Pith. sign in

Paper Citation Record · LEDGER

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code

As of 18 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2507.07498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07498 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:44:00.020826Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T08:40:40.910461Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:40:41.143168Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e97d3ca6-0613-403f-a167-1f9d2740b4e0 · outbound

This paper cites Program Synthesis with Large Language Models.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Program Synthesis with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:56.513180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:56.513180Z digest=sha256:f0f9c0ae3700168ca9864ca619e9c296d259a731d851ba77efed8dce86ce1909

Observation 3498fbf5-206a-4956-b0ee-4d4b3abd7f41 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Evaluating Large Language Models Trained on Code

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:56.595227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:56.595227Z digest=sha256:f86aa912b3d4088f1d5ce902310f8c81064602e98505a0a6c737736982d95883

Observation f7db6b0e-5d22-48a2-b67e-999a3b922f9f · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:56.693037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:56.693037Z digest=sha256:707ea6f190f5cbb7b2d78a94ef40dffc4e13580293af6e0990e88531d1861dc8

Observation f3e74ca6-911d-4584-919d-901d130cc690 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.702849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:56.789538Z digest=sha256:efef4ec13eebf3bfa788a35eece61b95a54ecdd8d09724f96dbbaac9da856b97

Observation a182342f-0939-4216-a5ad-c4014354a35f · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:56.867692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:56.867692Z digest=sha256:c5ac835987d172c1f10215fb8e9776ed2d05fd6850105756b5e955a27f680d79

Observation be213f86-508b-4eec-bd9d-7ce01221e1b0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:56.929622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:56.929622Z digest=sha256:4b02079d1e5a3d85017d29777d3dcf26e5521f379a99e54c40663327af03fff3

Observation d7be902d-f847-4705-b203-fc78ebf9373e · outbound

This paper cites Min, Gail E.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Min, Gail E

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.001314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.001314Z digest=sha256:5e4759392c0843dab84fa78bd10ea1bb5c7a3fad37724d32a7f07cc8f83e3f84

Observation 00ba49be-2db1-48ac-bad9-5f6480f21142 · outbound

This paper cites SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.092562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.092562Z digest=sha256:fcd8f6df91564d25eb28effbc693255860be07014a01f276e7b3c2ab2467af77

Observation c2b5160f-9463-4fb9-82bb-23e5fda2d77e · outbound

This paper cites Min, Gail E.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Min, Gail E

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:44:01.615383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:57.183728Z digest=sha256:cbfa09b91351078c711998982673b44fe21a7b5afa49cd060efa0983a9c8f7c0

Observation 63d3f85b-8bcb-49b3-acb4-4b9c2a70d255 · outbound

This paper cites Kaiser, Wei Le, and Baishakhi Ray.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Kaiser, Wei Le, and Baishakhi Ray

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.248831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.248831Z digest=sha256:5c47b28e909fd9709f177e1846c87db6b3b558460614dc587a4731d977381ab1

Observation f8999b04-aa9a-4d01-b602-c9f5584ecc5c · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.306817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.306817Z digest=sha256:60b9fe9bc4437c97445fdb354f4d58e123c0658ec508804488e2945f50310024

Observation 9933a054-1d12-4da6-8fde-87d0afd5cedc · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.538429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:57.360557Z digest=sha256:f7706038139628aa9971090b07c62c8e8a94a38ec4a3e3a28beccd9634160eb3

Observation 9ee7c40b-946e-46c1-9991-3dda6ac8a2ae · outbound

This paper cites Are We Done with MMLU?.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Are We Done with MMLU?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.410818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.410818Z digest=sha256:9e2567dc403d991baf6d45f7510585dbf8ae907e52d476fb9819087f87d0d39a

Observation 2524e37f-3834-445b-ad3f-2766ed9f426a · outbound

This paper cites Pseudocode-Injection Magic: Enabling LLMs to Tackle Graph Computational Tasks.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Pseudocode-Injection Magic: Enabling LLMs to Tackle Graph Computational Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.462689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.462689Z digest=sha256:288f6629513498721bec98150c08f037e7a4f8cd9c0f988a87b772369d85dbce

Observation 7ea1aca0-17cf-44ce-bab3-d6912a0b2bc1 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.449710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:57.510655Z digest=sha256:1b09f9387e69993703d410c8e5334736c76eff25c162fdb5df0c638e1e5a1a81

Observation 89b15af0-e0ae-43b1-876b-86fffe9e5918 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.355905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:57.558937Z digest=sha256:f4a7f2638f7b696a5d70557dc79c4a9b66a2e1ee8e6a88f448ef4bc8281457ed

Observation b770dbb4-18f9-4cd7-8879-0dc42f6fe197 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.608469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.608469Z digest=sha256:0ade5d9231ef918c155c26b46c5fb84f385297b1dc1a576b93997e7a69d54016

Observation e5066ef0-7bd0-4cd0-84fa-93f0e1316849 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.651700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.651700Z digest=sha256:5d6ba9aaee89bc6a6e1c98cb77880fe98e85873b4504a6dd3993bdf5af9cb3df

Observation 81f6f63a-f19c-4e28-9807-46b81b6a8dbd · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Measuring Mathematical Problem Solving With the MATH Dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.708756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.708756Z digest=sha256:99a3765ec71f8d1b7485a2c4bfca8239bcddf61cbf1b4303002d4e6807433804

Observation 4a38939c-a5c4-4976-952f-08e297d0c85a · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.785882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.785882Z digest=sha256:d41bccc88a46670ba610f5f779895430911b6c5e0b3ef3bbe15efbdab53541ff

Observation 0f54101d-8a92-4ba9-ae5e-9ccf6a4cad05 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Qwen2.5-Coder Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.839394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.839394Z digest=sha256:a7114bf13166b8fc78266946b53ce827916df6887d203bcc32d41b57e95942c6

Observation c01f51ce-a523-47e5-ba9f-cdb2e14be61a · outbound

This paper cites OpenAI o1 System Card.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code OpenAI o1 System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.895612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.895612Z digest=sha256:e11b36f14ce78e6fb6b35702b4b5070658c68f596e96c94cb8bccd1a1c97f1c4

Observation fc6c8bd9-bc18-447b-8b49-db539a07e775 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.969802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.969802Z digest=sha256:64c1d0f20b24ad4a71cce5dd6ded167016313bfdd42b4c9b2db9f7f5c58ef8c4

Observation f775d378-1c6f-41d0-93b1-f1171f34f13f · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.283436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:58.020560Z digest=sha256:ebab622cd0fb1550f869203fdd61955d2470f59c06b4b218473f8f1bc5c5f148

Observation 493bf569-5162-4e7c-a312-81eb33bb145a · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.069083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.069083Z digest=sha256:fd833ad1537862e4ee84a26cdda4cfac53429f8d01a16e28d4830b4ca9f836f7

Observation 5ca038d3-4070-4412-93fe-258d7d9f6090 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.194426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.194426Z digest=sha256:8a629918b733e4525602b11215e8a6c19c1326ddd9d8c153a26de1ad1b202452

Observation 705bca43-4958-4332-8364-62c8169b340e · outbound

This paper cites CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.301904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.301904Z digest=sha256:96139ff1846851a084ab731a14b9f59e281c92ca7127297948eb7291076528f5

Observation c8d37bcf-d2fd-4899-ae03-7805c3e87a0c · outbound

This paper cites MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.430688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.430688Z digest=sha256:dfe0d62ca2fc59554d7fb0aa73c29cf78a005d3d4611887792ff600577506e06

Observation abe9e047-01dc-4272-a012-e5869a9c0d20 · outbound

This paper cites HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.486504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.486504Z digest=sha256:d41a933432b0054b06e5f05e25878eaadce3887332b3ed5da82ace1001cecd27

Observation cf41a461-6ba7-4a9d-97a6-6c755177b0fc · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.158488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:58.628898Z digest=sha256:9bf2c0f86f3dcce875ca8715672d756a01a665f8fd119b953942520b99c2f1bb

Observation 351480f6-b042-44dc-a133-69877045730e · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.771660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.771660Z digest=sha256:d55906022d702fcf9a758ab611fc33f9fe2ead4b9fc563ee574a86ad9d205781

Observation 711345b0-5fcf-4d58-aae5-ae6f48db94f9 · outbound

This paper cites On the Impact of Fine-Tuning on Chain-of-Thought Reasoning.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code On the Impact of Fine-Tuning on Chain-of-Thought Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.869826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.869826Z digest=sha256:22540733461934d4147d59c3d6b3d387c87203f177f342c39337ba98c0d294ae

Observation 43a09bcd-2a34-4758-b151-876c4c53b183 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.929656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.929656Z digest=sha256:e57ecc838a02fa568e2e08083663eb8e22e57d5a350ae9d35c11d9e471ea911f

Observation b02a2f70-fb09-44ee-a3ad-ec2234bab7ff · outbound

This paper cites KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.987141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.987141Z digest=sha256:cd6f4150fa3ed363bd69ef7ab29a1432936ab8b3edba870be9d471bc9d1890f2

Observation 0270f618-f1c0-4595-84dd-2b347d0ec3fc · outbound

This paper cites Qwen2.5 Technical Report.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Qwen2.5 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.046822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.046822Z digest=sha256:5eada69fc61d34eeb8aa54edb75506aad311465c81adf29ca338eaca15e5c0c3

Observation 17feeeea-e0d2-4784-99ce-83f14916907e · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.109983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.109983Z digest=sha256:1b542b9a940a626e7a959c7392ed6015ed48ca37d069710e1bd6ad36ccf31225

Observation 67fb2b71-9929-4c31-8644-0cc18dac0ee4 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.155801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.155801Z digest=sha256:068979bf81027fa1b2bf44f03db33b450d5c570c352c8300155f74d9143c8851

Observation 2b48e226-7534-4af7-a9c1-1b5fcc0ea1a2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.205671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.205671Z digest=sha256:fa9b8c5b94d6f4b6fcd10d894c20c300e1b085d05130849610e55726b1945d56

Observation 62d1e1e1-8df0-4206-9df7-212aa93f435d · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.049327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:59.267078Z digest=sha256:fc0ca10b637af5f40b4ea82f4a9845fa73e7012dd1560584664337c230ade3d9

Observation 4c774638-4500-4e1d-beac-2ba61ae4e4b4 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.311348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.311348Z digest=sha256:39386d06ac736df9ce3b49e2acf35c3e6fe6631c38cbd1d176bc96765f37265d

Observation d7dde5ff-939b-4562-b0e6-95af6b34ff72 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.388088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.388088Z digest=sha256:0c5da1a1f13d2d353c7ee4613bfff41f501d691512cc48fa5a697452d05d3a94

Observation ff4a753f-d84a-4aaf-b0b8-e6b537bc62dc · outbound

This paper cites OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.471326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.471326Z digest=sha256:57b2ccf88523080d8809e9d6c52a0c6d9bb4c24840d08bd871efa60add497688

Observation 88366e24-9723-48b8-933d-b688da49350d · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:00.969329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:59.551743Z digest=sha256:078f950b682ab3256281ccf80aee3fc835e5c1c483d5e0ddd68dc8c971c3922e

Observation 38559dbf-f1ae-4979-aae0-abcdcdc53f6f · outbound

This paper cites MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.614488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.614488Z digest=sha256:9039d6a8dd665e0ba94bd022d2ad3dc07eb482f7d723498a015bdfe191c23445

Observation db14ff2f-6586-4fba-a676-ab1b1c522760 · outbound

This paper cites Le, Ed H.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Le, Ed H

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:44:00.892980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:59.669231Z digest=sha256:8813da1e07fbab5e1c8229aa7c7e55e9ccb62d3a035427520596ae82a44d3a06

Observation 2e0b8cd6-af3b-40a5-8efa-256c0de4201b · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:00.785136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:59.719102Z digest=sha256:8a8bf21379ab589275472f89f743b74ee5922ab5404bc770047ed89e3b21b8c1

Observation bf2044ff-eb66-4004-bf51-d582acf36600 · outbound

This paper cites Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:44:00.703241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:59.764041Z digest=sha256:d512f9894b1aea40838adad6b199226e44177e833163cd05058167f5a5d9e5a9

Observation 92874828-cc70-4d5e-b29b-aec626613d50 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:00.584735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:43:59.808781Z digest=sha256:bc9d83d42d8a5629ceb6d5b3efadafa4c972b2902f6e4f1244eaada4cfe803af

Observation b8b0e9da-da06-47ad-8c4d-09724995ea1a · outbound

This paper cites LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.848050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.848050Z digest=sha256:465f0c78081ca7e253cd68d9b04314bf00d80feb6fe390830adad6d61304434a

Observation 99a139d5-32ad-498a-880e-a8739508a026 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.893251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.893251Z digest=sha256:6b0389f28201bb9b70a363847c24617344ef63a4f1131fcfb3145ed34a83197a

Observation 498cf9e9-8b5a-42f6-9848-f9b4f31190ec · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.932862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.932862Z digest=sha256:0c78638640c0d9cfcf788fba87fb46e219f5ea793384b1970f50c7f3b1e398c6

Observation 536563bd-f0ae-4a38-ac31-fe7cc09dc943 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.959396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.959396Z digest=sha256:b2d180e3fc2a29b586dbfb508fdd7c4ca761ca7e4c6cc49e3f06d3186a94ef67

Observation 85f4a8c5-380c-40c7-ac76-6aedd63ed9e2 · outbound

This paper cites Learning to Execute.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Learning to Execute

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.965594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.965594Z digest=sha256:18acf190936ed5ff64fb008355d615b9f4c1b3d577126e535fe652667174da68

Observation f7351096-4071-4a91-8b50-e1129681d1ba · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.970269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.970269Z digest=sha256:8b8bd514a7877d213476baf6f9931c0206723039f07716e74ccb3e993794de90

Observation a5e5406b-ede2-49b5-b534-889a281da0ac · outbound

This paper cites online" 'onlinestring :=.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code online" 'onlinestring :=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.987191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.987191Z digest=sha256:638d9f721689feb0b537a507c5135ae38d1fa1cfa63c98c50a68b5524a07ecc4

Observation 9cff075c-7ed1-42d4-929a-396445ae561e · outbound

This paper cites write newline.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code write newline

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:44:00.020826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:44:00.020826Z digest=sha256:f0b8d78f913eaf78b5f6c458685748fac3cadd054bc7760c3301a2f5f62d9cc0

Pith citing papers

Observation bd083f39-b368-43e4-831b-cbe8ff80f22e · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.146090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:ff2d7972f2ce5ad4f900f73435fff77fdf6477675f8896d91e4e1e38ded86f72