Pith. sign in

Paper Citation Record · LEDGER

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 6 inbound Pith citation observations for arXiv:2507.08267.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08267 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:26:08.536798Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:18:39.658589Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T22:43:37.868176Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b57bf2b3-6c37-4caf-9e90-3603c8f86b2f · outbound

This paper cites write newline.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.461183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.461183Z digest=sha256:60451236b4363eecdc68a5eadc96fe23c4c542dc8688c56d00067603a6aa11cc

Observation 64f6567c-e674-4803-a3b2-9a077cd0b661 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.464611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.464611Z digest=sha256:9c76cfbccb12de865013b8083178c6245b434c87ed1f5b495bd7cbe9c8435117

Observation 92ed019e-744e-403f-88e5-6f7c2f51ed3d · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.467197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.467197Z digest=sha256:b134908e16515e0055a975ed544a7593a887b8d3d3e496c38cca0bd5afab3911

Observation 3fe5eb46-0af6-40a4-9cdd-03814bc37907 · outbound

This paper cites Alphamath almost zero: Process supervision without process.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Alphamath almost zero: Process supervision without process

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.949427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.470895Z digest=sha256:b733ef3fb3ebd327cadfb446180c433afd73a35b2c8500a0bf45b9ca1bf34b14

Observation 7e909581-2699-444c-a2d5-827300282a19 · outbound

This paper cites and Ngo, C.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning and Ngo, C

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.473351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.473351Z digest=sha256:d7e8c299e13be1d32e5a7d16f95e59d734fb813abb16e78b20ef84724ebb7fa4

Observation 37aca5f0-c980-4c84-ae41-f2358fc22c83 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.475643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.475643Z digest=sha256:643e7ead4258b4849e45a2e05a8ff342824c51a65d71aff54aaeca04f51ed87b

Observation 580f2210-02d4-4d7c-aa92-f4e60816498c · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.477985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.477985Z digest=sha256:0a230540b1e2519a9997dd7c6e5e54a0398736f4c2d80adf647d49eb22a5b865

Observation f6258d02-8f85-4857-8d2d-d5583d8f9fdc · outbound

This paper cites C., Buzzard, K., Gowers, T., Liu, P.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning C., Buzzard, K., Gowers, T., Liu, P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.938412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.480618Z digest=sha256:df29617b4f4e05e2bbb846e82d266f4978c23cd2e2f3bbbe8a0671fce6967007

Observation f6dfa5ca-d1dd-4f0a-88ba-63076368c673 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Measuring mathematical problem solving with the MATH dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.931008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.483094Z digest=sha256:f010b873bc60438a2c388875a235d3acce987f1db37809abacf0350ebc2e1cd6

Observation 71208300-83d6-435a-8ee8-eb585246e9b9 · outbound

This paper cites Training Compute-Optimal Large Language Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Training Compute-Optimal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.485151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.485151Z digest=sha256:163e09cc9c2627289d0f81d4bc81ab326b364b0af8808b551654be6945f429fc

Observation 3bf86600-963a-49d0-b315-4e4c9a5dd9a4 · outbound

This paper cites C3ot: Generating shorter chain-of-thought without compromising effectiveness.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning C3ot: Generating shorter chain-of-thought without compromising effectiveness

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.924347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.488176Z digest=sha256:8e69d0dc7e6fbea98d69ef52acaf6a3409144e5cbb618d4bd1e4d48b57a56bcd

Observation 34d686bd-584d-427a-a481-721473558fc3 · outbound

This paper cites Scaling Laws for Neural Language Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Scaling Laws for Neural Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.490527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.490527Z digest=sha256:cdcb8261ae4923b38cfb82adb1d0de1e7765aae752409b180928c2f87aedf328

Observation 97a5c783-791f-435b-8a72-10a9db1142a4 · outbound

This paper cites Solving quantitative reasoning problems with language models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Solving quantitative reasoning problems with language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.917123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.492780Z digest=sha256:47b7ee41e396876e96267c537f118920afb63f788f3e9a008737988b698af0ef

Observation 86cb55e2-32e5-41a4-917c-00cdee1a533a · outbound

This paper cites Competition-level code generation with alphacode.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Competition-level code generation with alphacode

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.494732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.494732Z digest=sha256:cefc42ddc641fd0f6161637ead57de52682ce7728b9c47e53a08fc7a615f4ee7

Observation d67c5727-2112-4cbd-8b15-c01819965be2 · outbound

This paper cites Can language models learn to skip steps? In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Can language models learn to skip steps? In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.906089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.496800Z digest=sha256:384461586cf8e31b7a6b7dc3ea49d0b7a0287ccf9bc22c7ed9d3e0fdb47d46b5

Observation b6a561d6-ea7c-40c9-9b64-6c5474e58b1c · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.499132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.499132Z digest=sha256:bdfcd5e184dc17244676d8cdd7b82790e5c914455a30067210e5298dfc932efa

Observation b34e44a7-7b90-4f77-b857-fadfa7d5e00a · outbound

This paper cites Wider or deeper? scaling llm inference-time compute with adaptive branching tree search.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Wider or deeper? scaling llm inference-time compute with adaptive branching tree search

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.501675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.501675Z digest=sha256:682ecf699845f5f7399bc0ad0179560d49a2b3991f180d2929a27ab0e9c2f1a0

Observation 540f1a2e-80f3-4915-ab92-4d4b1f879a47 · outbound

This paper cites s1: Simple test-time scaling.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning s1: Simple test-time scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.503965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.503965Z digest=sha256:7bdb28771ce7c544d7cb7f61bed16c5893c274f6145bc509e720fd65aa733810

Observation 2a50dcea-77aa-48c6-b002-3b3d8b4a68ee · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Self-Training Elicits Concise Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.506610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.506610Z digest=sha256:ed0cfddae02fefd31ec082e82e27cf471ca5fcaedc3a68e0bac56aa055a7cb9b

Observation 58b1e904-03b7-4c56-81ca-614ffd5da247 · outbound

This paper cites OpenAI o1 System Card.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning OpenAI o1 System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.509681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.509681Z digest=sha256:ef68fa61e7f4d6df6c955f676315432393131a8555a59384c0e448f1578bd754

Observation 8bd696b0-af35-491a-bd05-22225f5cf6a0 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Competitive Programming with Large Reasoning Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.512634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.512634Z digest=sha256:fa94cb26936b2ed11797425849f27a170f239a69a2cbb425f2681be205632ed8

Observation 02302a27-eca4-4170-b3c1-f439d6e8558b · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.898157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.515249Z digest=sha256:a733de3d37f006c8aba9599ba06f6153abc478451bc52c425ed39eb8bbabdc2b

Observation 8615f2d8-2c23-4121-9b1b-6851900bd419 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.517455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.517455Z digest=sha256:e386289bc0432c6902b9e3ebb787735ff9de92636a50eb0249358c004a293fe7

Observation c9e81936-ede2-4c91-8adf-4e591b2ceaf4 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.520090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.520090Z digest=sha256:62bd346c21009ccf589aa6e4cc83b8eb39887a9e18e66fbb16bf8c61225abc73

Observation aa33f8c4-dc5e-4f4f-b329-bcc927e8dd22 · outbound

This paper cites H., Le, Q.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning H., Le, Q

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.522615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.522615Z digest=sha256:20a5ec2c0d1e3bd224acaa3544c232dc9606bcf105740dec2e34bbd65a488d44

Observation 6c52c4c3-7204-4084-8ba8-a61935fa4f62 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.525387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.525387Z digest=sha256:b57f1e43524ff9533c8ba3aae05218a27501c1a51ac1c3d8610d7d6211b8a2a9

Observation 2df7098c-9852-4f97-8a62-9c8852132e61 · outbound

This paper cites Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.885946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.527957Z digest=sha256:c2656ef6ec4e5c96df6691e5149af5d12dc07e85072196c818bce7a32d63bb8d

Observation 21d4e395-d743-4d01-aa42-52e48ab6cfea · outbound

This paper cites T., Wang, W., and Li, W.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning T., Wang, W., and Li, W

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.530115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.530115Z digest=sha256:817ef9cd6e8e9f79eef6d825462972be051d1d8d86f1e6275649b0f7f67d955c

Observation 8d0d3622-a270-4a90-b1b8-f81966e32d0a · outbound

This paper cites LIMO: Less is More for Reasoning.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning LIMO: Less is More for Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.532302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.532302Z digest=sha256:4be915cd6b13cd8cdc3e9454806413373428bfe218d6ff50485d576fb2bbf529

Observation cb22ee4b-3048-4dc8-8a0d-33c50d00818a · outbound

This paper cites Demystifying long chain-of-thought reasoning in LLM s.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Demystifying long chain-of-thought reasoning in LLM s

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.878484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.534689Z digest=sha256:ba4edc0a908c0d53b510f837939c47253a09680f72e3b8c3e45c1054ee4bbf46

Observation 432f3e6a-5dd0-4835-afcf-14f230158f4b · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.536798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.536798Z digest=sha256:465f5be64ad38c33f01e358b564ee82c95cdf91afd15a6e7727cbd24b00f6aeb

Pith citing papers

Observation cc5cb283-56ab-47d4-a8ee-94f90b6acb05 · inbound

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning cites this paper.

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:43:37.870506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T22:43:01.937642Z digest=sha256:169ddf3c6faf14e64604cffb3fbc71a7d2893af86efa26c56f8ea35685c30d86

Observation 73d7339c-e9ef-4e58-b417-a13640e62122 · inbound

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings cites this paper.

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:45:37.299666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T12:44:50.752937Z digest=sha256:24beb7cb929d7e5c72303ec2c7f3690543d60be4f5fe6c2c95f8e6b4c9c6f79b

Observation d0d6a44c-6a40-42d8-bd33-517004870c9c · inbound

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning cites this paper.

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:53.376328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:45:40.915428Z digest=sha256:50b18534459f6452bc4bf3f0c52f95109e902925922e88bd4203b33a45e9b2b6

Observation 4757bf6b-42e1-436c-812b-69a03b3cd5cf · inbound

Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning cites this paper.

Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:06:00.003903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:49:26.829527Z digest=sha256:ea237df72fd3e8c74615acd5fd56492de092008e928e81a61465960069f4d550

Observation b7e02888-5225-4b44-9d33-6f18a8d8d872 · inbound

CRAFT: Learn the Schema, Execute the Plan cites this paper.

CRAFT: Learn the Schema, Execute the Plan A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T10:18:39.658589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:18:39.658589Z digest=sha256:2612cca32f357f3a17ef64d72c9568e3a10c093f0137cd53b161efebee871280

Observation 94b467e5-494d-45ae-ab3b-100edb5b663d · inbound

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding cites this paper.

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T21:13:06.724007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:13:06.724007Z digest=sha256:7c3f362a68e63bc26040d10222dce1cbdc3ba46b170861a15b093723826ca5b9