Pith. sign in

Paper Citation Record · LEDGER

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

As of 14 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 6 inbound Pith citation observations for arXiv:2507.08267.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08267 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:26:08.536798Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:18:39.658589Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T22:43:37.868176Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b57bf2b3-6c37-4caf-9e90-3603c8f86b2f · outbound

This paper cites write newline.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.461183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.461183Z digest=sha256:60451236b4363eecdc68a5eadc96fe23c4c542dc8688c56d00067603a6aa11cc

Observation 64f6567c-e674-4803-a3b2-9a077cd0b661 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.464611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.464611Z digest=sha256:9c76cfbccb12de865013b8083178c6245b434c87ed1f5b495bd7cbe9c8435117

Observation 92ed019e-744e-403f-88e5-6f7c2f51ed3d · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.467197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.467197Z digest=sha256:b134908e16515e0055a975ed544a7593a887b8d3d3e496c38cca0bd5afab3911

Observation 3fe5eb46-0af6-40a4-9cdd-03814bc37907 · outbound

This paper cites Alphamath almost zero: Process supervision without process.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Alphamath almost zero: Process supervision without process

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.949427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.470895Z digest=sha256:2f1b0235ed4e5aedc0ca6385c157e9d0158813939b8636f757dc603148b5e846

Observation 7e909581-2699-444c-a2d5-827300282a19 · outbound

This paper cites and Ngo, C.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning and Ngo, C

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.473351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.473351Z digest=sha256:d7e8c299e13be1d32e5a7d16f95e59d734fb813abb16e78b20ef84724ebb7fa4

Observation 37aca5f0-c980-4c84-ae41-f2358fc22c83 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.475643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.475643Z digest=sha256:4730d8b24634e4ddbe49305808a57a08af8e8898aebb2b10e3401240985044d6

Observation 580f2210-02d4-4d7c-aa92-f4e60816498c · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.477985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.477985Z digest=sha256:0a230540b1e2519a9997dd7c6e5e54a0398736f4c2d80adf647d49eb22a5b865

Observation f6258d02-8f85-4857-8d2d-d5583d8f9fdc · outbound

This paper cites C., Buzzard, K., Gowers, T., Liu, P.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning C., Buzzard, K., Gowers, T., Liu, P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.938412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.480618Z digest=sha256:670ee715591b20f4b05f355c28bedd8ce3618d2f840ed3f3fa1ab1db3d2e90e7

Observation f6dfa5ca-d1dd-4f0a-88ba-63076368c673 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Measuring mathematical problem solving with the MATH dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.931008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.483094Z digest=sha256:731a571dab4a50d32e71049f0aec38ba11ffe274805f28089ee428b8cb731848

Observation 71208300-83d6-435a-8ee8-eb585246e9b9 · outbound

This paper cites Training Compute-Optimal Large Language Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Training Compute-Optimal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.485151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.485151Z digest=sha256:65f28a4d8f27c530fd043864421b5a39078d17f9f96c562b99f61241afc28daa

Observation 3bf86600-963a-49d0-b315-4e4c9a5dd9a4 · outbound

This paper cites C3ot: Generating shorter chain-of-thought without compromising effectiveness.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning C3ot: Generating shorter chain-of-thought without compromising effectiveness

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.924347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.488176Z digest=sha256:4770e5a355499cbdf3f2421f5f1c5d16c28dd94870ebec273bfaae98c2bb04a5

Observation 34d686bd-584d-427a-a481-721473558fc3 · outbound

This paper cites Scaling Laws for Neural Language Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Scaling Laws for Neural Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.490527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.490527Z digest=sha256:8393bbda848de174ff2713b81ebbb23c90e1134a3eb762877a4a343229c765be

Observation 97a5c783-791f-435b-8a72-10a9db1142a4 · outbound

This paper cites Solving quantitative reasoning problems with language models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Solving quantitative reasoning problems with language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.917123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.492780Z digest=sha256:364d8704df193482d1dffe1c31c8451ae4d33c2c52f38b66d0d17ca7f27fa7c1

Observation 86cb55e2-32e5-41a4-917c-00cdee1a533a · outbound

This paper cites Competition-level code generation with alphacode.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Competition-level code generation with alphacode

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.494732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.494732Z digest=sha256:cefc42ddc641fd0f6161637ead57de52682ce7728b9c47e53a08fc7a615f4ee7

Observation d67c5727-2112-4cbd-8b15-c01819965be2 · outbound

This paper cites Can language models learn to skip steps? In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Can language models learn to skip steps? In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.906089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.496800Z digest=sha256:07de851f16a7a0d8231dae97dc798735d3b40078a48ca3dbdf7d97ad388b5b0d

Observation b6a561d6-ea7c-40c9-9b64-6c5474e58b1c · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.499132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.499132Z digest=sha256:f59e857ed3961f89764248b12e3fd23bdffb59347083de9b3e0c20683f536340

Observation b34e44a7-7b90-4f77-b857-fadfa7d5e00a · outbound

This paper cites Wider or deeper? scaling llm inference-time compute with adaptive branching tree search.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Wider or deeper? scaling llm inference-time compute with adaptive branching tree search

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.501675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.501675Z digest=sha256:682ecf699845f5f7399bc0ad0179560d49a2b3991f180d2929a27ab0e9c2f1a0

Observation 540f1a2e-80f3-4915-ab92-4d4b1f879a47 · outbound

This paper cites s1: Simple test-time scaling.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning s1: Simple test-time scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.503965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.503965Z digest=sha256:7bdb28771ce7c544d7cb7f61bed16c5893c274f6145bc509e720fd65aa733810

Observation 2a50dcea-77aa-48c6-b002-3b3d8b4a68ee · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Self-Training Elicits Concise Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.506610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.506610Z digest=sha256:ed0cfddae02fefd31ec082e82e27cf471ca5fcaedc3a68e0bac56aa055a7cb9b

Observation 58b1e904-03b7-4c56-81ca-614ffd5da247 · outbound

This paper cites OpenAI o1 System Card.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning OpenAI o1 System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.509681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.509681Z digest=sha256:ef68fa61e7f4d6df6c955f676315432393131a8555a59384c0e448f1578bd754

Observation 8bd696b0-af35-491a-bd05-22225f5cf6a0 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Competitive Programming with Large Reasoning Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.512634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.512634Z digest=sha256:1cbbb2ac7996d9a0672d87fb36c38dcf6cb998e6217167f365e23ab14d01d1a9

Observation 02302a27-eca4-4170-b3c1-f439d6e8558b · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.898157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.515249Z digest=sha256:c303a7c0db5cf2667c39113e16353b88c5f83c7f1f543db9c4c37f7e1648e0d7

Observation 8615f2d8-2c23-4121-9b1b-6851900bd419 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.517455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.517455Z digest=sha256:e386289bc0432c6902b9e3ebb787735ff9de92636a50eb0249358c004a293fe7

Observation c9e81936-ede2-4c91-8adf-4e591b2ceaf4 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.520090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.520090Z digest=sha256:5eb67ef5e12ae386be82ed03c71432f0a61aa2a6e5789674846f66be2a6d71b1

Observation aa33f8c4-dc5e-4f4f-b329-bcc927e8dd22 · outbound

This paper cites H., Le, Q.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning H., Le, Q

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.522615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.522615Z digest=sha256:20a5ec2c0d1e3bd224acaa3544c232dc9606bcf105740dec2e34bbd65a488d44

Observation 6c52c4c3-7204-4084-8ba8-a61935fa4f62 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.525387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.525387Z digest=sha256:2192242b07e3a44f7100d897bb441f42a17a715017e6f263650f0237ed657efa

Observation 2df7098c-9852-4f97-8a62-9c8852132e61 · outbound

This paper cites Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.885946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.527957Z digest=sha256:bd619c57a3dca285b4e482956b9dd6df55656bf3d7759b9f05d245cf58fea803

Observation 21d4e395-d743-4d01-aa42-52e48ab6cfea · outbound

This paper cites T., Wang, W., and Li, W.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning T., Wang, W., and Li, W

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.530115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.530115Z digest=sha256:817ef9cd6e8e9f79eef6d825462972be051d1d8d86f1e6275649b0f7f67d955c

Observation 8d0d3622-a270-4a90-b1b8-f81966e32d0a · outbound

This paper cites LIMO: Less is More for Reasoning.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning LIMO: Less is More for Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.532302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.532302Z digest=sha256:4be915cd6b13cd8cdc3e9454806413373428bfe218d6ff50485d576fb2bbf529

Observation cb22ee4b-3048-4dc8-8a0d-33c50d00818a · outbound

This paper cites Demystifying long chain-of-thought reasoning in LLM s.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Demystifying long chain-of-thought reasoning in LLM s

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.878484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.534689Z digest=sha256:bb85a0fd086234ab63776b1aca294f35a22f6a50236d7b481839e7b3f9013872

Observation 432f3e6a-5dd0-4835-afcf-14f230158f4b · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.536798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.536798Z digest=sha256:465f5be64ad38c33f01e358b564ee82c95cdf91afd15a6e7727cbd24b00f6aeb

Pith citing papers

Observation cc5cb283-56ab-47d4-a8ee-94f90b6acb05 · inbound

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning cites this paper.

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:43:37.870506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-16T22:43:01.937642Z digest=sha256:765724051692be4442306b4dfa2891d8fa7c590aca6c77ca5fedc03fdf00f4cd

Observation 73d7339c-e9ef-4e58-b417-a13640e62122 · inbound

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings cites this paper.

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:45:37.299666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T12:44:50.752937Z digest=sha256:df983747c7e8df6dd33ecc9f5983330872199430c0e0fd9669ba1c3311e04cda

Observation d0d6a44c-6a40-42d8-bd33-517004870c9c · inbound

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning cites this paper.

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:53.376328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T19:45:40.915428Z digest=sha256:35d4ddea8488a0501954ef043da350ed7c0420e8c588dde795f03f51fb7b6f75

Observation 4757bf6b-42e1-436c-812b-69a03b3cd5cf · inbound

Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning cites this paper.

Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:06:00.003903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:49:26.829527Z digest=sha256:36d1fdf3c8f54c00acad94015f3021a0ee8207da3a4a31499a045e1960f1b3c4

Observation b7e02888-5225-4b44-9d33-6f18a8d8d872 · inbound

CRAFT: Learn the Schema, Execute the Plan cites this paper.

CRAFT: Learn the Schema, Execute the Plan A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T10:18:39.658589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:18:39.658589Z digest=sha256:2612cca32f357f3a17ef64d72c9568e3a10c093f0137cd53b161efebee871280

Observation 94b467e5-494d-45ae-ab3b-100edb5b663d · inbound

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding cites this paper.

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T21:13:06.724007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:13:06.724007Z digest=sha256:7c3f362a68e63bc26040d10222dce1cbdc3ba46b170861a15b093723826ca5b9