Pith. sign in

Paper Citation Record · LEDGER

RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2406.14532.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.14532 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:47:31.286943Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:57:41.298545Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0317369c-a3d0-4623-969c-eaef666a455c · inbound

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters cites this paper.

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:16:23.877962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:16:23.842092Z digest=sha256:e86b357aa04d3578975f45b35730489d29cdb8966af0761e10c72d02fae0ceb4

Observation 7b6fed2f-15ec-4979-9ef2-4a60d3cf925a · inbound

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents cites this paper.

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T09:42:04.217081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T09:41:59.979595Z digest=sha256:2e619e63e9255c85d6126b31b2ffec641c9882432fe7e93df6340adfe081cad5

Observation eb91be93-fdbe-4e6a-9726-dbe0f96d824a · inbound

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning cites this paper.

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:42:19.113640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T01:42:19.004468Z digest=sha256:bf3c04d02d6fa73351a43e309ae5036aa487fdffd3701eb74cb0b693974e5fb1

Observation f90cc920-7429-4ba0-b076-b4969716c1e6 · inbound

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search cites this paper.

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:31.286943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:31.286943Z digest=sha256:c6103cb035368c62e079e6d6d4b31de1de6e74d08492ad769e2335b467251823

Observation 2a82e362-277a-4120-910f-1cc2d62ec641 · inbound

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling cites this paper.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.027707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.027707Z digest=sha256:5c55f9db33594209d6c06fc70fe7d02424331cea49ee2c4f246f340db3e4b5a8

Observation 0206bbf1-b54a-4626-bcb5-a65cff803769 · inbound

InSTA: Towards Internet-Scale Training For Agents cites this paper.

InSTA: Towards Internet-Scale Training For Agents RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:24:50.406158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:24:50.406158Z digest=sha256:98d49d34dc306d0280a9308eeb791426564595b66319d4a7ca8946ef2e9053bd

Observation e1279313-417a-47ed-a1a8-01a313703505 · inbound

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning cites this paper.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.967705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.967705Z digest=sha256:245ab56c8f545ed44941793cbad1cac5317c26254681aad74de7e7df396429a4

Observation 1c36be57-9ab2-463b-a971-a828d05c21ad · inbound

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning cites this paper.

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:28:05.703259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T16:24:04.132572Z digest=sha256:8c8b7afc8eae385c9042c7d58506686696ebce96f3133c177333ef2b03fe0d65

Observation 200d5800-cc90-4686-a480-1fc7deb21f6f · inbound

Reinforcement Learning Improves LLM Accuracy and Reasoning in Disease Classification from Radiology Reports cites this paper.

Reinforcement Learning Improves LLM Accuracy and Reasoning in Disease Classification from Radiology Reports RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:16:04.302777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:09:13.352298Z digest=sha256:2f218701bbc6cb0e751dd9a121f86dcc9852238a2947a4ca55b88835f07298cf

Observation e14860d0-4933-4bff-86ee-ea75944d0428 · inbound

Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning cites this paper.

Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:07:08.108704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T02:03:18.320462Z digest=sha256:e133e33ae7d5b3dbf3dfa4fe4b2d08df3934566cd7e88ca75e57dc122cf1c04e

Observation 5ab74adf-c224-48a9-b97d-76eda0cb26d4 · inbound

Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning cites this paper.

Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:45:07.399654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:42:36.023338Z digest=sha256:8a98b3794cace79ec692acab37f8bb5a369014ff5f5cc86ac483b5b43c1aee2d

Observation 2bc84b52-7025-4321-a992-b85bdbac54b7 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 204

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.300307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:4bbd4aa9a1e9e73c18eb5b3727809a69ac0cc023978137ba1ffed8eb90bda079

Observation 22ac0092-d162-4cae-82a9-e23c83dbb28c · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:55.203968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:55.203968Z digest=sha256:1577f37a7227b264d2ddd474850456b7217254a8096510169bf7c3ec42fe7d6e