Pith. sign in

Paper Citation Record · LEDGER

RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2406.14532.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.14532 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:47:31.286943Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:57:41.298545Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0317369c-a3d0-4623-969c-eaef666a455c · inbound

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters cites this paper.

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:16:23.877962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:16:23.842092Z digest=sha256:10253efc6f6c7969a7909231fe1c3d63c31e60f1b925214ff942a48109d903d9

Observation 7b6fed2f-15ec-4979-9ef2-4a60d3cf925a · inbound

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents cites this paper.

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T09:42:04.217081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T09:41:59.979595Z digest=sha256:fe74e6d6972162c551edeecb272e00e731573993d8f817608a864b5bf580e628

Observation eb91be93-fdbe-4e6a-9726-dbe0f96d824a · inbound

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning cites this paper.

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:42:19.113640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T01:42:19.004468Z digest=sha256:2023983a7e919f091e9c55017c38234293908951eea0f656e1b0306952bef996

Observation f90cc920-7429-4ba0-b076-b4969716c1e6 · inbound

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search cites this paper.

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:31.286943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:31.286943Z digest=sha256:c6103cb035368c62e079e6d6d4b31de1de6e74d08492ad769e2335b467251823

Observation 2a82e362-277a-4120-910f-1cc2d62ec641 · inbound

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling cites this paper.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.027707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.027707Z digest=sha256:5c55f9db33594209d6c06fc70fe7d02424331cea49ee2c4f246f340db3e4b5a8

Observation 0206bbf1-b54a-4626-bcb5-a65cff803769 · inbound

InSTA: Towards Internet-Scale Training For Agents cites this paper.

InSTA: Towards Internet-Scale Training For Agents RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:24:50.406158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:24:50.406158Z digest=sha256:98d49d34dc306d0280a9308eeb791426564595b66319d4a7ca8946ef2e9053bd

Observation e1279313-417a-47ed-a1a8-01a313703505 · inbound

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning cites this paper.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.967705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.967705Z digest=sha256:245ab56c8f545ed44941793cbad1cac5317c26254681aad74de7e7df396429a4

Observation 1c36be57-9ab2-463b-a971-a828d05c21ad · inbound

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning cites this paper.

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:28:05.703259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T16:24:04.132572Z digest=sha256:b21db9bc9850cfaf11bae670ddaff0125983edef158b6253cb8e0d09bc968f96

Observation 200d5800-cc90-4686-a480-1fc7deb21f6f · inbound

Reinforcement Learning Improves LLM Accuracy and Reasoning in Disease Classification from Radiology Reports cites this paper.

Reinforcement Learning Improves LLM Accuracy and Reasoning in Disease Classification from Radiology Reports RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:16:04.302777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:09:13.352298Z digest=sha256:49eb9b48f8275ecd8469706f57e86f7865259fe9196c3b3d6f7ec0525584926f

Observation e14860d0-4933-4bff-86ee-ea75944d0428 · inbound

Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning cites this paper.

Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:07:08.108704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T02:03:18.320462Z digest=sha256:e3943040d4374b33e8c02d2b1eb87ad03dee74f161408d245cbf83f9c2352fbb

Observation 5ab74adf-c224-48a9-b97d-76eda0cb26d4 · inbound

Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning cites this paper.

Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:45:07.399654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:42:36.023338Z digest=sha256:127850bae24450c4b122ca9bed631980191677e57acd9ea70ee52f34b69384c6

Observation 2bc84b52-7025-4321-a992-b85bdbac54b7 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 204

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.300307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:2a8f5cf9c0893afbcbcf08e50409959073d23f4f6f67a79dd776ccda11ae1b88

Observation 22ac0092-d162-4cae-82a9-e23c83dbb28c · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:55.203968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:55.203968Z digest=sha256:1577f37a7227b264d2ddd474850456b7217254a8096510169bf7c3ec42fe7d6e