Pith. sign in

Paper Citation Record · LEDGER

Rethinking Reflection in Pre-Training

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2504.04022.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.04022 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:49:49.869704Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-05T10:20:57.237283Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6e0d4383-d98e-4329-bae7-b28ca559b054 · inbound

Reinforcement Learning for Reasoning in Large Language Models with One Training Example cites this paper.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Rethinking Reflection in Pre-Training

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.097590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:4b910020de77ab80943f6a2a66cd20d8254a7aa69f72f7e4b1b743bba4c1f5ad

Observation da3a0101-9dc7-4161-a81f-dab338fd9a5a · inbound

Prompting Large Language Models with Partial Knowledge for Answering Questions with Unseen Entities cites this paper.

Prompting Large Language Models with Partial Knowledge for Answering Questions with Unseen Entities Rethinking Reflection in Pre-Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:49.869704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:49.869704Z digest=sha256:e925f5bf2e338e4ccd6829e2d76c2d742eb1dbae14390c8d09648131ad6fca9b

Observation c5f341ca-6c69-4efc-91c7-bdbe9aabc616 · inbound

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration cites this paper.

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration Rethinking Reflection in Pre-Training

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:36:53.752419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T22:33:01.074518Z digest=sha256:986518f0ecaa96e2dad93c686c784eb5994a1d1276bc45be65673ec823f3ed24

Observation 08718af6-5d1d-47ea-9c59-ae0a53e26207 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Rethinking Reflection in Pre-Training

Reference 150

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:38.823390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:38.823390Z digest=sha256:89abc43be556739adbce63ea8c1a62f7966ff8d9f62eea68408e877cf7bd6a37

Observation ecb10b1d-f835-47cd-87eb-b986a7e0759a · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Rethinking Reflection in Pre-Training

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:26:28.217277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T14:24:48.666197Z digest=sha256:c348def3fbad548700ef6945c839651943ce07e0f4ef6dd7a839d30eda935760

Observation 9dc31577-1fef-4ed5-8cd1-189b099b35cf · inbound

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning cites this paper.

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning Rethinking Reflection in Pre-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T12:52:27.595638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:52:27.595638Z digest=sha256:89dfc8f0381c45da97f775c6611ab81e09a7cb286c3e2ab8fe3e50153f57ec68

Observation 7a64fd7d-c393-4110-94f6-f1818ad00ff2 · inbound

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning cites this paper.

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning Rethinking Reflection in Pre-Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-05T10:20:57.239715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-05T10:18:18.717871Z digest=sha256:cb17827caa422cde9b2c90e27dc42bd855d8d04cb63ae3355de56067b8b0085c

Observation 283fed06-b34b-4ef7-9aa5-28b5431b5e8f · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients Rethinking Reflection in Pre-Training

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.418746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:094c48cdd99169916f81fad8ef516807b9c4d3951f670dce7372cd31ee563b2c

Observation cecee2b2-8406-4f3c-8593-80399a10b817 · inbound

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation cites this paper.

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation Rethinking Reflection in Pre-Training

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:56.961703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T02:17:30.974692Z digest=sha256:88d7aa69da4ac5c902f7f7e2234e8c00c5d8be48e83611298b09e900b73760fc