Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2504.04022.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:49:49.869704Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-05T10:20:57.237283Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 6e0d4383-d98e-4329-bae7-b28ca559b054 · inbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Rethinking Reflection in Pre-Training
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation da3a0101-9dc7-4161-a81f-dab338fd9a5a · inbound
Prompting Large Language Models with Partial Knowledge for Answering Questions with Unseen Entities Rethinking Reflection in Pre-Training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5f341ca-6c69-4efc-91c7-bdbe9aabc616 · inbound
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration Rethinking Reflection in Pre-Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 08718af6-5d1d-47ea-9c59-ae0a53e26207 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Rethinking Reflection in Pre-Training
Reference 150
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecb10b1d-f835-47cd-87eb-b986a7e0759a · inbound
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Rethinking Reflection in Pre-Training
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9dc31577-1fef-4ed5-8cd1-189b099b35cf · inbound
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning Rethinking Reflection in Pre-Training
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a64fd7d-c393-4110-94f6-f1818ad00ff2 · inbound
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning Rethinking Reflection in Pre-Training
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 283fed06-b34b-4ef7-9aa5-28b5431b5e8f · inbound
On Advantage Estimates for Max@K Policy Gradients Rethinking Reflection in Pre-Training
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cecee2b2-8406-4f3c-8593-80399a10b817 · inbound
OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation Rethinking Reflection in Pre-Training
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.