Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:16:03.097264Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 14 inbound Pith citation observations for arXiv:2506.17211.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:16:03.097264Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:58.173572Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:29:57.036901Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8b615c96-ca91-46e9-918f-b020b315d4df · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Phi-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bff033ec-fa38-4fdf-be7e-a22e310e3686 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ec446cf-31e8-43e2-9d83-7930ac67c78b · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Hindsight experience replay.Advances in neural information processing systems, 30, 2017
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40101499-a613-40ac-80cf-f556f124128c · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Thinking fast and slow with deep learning and tree search.Advances in neural information processing systems, 30, 2017
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 53069f5d-a754-4445-a10c-53637e29997d · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90e7283d-54d3-4a8e-a0c5-0a7f978140b2 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46184443-4e7b-44fa-b5d6-6cb689cf7460 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Cambridge university press, 2019
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b16e010d-9345-4da8-8280-419adc28b9b7 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning First return, then explore.Nature, 590(7847):580–586, 2021
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb6da3df-2c18-4ef9-b9d3-fb78a3111f55 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95f2eeb4-fba4-4c28-b47d-64dbf3543f6e · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning John Wiley & Sons, 1991
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dffc416-524d-4e91-abcb-c34974ad2785 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Gemini 2.0 flash thinking mode (gemini-2.0f lash-thinking-exp-1219), 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c63d430d-3448-4853-9228-fb9d71aa8f35 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Last updated 14 May 2025
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 60bd52a3-f3e7-4bf1-b07b-3ae0b83e56ce · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 887f31d0-d506-47fd-872b-fb9905231a5a · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Language model cascades: Token-level uncertainty and beyond
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8b12159c-15aa-4e5f-9e03-c6baea6a4a4a · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64f3349c-3e61-4df2-9aab-ae5b96d69797 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Training Compute-Optimal Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 593e18c5-091f-4cbc-bda4-82ccc16f7fdc · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning OpenAI o1 System Card
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ecc5a44-9a45-40c7-8ef7-f3a14d1e6ffd · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning When to trust your model: Model-based policy optimization.Advances in neural information processing systems, 32, 2019
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db10e8a0-6ee5-4b4d-9b18-f884620abb2d · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Gemini 2.5: Our most intelligent ai model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 561abd4e-f2d3-4c36-ba0e-16811971c35b · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Adam: A Method for Stochastic Optimization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8ebe66f-f197-404e-81f2-68354b35d613 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Numinamath
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a63cfc0-16ac-42ff-a5f1-6966c7208a85 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Small models struggle to learn from strong reasoners
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31d7cf98-6696-4645-9e55-af9e5cc7b8fa · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 750e1d5a-a226-47c7-9ccd-5fb331ba6c16 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning s1: Simple test-time scaling
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79ecc836-e87b-44cc-95b4-944d64316795 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Policy invariance under reward transforma- tions: Theory and application to reward shaping
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e71d2799-3c7c-4413-bd60-515ceabb671f · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f601ef-e550-4ff5-9a70-5cd29c1831fc · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Direct preference optimization: Your language model is secretly a reward model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b12fe50-36a1-4ed4-a833-24583098b27d · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Gpqa: A graduate-level google-proof q&a benchmark
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfce9b3b-e663-4831-87a3-f3aa9ae05b31 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning A reduction of imitation learning and structured prediction to no-regret online learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a39a7a7-8e3e-4c8f-a3e2-68d6bea889d8 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ac2be1a-098e-4b41-9f47-060728cf6955 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Reasoning with Latent Thoughts: On the Power of Looped Transformers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d093823e-5fac-45d9-906e-229a3be8aed3 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Kickstarting Deep Reinforcement Learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e4a7f8f-8989-4f57-8d08-2f8206a9bb91 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6081e96-a962-48bf-a4ab-1701783111e6 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebe15589-435c-4504-918b-7ce1d9456038 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b70788f5-3f86-4eb3-9235-0ec5698f3ca6 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10649a21-2386-426a-b0e4-a1094e740573 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5469da06-763c-47c3-8758-f199384891e6 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Dump: Automated distribution- level curriculum learning for rl-based llm post-training.arXiv preprint arXiv:2504.09710, 2025
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0876de9d-3300-4349-b82f-43b06698d209 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Chain-of-thought prompting elicits reasoning in large language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18c17d81-78aa-4fc5-9c1d-c31bce1bb703 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f67e7e05-8d2a-4de4-942c-4630b9cd4e06 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f9b2c27-53a9-4594-988e-304daa042a40 · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6820098a-2d10-4b21-ae17-00996f34dbae · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0536afb-af6a-499c-8e9f-d05b9d33ac9d · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning ” or“\n
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c520faaf-ed47-4b73-bd70-c212a9150b9e · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 163ff7ed-6a39-4c59-a15b-3d761b97bdcb · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 46d42dce-87aa-4f7d-9e79-b9470f81b9ae · outbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8858a5b0-372f-42e1-a199-4c9d910cbcb6 · inbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce6002ff-3b9b-4330-81cc-1e7f8f8433b4 · inbound
Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f7d246d-fdc6-4496-bda4-305a41ec8b45 · inbound
CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1f90e798-f156-4265-bbf2-6bb762263edc · inbound
Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cfa911e6-f72e-4014-9b58-18e0a0cdc9d4 · inbound
Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac947f1f-e450-4296-bbf3-914c59795de6 · inbound
Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4868544b-4871-43fe-87d0-83e5fc4872bf · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 176
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c3895e82-69c5-49c9-b926-1d6b03232eeb · inbound
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 75445301-2356-42f7-b965-df56af58a95b · inbound
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6a6d3607-e3c6-4c8e-bcca-e0ce1ebe180d · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8a73fa41-b0de-4b66-9516-1b598333ed61 · inbound
Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f74b1d37-d4a9-4f65-a79e-667f30daddb4 · inbound
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4499d63a-74c1-46f9-9299-604148447797 · inbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cabe3706-fffd-4c42-8540-0b7c87b376a6 · inbound
Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.