Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:30:38.148313Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 30 inbound Pith citation observations for arXiv:2506.19767.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:30:38.148313Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:52:15.303331Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
25 of 25 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation cbe01040-8176-406a-b45b-134de8acc13c · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning The magnitude of the decrease for each non-target tokenvis proportional to its current probabilityπ θ(v|x,y<t)
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 19f55063-3ef4-4662-b74f-643205027fcd · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7775247c-7ce1-4249-84ee-9ace25ee397e · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a57d2fe3-68d0-4906-9adc-01e65938bc09 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Improving RL Exploration for LLM Reasoning through Retrospective Replay
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cf16aa3-009e-4a19-9e08-92cbd05491ec · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dc7a3d8-dbd3-47b7-8823-1fc9362047f5 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65877c37-2906-48b1-8829-8017d612a744 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Learning to Reason under Off-Policy Guidance
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0783268-72c6-4f8e-8038-25ed0fae6a5b · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e18dc3a-2a3c-4b87-9f02-c016dfab8075 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning UFT: Unifying supervised and reinforcement fine-tuning.arXiv preprint arXiv:2505.16984, 2025a
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad07a096-01e3-40e2-abf5-7afca31de710 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning 5, 22 Lu Ma, Hao Liang, Meiyi Qiang, Lexiang Tang, Xiaochen Ma, Zhen Hao Wong, Junbo Niu, Chengyu Shen, Runming He, Bin Cui, et al
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeeb7730-bc28-48cb-8a7d-8a72a80e125d · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84bf90af-a3eb-4d56-8eea-75b646a479d1 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e80430f-21f4-4522-9a3c-1321d14f7aeb · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426a2e7a-f3f8-4cec-8b43-72664517a6fa · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Process Reinforcement through Implicit Rewards
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b4cafe-138b-4da1-b556-255e515f2772 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c0839fc-e534-46bd-b906-90277b620ff8 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f77ef34b-767b-4b0f-8fd2-74e212b73360 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning For datasets with limited sample sizes (AIME24 and AMC), we report the avg@32 metric; for the remaining three datasets, we adopt pass@1 as the evaluation criterion
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 823de888-79d3-4a9f-bb10-161813419857 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning <think>\n {thoughts}</think>\n
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 147acd18-44e4-4828-abf0-79a993c358e8 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Google-proof
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 41f8982b-c858-4a83-8a96-bd5958b82045 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf1c6665-0cd0-457f-aa09-037c426ab0b7 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe7424f0-86da-4f18-b9f3-6817648067c8 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f53627-385a-43b0-84cc-c2ce833bc023 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning 17 A.2 Entropy-aware Gradient Clipping
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ff13ae6f-403c-4c05-8609-00d8adfe5d6f · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Proximal Policy Optimization Algorithms
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f5e7844-bbd9-4cb7-9fa5-b754265c7a04 · outbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ce80c5e-b8d2-4b1e-8cb6-e72c91dd252f · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 193
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6c3f19e7-c0d6-459f-924c-eb6dc8dd69a3 · inbound
Post-Completion Learning for Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f2dbca8-2b50-4732-998c-73f84649e9bb · inbound
EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f9961d65-2c5c-4622-923a-f528de674486 · inbound
Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e36dad10-0739-4a45-b78b-9b7028f14b24 · inbound
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2d21066e-700b-4b6d-a68f-377c3c099a3a · inbound
A Survey of Reinforcement Learning for Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 148
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 975e4c8d-006c-480c-b21b-c252f9e63794 · inbound
CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5c519e1-9d20-4eb9-af14-6b057420aa87 · inbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a44f1f16-1fbc-4dc5-bee6-5a444af496a9 · inbound
Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a5c8a6-8a44-46ff-b8ce-25da8195f68b · inbound
RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 336fcfdd-e5d7-437a-87d2-d788895f128a · inbound
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50b95321-bffc-40e8-867f-d8c974c23178 · inbound
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5853afe1-a468-491f-b111-8e395ab4a9a5 · inbound
Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00e54c2-0f98-4160-835d-a032820f5cf6 · inbound
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 587b056d-a5f2-41cc-93fa-1a211b8b1ed3 · inbound
$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3b331609-45b2-472c-b100-ac69b4472109 · inbound
Near-Future Policy Optimization SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7c4bd75f-8c56-4966-a086-5eedfcbb4d73 · inbound
Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 285b16b3-efde-4ca8-89b8-1e5350cd1871 · inbound
AIPO: Learning to Reason from Active Interaction SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 83613a1d-44e0-467f-8c7c-ef40c35fc3ac · inbound
AIPO: Learning to Reason from Active Interaction SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 859ffaa4-7137-4c9d-bc2d-2b543d12052e · inbound
Learning Agentic Policy from Action Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 96a80a11-f43d-46f4-a21d-eb95fc43ec92 · inbound
LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c60b7a7d-435d-4f00-bc9e-a0ce27f0fe67 · inbound
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 21f8c1db-3f48-42ea-9da4-fb704c60b231 · inbound
Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 967da951-9120-4295-8a7c-68731a45b4d7 · inbound
What are Key Factors for Updates in RL for LLM Reasoning? SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ee13d56a-e9c9-4aa0-9614-253e730673ff · inbound
Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 94116e2e-ed12-4903-acf6-fe4865d5320d · inbound
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 151eb895-fbf2-4d7d-ba72-a03c726fcb31 · inbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7b7907b-ccc5-4b9b-be3f-b7943ba9f25f · inbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b13a36df-6e61-4002-bd65-f7408ccdc670 · inbound
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66463944-db0e-4f81-8a74-1a8135b269ac · inbound
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.