Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2506.19767.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:13:48.101782Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 9ce80c5e-b8d2-4b1e-8cb6-e72c91dd252f · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 193
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f2dbca8-2b50-4732-998c-73f84649e9bb · inbound
EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9961d65-2c5c-4622-923a-f528de674486 · inbound
Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e36dad10-0739-4a45-b78b-9b7028f14b24 · inbound
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2d21066e-700b-4b6d-a68f-377c3c099a3a · inbound
A Survey of Reinforcement Learning for Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 148
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 975e4c8d-006c-480c-b21b-c252f9e63794 · inbound
CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a44f1f16-1fbc-4dc5-bee6-5a444af496a9 · inbound
Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a5c8a6-8a44-46ff-b8ce-25da8195f68b · inbound
RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 336fcfdd-e5d7-437a-87d2-d788895f128a · inbound
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50b95321-bffc-40e8-867f-d8c974c23178 · inbound
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5853afe1-a468-491f-b111-8e395ab4a9a5 · inbound
Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00e54c2-0f98-4160-835d-a032820f5cf6 · inbound
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 587b056d-a5f2-41cc-93fa-1a211b8b1ed3 · inbound
$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b331609-45b2-472c-b100-ac69b4472109 · inbound
Near-Future Policy Optimization SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c4bd75f-8c56-4966-a086-5eedfcbb4d73 · inbound
Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 285b16b3-efde-4ca8-89b8-1e5350cd1871 · inbound
AIPO: Learning to Reason from Active Interaction SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83613a1d-44e0-467f-8c7c-ef40c35fc3ac · inbound
AIPO: Learning to Reason from Active Interaction SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 859ffaa4-7137-4c9d-bc2d-2b543d12052e · inbound
Learning Agentic Policy from Action Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96a80a11-f43d-46f4-a21d-eb95fc43ec92 · inbound
LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c60b7a7d-435d-4f00-bc9e-a0ce27f0fe67 · inbound
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21f8c1db-3f48-42ea-9da4-fb704c60b231 · inbound
Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 967da951-9120-4295-8a7c-68731a45b4d7 · inbound
What are Key Factors for Updates in RL for LLM Reasoning? SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee13d56a-e9c9-4aa0-9614-253e730673ff · inbound
Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94116e2e-ed12-4903-acf6-fe4865d5320d · inbound
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 151eb895-fbf2-4d7d-ba72-a03c726fcb31 · inbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7b7907b-ccc5-4b9b-be3f-b7943ba9f25f · inbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b13a36df-6e61-4002-bd65-f7408ccdc670 · inbound
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.