Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 61 inbound Pith citation observations for arXiv:2503.10460.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:28.847092Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 3a06483f-e30c-4c86-a0d8-fc1f430aa4c1 · inbound
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 191
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 61e3575d-57e1-43c4-ac9a-d246f320caf0 · inbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 616e83c4-3352-43ef-b983-289c2dbd6c2a · inbound
Learning to Reason under Off-Policy Guidance Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 03557996-dbda-41b0-8670-3f7c8f9d161f · inbound
DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89e88595-6b82-4a51-ba4e-fabbdd4e0f50 · inbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9c66fe2a-5ebb-41b7-b8f2-e87660e5ff04 · inbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 131
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce99a49-f25f-48f4-95fe-c7949ed87f8b · inbound
Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d6df71d-8774-429f-9625-23df6d8e680a · inbound
Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f4767c9-cb2f-4298-bdf5-5277d918943c · inbound
MoL for LLMs: Dual-Loss Optimization to Enhance Domain Expertise While Preserving General Capabilities Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c883251e-66ba-41a3-ba3f-4962c134865a · inbound
Towards A Generalist Code Embedding Model Based On Massive Data Synthesis Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51031af8-96fb-4a71-ae2a-093c23cca5ee · inbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a891a60-d82e-4554-820c-e3eec9d5f64f · inbound
Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b3a869c-772d-4d03-b820-8926f46311ae · inbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f6d38db-9d42-4123-b1f9-235377011f6b · inbound
Stable Reinforcement Learning for Efficient Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58bcbb80-eed5-449a-9c08-cdb14c4ad7e0 · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d83f6548-a7f5-4fd5-aba7-4754a97c478f · inbound
MMATH: A Multilingual Benchmark for Mathematical Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 857e40f0-629c-4ec6-a05c-83cfacac4271 · inbound
Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1d4155d-0ead-46e9-a738-5be0cf0490c0 · inbound
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c8782f-2b2e-43bf-b38f-0383ac676e04 · inbound
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eb165d9-c577-4497-85c8-12b25fff21a5 · inbound
Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a85ba105-22b8-4869-ad5a-b8688d9142cc · inbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d906625e-03ee-4bc7-90cf-733794c7e472 · inbound
Skywork Open Reasoner 1 Technical Report Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6d56d6f1-1820-4fc9-865c-f168dcb42c1c · inbound
Decomposing Elements of Problem Solving: What "Math" Does RL Teach? Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40d1d414-ca7a-4e9a-8d5a-2ee63473299f · inbound
Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec882aa4-946f-4bb0-924e-0ca1319546e9 · inbound
StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 235ad0e0-a4d8-4234-93ea-98d1f2f340ca · inbound
Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf457c15-fa3a-4d33-a7ad-bcdc482e83be · inbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39536fc5-73e8-4b6e-b543-381f67e043d0 · inbound
Schema-R1: A reasoning training approach for schema linking in Text-to-SQL Task Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71a6e346-aac4-4dea-a418-2222643eb1ae · inbound
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c88e462-fd60-4c69-a1e3-02b0b28a2579 · inbound
Enhancing Large Language Models through Structured Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cb3ee36-7012-4e15-b16c-47b1fc5031b4 · inbound
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c17a0cd6-9280-4801-9636-f0b99cb58c8d · inbound
MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d0928f-6f77-40d6-9eaf-e34070df6e62 · inbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c52c4c3-7204-4084-8ba8-a61935fa4f62 · inbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4364aea-4370-4dbf-bd1e-802911a48418 · inbound
Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 200
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b60c03f1-14c4-4a95-a520-91ce8f4fccf3 · inbound
REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f940d917-eda9-44d6-8168-e3c54efb1326 · inbound
Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba84593e-9ae9-4427-8d11-3e07adbfd0b4 · inbound
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c5b6fc67-3e29-4c6a-bd9e-45e90dcbf817 · inbound
ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d205813b-567a-43ec-9793-c23f36ab7807 · inbound
Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08fe2101-9ae3-4671-9a47-74cd828efa0a · inbound
Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df28670-0e00-44f9-a3c3-8772a2db9d8a · inbound
Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1c883b7-520d-4b4a-97f1-d8b690a827bc · inbound
Domain-Aware RAG: MoL-Enhanced RL for Efficient Training and Scalable Retrieval Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 400fda6c-c35b-4b51-92e0-b5081c768edb · inbound
Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b7b7f24b-3eea-4265-9b27-09e63c6e8bb5 · inbound
The Signal is in the Steps: Local Scoring for Reasoning Data Selection Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6eb96528-8a20-4626-b3e4-0e81d6459750 · inbound
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 15289bb2-e131-49f9-ae78-6e8beb41fd51 · inbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation be7e2720-b16e-4fdb-ae74-588e4da07b09 · inbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a4abdab7-eff5-47e4-8935-3ee46b435881 · inbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d14952ea-21c3-4e8c-8ef4-34e0280d2194 · inbound
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1bf0c8de-d4d5-449c-998f-c000e8621a0d · inbound
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 09a06dac-043b-4ecc-bfb4-0a1933af2e20 · inbound
Characterizing Model-Native Skills Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0e9be8f2-6c0b-4de7-98bb-6803cffdb2a4 · inbound
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 159
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5a49ae54-1398-4a86-9a44-d95a03d6c9a1 · inbound
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e229cccd-8685-440f-a136-6c54e5289bd3 · inbound
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1d6efd00-eb3d-4326-8317-629164fefbbe · inbound
Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d4a5b6d9-9cbb-40e4-9b29-12e08523d53f · inbound
LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ae67e4f2-82f4-4c79-bc4f-db0a819cbd93 · inbound
Thinking Economically: A Hierarchical Framework for Adaptive-Complexity Reasoning in LLMs Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 221004fe-29e5-487f-8700-2e681c0208f9 · inbound
TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ec1f147d-143b-4586-bbec-816290c2e37b · inbound
Purified OPSD: On-Policy Self-Distillation Without Losing How to Think Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1eafb797-f179-47da-97b1-418540031a12 · inbound
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.