Pith. sign in

Paper Citation Record · LEDGER

Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 61 inbound Pith citation observations for arXiv:2503.10460.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.10460 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 61 of 61 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:28.847092Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3a06483f-e30c-4c86-a0d8-fc1f430aa4c1 · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 191

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:29:57.071592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:fa39c32c181c16d60e1d580a60bed2b6c398edd060d3bf95573bf022dcbfd80b

Observation 61e3575d-57e1-43c4-ac9a-d246f320caf0 · inbound

Improving RL Exploration for LLM Reasoning through Retrospective Replay cites this paper.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.847092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.847092Z digest=sha256:fbd1ff987ab84ccf9fb2df82872e8cc9a49e09064084dc5cbf82c59456be13d2

Observation 616e83c4-3352-43ef-b983-289c2dbd6c2a · inbound

Learning to Reason under Off-Policy Guidance cites this paper.

Learning to Reason under Off-Policy Guidance Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:17:02.869340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T23:17:02.701393Z digest=sha256:ea1242675aca743b419e087d8e43a4ec804dcdf94c1484677f516d1398fc9538

Observation 03557996-dbda-41b0-8670-3f7c8f9d161f · inbound

DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training cites this paper.

DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T10:39:31.300027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:39:31.300027Z digest=sha256:48983905b06007b3dc7af7234be04f7b09fd94de71e43d6378971558dad5e485

Observation 89e88595-6b82-4a51-ba4e-fabbdd4e0f50 · inbound

Reinforcement Learning for Reasoning in Large Language Models with One Training Example cites this paper.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:05.002577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:167d167a86c8be709adec54cad79ba54fd69ac887352b35ed8f6afeab82da6ed

Observation 9c66fe2a-5ebb-41b7-b8f2-e87660e5ff04 · inbound

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models cites this paper.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.576915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.576915Z digest=sha256:7c74ff96f4de6521d6fbeb084ace47a8110319ed199b11d9a7a03d6ae8595f3b

Observation 7ce99a49-f25f-48f4-95fe-c7949ed87f8b · inbound

Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models cites this paper.

Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:46.331840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:56:46.331840Z digest=sha256:433cfaa4905138699d668feb6b78d94845e8a03186f8c6f90304f8ccdc89207f

Observation 5d6df71d-8774-429f-9625-23df6d8e680a · inbound

Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations cites this paper.

Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:08.171603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:08.171603Z digest=sha256:8ab077af3db704cbb67609f990f24000e450fe952a3423b42c298be33d9c64d3

Observation 4f4767c9-cb2f-4298-bdf5-5277d918943c · inbound

MoL for LLMs: Dual-Loss Optimization to Enhance Domain Expertise While Preserving General Capabilities cites this paper.

MoL for LLMs: Dual-Loss Optimization to Enhance Domain Expertise While Preserving General Capabilities Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:44.462106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:48:44.462106Z digest=sha256:7f2014846d8c70855e99d3edd61c4c25aaa8e30d29cecdc5c17c7540471019a3

Observation c883251e-66ba-41a3-ba3f-4962c134865a · inbound

Towards A Generalist Code Embedding Model Based On Massive Data Synthesis cites this paper.

Towards A Generalist Code Embedding Model Based On Massive Data Synthesis Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:54.346546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:54.346546Z digest=sha256:6bf6666b2b745f4cf2181aafab8ef9d8d6dbbb5f90b10aa80640e6cc7ce6754e

Observation 51031af8-96fb-4a71-ae2a-093c23cca5ee · inbound

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning cites this paper.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.469350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.469350Z digest=sha256:b16f64032ae4847c655108178311d295178fa9ea73c9eaea4ece9764712eff78

Observation 8a891a60-d82e-4554-820c-e3eec9d5f64f · inbound

Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN cites this paper.

Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:18.720509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:05:18.720509Z digest=sha256:5f13699c596f840810d41a8d77b2696629c9287a302855662ddfed88acad3eac

Observation 9b3a869c-772d-4d03-b820-8926f46311ae · inbound

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning cites this paper.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.110674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.110674Z digest=sha256:c200b603660ececb361a79b2d0908eb639d91def3910c7b1dd91b5a4a11414e2

Observation 8f6d38db-9d42-4123-b1f9-235377011f6b · inbound

Stable Reinforcement Learning for Efficient Reasoning cites this paper.

Stable Reinforcement Learning for Efficient Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.973811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.973811Z digest=sha256:d210d0b25dd420f6f7ec894636d14c8ad0302f72d6d52e95111cc68ccc87eb9c

Observation 58bcbb80-eed5-449a-9c08-cdb14c4ad7e0 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.891931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.891931Z digest=sha256:a1c9bd8ce3af029812c8925c51b436b484d8dc94ea9d189ca3de46dee4f78c3e

Observation d83f6548-a7f5-4fd5-aba7-4754a97c478f · inbound

MMATH: A Multilingual Benchmark for Mathematical Reasoning cites this paper.

MMATH: A Multilingual Benchmark for Mathematical Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.386204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:23:09.386204Z digest=sha256:9d4877ef60d73bf567b34bf7ed8e4af8296a7f0d9ffa010a05f8d09fca01fa05

Observation 857e40f0-629c-4ec6-a05c-83cfacac4271 · inbound

Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting cites this paper.

Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:19.947222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:19.947222Z digest=sha256:cf352b8c8d3df50fc98ac3d1fa19714b27686de2e518bd68c02d7b7d5430caad

Observation d1d4155d-0ead-46e9-a738-5be0cf0490c0 · inbound

Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective cites this paper.

Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:11.080617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:11.080617Z digest=sha256:b1861bdf85ff7167fe155b17bb3c05535a83179552894527388df8c00cdc4541

Observation 28c8782f-2b2e-43bf-b38f-0383ac676e04 · inbound

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles cites this paper.

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:08.796570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:08.796570Z digest=sha256:4bd567accf3ed46289ffc6a52bbe25eb682c7b0ab72cf74e21196e6cfb1a38a8

Observation 6eb165d9-c577-4497-85c8-12b25fff21a5 · inbound

Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions cites this paper.

Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:49.365734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:49.365734Z digest=sha256:bec8813fafef14c7b99bb74e51efa977b52d029f4d2794e863f1f267f8534118

Observation a85ba105-22b8-4869-ad5a-b8688d9142cc · inbound

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning cites this paper.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:30.443487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:30.443487Z digest=sha256:09dd133e1eab1fc142ebe0661344418f184acc1b642cd99c2354f71d0ab8abd7

Observation d906625e-03ee-4bc7-90cf-733794c7e472 · inbound

Skywork Open Reasoner 1 Technical Report cites this paper.

Skywork Open Reasoner 1 Technical Report Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:26:47.350497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T04:26:47.283983Z digest=sha256:be4f0f54c097b7cf46ab5229ad97dec389f6599b9f2ac528e94a82b0f10efb1c

Observation 6d56d6f1-1820-4fc9-865c-f168dcb42c1c · inbound

Decomposing Elements of Problem Solving: What "Math" Does RL Teach? cites this paper.

Decomposing Elements of Problem Solving: What "Math" Does RL Teach? Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:27.038237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:27.038237Z digest=sha256:958de69450bde13dbeaae056065df23bdf2b942f3acd817617905fb2ab4913d2

Observation 40d1d414-ca7a-4e9a-8d5a-2ee63473299f · inbound

Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models cites this paper.

Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:21.956041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:21.956041Z digest=sha256:4435c86fc480204e736f35d4af75846e8b87065456f3b8a7a940b949782d297e

Observation ec882aa4-946f-4bb0-924e-0ca1319546e9 · inbound

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs cites this paper.

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:37.235199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:17:37.235199Z digest=sha256:2e8c7ec04c912b89dd4fa591d3f89e78f25f690c1b917021b1b9e885caabc02b

Observation 235ad0e0-a4d8-4234-93ea-98d1f2f340ca · inbound

Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning cites this paper.

Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:07.838570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:07.838570Z digest=sha256:7d4d0cb8162bbe13ebfe6a379e48f6c2e8c1cf7885290d4be1933bde7bb65ef8

Observation bf457c15-fa3a-4d33-a7ad-bcdc482e83be · inbound

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning cites this paper.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.233618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.233618Z digest=sha256:2a27872b8471f473b6badc344676c54d3f94a3f26a8ed1284984af22c6460b25

Observation 39536fc5-73e8-4b6e-b543-381f67e043d0 · inbound

Schema-R1: A reasoning training approach for schema linking in Text-to-SQL Task cites this paper.

Schema-R1: A reasoning training approach for schema linking in Text-to-SQL Task Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:26.848970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:26.848970Z digest=sha256:2b2d1e9f5a681ff12adf6e6abff79894450226dc50d2ece81eddb565f2cb7433

Observation 71a6e346-aac4-4dea-a418-2222643eb1ae · inbound

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy cites this paper.

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:36.113617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:36.113617Z digest=sha256:749e8434ad64409b243ddf3e938c5efaf921734238d5b9797f93971ace446c69

Observation 2c88e462-fd60-4c69-a1e3-02b0b28a2579 · inbound

Enhancing Large Language Models through Structured Reasoning cites this paper.

Enhancing Large Language Models through Structured Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:59:47.990084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:59:47.990084Z digest=sha256:8a4706b1e13723d73513b1dade25917c954a4848f86aeed48e870bf9e4842349

Observation 8cb3ee36-7012-4e15-b16c-47b1fc5031b4 · inbound

CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning cites this paper.

CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T18:39:18.971569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:39:18.971569Z digest=sha256:ac37c43c81feca25b60ca1915ae084b20d730b0f208a66d6d58740af39d77ccd

Observation c17a0cd6-9280-4801-9636-f0b99cb58c8d · inbound

MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants cites this paper.

MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:22.631170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:22.631170Z digest=sha256:6fac475c0742345e60f6261f8459b3456e471fd1be3c0b89ff0353fa5d8677c1

Observation f4d0928f-6f77-40d6-9eaf-e34070df6e62 · inbound

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model cites this paper.

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:38.720506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:57:38.720506Z digest=sha256:a53233d6f052dbc837b06d8d16d112deb06fa9ca4fa92c8fbdb072b037dce372

Observation 6c52c4c3-7204-4084-8ba8-a61935fa4f62 · inbound

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning cites this paper.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.525387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.525387Z digest=sha256:4430e4ceba5edb866707ea793f06bf49f647a0478a4222bcb23a7c3b91c7b85b

Observation d4364aea-4370-4dbf-bd1e-802911a48418 · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.847748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.847748Z digest=sha256:8995e6ae53d336e1814608ef90c22ac9ac9ac2e37e75ef189559244811a9d0b0

Observation b60c03f1-14c4-4a95-a520-91ce8f4fccf3 · inbound

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once cites this paper.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.692436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.692436Z digest=sha256:6f3b3c7ae7ed636eba03e29d490fdf28a06f6dc46723a3695babe5da30fbac20

Observation f940d917-eda9-44d6-8168-e3c54efb1326 · inbound

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning cites this paper.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.599526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.599526Z digest=sha256:d01c8676411db8a9b9b2421dd315ef807f2e3362a67f7dbc8d53bcdc416df37d

Observation ba84593e-9ae9-4427-8d11-3e07adbfd0b4 · inbound

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling cites this paper.

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T22:34:23.988237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T22:33:09.674822Z digest=sha256:f1f781f78976b52dfd0d60ef8a462032bf49d7dcc57923e82e12f5f0a4c1314a

Observation c5b6fc67-3e29-4c6a-bd9e-45e90dcbf817 · inbound

ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models cites this paper.

ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T16:19:53.921525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:19:53.921525Z digest=sha256:03921f4a8c8a35fe2ae9ca632920f94a466ab8b8415ed14907c0792e56cd58e8

Observation d205813b-567a-43ec-9793-c23f36ab7807 · inbound

Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM cites this paper.

Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:11:01.110314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:11:01.110314Z digest=sha256:e44bf30fcbfaa0d22fdbf5ad0547fa5409493a4d92f0d5d865a641ca0f49f4af

Observation 08fe2101-9ae3-4671-9a47-74cd828efa0a · inbound

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework cites this paper.

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T05:45:04.041310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:45:04.041310Z digest=sha256:2e797c24b0f339d81b0d2da23498b6c7f712c5e67cb456c92f7fe4d7968b185a

Observation 2df28670-0e00-44f9-a3c3-8772a2db9d8a · inbound

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework cites this paper.

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T05:45:04.085149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:45:04.085149Z digest=sha256:1634990f7f0a39f53ca2703ff1abd21f00919af5fac52f848580741065d3c9bb

Observation e1c883b7-520d-4b4a-97f1-d8b690a827bc · inbound

Domain-Aware RAG: MoL-Enhanced RL for Efficient Training and Scalable Retrieval cites this paper.

Domain-Aware RAG: MoL-Enhanced RL for Efficient Training and Scalable Retrieval Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:41.779153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:41.779153Z digest=sha256:c60cbdb5843552a3580d319564d320a9ca9eff4d70e90748dd0078e6de37a977

Observation 400fda6c-c35b-4b51-92e0-b5081c768edb · inbound

Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training cites this paper.

Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:31:24.978441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T13:28:32.093512Z digest=sha256:327b2608531d11f223b51990a83f9eebe1fc467532d1bd710a7b26a743c89284

Observation b7b7f24b-3eea-4265-9b27-09e63c6e8bb5 · inbound

The Signal is in the Steps: Local Scoring for Reasoning Data Selection cites this paper.

The Signal is in the Steps: Local Scoring for Reasoning Data Selection Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:16:14.172268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T10:14:27.739531Z digest=sha256:d0995cc07101c277edb1151fdb718fb1430ab75fcd8f8b36f9364db28497d458

Observation 6eb96528-8a20-4626-b3e4-0e81d6459750 · inbound

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation cites this paper.

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:06:13.802986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T10:04:39.223895Z digest=sha256:0e9609da907331c5d77f84c78ead7ee04aa2e202e8c0d9b52a97775dc5728a0f

Observation 15289bb2-e131-49f9-ae78-6e8beb41fd51 · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:12:23.662328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T05:11:29.205366Z digest=sha256:3f56d528dcf448527f0f6c2d120c044a859494ea0407278869e635541cad0bdf

Observation be7e2720-b16e-4fdb-ae74-588e4da07b09 · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T20:10:34.891829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T20:06:16.172916Z digest=sha256:861512d8b1ed2189449d48a357dec85d6af70845c86885b4ee2c0e94089d78a4

Observation a4abdab7-eff5-47e4-8935-3ee46b435881 · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.747822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.747822Z digest=sha256:62516d5c40444648c521764b590f9b5fd4c91774c0d7c608c0da35ca03b2db02

Observation d14952ea-21c3-4e8c-8ef4-34e0280d2194 · inbound

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training cites this paper.

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:48:50.926478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T01:46:21.744857Z digest=sha256:3b3035f6fbfe5ce266749083146594fda4cd6747ba22ec6eb752f333573bc0ab

Observation 1bf0c8de-d4d5-449c-998f-c000e8621a0d · inbound

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models cites this paper.

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T14:10:13.173355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T14:09:26.842696Z digest=sha256:d304fe135ae775d942bbf63fa6603b4f2f16bdb29b4fbc15c7d9a500645f6f55

Observation 09a06dac-043b-4ecc-bfb4-0a1933af2e20 · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.521433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:4e83626bc8237bc5ad766719ebb04ee03ad57ee6d87159e8c1c808992153ace7

Observation 0e9be8f2-6c0b-4de7-98bb-6803cffdb2a4 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 159

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.755531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:57ba0d93ad5a3de2b57525aaf6384461c0c8704c381e6645287352b27878ae05

Observation 5a49ae54-1398-4a86-9a44-d95a03d6c9a1 · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:46:28.509523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:44aa47c192baacd216e9e66a7f3ef44a7bd6b8523f8b79eaa634cde4efcf8e09

Observation e229cccd-8685-440f-a136-6c54e5289bd3 · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:46:18.839590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:01a3799eed062fb79f4445f9d272191abe49d34f3953fa5a1eac71ad18d1489e

Observation 1d6efd00-eb3d-4326-8317-629164fefbbe · inbound

Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning cites this paper.

Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:15:02.865013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T05:14:28.256032Z digest=sha256:e8d5d90b2e5b48e08ddfb728291b0479d96f818374908c055a11f72cbee923fc

Observation d4a5b6d9-9cbb-40e4-9b29-12e08523d53f · inbound

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance cites this paper.

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:09.548667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-22T06:19:44.377733Z digest=sha256:b6d62d89c9aea65da35c802822bdf1aa56286b5626d73d57635a1e209e39df5b

Observation ae67e4f2-82f4-4c79-bc4f-db0a819cbd93 · inbound

Thinking Economically: A Hierarchical Framework for Adaptive-Complexity Reasoning in LLMs cites this paper.

Thinking Economically: A Hierarchical Framework for Adaptive-Complexity Reasoning in LLMs Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:12:25.200187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T17:05:48.244094Z digest=sha256:598a1d6adf36e8f33f02e209fea3109bc9b37aa61d888bccb3b9f12cee63cca8

Observation 221004fe-29e5-487f-8700-2e681c0208f9 · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.717007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:aae554c33c0d9f0914bc10abd25458724a6830b6e6360973c8d85d3bc56542ee

Observation ec1f147d-143b-4586-bbec-816290c2e37b · inbound

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think cites this paper.

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:58:21.051491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T13:56:13.827493Z digest=sha256:d95d9d81127f7cf12fe3a93a0a4f39bdaa1c78110cb5a0017d7ddf88775b7556

Observation 1eafb797-f179-47da-97b1-418540031a12 · inbound

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility cites this paper.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.983309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.983309Z digest=sha256:6dabe9f01f7ab13df536c2fc9acb374291f091e560e987a276c336ce9d4d30f0