Pith. sign in

Paper Citation Record · LEDGER

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

As of 18 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 30 inbound Pith citation observations for arXiv:2506.19767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19767 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:30:38.148313Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:52:15.303331Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation cbe01040-8176-406a-b45b-134de8acc13c · outbound

This paper cites The magnitude of the decrease for each non-target tokenvis proportional to its current probabilityπ θ(v|x,y<t).

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning The magnitude of the decrease for each non-target tokenvis proportional to its current probabilityπ θ(v|x,y<t)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:38.558327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:38.148313Z digest=sha256:ed89bcf1e6c56a0bd12e030721f1e2476b8dde6f118f2c3b682eb22ff3487a22

Observation 19f55063-3ef4-4662-b74f-643205027fcd · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.045826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.045826Z digest=sha256:c8780bb8b4da2891d1654a76635e72c46c15c5399968db74f7134384d9f91514

Observation 7775247c-7ce1-4249-84ee-9ace25ee397e · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.050472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.050472Z digest=sha256:279ee61e199d3480ae7a72fe323e5210b37fb57e4a3b5cfeeb7b6fa50d130d80

Observation a57d2fe3-68d0-4906-9adc-01e65938bc09 · outbound

This paper cites Improving RL Exploration for LLM Reasoning through Retrospective Replay.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.055667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.055667Z digest=sha256:a23e857f6860e65b05760b3860df5ffb707262c9e44ab9cb07c6e1ce701e55f6

Observation 5cf16aa3-009e-4a19-9e08-92cbd05491ec · outbound

This paper cites LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.060670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.060670Z digest=sha256:37468c6c6200014606b31fb015ed7a43cb3aa8c070bf035b9c8065b9768b3f24

Observation 5dc7a3d8-dbd3-47b7-8823-1fc9362047f5 · outbound

This paper cites How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.065791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.065791Z digest=sha256:99a8b6444f51b2d3754cf631cfec0181dd7bfa66f1188bfb3f9d2201f5b1e523

Observation 65877c37-2906-48b1-8829-8017d612a744 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Learning to Reason under Off-Policy Guidance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.070952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.070952Z digest=sha256:e2f6d9ef08dc18178383546ad31949e55218255df149ded6e3b08d742deae70b

Observation c0783268-72c6-4f8e-8038-25ed0fae6a5b · outbound

This paper cites TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.074895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.074895Z digest=sha256:30a933fea81447e20c9f130ba5be700dd08a3ccc1ee7e6caa6c35960d496ed4a

Observation 1e18dc3a-2a3c-4b87-9f02-c016dfab8075 · outbound

This paper cites UFT: Unifying supervised and reinforcement fine-tuning.arXiv preprint arXiv:2505.16984, 2025a.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning UFT: Unifying supervised and reinforcement fine-tuning.arXiv preprint arXiv:2505.16984, 2025a

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.078831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.078831Z digest=sha256:136ad8bc56cbfb9a22b978d055fd02ea25e7545cbd6824186ee29c73089d3052

Observation ad07a096-01e3-40e2-abf5-7afca31de710 · outbound

This paper cites 5, 22 Lu Ma, Hao Liang, Meiyi Qiang, Lexiang Tang, Xiaochen Ma, Zhen Hao Wong, Junbo Niu, Chengyu Shen, Runming He, Bin Cui, et al.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning 5, 22 Lu Ma, Hao Liang, Meiyi Qiang, Lexiang Tang, Xiaochen Ma, Zhen Hao Wong, Junbo Niu, Chengyu Shen, Runming He, Bin Cui, et al

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.092173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.092173Z digest=sha256:ec8d304141c194973fdea6b983d8ae2e40c0b91e92926fa1122e7b2771db4b47

Observation aeeb7730-bc28-48cb-8a7d-8a72a80e125d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.100595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.100595Z digest=sha256:93c4895b69f175a07e13852cf9882c1387084ab736b6bf879e18f32aa826f503

Observation 84bf90af-a3eb-4d56-8eea-75b646a479d1 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.109768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.109768Z digest=sha256:f69b57c7dd4ddbd132f865866369fcce7efa966ac4f4f1ef7be24784f7737d70

Observation 1e80430f-21f4-4522-9a3c-1321d14f7aeb · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.114032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.114032Z digest=sha256:ee1c3a6fe2e27d413f7dc9b34f1a7843389d4861c4cfafc3a3d3934eaecbd93a

Observation 426a2e7a-f3f8-4cec-8b43-72664517a6fa · outbound

This paper cites Process Reinforcement through Implicit Rewards.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Process Reinforcement through Implicit Rewards

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.118345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.118345Z digest=sha256:da9ddbdb26bda6503dcf5432b308b2bfc288f55bdb9667910c5550a6ad1c12e8

Observation d4b4cafe-138b-4da1-b556-255e515f2772 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.122778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.122778Z digest=sha256:74be3e839b3c138a7831f2b06d6f501bbd4e568ecdb6693e837ea120ac0b422d

Observation 7c0839fc-e534-46bd-b906-90277b620ff8 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.127553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.127553Z digest=sha256:0c45e9b98f2a4870c14578fc368252248e8726e16f676baa0c0cc1c9d48a5465

Observation f77ef34b-767b-4b0f-8fd2-74e212b73360 · outbound

This paper cites For datasets with limited sample sizes (AIME24 and AMC), we report the avg@32 metric; for the remaining three datasets, we adopt pass@1 as the evaluation criterion.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning For datasets with limited sample sizes (AIME24 and AMC), we report the avg@32 metric; for the remaining three datasets, we adopt pass@1 as the evaluation criterion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:38.601772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:38.136077Z digest=sha256:049d3bfe6ced001dcd4c51d6515adc1102212aad223c4b7d2d5a27f344d9fd3d

Observation 823de888-79d3-4a9f-bb10-161813419857 · outbound

This paper cites <think>\n {thoughts}</think>\n.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning <think>\n {thoughts}</think>\n

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:38.586601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:38.140023Z digest=sha256:6e8040b9a9e8cd7b2da8010d572ee47320a2214825cfcac45594bca6ce0352c1

Observation 147acd18-44e4-4828-abf0-79a993c358e8 · outbound

This paper cites Google-proof.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Google-proof

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:38.572625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:38.144291Z digest=sha256:e60529f7333377452f85009e3f38c9c473fa73c5f21d1de89359702ede7c6c98

Observation 41f8982b-c858-4a83-8a96-bd5958b82045 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.083143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.083143Z digest=sha256:088130758449b0b36f4c656a7f6f871290440e6b17f35cdbad73a28bcc311198

Observation cf1c6665-0cd0-457f-aa09-037c426ab0b7 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.105293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.105293Z digest=sha256:352114001334ddbc77684fdacd053d36a11f073abb6495f22a3421729df83c3d

Observation fe7424f0-86da-4f18-b9f3-6817648067c8 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.096229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.096229Z digest=sha256:bfe446a7d1bd199e3d9cda469c6760025a550ace93ffc8f95dae8567de093803

Observation 28f53627-385a-43b0-84cc-c2ce833bc023 · outbound

This paper cites 17 A.2 Entropy-aware Gradient Clipping.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning 17 A.2 Entropy-aware Gradient Clipping

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:38.616635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:38.131784Z digest=sha256:371fbcbf3d45fbe3a0911acb6a248d82c32afe33bacb5e752190ac9cbb1cd6de

Observation ff13ae6f-403c-4c05-8609-00d8adfe5d6f · outbound

This paper cites Proximal Policy Optimization Algorithms.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Proximal Policy Optimization Algorithms

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.087653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.087653Z digest=sha256:0d0495f018c5f12fbe9e5477bae56f7b19b29a62d1c97559bbc59f0eee41f229

Observation 0f5e7844-bbd9-4cb7-9fa5-b754265c7a04 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.040628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.040628Z digest=sha256:c3711b14a01ac3874656bff11c5c6c3823c61fb638e264b8c41aec85fc72eef0

Pith citing papers

Observation 9ce80c5e-b8d2-4b1e-8cb6-e72c91dd252f · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 193

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:41:23.526899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:a7d853dcbbd89c5b967b800bdbe3c7931f73fe5016d2a2acd53fcbffd1a262a3

Observation 6c3f19e7-c0d6-459f-924c-eb6dc8dd69a3 · inbound

Post-Completion Learning for Language Models cites this paper.

Post-Completion Learning for Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:52:15.303331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:52:15.303331Z digest=sha256:d538973b90089108afea7eb5af513cc826f5f2234a32c9406d903de644c923d9

Observation 3f2dbca8-2b50-4732-998c-73f84649e9bb · inbound

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning cites this paper.

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:56:55.045505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T23:56:47.274169Z digest=sha256:59f8cd7d523d940bef71cac6f1921886349faf70bce292d08b9e633efa782b97

Observation f9961d65-2c5c-4622-923a-f528de674486 · inbound

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning cites this paper.

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T20:49:33.169961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:49:33.169961Z digest=sha256:851a186ce1da4d61db5627b5a578727629132374b8171f52f149f57135cf2b21

Observation e36dad10-0739-4a45-b78b-9b7028f14b24 · inbound

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration cites this paper.

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:36:53.701052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T22:33:01.074518Z digest=sha256:f90771e83f53df00bfc585e9bc09998e35ae502ec874a7467517ead34e72a6e4

Observation 2d21066e-700b-4b6d-a68f-377c3c099a3a · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 148

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:25.405749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:44459f263082b52bf50a06e856e70aad858289115f546345d687c9b8905abbb0

Observation 975e4c8d-006c-480c-b21b-c252f9e63794 · inbound

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems cites this paper.

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:33.241269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:42:33.241269Z digest=sha256:538f29f1439cff71dd0e66e9d8b53523ffb2552b7bd91225cd35ba810dc30b18

Observation a5c519e1-9d20-4eb9-af14-6b057420aa87 · inbound

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models cites this paper.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.176602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.176602Z digest=sha256:2754bc3286d0efe0fff8a5d9b15c7b953f6630ff9dc2cd026b2f969aa33e1dba

Observation a44f1f16-1fbc-4dc5-bee6-5a444af496a9 · inbound

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts cites this paper.

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:39:26.597119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:39:26.597119Z digest=sha256:6103d57f750e22a2cc834c68f9b3cdc240051be04b08d57d0757cf8aea52633f

Observation 95a5c8a6-8a44-46ff-b8ce-25da8195f68b · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:38.707015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:38.707015Z digest=sha256:6b2d257cc0e6f4a68178f2cecb7a735a05688bfccf124f6416b696a8a30983f8

Observation 336fcfdd-e5d7-437a-87d2-d788895f128a · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:50.322424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:50.322424Z digest=sha256:e96eb8af3a52caa75fa8abe8b0572174281a42587569f89f9b8dd1b4d1b6bbb7

Observation 50b95321-bffc-40e8-867f-d8c974c23178 · inbound

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning cites this paper.

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:43:37.638943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T22:43:01.937642Z digest=sha256:b10634ba872efb7fb35f7cbd6996ffbb9c02a325b427ce404795d82d00620a08

Observation 5853afe1-a468-491f-b111-8e395ab4a9a5 · inbound

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models cites this paper.

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T05:28:11.662347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:28:11.662347Z digest=sha256:797cb56536f0fa95c020b64495cd4b919526dea45c5bb594b884d70f541d51e2

Observation b00e54c2-0f98-4160-835d-a032820f5cf6 · inbound

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings cites this paper.

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:45:37.356079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T12:44:50.752937Z digest=sha256:cee0046834891344d6afc64d6a7c34a45df807032d28b48917db848ad36f27f5

Observation 587b056d-a5f2-41cc-93fa-1a211b8b1ed3 · inbound

$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data cites this paper.

$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:27.829704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T13:43:09.581054Z digest=sha256:0a52e563831c6e9326df456c0a8b45782dce2b9c8690d129aa731eb84f0c977c

Observation 3b331609-45b2-472c-b100-ac69b4472109 · inbound

Near-Future Policy Optimization cites this paper.

Near-Future Policy Optimization SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:54:48.662477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T00:51:36.580600Z digest=sha256:64e9583dc00fc1404a7161ea5cc9c466bbb3bd7da6a27905091fba617be68ffa

Observation 7c4bd75f-8c56-4966-a086-5eedfcbb4d73 · inbound

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors cites this paper.

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:44.011351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-09T19:05:51.423427Z digest=sha256:cc7f1e11bc7602ffdbcd7193a016adb80185bc457b2704c3fabff47f79264db2

Observation 285b16b3-efde-4ca8-89b8-1e5350cd1871 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:06:31.724069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:4e058f1d25d4f3d6f1dcc9ba4f0d7de3c390453a1fabb5f5ba99ae4c3c93bc49

Observation 83613a1d-44e0-467f-8c7c-ef40c35fc3ac · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:07:42.235985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:7ead29fb70d00c3bddcfe913713bc446a1e8d686bd77afeaed6095ed594964b6

Observation 859ffaa4-7137-4c9d-bc2d-2b543d12052e · inbound

Learning Agentic Policy from Action Guidance cites this paper.

Learning Agentic Policy from Action Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.645623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:a1a7c59d57ff34e512ed84cfc3ae762f73876fbc264c9c4d4e0f2f70820dd5ab

Observation 96a80a11-f43d-46f4-a21d-eb95fc43ec92 · inbound

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance cites this paper.

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.377765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-22T06:19:44.377733Z digest=sha256:3327b730c2d23e0d21f0d8e4cfc6b02f5542c2a6654aeeabc0b34a18028ae355

Observation c60b7a7d-435d-4f00-bc9e-a0ce27f0fe67 · inbound

GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training cites this paper.

GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.719169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-29T22:29:05.467252Z digest=sha256:20544abc2130080f4a55a96eb20d1c9ff64b6ce4590f534511eca736b5a57104

Observation 21f8c1db-3f48-42ea-9da4-fb704c60b231 · inbound

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models cites this paper.

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.496547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-29T08:01:39.412431Z digest=sha256:c4d01afb79a4b8254057fb9414a35bda658cd8b6dc3ef42e3bb266df2459cdbd

Observation 967da951-9120-4295-8a7c-68731a45b4d7 · inbound

What are Key Factors for Updates in RL for LLM Reasoning? cites this paper.

What are Key Factors for Updates in RL for LLM Reasoning? SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.590285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:260771a3082c10e0c3492c9acb6c7854370060cab44298f013fcd56ba326e909

Observation ee13d56a-e9c9-4aa0-9614-253e730673ff · inbound

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning cites this paper.

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:29:56.991869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T00:39:29.703711Z digest=sha256:0196b9bff4d8c2b1a6acfc8ed035b45598968cb29f07d7d1da469a308790a8a6

Observation 94116e2e-ed12-4903-acf6-fe4865d5320d · inbound

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning cites this paper.

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:26:58.442808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-02T13:24:17.538850Z digest=sha256:5768b8ce582ba467e05cc8ee0349105c29ff6d5231a94a92d34d37ad159ab6b7

Observation 151eb895-fbf2-4d7d-ba72-a03c726fcb31 · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T11:23:30.063231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:23:30.063231Z digest=sha256:90e1f6e5f96bd59a5aee5f7ce2cfcdda4d4e61b125baf8825ff1e0e6f8098c60

Observation e7b7907b-ccc5-4b9b-be3f-b7943ba9f25f · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.545828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.545828Z digest=sha256:5a828cd4610957997958cabcc18e2620145caf9150c1cc60879a18a180228ecc

Observation b13a36df-6e61-4002-bd65-f7408ccdc670 · inbound

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs cites this paper.

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:48.101782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:48.101782Z digest=sha256:f9fd5d222c98eecb5880999008c6c9372ee87914ba3a839aba7fa5b7716cf523

Observation 66463944-db0e-4f81-8a74-1a8135b269ac · inbound

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs cites this paper.

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T00:51:24.937566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:51:24.937566Z digest=sha256:5afc79e992a2f8b7bc03207b7abeb9dcb6726148a2a699e85f9f23cca09169ce