Pith. sign in

Paper Citation Record · LEDGER

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

As of 9 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 12 inbound Pith citation observations for arXiv:2507.12856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12856 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:44:51.047658Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T17:03:13.686066Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact4
  • verified fuzzy25
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation afb97173-c846-4dd5-aa11-24107e077d56 · outbound

This paper cites Training language models to follow instructions with human feedback.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Training language models to follow instructions with human feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:45.401136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:45.401136Z digest=sha256:5e5337f22bb0f163d2b676a16bac3d8a12fbfb37b0d54e492276529b36538207

Observation 2ae96b4f-752c-4734-b189-ad82e388fa32 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Fine-Tuning Language Models from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:45.500663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:45.500663Z digest=sha256:e2936fb013e29683f5c7d3bf01a7bca52890416d6be83c99d85467173d735aea

Observation 2dd8d10b-7d19-4cf4-837d-47a1138126a0 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:45.719179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:45.719179Z digest=sha256:0cdb21b0e2dbe2dc2290750d6dfcf83986bb6090b6b0c0aaf613468e28262150

Observation e196e2bc-700e-4987-ab05-180d411725d2 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:45.773825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:45.773825Z digest=sha256:dcb8487c22e6cc79fa29c6f33835aae8fcefedc88502faed29dc5ff124494478

Observation fde1362d-60b9-4d0e-8d05-604669c47cd2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:45.860300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:45.860300Z digest=sha256:b5a5acb3a432cc7b72ef66ba371d5156f0d113c2b6c87e57b2a7b35b108f35c9

Observation 8300ce0c-2d2a-4072-8b57-4d8ea246b60a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:45.911559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:45.911559Z digest=sha256:b46da45849fd72dc72960cce6f91ecb660a7ef1c764d345d2249196eb4614dbd

Observation 7c232c5e-c530-4ae6-92d8-14ba6956428d · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.002490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.002490Z digest=sha256:5eb91f3c2c18f559836ba9e7cf4e9af1b5882667410984df6db8b8e96c054c44

Observation 62ef0cf6-7c98-4fcd-90de-0686f7a6e8d3 · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.084746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.084746Z digest=sha256:8e0ae02c5f711a1df244d43ec9ece7876f95eadfaacdd0434ac18739bb96159f

Observation b86cc15b-f822-4864-ae8e-5ebee5ffefb5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Direct preference optimization: Your language model is secretly a reward model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.142290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.142290Z digest=sha256:3bd4f4618985a8602f1c7958bea1e198a1b79802053348eab3dae06a02242401

Observation bfb073a7-0f9e-4f28-a85d-514a64fbb332 · outbound

This paper cites Learning from negative feedback, or positive feedback or both.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Learning from negative feedback, or positive feedback or both

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:44:52.510853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:46.237472Z digest=sha256:f468ecaf2f46c5fd5e491aac41b07ac0239f944431ac9f1ee10e3efe3d762217

Observation 5640087d-64d5-4b35-ab97-d304c031a795 · outbound

This paper cites s1: Simple test-time scaling.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) s1: Simple test-time scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.335618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.335618Z digest=sha256:aba7c7f5fc6a5676b7abd6ded149349b2cd7292224efe61291d5d04ef0ccf6c6

Observation 4f91e8dd-5014-4aa0-afd3-33baf1614ff9 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.435921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.435921Z digest=sha256:1cc06ae18908fc7a88fe2c0f5c421dd7c203a2f9ec7d3a6f953253bcd977873d

Observation efee556c-471a-43b1-a52e-e3c1a6e1fe78 · outbound

This paper cites Peters and S.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Peters and S

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.861232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:46.568818Z digest=sha256:010dc0ccd9ec23bedb840b29c7008b410f3b10327de595dbf9c4d589331c447d

Observation b741377d-402c-46d2-8377-8ae21207358e · outbound

This paper cites Policy search for motor primitives in robotics.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Policy search for motor primitives in robotics

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-06T16:44:52.290791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:46.673807Z digest=sha256:cb2a85d8f0983560d40d7d438dc821ad42cf688e198f122d1760212fe3a66da1

Observation 57f29dac-888b-4f4d-9cd4-52053917f724 · outbound

This paper cites Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.752711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.752711Z digest=sha256:f2dc104744003b4721a084f869b916ce90fc0e7fc64d84c1ec4b435bb6e6ebf1

Observation 6c0d0215-1793-4c1c-b3cc-6c7d45288ae8 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.837944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.837944Z digest=sha256:26654f5b229578558744122434cff4e51215235271704930b0990ccb51b859ef

Observation 16728588-cee1-4bd4-a0ea-133600395631 · outbound

This paper cites D4rl: Datasets for deep data-driven reinforcement learning, 2020.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) D4rl: Datasets for deep data-driven reinforcement learning, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.734986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:46.919989Z digest=sha256:acc1282a03e04fdb9f273d0703a2cba0eaec97c66675579bcd35fde5e59f654b

Observation 8e88bc7c-02fd-4134-8a20-025d8f5fa7ca · outbound

This paper cites an unresolved cited work.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Unresolved cited work

Reference 18

Resolution
verified exact
doi, observed 2026-08-06T16:44:51.237564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.018101Z digest=sha256:b1f8adbc5417fe773a2eebc4ae51927258531a7682c19329064ff884d2127a18

Observation 7ab64d06-7140-4cfd-99d0-38a5b645dc80 · outbound

This paper cites Learning math reasoning from self-sampled correct and partially-correct solutions.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Learning math reasoning from self-sampled correct and partially-correct solutions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.594199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.085028Z digest=sha256:7021568f390c3b5c9d1991aa8fa554ca57b56a706c393793a74a215eeaadfa95

Observation 83d709d3-e8cb-4135-8314-b8b5f1ab78a7 · outbound

This paper cites Learning to generalize from sparse and underspecified rewards.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Learning to generalize from sparse and underspecified rewards

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.440699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.185820Z digest=sha256:9c3ede481580b1362a149a27869f526aacf484edf5a7e265a91d0907d023134b

Observation d419a748-668a-4a22-8bd7-aaa10210dcb6 · outbound

This paper cites Reward augmented maximum likelihood for neural structured prediction.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Reward augmented maximum likelihood for neural structured prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.339577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.286124Z digest=sha256:bfbe84ff730afa8e8c3ad0b209d343e420e83532e48c07cfb4736ab75ac51490

Observation c2cdac58-84db-41e2-ad62-c8f93657213c · outbound

This paper cites Scaling relationship on learning mathematical reasoning with large language models, 2024.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Scaling relationship on learning mathematical reasoning with large language models, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.228534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.353089Z digest=sha256:d3145cbc6ac5445c9d6daa1ac4132e90e97a589b21c6bb8c90967b4f0adcb01f

Observation 4a3e8de1-e1a8-473a-9616-c3df12e2d901 · outbound

This paper cites Simplify rlhf as reward-weighted sft: A variational method, 2025.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Simplify rlhf as reward-weighted sft: A variational method, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:47.432332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:47.432332Z digest=sha256:1b2dc66debdd15dbe27ab0c264d99c0039ea5a30a02e443e5b79e8ed1cfbf537

Observation 257f81c1-d435-4a4a-875b-84f7e81c6f2d · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Reinforced Self-Training (ReST) for Language Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:47.542530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:47.542530Z digest=sha256:66fe2941f59a6c2e0e7f9f1ac65dd002f81350f538e11c601d646326c159506d

Observation 771076c3-756d-4b82-bf86-b02498f3c9c5 · outbound

This paper cites Beyond human data: Scaling self-training for problem-solving with language models.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Beyond human data: Scaling self-training for problem-solving with language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.075852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.619956Z digest=sha256:5bc2f91605f8ae874da6e9b294351aedb67788c3a9c567356072ef8d8f7a18b0

Observation fa584de8-e969-4035-84af-959a13e2f10c · outbound

This paper cites Peters, K.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Peters, K

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:55.960423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.751813Z digest=sha256:a8d260ed28fafeca3b0bfbbf8d2b9fe847242215ef0958c61e91e174872d8006

Observation 0c3b15e6-3dbd-407f-982c-21099dd00054 · outbound

This paper cites Efficient iterative policy optimization.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Efficient iterative policy optimization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:44:51.818563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.819948Z digest=sha256:7f55a9b7263d2a86287bf01fe40433a811714ca28edecc1ea14619c0da8625bf

Observation 4c08c25d-820a-4ef7-9dae-a1be44beee82 · outbound

This paper cites Tighter bounds lead to improved classifiers.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Tighter bounds lead to improved classifiers

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:44:51.678790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.885875Z digest=sha256:fef06c3e628bdc5d539ec6f58bde3bf369488b4fb369cf2c7cd4c125ed2b6f66

Observation 7997a69a-b90a-4623-8aca-0ab0520792b4 · outbound

This paper cites Process for adapting language models to society (palms) with values-targeted datasets.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Process for adapting language models to society (palms) with values-targeted datasets

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:55.847173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.960817Z digest=sha256:6d8570be6d2e056097c2fbeabed6f1552816cb0fdc031dd4075e014951e9c47c

Observation c11d96f1-3267-4067-9e90-17b4b8c3e13d · outbound

This paper cites Self-consuming generative models with curated data provably optimize human preferences.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Self-consuming generative models with curated data provably optimize human preferences

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:55.709899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:48.055088Z digest=sha256:c2d0f677c3f4c44e33e2dfa4e6b4a67252b907a23506aa87e970838beb607aab

Observation a97c8070-9390-4a67-b0ff-4be30d82d30a · outbound

This paper cites Maximum a posteriori policy optimisation.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Maximum a posteriori policy optimisation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:55.501691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:48.153226Z digest=sha256:fcbe9e6fa4d8e4d2972e9741cb81151343ebe7232d50de5cc3000e1ec21f495d

Observation 45f416fc-0333-48e9-97be-f7b45475e876 · outbound

This paper cites On Multi-objective Policy Optimization as a Tool for Reinforcement Learning: Case Studies in Offline RL and Finetuning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) On Multi-objective Policy Optimization as a Tool for Reinforcement Learning: Case Studies in Offline RL and Finetuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:48.223183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:48.223183Z digest=sha256:584e94eb7374affb60a4d3626ee141302057174a59d3fe2063dd054430aa4db4

Observation 29d5868e-4a55-4112-ad0a-c0eea351d23a · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:48.304660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:48.304660Z digest=sha256:8f55f0c1757c2a842fa4c4e7fa3435f79f8461c27b83661dc18685c70d72b6b8

Observation c4f35b25-c428-4649-b9d0-a82cccd0435b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:48.492080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:48.492080Z digest=sha256:a39db77a5dcc7549d888a291464e3de69dbe61d73989aa747c1a17839009961f

Observation 7d4a7ba3-4f64-4919-8e9d-09a6ea8441d3 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Offline Reinforcement Learning with Implicit Q-Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:48.588072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:48.588072Z digest=sha256:747f3cd477a5a4693d7b434f4712064a7f58264dd79a223ad8e75ce9d2144bdf

Observation 6b54d76e-aa16-4dc0-9478-4302dc1c0d4b · outbound

This paper cites When does return-conditioned supervised learning work for offline reinforcement learning? In S.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) When does return-conditioned supervised learning work for offline reinforcement learning? In S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:55.354255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:48.689749Z digest=sha256:bf450c8a1efec4c98447ce6cc4dc1797b783e6a4811bb346c74db61f1dafc324

Observation 57d6da3f-8e75-4292-9d15-21a48fb37d5f · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Decision transformer: Reinforcement learning via sequence modeling

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:55.137459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:48.748441Z digest=sha256:4d1794624e9c1472bb253c99dc79cd2e876c300adca69121fba9e35510a8224e

Observation fdd94b3f-bdf5-4ec4-a8a4-7e990cac77b1 · outbound

This paper cites Kahn and Andrew W.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Kahn and Andrew W

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:54.952638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:48.854732Z digest=sha256:e2defe379954e9c04b75b98a1b5c3adf54c2517138fee72277fc9a9a08f2b774

Observation a0d9b5b8-8fbb-4d1f-baef-1e95294083a1 · outbound

This paper cites Rubinstein and Dirk P.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Rubinstein and Dirk P

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:54.740440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:48.934050Z digest=sha256:c4c0d0f08034dc3396fb55a9dfe1cd5b243daf85e38f50a6efa68f32538d86be

Observation 88f4b4b9-f99e-4300-aeea-dd32ec107cd3 · outbound

This paper cites Stochastic simulation.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Stochastic simulation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:54.569844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.002793Z digest=sha256:247e36e389372350d9423acffc72018e3eb6bb587516a6c02b1fe0bd1d86f582

Observation d1290aa6-45e6-4f4f-bdb3-f7946219a033 · outbound

This paper cites Doubly robust off-policy value evaluation for reinforcement learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Doubly robust off-policy value evaluation for reinforcement learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:49.093487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:49.093487Z digest=sha256:9c1a50e81bbec05b3b1e91c18e598991622c8b87943405224e9d4dfaeff83426

Observation a26b3599-c974-47b4-808f-22057a1eaf14 · outbound

This paper cites Policy optimization via importance sampling.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Policy optimization via importance sampling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:54.382252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.206149Z digest=sha256:6d21d29b18acb260454828b1712327357153d416af808229163bdd1c693f9c80

Observation f2d7eba5-97db-40d1-8f10-b4034d166bbe · outbound

This paper cites Andrad\'ottir, D.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Andrad\'ottir, D

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:54.218210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.260480Z digest=sha256:ac1d2eb4db9ceba4c4581ccfefe5ade4f23256393b39c0b44e07c388f505a2a6

Observation cced70b8-1f35-4901-b3c6-883f2bf2c9db · outbound

This paper cites Qwen2.5 Technical Report.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Qwen2.5 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:49.354945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:49.354945Z digest=sha256:00f177d8361754ebc0e80e87cffe385905c056dd45424a6f976387f3d654225f

Observation 9d50992c-f704-45d8-853f-9a1058b752be · outbound

This paper cites Numinamath, 2024.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Numinamath, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:54.057836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.442132Z digest=sha256:840a71fa1a7f6d699cf2d285ca4a85e2a5706a341265929c37c5fd6f142f2257

Observation a673c9d4-9d75-4660-8609-a19683a7ce93 · outbound

This paper cites Aime, February 2024.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Aime, February 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:53.913396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.572838Z digest=sha256:a6ea11fe11c9ecad173446bb93a76ba1ef66bd0e5c4f021119cad253e5d86dcb

Observation 0c4b594e-f0a2-43cc-8651-0f6f4b6a0a1d · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:49.647176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:49.647176Z digest=sha256:42e6c3b76ccb564d2b6e89359fee893f7fc78b30aa2ca37cb2242c833e11bf00

Observation 336f7a07-a1a9-4885-9ce8-204c9c234a51 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:49.735291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:49.735291Z digest=sha256:879279c75b068fc33658a9f691b2293ca9ea8480452229b6e96f1c59a1021c79

Observation ae6dc2d4-66ed-44d4-9409-536ced3b8c14 · outbound

This paper cites Gemini 2.0 flash thinking mode (gemini-2.0-flash-thinking-exp-1219), December 2024.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Gemini 2.0 flash thinking mode (gemini-2.0-flash-thinking-exp-1219), December 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:53.691885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.816139Z digest=sha256:84ca5938ee8898d35b96916b574ee91233c91daef5feead1d22fe0e71b459a49

Observation 4ee868a3-7f94-42d6-b3e0-f291b78a9a62 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:49.880551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:49.880551Z digest=sha256:e31dcd731622489bda1e8f2e8f031c9156c26cd790f9fb7fc85144a301f8fcce

Observation f4d49f26-0a3f-4942-8c82-e2dd23e1a758 · outbound

This paper cites Learning to reason with llms, September 2024.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Learning to reason with llms, September 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:53.487524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.948834Z digest=sha256:70c7329f75eb8a5f2cc88eafd6c018710b8c0689393f32ab6defe474669caf78

Observation 7470890e-5934-4df4-85ca-ca63db0f7f60 · outbound

This paper cites Bespoke-stratos: The unreasonable effectiveness of reasoning distillation, 2025.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Bespoke-stratos: The unreasonable effectiveness of reasoning distillation, 2025

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:53.298736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:50.031931Z digest=sha256:00884e94185449f2139d53cfcfdbaf482e3731002a639d9294b236e2d54eb0ae

Observation 668ec223-1aa3-4844-b3d9-059bb8fb9e58 · outbound

This paper cites Sky-t1: Fully open-source reasoning model with o1-preview performance in \ 450 budget, 2025.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Sky-t1: Fully open-source reasoning model with o1-preview performance in \ 450 budget, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:53.146540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:50.135012Z digest=sha256:24dbe38fa47939b16f7000a845341335c6a581206efbb6e2e5d78d2cb6348f4c

Observation a47ffc63-3e48-4a2b-a199-8c7edfecea5b · outbound

This paper cites Gemini 2.5: Our most intelligent ai model, March 2025.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Gemini 2.5: Our most intelligent ai model, March 2025

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:52.995163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:44:50.200843Z digest=sha256:d64c0267ba8dd5bfe7dff87672902598928ff9e29c582b5b8e60839872dc8d5d

Observation d4722487-92ec-4734-942d-1a862edfe79b · outbound

This paper cites Decoupled Weight Decay Regularization.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Decoupled Weight Decay Regularization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.258889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.258889Z digest=sha256:41ccd4c26b17b2e55a15dcebbe9ab9b07f8fb89e869d9deaaf9f5609240909ae

Observation 2ee6a2d9-e86e-4377-8003-f89dea17118c · outbound

This paper cites Continuous control with deep reinforcement learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Continuous control with deep reinforcement learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.353512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.353512Z digest=sha256:60b80c0580dbc5a7ebabe866117668919890e16373ed810b3d78e9c5e3ed71b7

Observation 151b4ce6-8c6f-4147-9149-2611c709b90f · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) A minimalist approach to offline reinforcement learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.436962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.436962Z digest=sha256:cf4ea2130967d892292046c37d048ad8759af77fb00db7cef10bd441450fb740

Observation 5f45d04a-2659-4dcf-a7de-a64126f29578 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.531632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.531632Z digest=sha256:3c8480d8d54c2417a08a46892b4502f606fc1603146d15a86fe71a80c0e99813

Observation 9adab734-26de-496a-b29b-9eb92f25db63 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Conservative q-learning for offline reinforcement learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.671535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.671535Z digest=sha256:b448e88cae3632edae6467ce11aa877f03c21cfe52c88fa04b8e4f99681e518a

Observation 5eb94329-e8c1-470b-b550-a4b008815cab · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.738027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.738027Z digest=sha256:af52b69473c573b0e63fd13ae5ad16d498310ec957223e0b57f71a7c0cd69a2a

Observation f70e476b-efaf-42b7-a1db-205748ae26fa · outbound

This paper cites Accelerate: Training and inference at scale made simple, efficient and adaptable.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Accelerate: Training and inference at scale made simple, efficient and adaptable

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.830477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.830477Z digest=sha256:3ee5f9b8a32f6c35885fe24d7af717d516c17a186c6643d21463193f53f58506

Observation c5d0cf3f-3ad7-4711-a8e4-f2ea0bb0b5a9 · outbound

This paper cites SGDR : Stochastic gradient descent with warm restarts.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) SGDR : Stochastic gradient descent with warm restarts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.909102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.909102Z digest=sha256:30553b16bf718f60d4fdd41dd7118b2bdb4f84b82cb2281d09faa855cf12e625

Observation c37973e0-39b8-43df-a3d4-68712b60c06e · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Adam: A Method for Stochastic Optimization

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:51.047658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:51.047658Z digest=sha256:0885bfdfceb84e1314e1d8b9abaad76e8d9e0a42c2e8313eb953afcb0e5d62ef

Pith citing papers

Observation 54a57a42-7d6b-497a-b993-a839ecc1914d · inbound

Proximal Supervised Fine-Tuning cites this paper.

Proximal Supervised Fine-Tuning Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:42:50.932782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T20:42:14.423836Z digest=sha256:8bfd0a544d1984e3d08252a4cccd28bc26e5a5ef79f9ade5c0d14e800e1e2f3d

Observation 3f61dc23-2063-46c8-8d16-a499a3e67153 · inbound

Multi-LLM Orchestration for High-Quality Code Generation: Exploiting Complementary Model Strengths cites this paper.

Multi-LLM Orchestration for High-Quality Code Generation: Exploiting Complementary Model Strengths Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:22:33.306633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T10:22:09.037783Z digest=sha256:f91e7d880ee198978fafc4ea0a83d043cc71f8b93ef0809dbbb629e16b1b6306

Observation 1c44e9b4-86b8-4c9d-9515-fd5da3410e5d · inbound

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors cites this paper.

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:51:44.181536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T19:05:51.423427Z digest=sha256:b1c3c811479850dcf80af7929c4fdc8050f97accc461bbaa23fc2c04a636827b

Observation 5a99cba9-220f-481b-9f34-85429aad88fb · inbound

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent cites this paper.

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:23.212875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T19:34:22.546508Z digest=sha256:3a1d0c6996b6c6396e2945c129e4206db270e9c46b67db770148a9f3fb03b64f

Observation c53f467c-925c-4fd4-b0f5-bd44f83690f3 · inbound

Rotation-Preserving Supervised Fine-Tuning cites this paper.

Rotation-Preserving Supervised Fine-Tuning Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:27:24.422736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T06:26:20.393476Z digest=sha256:da5e60804ac674dd1d04927cc8d9dc5688ea14a61343f81b2f2fddc4d88186c4

Observation b6439b38-6462-46eb-b7c8-88e767908f8d · inbound

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models cites this paper.

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.954844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T08:01:39.412431Z digest=sha256:d813fa26d4f7763909c874ac407a65f1edcd9b3a3e049faf206c01b1f46f65a9

Observation 353feea3-2ed2-4330-92c0-23c886e339e7 · inbound

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization cites this paper.

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:22:47.435668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:16:49.358792Z digest=sha256:0911218ee17058bda2236dce1c6db216fc87db6ba935311bebdfea349547c4c1

Observation 4d8fce38-062b-45f3-8b32-d81b7b5715d8 · inbound

PriFT: Prior-Support Guided Supervised Fine-Tuning cites this paper.

PriFT: Prior-Support Guided Supervised Fine-Tuning Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:57:29.908830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T16:58:09.015766Z digest=sha256:d4f70e04f8a863674ccbd0c9525117a7f39e6aa6cbab983dff65571effb8ee32

Observation 197441d6-2553-4e83-98c0-184d73a79727 · inbound

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff cites this paper.

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:27:26.355797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T18:49:47.876179Z digest=sha256:c2920748c05fc9526f9e3324f7420bc348e75c466081ee84c6d28c9947b4cd2f

Observation 4e3531e1-fe33-4e57-97ba-b04f0faf917c · inbound

Compatibility-Aware Dynamic Fine-Tuning for Large Language Models cites this paper.

Compatibility-Aware Dynamic Fine-Tuning for Large Language Models Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-05T02:30:40.844795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-05T02:27:50.369374Z digest=sha256:ce8ea466e91e8959aa2cff4cdda5277d940525921e3c537ece4d409463bc303d

Observation 01422f85-b83c-4c76-ab8e-e3f2056d0f89 · inbound

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning cites this paper.

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:29:57.043899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T00:39:29.703711Z digest=sha256:5d715758306097ee56d740604954797ac5e88cc873146bed80c7ecd8569cb770

Observation 6f56a063-4987-42b8-8411-20189f8b2049 · inbound

A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training cites this paper.

A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T17:03:13.686066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T17:03:13.686066Z digest=sha256:7dfd80a53934f774bc66583cbac3f97dc7ccb5b3acb85552b97911c6cc7fc35b