Pith. sign in

Paper Citation Record · LEDGER

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

As of 17 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 12 inbound Pith citation observations for arXiv:2507.12856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12856 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:44:51.047658Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T17:03:13.686066Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact4
  • verified fuzzy25
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation afb97173-c846-4dd5-aa11-24107e077d56 · outbound

This paper cites Training language models to follow instructions with human feedback.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Training language models to follow instructions with human feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:45.401136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:45.401136Z digest=sha256:05c39ca5c9a934d266d7bfd2c9f4c31b1d6b7b64e283f1a127d01bf07eecc112

Observation 2ae96b4f-752c-4734-b189-ad82e388fa32 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Fine-Tuning Language Models from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:45.500663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:45.500663Z digest=sha256:afcdf7ceb45642975212a4e0ab5c96e3f368bc9522ba5be07df63edab7472ad5

Observation 2dd8d10b-7d19-4cf4-837d-47a1138126a0 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:45.719179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:45.719179Z digest=sha256:ba667cf0187f40b3cffb453374f7bb626b0694f47ce93752cbf61df6457fc9da

Observation e196e2bc-700e-4987-ab05-180d411725d2 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:45.773825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:45.773825Z digest=sha256:9be52fb1bfaa01f10ab43d2e72ebecd8cfbcfd1da41b66fea170c047ec6c7c9c

Observation fde1362d-60b9-4d0e-8d05-604669c47cd2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:45.860300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:45.860300Z digest=sha256:8b4c9af3e2aad823fe28250fe58a01b42657ae81a77e00dabfaa7b3dfa16885d

Observation 8300ce0c-2d2a-4072-8b57-4d8ea246b60a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:45.911559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:45.911559Z digest=sha256:1a61709d7c5c81ec724ac0579d510266eddafd497bebd2331eb4d66ce416eb86

Observation 7c232c5e-c530-4ae6-92d8-14ba6956428d · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.002490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.002490Z digest=sha256:b13e184398dfef4bd4ff52517704f9b0b48f5dac73981e276b44b933b854912e

Observation 62ef0cf6-7c98-4fcd-90de-0686f7a6e8d3 · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.084746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.084746Z digest=sha256:2dff616c76e8a8be30fdae399690ff7349996c50a9b8d8495f35d1ca885bb285

Observation b86cc15b-f822-4864-ae8e-5ebee5ffefb5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Direct preference optimization: Your language model is secretly a reward model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.142290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.142290Z digest=sha256:b238406b5ebc56c2208ed2b3a7103818f4c9de5f483b863375645f84ae689bc2

Observation bfb073a7-0f9e-4f28-a85d-514a64fbb332 · outbound

This paper cites Learning from negative feedback, or positive feedback or both.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Learning from negative feedback, or positive feedback or both

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:44:52.510853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:46.237472Z digest=sha256:758cec0e08d85e66f0fc98a41cd479bee2f7cac35f3a051177ebb36089eb4c7a

Observation 5640087d-64d5-4b35-ab97-d304c031a795 · outbound

This paper cites s1: Simple test-time scaling.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) s1: Simple test-time scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.335618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.335618Z digest=sha256:a7843095e64e19534951ce6f6be7fe9c8e34145d199d92f3da74137453886dda

Observation 4f91e8dd-5014-4aa0-afd3-33baf1614ff9 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.435921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.435921Z digest=sha256:b3e59447a78cf0bce2e964d3f443cdd90a3e29b70bb34f077c6154f9774eba31

Observation efee556c-471a-43b1-a52e-e3c1a6e1fe78 · outbound

This paper cites Peters and S.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Peters and S

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.861232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:46.568818Z digest=sha256:b521508cc6df44c04e65b424454d5c1df5d536f33434d15c50899c6f63044bfa

Observation b741377d-402c-46d2-8377-8ae21207358e · outbound

This paper cites Policy search for motor primitives in robotics.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Policy search for motor primitives in robotics

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-06T16:44:52.290791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:46.673807Z digest=sha256:4da48aa4d9d16650854e363458abaffefd3962b1a03b6d173372527da281a842

Observation 57f29dac-888b-4f4d-9cd4-52053917f724 · outbound

This paper cites Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.752711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.752711Z digest=sha256:08c769b5da1445dffd66a2409f5f89b3e9780ac37da10ece2571af41eb1154d6

Observation 6c0d0215-1793-4c1c-b3cc-6c7d45288ae8 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.837944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.837944Z digest=sha256:032fb0e5534cf44b2de4db9a340e83e6867fec3874ee790514dffdf19e179377

Observation 16728588-cee1-4bd4-a0ea-133600395631 · outbound

This paper cites D4rl: Datasets for deep data-driven reinforcement learning, 2020.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) D4rl: Datasets for deep data-driven reinforcement learning, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.734986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:46.919989Z digest=sha256:6987f44a31f556d070f4f7cabfab7cb5c1e28c7ec23079aeee15ecb7c1f74fca

Observation 8e88bc7c-02fd-4134-8a20-025d8f5fa7ca · outbound

This paper cites an unresolved cited work.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Unresolved cited work

Reference 18

Resolution
verified exact
doi, observed 2026-08-06T16:44:51.237564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.018101Z digest=sha256:67965ceef41d6544d457a87e3f49ce8c0594e48ee058310dd043c8193722be24

Observation 7ab64d06-7140-4cfd-99d0-38a5b645dc80 · outbound

This paper cites Learning math reasoning from self-sampled correct and partially-correct solutions.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Learning math reasoning from self-sampled correct and partially-correct solutions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.594199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.085028Z digest=sha256:dfece5194c28980ee1a6584f582a86e5f9eda8fc45c92ecf679820907ec7a2de

Observation 83d709d3-e8cb-4135-8314-b8b5f1ab78a7 · outbound

This paper cites Learning to generalize from sparse and underspecified rewards.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Learning to generalize from sparse and underspecified rewards

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.440699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.185820Z digest=sha256:95cb1aab21aa83bf0bad97adbdcec04fcc89615632451a38a3cee066da701ac6

Observation d419a748-668a-4a22-8bd7-aaa10210dcb6 · outbound

This paper cites Reward augmented maximum likelihood for neural structured prediction.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Reward augmented maximum likelihood for neural structured prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.339577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.286124Z digest=sha256:8a2b4417085791a37f2866d3afe22f3745776d53ea65df5eac4c879fc7477dfb

Observation c2cdac58-84db-41e2-ad62-c8f93657213c · outbound

This paper cites Scaling relationship on learning mathematical reasoning with large language models, 2024.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Scaling relationship on learning mathematical reasoning with large language models, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.228534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.353089Z digest=sha256:1f1032c3e14def01d4eb8159810922b9d299ab1dda4fbfac9d3aedcaf5eeff2c

Observation 4a3e8de1-e1a8-473a-9616-c3df12e2d901 · outbound

This paper cites Simplify rlhf as reward-weighted sft: A variational method, 2025.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Simplify rlhf as reward-weighted sft: A variational method, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:47.432332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:47.432332Z digest=sha256:1c01b2464613333b917894d19fc34b6c802ef26e85f447ebac28a39dc28c7063

Observation 257f81c1-d435-4a4a-875b-84f7e81c6f2d · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Reinforced Self-Training (ReST) for Language Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:47.542530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:47.542530Z digest=sha256:5cded78dc585969d6d431c66483e93608c3c9ddc283bb3966e064a68aa6a8184

Observation 771076c3-756d-4b82-bf86-b02498f3c9c5 · outbound

This paper cites Beyond human data: Scaling self-training for problem-solving with language models.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Beyond human data: Scaling self-training for problem-solving with language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:56.075852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.619956Z digest=sha256:b11ac0758581be763fa932d499081e9036e14a028ca843ccc02f1b8bd40ca2da

Observation fa584de8-e969-4035-84af-959a13e2f10c · outbound

This paper cites Peters, K.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Peters, K

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:55.960423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.751813Z digest=sha256:6729f39c09949cd8d22fee23b225a1ac3eedb53fd69a9e4e8078dc166713d6a5

Observation 0c3b15e6-3dbd-407f-982c-21099dd00054 · outbound

This paper cites Efficient iterative policy optimization.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Efficient iterative policy optimization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:44:51.818563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.819948Z digest=sha256:cf019ef7ad54f7624519bab4ca84e0ff15c0290923af8579cafa562c8815eea6

Observation 4c08c25d-820a-4ef7-9dae-a1be44beee82 · outbound

This paper cites Tighter bounds lead to improved classifiers.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Tighter bounds lead to improved classifiers

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:44:51.678790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.885875Z digest=sha256:5f8b0b2302f8c2dcd097baf1500a7a1dc6bf71857b899884278fcfc61ae83401

Observation 7997a69a-b90a-4623-8aca-0ab0520792b4 · outbound

This paper cites Process for adapting language models to society (palms) with values-targeted datasets.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Process for adapting language models to society (palms) with values-targeted datasets

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:55.847173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:47.960817Z digest=sha256:25ca23b78464d8ce003e4257cfc0f0c7a5c397bfa9a92fde8cf722e0d63b3272

Observation c11d96f1-3267-4067-9e90-17b4b8c3e13d · outbound

This paper cites Self-consuming generative models with curated data provably optimize human preferences.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Self-consuming generative models with curated data provably optimize human preferences

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:55.709899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:48.055088Z digest=sha256:74f8371932063b8240c13d7ec202a1fc79f8ecb74db6195614549577cf79f720

Observation a97c8070-9390-4a67-b0ff-4be30d82d30a · outbound

This paper cites Maximum a posteriori policy optimisation.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Maximum a posteriori policy optimisation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:55.501691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:48.153226Z digest=sha256:da4e2604842bd4c66d7299381a02c68af473e567805a5766e69417faf79b11dc

Observation 45f416fc-0333-48e9-97be-f7b45475e876 · outbound

This paper cites On Multi-objective Policy Optimization as a Tool for Reinforcement Learning: Case Studies in Offline RL and Finetuning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) On Multi-objective Policy Optimization as a Tool for Reinforcement Learning: Case Studies in Offline RL and Finetuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:48.223183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:48.223183Z digest=sha256:450813219c1331a5488f0dd684ab7844395c569cf6ab985928e902892244025a

Observation 29d5868e-4a55-4112-ad0a-c0eea351d23a · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:48.304660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:48.304660Z digest=sha256:1d19056cedb22c042ec70176e9608a6aeb8f4cc898461c40c83ff865ae87d367

Observation c4f35b25-c428-4649-b9d0-a82cccd0435b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:48.492080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:48.492080Z digest=sha256:6ae6c91d98ecfeb3d8065d214b9130b1a6a66153be9586b92b8dbfd18be26aca

Observation 7d4a7ba3-4f64-4919-8e9d-09a6ea8441d3 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Offline Reinforcement Learning with Implicit Q-Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:48.588072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:48.588072Z digest=sha256:632da3e12597f75227aa7c1ab43ba6c32386313d3c870a4a968e5fa59d2fc556

Observation 6b54d76e-aa16-4dc0-9478-4302dc1c0d4b · outbound

This paper cites When does return-conditioned supervised learning work for offline reinforcement learning? In S.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) When does return-conditioned supervised learning work for offline reinforcement learning? In S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:55.354255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:48.689749Z digest=sha256:de72b3d5f5634909cd1a26201a91c598b14a1491003808dae3ac3f54de321116

Observation 57d6da3f-8e75-4292-9d15-21a48fb37d5f · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Decision transformer: Reinforcement learning via sequence modeling

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:55.137459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:48.748441Z digest=sha256:dffadc98cb7ee9e5ea1303ba438f2e4265238a66f6182bc55147a2ae8cca6677

Observation fdd94b3f-bdf5-4ec4-a8a4-7e990cac77b1 · outbound

This paper cites Kahn and Andrew W.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Kahn and Andrew W

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:54.952638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:48.854732Z digest=sha256:b86893cf1e135af22fc5407c1cb38e49d4b41b6e090dd9a50287813801abae89

Observation a0d9b5b8-8fbb-4d1f-baef-1e95294083a1 · outbound

This paper cites Rubinstein and Dirk P.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Rubinstein and Dirk P

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:54.740440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:48.934050Z digest=sha256:c7ade83364d7359fc707d0bc375a689197588ea1f46239d9c8a6ded6804bb320

Observation 88f4b4b9-f99e-4300-aeea-dd32ec107cd3 · outbound

This paper cites Stochastic simulation.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Stochastic simulation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:54.569844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.002793Z digest=sha256:bda9bd1606847083ddddc9932772d82af377bddfc5a0f0e333d90176f758b9c5

Observation d1290aa6-45e6-4f4f-bdb3-f7946219a033 · outbound

This paper cites Doubly robust off-policy value evaluation for reinforcement learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Doubly robust off-policy value evaluation for reinforcement learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:49.093487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:49.093487Z digest=sha256:5f23d27624129270ff8442f709b1bfd676137580843f0155ce93322496685d29

Observation a26b3599-c974-47b4-808f-22057a1eaf14 · outbound

This paper cites Policy optimization via importance sampling.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Policy optimization via importance sampling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:54.382252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.206149Z digest=sha256:42943c0e6889b0b6be612f121b2d82adde29859922a85475bb12ce86e7234b1e

Observation f2d7eba5-97db-40d1-8f10-b4034d166bbe · outbound

This paper cites Andrad\'ottir, D.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Andrad\'ottir, D

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:54.218210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.260480Z digest=sha256:540ae86773cd50e342b5c4e87e30de9d48c9b46660b155f868b2c80f3642fcda

Observation cced70b8-1f35-4901-b3c6-883f2bf2c9db · outbound

This paper cites Qwen2.5 Technical Report.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Qwen2.5 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:49.354945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:49.354945Z digest=sha256:afd369a40e85bd9f8c4e467bb4ec9f4ea74459ae20f2faed32a42698bd9abe59

Observation 9d50992c-f704-45d8-853f-9a1058b752be · outbound

This paper cites Numinamath, 2024.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Numinamath, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:54.057836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.442132Z digest=sha256:6c83d1708e365affecce8917b82dbb1ca217c1794f9d98f82fa06021e45c3097

Observation a673c9d4-9d75-4660-8609-a19683a7ce93 · outbound

This paper cites Aime, February 2024.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Aime, February 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:53.913396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.572838Z digest=sha256:ff7635bc4000508273f2124d78967a0e81bfa3675d0081060a6613e8754cfd3a

Observation 0c4b594e-f0a2-43cc-8651-0f6f4b6a0a1d · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:49.647176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:49.647176Z digest=sha256:bdce80e1e5a448770b78dc0708c7fd61272969d8663290984c82aa9896aabd9e

Observation 336f7a07-a1a9-4885-9ce8-204c9c234a51 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:49.735291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:49.735291Z digest=sha256:01167f7732227bbdfe5500328fa7a823cf6fab974625ccd49aef51abf222e443

Observation ae6dc2d4-66ed-44d4-9409-536ced3b8c14 · outbound

This paper cites Gemini 2.0 flash thinking mode (gemini-2.0-flash-thinking-exp-1219), December 2024.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Gemini 2.0 flash thinking mode (gemini-2.0-flash-thinking-exp-1219), December 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:53.691885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.816139Z digest=sha256:6e3c7ad21389cd960787b5e44f3e40900e25f4710090ab18a3350d8b42ce0fd5

Observation 4ee868a3-7f94-42d6-b3e0-f291b78a9a62 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:49.880551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:49.880551Z digest=sha256:3d36fc4425c64bd3eb925216b200deb13a5c6d4754b2b737c6743225c1de6a90

Observation f4d49f26-0a3f-4942-8c82-e2dd23e1a758 · outbound

This paper cites Learning to reason with llms, September 2024.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Learning to reason with llms, September 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:53.487524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:49.948834Z digest=sha256:6c27f1065128c55398bffba5c05f48c60b50f2832bef5ecdc1e5097faec62de0

Observation 7470890e-5934-4df4-85ca-ca63db0f7f60 · outbound

This paper cites Bespoke-stratos: The unreasonable effectiveness of reasoning distillation, 2025.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Bespoke-stratos: The unreasonable effectiveness of reasoning distillation, 2025

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:53.298736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:50.031931Z digest=sha256:8373673d13883d0f019237e7e7dca8760e295d4de28a8f5c3e47b941e4064b62

Observation 668ec223-1aa3-4844-b3d9-059bb8fb9e58 · outbound

This paper cites Sky-t1: Fully open-source reasoning model with o1-preview performance in \ 450 budget, 2025.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Sky-t1: Fully open-source reasoning model with o1-preview performance in \ 450 budget, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:53.146540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:50.135012Z digest=sha256:97cbaec866832e00db5507e42a21851b6e2a28a05dbf2293347a0f233ea634e3

Observation a47ffc63-3e48-4a2b-a199-8c7edfecea5b · outbound

This paper cites Gemini 2.5: Our most intelligent ai model, March 2025.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Gemini 2.5: Our most intelligent ai model, March 2025

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:44:52.995163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T16:44:50.200843Z digest=sha256:076365180914e23238f7897993dd87799fb8299953b063e8432e316d24c334f5

Observation d4722487-92ec-4734-942d-1a862edfe79b · outbound

This paper cites Decoupled Weight Decay Regularization.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Decoupled Weight Decay Regularization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.258889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.258889Z digest=sha256:1abe28bc743ac761bafb9042d61bdbe85263de9ff3f9a38f57f78d9821b5089b

Observation 2ee6a2d9-e86e-4377-8003-f89dea17118c · outbound

This paper cites Continuous control with deep reinforcement learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Continuous control with deep reinforcement learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.353512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.353512Z digest=sha256:25724c78f62c8f6ca5d629dfa19fa228d2f025e72abaf606d3538404cc96d9e9

Observation 151b4ce6-8c6f-4147-9149-2611c709b90f · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) A minimalist approach to offline reinforcement learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.436962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.436962Z digest=sha256:0e0bffcf5c0f1172234e2fe1a3b1944821656db7e0266dbfa517e17a6213882e

Observation 5f45d04a-2659-4dcf-a7de-a64126f29578 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.531632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.531632Z digest=sha256:24add6c69ef84490849924cffcb50dbe5e6a422dabb6ac941caea1469f9bef2c

Observation 9adab734-26de-496a-b29b-9eb92f25db63 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Conservative q-learning for offline reinforcement learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.671535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.671535Z digest=sha256:b2516cbe3d492a589e062fbf9b5db3074eb792e4b757f8173ffbb84fe91d1d1c

Observation 5eb94329-e8c1-470b-b550-a4b008815cab · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.738027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.738027Z digest=sha256:fd2547328a72e4964340d3c1cda3e2c24123641f3fc20690521618661c8dd5a9

Observation f70e476b-efaf-42b7-a1db-205748ae26fa · outbound

This paper cites Accelerate: Training and inference at scale made simple, efficient and adaptable.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Accelerate: Training and inference at scale made simple, efficient and adaptable

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.830477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.830477Z digest=sha256:896108aaf52bae23311d9e83cfe77c271cebc8aab11e7d97877ded3abad8cc0c

Observation c5d0cf3f-3ad7-4711-a8e4-f2ea0bb0b5a9 · outbound

This paper cites SGDR : Stochastic gradient descent with warm restarts.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) SGDR : Stochastic gradient descent with warm restarts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:50.909102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:50.909102Z digest=sha256:279c96e24a0d0b4834598ec423d90442fb531508199aac2b6f1b73626d32788e

Observation c37973e0-39b8-43df-a3d4-68712b60c06e · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Adam: A Method for Stochastic Optimization

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:51.047658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:51.047658Z digest=sha256:39211f8b1be9bd1db33e33d73c748699ba506d09b4ed3c2c69707781cace15b3

Pith citing papers

Observation 54a57a42-7d6b-497a-b993-a839ecc1914d · inbound

Proximal Supervised Fine-Tuning cites this paper.

Proximal Supervised Fine-Tuning Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:42:50.932782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T20:42:14.423836Z digest=sha256:580418564791a0f942735a87f4da5eba6688e461baca4a5e0757857333d29dd9

Observation 3f61dc23-2063-46c8-8d16-a499a3e67153 · inbound

Multi-LLM Orchestration for High-Quality Code Generation: Exploiting Complementary Model Strengths cites this paper.

Multi-LLM Orchestration for High-Quality Code Generation: Exploiting Complementary Model Strengths Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:22:33.306633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T10:22:09.037783Z digest=sha256:8af1b97687b4665d77243b65150eabc31899e0703fd3a98bc8396b7137ddd173

Observation 1c44e9b4-86b8-4c9d-9515-fd5da3410e5d · inbound

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors cites this paper.

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:51:44.181536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T19:05:51.423427Z digest=sha256:f016230fb938a4ea13225a5c9862adb0f6c8ae7bedffcdd89d0a0753932c62d3

Observation 5a99cba9-220f-481b-9f34-85429aad88fb · inbound

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent cites this paper.

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:23.212875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T19:34:22.546508Z digest=sha256:f008a980af8738d1d15af8cf593155559d3fa16c594a2b6e9b3a25d75ee2c782

Observation c53f467c-925c-4fd4-b0f5-bd44f83690f3 · inbound

Rotation-Preserving Supervised Fine-Tuning cites this paper.

Rotation-Preserving Supervised Fine-Tuning Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:27:24.422736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T06:26:20.393476Z digest=sha256:ce42ebe16708255c40e422f063016da2357217c5fd4c0543cf29418ec5ba7047

Observation b6439b38-6462-46eb-b7c8-88e767908f8d · inbound

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models cites this paper.

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.954844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-29T08:01:39.412431Z digest=sha256:f7f2c052262cf525ed7ab00310c9401e10f56ec022d9800baf3b02e44d7981a9

Observation 353feea3-2ed2-4330-92c0-23c886e339e7 · inbound

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization cites this paper.

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:22:47.435668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T23:16:49.358792Z digest=sha256:57e05c2701c9ff42f784bff00a722002de193c1ffb28815554f349e100e175ec

Observation 4d8fce38-062b-45f3-8b32-d81b7b5715d8 · inbound

PriFT: Prior-Support Guided Supervised Fine-Tuning cites this paper.

PriFT: Prior-Support Guided Supervised Fine-Tuning Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:57:29.908830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T16:58:09.015766Z digest=sha256:2f16652f623a74acbd981ce35aef0abb52ac19f0f2e76a7a775099bc51aedc24

Observation 197441d6-2553-4e83-98c0-184d73a79727 · inbound

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff cites this paper.

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:27:26.355797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T18:49:47.876179Z digest=sha256:08ece810d807c535fedc6b14328b0ef3361899682498f628fba2df48b55ea0ab

Observation 4e3531e1-fe33-4e57-97ba-b04f0faf917c · inbound

Compatibility-Aware Dynamic Fine-Tuning for Large Language Models cites this paper.

Compatibility-Aware Dynamic Fine-Tuning for Large Language Models Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-05T02:30:40.844795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-05T02:27:50.369374Z digest=sha256:b9d179391585f7c5b0b5f4d36c8eeb91863ad5a7106973c8c3ef8d85a67ec176

Observation 01422f85-b83c-4c76-ab8e-e3f2056d0f89 · inbound

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning cites this paper.

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:29:57.043899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T00:39:29.703711Z digest=sha256:936e5da1eba6995fcc6853ef2e68dd483c594dbb9a129a8194e3d27eea4fb141

Observation 6f56a063-4987-42b8-8411-20189f8b2049 · inbound

A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training cites this paper.

A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T17:03:13.686066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T17:03:13.686066Z digest=sha256:4faad6c9a45a02051fb247230bc68172c8b608aea472c37966572944079becf9