Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:44:51.047658Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 12 inbound Pith citation observations for arXiv:2507.12856.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:44:51.047658Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T17:03:13.686066Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
63 of 63 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation afb97173-c846-4dd5-aa11-24107e077d56 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Training language models to follow instructions with human feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae96b4f-752c-4734-b189-ad82e388fa32 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Fine-Tuning Language Models from Human Preferences
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dd8d10b-7d19-4cf4-837d-47a1138126a0 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e196e2bc-700e-4987-ab05-180d411725d2 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Constitutional AI: Harmlessness from AI Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde1362d-60b9-4d0e-8d05-604669c47cd2 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8300ce0c-2d2a-4072-8b57-4d8ea246b60a · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c232c5e-c530-4ae6-92d8-14ba6956428d · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62ef0cf6-7c98-4fcd-90de-0686f7a6e8d3 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86cc15b-f822-4864-ae8e-5ebee5ffefb5 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Direct preference optimization: Your language model is secretly a reward model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb073a7-0f9e-4f28-a85d-514a64fbb332 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Learning from negative feedback, or positive feedback or both
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5640087d-64d5-4b35-ab97-d304c031a795 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) s1: Simple test-time scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f91e8dd-5014-4aa0-afd3-33baf1614ff9 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efee556c-471a-43b1-a52e-e3c1a6e1fe78 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Peters and S
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b741377d-402c-46d2-8377-8ae21207358e · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Policy search for motor primitives in robotics
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 57f29dac-888b-4f4d-9cd4-52053917f724 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c0d0215-1793-4c1c-b3cc-6c7d45288ae8 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Measuring Mathematical Problem Solving With the MATH Dataset
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16728588-cee1-4bd4-a0ea-133600395631 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) D4rl: Datasets for deep data-driven reinforcement learning, 2020
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8e88bc7c-02fd-4134-8a20-025d8f5fa7ca · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7ab64d06-7140-4cfd-99d0-38a5b645dc80 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Learning math reasoning from self-sampled correct and partially-correct solutions
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 83d709d3-e8cb-4135-8314-b8b5f1ab78a7 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Learning to generalize from sparse and underspecified rewards
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d419a748-668a-4a22-8bd7-aaa10210dcb6 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Reward augmented maximum likelihood for neural structured prediction
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c2cdac58-84db-41e2-ad62-c8f93657213c · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Scaling relationship on learning mathematical reasoning with large language models, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a3e8de1-e1a8-473a-9616-c3df12e2d901 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Simplify rlhf as reward-weighted sft: A variational method, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 257f81c1-d435-4a4a-875b-84f7e81c6f2d · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Reinforced Self-Training (ReST) for Language Modeling
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 771076c3-756d-4b82-bf86-b02498f3c9c5 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Beyond human data: Scaling self-training for problem-solving with language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fa584de8-e969-4035-84af-959a13e2f10c · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Peters, K
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c3b15e6-3dbd-407f-982c-21099dd00054 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Efficient iterative policy optimization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c08c25d-820a-4ef7-9dae-a1be44beee82 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Tighter bounds lead to improved classifiers
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7997a69a-b90a-4623-8aca-0ab0520792b4 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Process for adapting language models to society (palms) with values-targeted datasets
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c11d96f1-3267-4067-9e90-17b4b8c3e13d · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Self-consuming generative models with curated data provably optimize human preferences
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a97c8070-9390-4a67-b0ff-4be30d82d30a · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Maximum a posteriori policy optimisation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 45f416fc-0333-48e9-97be-f7b45475e876 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) On Multi-objective Policy Optimization as a Tool for Reinforcement Learning: Case Studies in Offline RL and Finetuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29d5868e-4a55-4112-ad0a-c0eea351d23a · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4f35b25-c428-4649-b9d0-a82cccd0435b · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Proximal Policy Optimization Algorithms
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d4a7ba3-4f64-4919-8e9d-09a6ea8441d3 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Offline Reinforcement Learning with Implicit Q-Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b54d76e-aa16-4dc0-9478-4302dc1c0d4b · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) When does return-conditioned supervised learning work for offline reinforcement learning? In S
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 57d6da3f-8e75-4292-9d15-21a48fb37d5f · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Decision transformer: Reinforcement learning via sequence modeling
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fdd94b3f-bdf5-4ec4-a8a4-7e990cac77b1 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Kahn and Andrew W
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a0d9b5b8-8fbb-4d1f-baef-1e95294083a1 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Rubinstein and Dirk P
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 88f4b4b9-f99e-4300-aeea-dd32ec107cd3 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Stochastic simulation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d1290aa6-45e6-4f4f-bdb3-f7946219a033 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Doubly robust off-policy value evaluation for reinforcement learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a26b3599-c974-47b4-808f-22057a1eaf14 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Policy optimization via importance sampling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f2d7eba5-97db-40d1-8f10-b4034d166bbe · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Andrad\'ottir, D
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cced70b8-1f35-4901-b3c6-883f2bf2c9db · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Qwen2.5 Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d50992c-f704-45d8-853f-9a1058b752be · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Numinamath, 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a673c9d4-9d75-4660-8609-a19683a7ce93 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Aime, February 2024
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c4b594e-f0a2-43cc-8651-0f6f4b6a0a1d · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 336f7a07-a1a9-4885-9ce8-204c9c234a51 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae6dc2d4-66ed-44d4-9409-536ced3b8c14 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Gemini 2.0 flash thinking mode (gemini-2.0-flash-thinking-exp-1219), December 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4ee868a3-7f94-42d6-b3e0-f291b78a9a62 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d49f26-0a3f-4942-8c82-e2dd23e1a758 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Learning to reason with llms, September 2024
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7470890e-5934-4df4-85ca-ca63db0f7f60 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Bespoke-stratos: The unreasonable effectiveness of reasoning distillation, 2025
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 668ec223-1aa3-4844-b3d9-059bb8fb9e58 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Sky-t1: Fully open-source reasoning model with o1-preview performance in \ 450 budget, 2025
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a47ffc63-3e48-4a2b-a199-8c7edfecea5b · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Gemini 2.5: Our most intelligent ai model, March 2025
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d4722487-92ec-4734-942d-1a862edfe79b · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Decoupled Weight Decay Regularization
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ee6a2d9-e86e-4377-8003-f89dea17118c · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Continuous control with deep reinforcement learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 151b4ce6-8c6f-4147-9149-2611c709b90f · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) A minimalist approach to offline reinforcement learning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f45d04a-2659-4dcf-a7de-a64126f29578 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9adab734-26de-496a-b29b-9eb92f25db63 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Conservative q-learning for offline reinforcement learning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eb94329-e8c1-470b-b550-a4b008815cab · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) PyTorch: An Imperative Style, High-Performance Deep Learning Library
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f70e476b-efaf-42b7-a1db-205748ae26fa · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Accelerate: Training and inference at scale made simple, efficient and adaptable
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5d0cf3f-3ad7-4711-a8e4-f2ea0bb0b5a9 · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) SGDR : Stochastic gradient descent with warm restarts
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c37973e0-39b8-43df-a3d4-68712b60c06e · outbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Adam: A Method for Stochastic Optimization
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54a57a42-7d6b-497a-b993-a839ecc1914d · inbound
Proximal Supervised Fine-Tuning Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3f61dc23-2063-46c8-8d16-a499a3e67153 · inbound
Multi-LLM Orchestration for High-Quality Code Generation: Exploiting Complementary Model Strengths Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1c44e9b4-86b8-4c9d-9515-fd5da3410e5d · inbound
Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5a99cba9-220f-481b-9f34-85429aad88fb · inbound
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c53f467c-925c-4fd4-b0f5-bd44f83690f3 · inbound
Rotation-Preserving Supervised Fine-Tuning Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b6439b38-6462-46eb-b7c8-88e767908f8d · inbound
Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 353feea3-2ed2-4330-92c0-23c886e339e7 · inbound
DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4d8fce38-062b-45f3-8b32-d81b7b5715d8 · inbound
PriFT: Prior-Support Guided Supervised Fine-Tuning Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 197441d6-2553-4e83-98c0-184d73a79727 · inbound
When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4e3531e1-fe33-4e57-97ba-b04f0faf917c · inbound
Compatibility-Aware Dynamic Fine-Tuning for Large Language Models Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 01422f85-b83c-4c76-ab8e-e3f2056d0f89 · inbound
Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6f56a063-4987-42b8-8411-20189f8b2049 · inbound
A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.