Pith. sign in

Paper Citation Record · LEDGER

RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

As of 9 August 2026, this Paper Citation Record lists 2 of 2 outbound references and 6 inbound Pith citation observations for arXiv:2508.16546.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16546 v1

Coverage vector

measured 2 of 2 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:21:25.280778Z

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:51:24.929359Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:08:57.813685Z

Reference resolution

2 of 2 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf3bf01c-0ebf-45bb-ae2d-bdd292bf06fc · outbound

This paper cites Learning Dynamics of LLM Finetuning.

RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs Learning Dynamics of LLM Finetuning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:21:25.280778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:21:25.280778Z digest=sha256:26f65a4f209bf98129c9b5aed4d5b442ba4b8259deca97200104347c480a5fed

Observation 25b08014-c2c5-4114-9a8a-66e19e7ab620 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T17:21:25.215188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:21:25.215188Z digest=sha256:bba8719da907046b477168d2f1cb956c4aca78723057f2c24d8bc1a0a4104e0d

Pith citing papers

Observation 2bd11310-06fa-447c-b220-ffbeb8dc6875 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 247

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.606031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:a148e8afd3fdd53608e9f63a321d88abe647d0a046751f50c69d6c7edbfde034

Observation 8e48779b-a561-4402-84fd-9a89f9097237 · inbound

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability cites this paper.

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.629820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:53:22.553430Z digest=sha256:949210c85b1b0d55feec46be191ce732b6c62124a8fa52546786bb57e9a143f6

Observation 3fe11734-8d35-4637-8da6-aaad2aaf8d73 · inbound

Visual Reasoning through Tool-supervised Reinforcement Learning cites this paper.

Visual Reasoning through Tool-supervised Reinforcement Learning RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:03.023471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T03:05:21.688216Z digest=sha256:19c543c8031e54b8741aac01e7b63f1e114e16001d002ab9cfb6a15d75419712

Observation bd565c28-485f-4cb0-9e7a-9263ae2c2ae4 · inbound

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff cites this paper.

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:27:26.418219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T18:49:47.876179Z digest=sha256:87e8f031cffc5b51aa37ca75c1756e70a2590488e208990d0877dc30d528e293

Observation 26b368a4-02fb-414a-8bd5-4864e047040c · inbound

Sparsity Curse: Understanding RLVR Model Parameter Space from Model Merging cites this paper.

Sparsity Curse: Understanding RLVR Model Parameter Space from Model Merging RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:57.816892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T00:57:02.938997Z digest=sha256:297f02bd01f19fc025aaea8c66fc4a2d2ad39d01283519e2c28ac2ffce9b99a5

Observation b7ae546c-595f-4e1e-a732-6e2679834658 · inbound

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs cites this paper.

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T00:51:24.929359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:51:24.929359Z digest=sha256:ff7419bb283b503ab0399ae36009929d36fc33dadea472ba6ca1e18234945323