Pith. sign in

Paper Citation Record · LEDGER

RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

As of 19 August 2026, this Paper Citation Record lists 2 of 2 outbound references and 6 inbound Pith citation observations for arXiv:2508.16546.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16546 v1

Coverage vector

measured 2 of 2 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:21:25.280778Z

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:51:24.929359Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:08:57.813685Z

Reference resolution

2 of 2 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf3bf01c-0ebf-45bb-ae2d-bdd292bf06fc · outbound

This paper cites Learning Dynamics of LLM Finetuning.

RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs Learning Dynamics of LLM Finetuning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:21:25.280778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:21:25.280778Z digest=sha256:f351ea45bb029a4e932363b55da7d7ba345b6a4af03fa6bbb8dd544e2def001e

Observation 25b08014-c2c5-4114-9a8a-66e19e7ab620 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T17:21:25.215188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:21:25.215188Z digest=sha256:7ccae6dd56b11d99add415c3d73fb49a535178c4b3c4a57e73eb5ef215259453

Pith citing papers

Observation 2bd11310-06fa-447c-b220-ffbeb8dc6875 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 247

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.606031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:ca18058220e429e2aabc852a141a8dfc7828dbadc5519fdef23b9107b6632a14

Observation 8e48779b-a561-4402-84fd-9a89f9097237 · inbound

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability cites this paper.

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.629820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:53:22.553430Z digest=sha256:4b37739f49e14c2dd3c10cc629865f3ca67d9ccb7499f418bcf87296eaee4ab5

Observation 3fe11734-8d35-4637-8da6-aaad2aaf8d73 · inbound

Visual Reasoning through Tool-supervised Reinforcement Learning cites this paper.

Visual Reasoning through Tool-supervised Reinforcement Learning RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:03.023471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T03:05:21.688216Z digest=sha256:544471188ffe0ee19738c793980d64a274779874213eb2ec010d7060ad5e0440

Observation bd565c28-485f-4cb0-9e7a-9263ae2c2ae4 · inbound

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff cites this paper.

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:27:26.418219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T18:49:47.876179Z digest=sha256:61f0cdf1d65eca402170f7aa6287b098115ab057e6df889dcafc5fb1e072d036

Observation 26b368a4-02fb-414a-8bd5-4864e047040c · inbound

Sparsity Curse: Understanding RLVR Model Parameter Space from Model Merging cites this paper.

Sparsity Curse: Understanding RLVR Model Parameter Space from Model Merging RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:57.816892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T00:57:02.938997Z digest=sha256:9af8ef91ee04a0d45bd7a211396ef821baf8acf3cc15f3ce9a717198862a9001

Observation b7ae546c-595f-4e1e-a732-6e2679834658 · inbound

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs cites this paper.

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T00:51:24.929359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:51:24.929359Z digest=sha256:1646f4436abbed0b13fb73342e28b9237cfed568e0abfcea78184b31ec4c298e