Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T11:57:56.105261Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 0 inbound Pith citation observations for arXiv:2601.04805.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T11:57:56.105261Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
6 of 6 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c3668722-a034-4959-8801-38af745e7961 · outbound
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca149062-f3c4-4f18-b192-c05c9a81c0c0 · outbound
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning Feedback Loops With Language Models Drive In-Context Reward Hacking
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f52f6b69-0b3a-4462-a845-512b93b01a08 · outbound
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92821b1a-0390-4fcf-8b03-debe79311163 · outbound
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning Let's Verify Step by Step
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac5ae07d-7a07-4554-bcb2-15a2561fdcd8 · outbound
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 395467f8-3c9a-47b1-adba-a8327d70034c · outbound
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.