Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T11:15:06.569421Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2607.07435.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T11:15:06.569421Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bae96de6-442c-45a9-9929-c2d07073cf79 · outbound
RLVP: Penalize the Path, Reward the Outcome VeriGate: Verifier-Gated Step-Level Supervision for GRPO
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8127ab77-f987-4b2f-96ae-af671848b9ac · outbound
RLVP: Penalize the Path, Reward the Outcome Constitutional AI: Harmlessness from AI Feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dec4b046-10b2-4450-98e7-4610de46580d · outbound
RLVP: Penalize the Path, Reward the Outcome $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 277e9b28-042c-4fb3-bd60-d5bc20a1880a · outbound
RLVP: Penalize the Path, Reward the Outcome Stop summation: Min-form credit assignment is all process reward model needs for reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 23faed5f-88eb-4f72-ac68-f495bf80106c · outbound
RLVP: Penalize the Path, Reward the Outcome DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 394678d5-61c0-451d-8924-8e7731e4d3e1 · outbound
RLVP: Penalize the Path, Reward the Outcome Group-in-Group Policy Optimization for LLM Agent Training
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 63e6b46b-70b4-4566-8f46-41b62353644a · outbound
RLVP: Penalize the Path, Reward the Outcome A sober look at progress in language model reasoning: Pitfalls and paths to repro- ducibility.arXiv preprint arXiv:2504.07086
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 197a620a-923f-4a98-a041-77e89dc92f05 · outbound
RLVP: Penalize the Path, Reward the Outcome Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cb0895ee-fa12-415a-95cf-58c0b8c7ecdc · outbound
RLVP: Penalize the Path, Reward the Outcome VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f943a7d1-3918-4d6f-a52e-8422a5d57d79 · outbound
RLVP: Penalize the Path, Reward the Outcome Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 353c53fe-8fb3-4a4e-ad63-ec7f4ebd72be · outbound
RLVP: Penalize the Path, Reward the Outcome V., Jeon, M., Vu, K., Lai, V., and Yang, E
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 931625b7-2dc1-4c9e-9341-909ba77a86c8 · outbound
RLVP: Penalize the Path, Reward the Outcome RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 423506aa-75f6-42fd-9526-0049c53259f3 · outbound
RLVP: Penalize the Path, Reward the Outcome Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f2614bf8-64dd-4a0d-bceb-15ae005edd3d · outbound
RLVP: Penalize the Path, Reward the Outcome Agentic reinforcement learning with implicit step rewards
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 75200a37-0b2d-4dd6-9cd9-18eb17d7ee96 · outbound
RLVP: Penalize the Path, Reward the Outcome Ngrpo: Negative-enhanced group relative policy optimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d992098a-e9c6-4fb6-a389-bac3dbbb4b13 · outbound
RLVP: Penalize the Path, Reward the Outcome ToolRL: Reward is All Tool Learning Needs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 520d274b-f92c-4b51-a10f-b89807016227 · outbound
RLVP: Penalize the Path, Reward the Outcome Spurious Rewards: Rethinking Training Signals in RLVR
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b1aff8df-40b0-4f09-9a91-27824b376e08 · outbound
RLVP: Penalize the Path, Reward the Outcome DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4e287174-7fa2-40a7-b574-76043c248d4e · outbound
RLVP: Penalize the Path, Reward the Outcome InThe Thirty-ninth Annual Conference on Neural Information Process- ing Systems
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 10e0078b-dea0-4119-b6ce-119e875d93c1 · outbound
RLVP: Penalize the Path, Reward the Outcome Qwen3 Technical Report
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8148a7e5-e65c-4119-b98f-7bc318c68640 · outbound
RLVP: Penalize the Path, Reward the Outcome SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 20eb2f64-cb0c-4e91-b943-b0c94bba18ae · outbound
RLVP: Penalize the Path, Reward the Outcome Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5b01ef1b-9cf8-4dbf-9158-8d55ba54e72a · outbound
RLVP: Penalize the Path, Reward the Outcome Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3f847dfc-dd58-4782-98ab-0a6cb32887bc · outbound
RLVP: Penalize the Path, Reward the Outcome Learning to Reason under Off-Policy Guidance
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a93dd69b-62cc-4eb5-861c-0cde563cb671 · outbound
RLVP: Penalize the Path, Reward the Outcome SWE-smith: Scaling Data for Software Engineering Agents
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8c9409aa-ee97-427b-a6e2-af585555535e · outbound
RLVP: Penalize the Path, Reward the Outcome $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ba6189f5-e6f4-4a7d-bb0b-5138b8c9a399 · outbound
RLVP: Penalize the Path, Reward the Outcome DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ae02f8ba-5a21-4114-bb82-ad1984456372 · outbound
RLVP: Penalize the Path, Reward the Outcome StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 79109ab6-23c1-4cf8-92d0-1d657aea6bb6 · outbound
RLVP: Penalize the Path, Reward the Outcome Free Process Rewards without Process Labels
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9d92d757-4540-435a-b8fd-b70cdcc849c2 · outbound
RLVP: Penalize the Path, Reward the Outcome From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4cbc42ed-6ff0-4f90-9a42-1cc2f6b83698 · outbound
RLVP: Penalize the Path, Reward the Outcome Group Sequence Policy Optimization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ab5052ff-f60f-462d-a3fe-4fb87f104a87 · outbound
RLVP: Penalize the Path, Reward the Outcome Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 414a6a04-8634-45b8-8de5-fce0ac8c5b3d · outbound
RLVP: Penalize the Path, Reward the Outcome The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 886e91d1-f6b4-4582-ade4-f5e7c488cfdd · outbound
RLVP: Penalize the Path, Reward the Outcome This appendix gives the full experiments
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7edaf817-368b-426d-a2dc-dfe408ef7fbf · outbound
RLVP: Penalize the Path, Reward the Outcome On SWE-bench software repair [Jimenez et al., 2024] the natural potential is the fraction of hidden FAIL_TO_PASS tests an episode makes pass
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8555e251-bec1-4240-82e7-c49646413fba · outbound
RLVP: Penalize the Path, Reward the Outcome Benefit appears only past a reachability threshold
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3ec87f6f-a0c5-431c-bdf3-1c45ae51a97b · outbound
RLVP: Penalize the Path, Reward the Outcome fragility,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 133097ed-8ffc-43ae-9a97-925b6d60e5ba · outbound
RLVP: Penalize the Path, Reward the Outcome any- valid-tactic
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0c7ce4fd-3d7b-4cd1-adec-d7b0caeb0e9e · outbound
RLVP: Penalize the Path, Reward the Outcome The learned critic is a poor training reward in every quadrant; its only robust use is offline diagnosis at the intent frontier
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.