Pith. sign in

Paper Citation Record · LEDGER

RLVP: Penalize the Path, Reward the Outcome

As of 18 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2607.07435.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07435 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T11:15:06.569421Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact29
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bae96de6-442c-45a9-9929-c2d07073cf79 · outbound

This paper cites VeriGate: Verifier-Gated Step-Level Supervision for GRPO.

RLVP: Penalize the Path, Reward the Outcome VeriGate: Verifier-Gated Step-Level Supervision for GRPO

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.348368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:bf7e85c1a01564acdfbf76217eba0ee25e1621b8ba13df3a09135bfff008e5e1

Observation 8127ab77-f987-4b2f-96ae-af671848b9ac · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

RLVP: Penalize the Path, Reward the Outcome Constitutional AI: Harmlessness from AI Feedback

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.338929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:7124e420167227f31ecfe1abdbbf292987cedc83d3ce565d15c9d5be13026805

Observation dec4b046-10b2-4450-98e7-4610de46580d · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

RLVP: Penalize the Path, Reward the Outcome $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.370402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:8ce6e61e467edd10e626a2efa25ccbd8ecdd3ff07745cf4748e1f0466d51a06f

Observation 277e9b28-042c-4fb3-bd60-d5bc20a1880a · outbound

This paper cites Stop summation: Min-form credit assignment is all process reward model needs for reasoning.

RLVP: Penalize the Path, Reward the Outcome Stop summation: Min-form credit assignment is all process reward model needs for reasoning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-09T11:16:11.374420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:8ddf0875e34b497ac44d8a7e905c52e44c0b7b2f1f01c67d6336656b81b175eb

Observation 23faed5f-88eb-4f72-ac68-f495bf80106c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RLVP: Penalize the Path, Reward the Outcome DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.385558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:9722f7715f350208d3f3d86b51fc2dc01926128ac5e46d7611eeb1f6e2559f50

Observation 394678d5-61c0-451d-8924-8e7731e4d3e1 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

RLVP: Penalize the Path, Reward the Outcome Group-in-Group Policy Optimization for LLM Agent Training

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.375765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:1dac1cc400da7895f1dac3627558a1bc15bba7432c7029b494e0c73076dbe006

Observation 63e6b46b-70b4-4566-8f46-41b62353644a · outbound

This paper cites A sober look at progress in language model reasoning: Pitfalls and paths to repro- ducibility.arXiv preprint arXiv:2504.07086.

RLVP: Penalize the Path, Reward the Outcome A sober look at progress in language model reasoning: Pitfalls and paths to repro- ducibility.arXiv preprint arXiv:2504.07086

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-09T11:16:11.347992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:041fccb32679ae616e198f3c4100cc926afb900edbad76280d4dae0dd3837b29

Observation 197a620a-923f-4a98-a041-77e89dc92f05 · outbound

This paper cites Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

RLVP: Penalize the Path, Reward the Outcome Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T11:16:11.516544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:e8c0f9ba99765b30e4efe6fa488af8cc8e459848c8da82882143ba7799e5a28c

Observation cb0895ee-fa12-415a-95cf-58c0b8c7ecdc · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

RLVP: Penalize the Path, Reward the Outcome VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T11:16:11.298001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:a9f5f805e6878b46f677db1a2124897958abd432ea441e986c1d10d0d6e9fd29

Observation f943a7d1-3918-4d6f-a52e-8422a5d57d79 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

RLVP: Penalize the Path, Reward the Outcome Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.378129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:736f30d18b474b9ad5c7e2a4921c3abab0d45556a9faa986e7c34237784704a6

Observation 353c53fe-8fb3-4a4e-ad63-ec7f4ebd72be · outbound

This paper cites V., Jeon, M., Vu, K., Lai, V., and Yang, E.

RLVP: Penalize the Path, Reward the Outcome V., Jeon, M., Vu, K., Lai, V., and Yang, E

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-09T11:16:11.377458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:a1980807628154719114323bfe13e71ab8692d81d7fc59bc23c2b5f654802334

Observation 931625b7-2dc1-4c9e-9341-909ba77a86c8 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

RLVP: Penalize the Path, Reward the Outcome RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.382159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:78eb471ab0c6cab8c2117fca55653b71812f7ad2f0b96da4f43cc681d157c880

Observation 423506aa-75f6-42fd-9526-0049c53259f3 · outbound

This paper cites Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning.

RLVP: Penalize the Path, Reward the Outcome Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.342496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:4170d3e4a0de3751b9be97943fa96e71fbf21ffa61b55cdd77759f8492e8d49c

Observation f2614bf8-64dd-4a0d-bceb-15ae005edd3d · outbound

This paper cites Agentic reinforcement learning with implicit step rewards.

RLVP: Penalize the Path, Reward the Outcome Agentic reinforcement learning with implicit step rewards

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-09T11:16:11.364593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:1b52bb9b7c578da9769ea6720001ee80c4ee0a46e67d0106a0c9819bed836b29

Observation 75200a37-0b2d-4dd6-9cd9-18eb17d7ee96 · outbound

This paper cites Ngrpo: Negative-enhanced group relative policy optimization.

RLVP: Penalize the Path, Reward the Outcome Ngrpo: Negative-enhanced group relative policy optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-09T11:16:11.385216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:039ec52be272850bed4a048af17e61ed72f3689907de670c35007851b1927e14

Observation d992098a-e9c6-4fb6-a389-bac3dbbb4b13 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

RLVP: Penalize the Path, Reward the Outcome ToolRL: Reward is All Tool Learning Needs

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.383178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:f222962101a8e3f182f26425fd0c3a661d54cd1d05dc8971faaee8d66d23da34

Observation 520d274b-f92c-4b51-a10f-b89807016227 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

RLVP: Penalize the Path, Reward the Outcome Spurious Rewards: Rethinking Training Signals in RLVR

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.379832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:c622f9948c5db2865438f13929d143f30e632413053e988d6fda0c34091e639d

Observation b1aff8df-40b0-4f09-9a91-27824b376e08 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RLVP: Penalize the Path, Reward the Outcome DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.390498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:9391ab5b419208ca42d0d393cc0c67bd1853f4a6254f598641cb87214a380c03

Observation 4e287174-7fa2-40a7-b574-76043c248d4e · outbound

This paper cites InThe Thirty-ninth Annual Conference on Neural Information Process- ing Systems.

RLVP: Penalize the Path, Reward the Outcome InThe Thirty-ninth Annual Conference on Neural Information Process- ing Systems

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T11:16:11.304849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:7fe7b8a113f59079a5b2234ab0e0b696ee820f0f07c39af995a822704d85b530

Observation 10e0078b-dea0-4119-b6ce-119e875d93c1 · outbound

This paper cites Qwen3 Technical Report.

RLVP: Penalize the Path, Reward the Outcome Qwen3 Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.388106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:8c22301653a7a9946bf72d9add0f94714e5c1364093a7f46263e24842d27dcae

Observation 8148a7e5-e65c-4119-b98f-7bc318c68640 · outbound

This paper cites SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution.

RLVP: Penalize the Path, Reward the Outcome SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.361536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:39607d3cdbbb0dfb8f5b1581d48b801c21bb1946147e6c55a00ec008b348f362

Observation 20eb2f64-cb0c-4e91-b943-b0c94bba18ae · outbound

This paper cites Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi.

RLVP: Penalize the Path, Reward the Outcome Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T11:16:11.393946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:c8234d847c83b14f5a7b8bdf79f80906ab07b9df7259d644c0c2fef780e13a52

Observation 5b01ef1b-9cf8-4dbf-9158-8d55ba54e72a · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

RLVP: Penalize the Path, Reward the Outcome Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.387942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:bb3a64fa2795eb69dd841b62d95c31234c11ca2b9bd8e89af1ac5e1c34cfd76e

Observation 3f847dfc-dd58-4782-98ab-0a6cb32887bc · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

RLVP: Penalize the Path, Reward the Outcome Learning to Reason under Off-Policy Guidance

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.366321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:550230496f289637e6e6e8ed5add10dcb1045d67bc3d74a6356a4ca772b49f18

Observation a93dd69b-62cc-4eb5-861c-0cde563cb671 · outbound

This paper cites SWE-smith: Scaling Data for Software Engineering Agents.

RLVP: Penalize the Path, Reward the Outcome SWE-smith: Scaling Data for Software Engineering Agents

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.318401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:4cb0057d5a3e4f51cb2d4a36be0f383d90b9617f483252d8a7c7aa977d9030ed

Observation 8c9409aa-ee97-427b-a6e2-af585555535e · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

RLVP: Penalize the Path, Reward the Outcome $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.360782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:3b65d1aa0c662c7ba27d64d66bf0b97fa2f3fd913147f9de6cb981231b20054a

Observation ba6189f5-e6f4-4a7d-bb0b-5138b8c9a399 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

RLVP: Penalize the Path, Reward the Outcome DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.371468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:3151d0ebeb13b5c8318293f6063ef1ba154585b531e23c5baef37b7ad6c2da27

Observation ae02f8ba-5a21-4114-bb82-ad1984456372 · outbound

This paper cites StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning.

RLVP: Penalize the Path, Reward the Outcome StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.390626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:c85e206ecb261fe4a8525cbc4786b73a300fa4d523bb484b32ebfc5a8963e8f5

Observation 79109ab6-23c1-4cf8-92d0-1d657aea6bb6 · outbound

This paper cites Free Process Rewards without Process Labels.

RLVP: Penalize the Path, Reward the Outcome Free Process Rewards without Process Labels

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.354882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:a4f9eff55a54a7bd8936d3c366c3b31d73d4cc053ead1f9383c3475104370f9e

Observation 9d92d757-4540-435a-b8fd-b70cdcc849c2 · outbound

This paper cites From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models.

RLVP: Penalize the Path, Reward the Outcome From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.380852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:53f42085179df5b1b8cfa63e58d97f55fb31f485515e0bebd3fa77e4059089df

Observation 4cbc42ed-6ff0-4f90-9a42-1cc2f6b83698 · outbound

This paper cites Group Sequence Policy Optimization.

RLVP: Penalize the Path, Reward the Outcome Group Sequence Policy Optimization

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.322505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:9bd828a39bd5f38b36f96de2110da817b0597f400b0ebb5a84da2521e24ac271

Observation ab5052ff-f60f-462d-a3fe-4fb87f104a87 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

RLVP: Penalize the Path, Reward the Outcome Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.341579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:a7d94527d22f168dc5ea615347294e313a5fab8b4537f1a539eb3e2769f0f589

Observation 414a6a04-8634-45b8-8de5-fce0ac8c5b3d · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347.

RLVP: Penalize the Path, Reward the Outcome The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-09T11:16:11.339361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:5b4e32fbc3573185fd1a45c57f462374d0319e3db7b02c2dfb83ab53b9b13b8c

Observation 886e91d1-f6b4-4582-ade4-f5e7c488cfdd · outbound

This paper cites This appendix gives the full experiments.

RLVP: Penalize the Path, Reward the Outcome This appendix gives the full experiments

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T11:16:11.523517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:4f0c13e1337e37146d1072f475ba5da7d8537ffeebd8b97593dfa81e913b3be9

Observation 7edaf817-368b-426d-a2dc-dfe408ef7fbf · outbound

This paper cites On SWE-bench software repair [Jimenez et al., 2024] the natural potential is the fraction of hidden FAIL_TO_PASS tests an episode makes pass.

RLVP: Penalize the Path, Reward the Outcome On SWE-bench software repair [Jimenez et al., 2024] the natural potential is the fraction of hidden FAIL_TO_PASS tests an episode makes pass

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T11:16:11.528376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:5b054f2b18beacb75ee0ac0ebd1738cb446bb6c2aeb2c76f89033db9e55a65ab

Observation 8555e251-bec1-4240-82e7-c49646413fba · outbound

This paper cites Benefit appears only past a reachability threshold.

RLVP: Penalize the Path, Reward the Outcome Benefit appears only past a reachability threshold

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T11:16:11.521743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:60c716cd086b0b1551cd948af9655832c0fe7f79c8b1c2e801ccca7264a6e52f

Observation 3ec87f6f-a0c5-431c-bdf3-1c45ae51a97b · outbound

This paper cites fragility,.

RLVP: Penalize the Path, Reward the Outcome fragility,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T11:16:11.513997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:c08518397416810ce8cce01a168dc3baff7025d87e4756947db4381a4316b234

Observation 133097ed-8ffc-43ae-9a97-925b6d60e5ba · outbound

This paper cites any- valid-tactic.

RLVP: Penalize the Path, Reward the Outcome any- valid-tactic

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-07-09T11:16:11.526298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:672a7380cab847693d49c050034074db006de11a4991db287b351fe61c2c174e

Observation 0c7ce4fd-3d7b-4cd1-adec-d7b0caeb0e9e · outbound

This paper cites The learned critic is a poor training reward in every quadrant; its only robust use is offline diagnosis at the intent frontier.

RLVP: Penalize the Path, Reward the Outcome The learned critic is a poor training reward in every quadrant; its only robust use is offline diagnosis at the intent frontier

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T11:16:11.526004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:986105eead23e3686240593094a83c881ea40c32283e5f33d9b892781277c98e

Pith citing papers

No inbound Pith citation observations are available.