Pith. sign in

Paper Citation Record · LEDGER

RLVP: Penalize the Path, Reward the Outcome

As of 19 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2607.07435.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07435 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T11:15:06.569421Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact29
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bae96de6-442c-45a9-9929-c2d07073cf79 · outbound

This paper cites VeriGate: Verifier-Gated Step-Level Supervision for GRPO.

RLVP: Penalize the Path, Reward the Outcome VeriGate: Verifier-Gated Step-Level Supervision for GRPO

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.348368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:911badd40b24ab4de21aa6fb5019dae0c32b13779e6281c751f876ecf73f8148

Observation 8127ab77-f987-4b2f-96ae-af671848b9ac · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

RLVP: Penalize the Path, Reward the Outcome Constitutional AI: Harmlessness from AI Feedback

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.338929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:ed220f0468e57f7b0e7c1a9275338f5e723d226ce6c6b46b17f33037d74109af

Observation dec4b046-10b2-4450-98e7-4610de46580d · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

RLVP: Penalize the Path, Reward the Outcome $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.370402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:661fd349c7f04cb075d968d494b7022b7dbd012e00eb887e71a095b7d350e6ed

Observation 277e9b28-042c-4fb3-bd60-d5bc20a1880a · outbound

This paper cites Stop summation: Min-form credit assignment is all process reward model needs for reasoning.

RLVP: Penalize the Path, Reward the Outcome Stop summation: Min-form credit assignment is all process reward model needs for reasoning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-09T11:16:11.374420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:46b7692298ac86935d65eccf0b7fc0998aea7099162bee84edfd320da7feb26b

Observation 23faed5f-88eb-4f72-ac68-f495bf80106c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RLVP: Penalize the Path, Reward the Outcome DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.385558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:9398416be505e0c230fda03e88e9635cf06c9bfd6a074b0c71ca71506a74273b

Observation 394678d5-61c0-451d-8924-8e7731e4d3e1 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

RLVP: Penalize the Path, Reward the Outcome Group-in-Group Policy Optimization for LLM Agent Training

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.375765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:7819051627424fa734129dddb506f46e8b6a44eea63b37407d4ce5fd05fa51ae

Observation 63e6b46b-70b4-4566-8f46-41b62353644a · outbound

This paper cites A sober look at progress in language model reasoning: Pitfalls and paths to repro- ducibility.arXiv preprint arXiv:2504.07086.

RLVP: Penalize the Path, Reward the Outcome A sober look at progress in language model reasoning: Pitfalls and paths to repro- ducibility.arXiv preprint arXiv:2504.07086

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-09T11:16:11.347992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:5b75c233fdb87100547597be52b8dcaa5de598c9a10a74579cfe8017b35a5899

Observation 197a620a-923f-4a98-a041-77e89dc92f05 · outbound

This paper cites Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

RLVP: Penalize the Path, Reward the Outcome Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T11:16:11.516544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:9bd5dde490caa23ba7f47373c671386901b332a4017b9142fdc1137d8b49cb38

Observation cb0895ee-fa12-415a-95cf-58c0b8c7ecdc · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

RLVP: Penalize the Path, Reward the Outcome VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T11:16:11.298001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:06ec0eba10cf2fad217e58c680921cd5ad214a06cbbf1694d9cea484f10dd74a

Observation f943a7d1-3918-4d6f-a52e-8422a5d57d79 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

RLVP: Penalize the Path, Reward the Outcome Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.378129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:b61742ce4f74eb9e172ac70188d810ef66ff283056383b5c9508e83b409e68f2

Observation 353c53fe-8fb3-4a4e-ad63-ec7f4ebd72be · outbound

This paper cites V., Jeon, M., Vu, K., Lai, V., and Yang, E.

RLVP: Penalize the Path, Reward the Outcome V., Jeon, M., Vu, K., Lai, V., and Yang, E

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-09T11:16:11.377458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:b8d9526fd405f7a4dff198cee1667caa1afdbc69988b569141d0207260429de9

Observation 931625b7-2dc1-4c9e-9341-909ba77a86c8 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

RLVP: Penalize the Path, Reward the Outcome RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.382159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:c588607c83bbaaf1c063234a1e67f99d713f0ee65874be97a7cc961f3eaff58c

Observation 423506aa-75f6-42fd-9526-0049c53259f3 · outbound

This paper cites Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning.

RLVP: Penalize the Path, Reward the Outcome Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.342496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:1733b6970d4e7804a5a3d3de3587457171d01df095afc563a4b78bc6c629608f

Observation f2614bf8-64dd-4a0d-bceb-15ae005edd3d · outbound

This paper cites Agentic reinforcement learning with implicit step rewards.

RLVP: Penalize the Path, Reward the Outcome Agentic reinforcement learning with implicit step rewards

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-09T11:16:11.364593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:6f3631a95e38dd049510c98324aab9bb2e6cfae4774be1486dd0d80bfd774853

Observation 75200a37-0b2d-4dd6-9cd9-18eb17d7ee96 · outbound

This paper cites Ngrpo: Negative-enhanced group relative policy optimization.

RLVP: Penalize the Path, Reward the Outcome Ngrpo: Negative-enhanced group relative policy optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-09T11:16:11.385216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:22e4658d0029154b299739e8b8111e1bfb1b56c8996546004f308c812abeb1c5

Observation d992098a-e9c6-4fb6-a389-bac3dbbb4b13 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

RLVP: Penalize the Path, Reward the Outcome ToolRL: Reward is All Tool Learning Needs

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.383178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:2d2a1e8cc56fd827216b3014481e2bbc135a0d94c796682be47112fb1a8b7682

Observation 520d274b-f92c-4b51-a10f-b89807016227 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

RLVP: Penalize the Path, Reward the Outcome Spurious Rewards: Rethinking Training Signals in RLVR

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.379832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:c6691cf31db03b57bd57970cd2d1f2878df6daa1d829823d5c197d459d528733

Observation b1aff8df-40b0-4f09-9a91-27824b376e08 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RLVP: Penalize the Path, Reward the Outcome DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.390498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:65a01c6b990d46a8cb8d9cd9b138338ae3770c4e0924f7bf567dd2fc86e9adb1

Observation 4e287174-7fa2-40a7-b574-76043c248d4e · outbound

This paper cites InThe Thirty-ninth Annual Conference on Neural Information Process- ing Systems.

RLVP: Penalize the Path, Reward the Outcome InThe Thirty-ninth Annual Conference on Neural Information Process- ing Systems

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T11:16:11.304849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:331c024e09b3fb6d65a70698db01b66d54fa5846fe7caf90efb85cbda2709f91

Observation 10e0078b-dea0-4119-b6ce-119e875d93c1 · outbound

This paper cites Qwen3 Technical Report.

RLVP: Penalize the Path, Reward the Outcome Qwen3 Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.388106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:d23ae0045d724b278cf18bc067021aedb04b8c7e47940a058c53d964855c724e

Observation 8148a7e5-e65c-4119-b98f-7bc318c68640 · outbound

This paper cites SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution.

RLVP: Penalize the Path, Reward the Outcome SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.361536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:031defb2fb52fd8cf632b83450482d2a5bb72ba68f08dab80de080ebaf44c43f

Observation 20eb2f64-cb0c-4e91-b943-b0c94bba18ae · outbound

This paper cites Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi.

RLVP: Penalize the Path, Reward the Outcome Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T11:16:11.393946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:255cb5aa4c959e592e0fdce1c1d7a800bbaa9a5869a575816568f6ae734248a2

Observation 5b01ef1b-9cf8-4dbf-9158-8d55ba54e72a · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

RLVP: Penalize the Path, Reward the Outcome Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.387942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:3f51481989e57b513304c0eb013faa066d3754ac94a0fa79a200b7c3f14058f8

Observation 3f847dfc-dd58-4782-98ab-0a6cb32887bc · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

RLVP: Penalize the Path, Reward the Outcome Learning to Reason under Off-Policy Guidance

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.366321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:745a44c31e2b0b007127e92aa9bff02c8f458aa5d225cab942033df383ea8c62

Observation a93dd69b-62cc-4eb5-861c-0cde563cb671 · outbound

This paper cites SWE-smith: Scaling Data for Software Engineering Agents.

RLVP: Penalize the Path, Reward the Outcome SWE-smith: Scaling Data for Software Engineering Agents

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.318401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:3561a03bcf07d9426c7af2a17f1b6e27d04013120efec83a6f828d1ee94eb55f

Observation 8c9409aa-ee97-427b-a6e2-af585555535e · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

RLVP: Penalize the Path, Reward the Outcome $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.360782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:fe43abd6dc9868621eb6e0c3653221ff2aa699638723d81cbf414edefd12aa63

Observation ba6189f5-e6f4-4a7d-bb0b-5138b8c9a399 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

RLVP: Penalize the Path, Reward the Outcome DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.371468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:c6e0e963432ab8e557ba5dd5473747a60da7ea227385f95f5030eee9bdf614a0

Observation ae02f8ba-5a21-4114-bb82-ad1984456372 · outbound

This paper cites StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning.

RLVP: Penalize the Path, Reward the Outcome StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.390626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:7b4ee2aa14986d3d13d607d97b5d3f21d038afefa28925aca67d159c4e68ade4

Observation 79109ab6-23c1-4cf8-92d0-1d657aea6bb6 · outbound

This paper cites Free Process Rewards without Process Labels.

RLVP: Penalize the Path, Reward the Outcome Free Process Rewards without Process Labels

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.354882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:f071bc0094ea163f5f600502b67b92af944e8451d6d66fdc6a38edc2ecfcddbc

Observation 9d92d757-4540-435a-b8fd-b70cdcc849c2 · outbound

This paper cites From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models.

RLVP: Penalize the Path, Reward the Outcome From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.380852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:fb38a98b187e17f71a84684589097910ded61170b8cfbd697220fc1312571a68

Observation 4cbc42ed-6ff0-4f90-9a42-1cc2f6b83698 · outbound

This paper cites Group Sequence Policy Optimization.

RLVP: Penalize the Path, Reward the Outcome Group Sequence Policy Optimization

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.322505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:5ff8553126c57bd90116ad46fd3665a1c0364a3a7c91d95809472a6a07432cda

Observation ab5052ff-f60f-462d-a3fe-4fb87f104a87 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

RLVP: Penalize the Path, Reward the Outcome Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.341579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:65f38951e9fd591b63b0f505329b2136d903cfd4d1d6a3b13ee2711a3a3acb83

Observation 414a6a04-8634-45b8-8de5-fce0ac8c5b3d · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347.

RLVP: Penalize the Path, Reward the Outcome The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-09T11:16:11.339361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:cdf5958b3a1389e0eb92fd87b85d96dc4aa694876a4143b6de0683142e4a480a

Observation 886e91d1-f6b4-4582-ade4-f5e7c488cfdd · outbound

This paper cites This appendix gives the full experiments.

RLVP: Penalize the Path, Reward the Outcome This appendix gives the full experiments

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T11:16:11.523517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:dac7e94024c0f192aea8c439929880f46d1d8acc7d1394bee2d42e72e0231fb2

Observation 7edaf817-368b-426d-a2dc-dfe408ef7fbf · outbound

This paper cites On SWE-bench software repair [Jimenez et al., 2024] the natural potential is the fraction of hidden FAIL_TO_PASS tests an episode makes pass.

RLVP: Penalize the Path, Reward the Outcome On SWE-bench software repair [Jimenez et al., 2024] the natural potential is the fraction of hidden FAIL_TO_PASS tests an episode makes pass

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T11:16:11.528376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:1b9268c0445511774a3fb287b4a735fe3421561d74dabe94fb2e00a007e19d50

Observation 8555e251-bec1-4240-82e7-c49646413fba · outbound

This paper cites Benefit appears only past a reachability threshold.

RLVP: Penalize the Path, Reward the Outcome Benefit appears only past a reachability threshold

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T11:16:11.521743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:1cced27c75b747c930751097ecf26e7320e80cfcd3427ed1e7f0cb183441bede

Observation 3ec87f6f-a0c5-431c-bdf3-1c45ae51a97b · outbound

This paper cites fragility,.

RLVP: Penalize the Path, Reward the Outcome fragility,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T11:16:11.513997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:d69cb8b377e51b50916a94f0d483405dc887f93bc1f3a758f1ef7d293333ad7c

Observation 133097ed-8ffc-43ae-9a97-925b6d60e5ba · outbound

This paper cites any- valid-tactic.

RLVP: Penalize the Path, Reward the Outcome any- valid-tactic

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-07-09T11:16:11.526298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:186e75caaf4e0647363725e5249530dcd3d350fa8a149a2908b38ce55f7e3976

Observation 0c7ce4fd-3d7b-4cd1-adec-d7b0caeb0e9e · outbound

This paper cites The learned critic is a poor training reward in every quadrant; its only robust use is offline diagnosis at the intent frontier.

RLVP: Penalize the Path, Reward the Outcome The learned critic is a poor training reward in every quadrant; its only robust use is offline diagnosis at the intent frontier

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T11:16:11.526004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:673db389daa98d80236d3a05a89ea6ccc8f6540329a6d0197ce10b34f981c72b

Pith citing papers

No inbound Pith citation observations are available.