Pith. sign in

Paper Citation Record · LEDGER

Process Reward Models for LLM Agents: Practical Framework and Directions

As of 14 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 25 inbound Pith citation observations for arXiv:2502.10325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10325 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:39:38.006409Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:40:55.391008Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:40:07.852377Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4a19a37-a103-41c0-9aa5-3e89233f0f4a · outbound

This paper cites Step: Stacked llm policies for web actions.

Process Reward Models for LLM Agents: Practical Framework and Directions Step: Stacked llm policies for web actions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.613373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.769548Z digest=sha256:53b5dcd5e0ed769412c7090c4e8fb5297694c73e019a6951c0e6b16e670d8a4f

Observation 544ebe5c-dfac-470b-8f68-28bfd3b71537 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Process Reward Models for LLM Agents: Practical Framework and Directions $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.773506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.773506Z digest=sha256:a2d0dae4c89151d1d0f1c7118d30c5be81d02558356a996148982e453a87c564

Observation 3afccf7a-92d7-478e-9c8c-92f9e4512480 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Process Reward Models for LLM Agents: Practical Framework and Directions SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.778546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.778546Z digest=sha256:2b446a84aa62c40c0be795e6f4aa92e679939c61b4abe220cf957cba9c87b371

Observation 42548ea1-5dc7-40ea-8817-b7bb4fddd832 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Process Reward Models for LLM Agents: Practical Framework and Directions ReAct: Synergizing Reasoning and Acting in Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.783337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.783337Z digest=sha256:c3c7b7a31fbf3454b8413d9c759836b20d133b9da82d71cd224e010477332bbf

Observation 47b6553b-c943-4114-b24b-5cfea61ec524 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Process Reward Models for LLM Agents: Practical Framework and Directions Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.788032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.788032Z digest=sha256:00626ea9e1319fce096ab761b8066d3ed3c7a0aa58304be4ef777b135acb9951

Observation 72bcb666-d5b4-4310-a92d-14213dc0e295 · outbound

This paper cites Fireact: Toward language agent fine-tuning, 2023.

Process Reward Models for LLM Agents: Practical Framework and Directions Fireact: Toward language agent fine-tuning, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.793329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.793329Z digest=sha256:0eda19f3f442ea319ba56f136670c96d6e1d5bfec2a49b60d2470f8821f27ce8

Observation 4f80b961-a6b3-4f94-90de-6b1eed3ba107 · outbound

This paper cites Decomposed Prompting: A Modular Approach for Solving Complex Tasks.

Process Reward Models for LLM Agents: Practical Framework and Directions Decomposed Prompting: A Modular Approach for Solving Complex Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.797982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.797982Z digest=sha256:33618446d1d5b08623fb9ae173bff33136487188b8250d76da6be42b847f4212

Observation abb81628-ae03-4ac8-a6f1-c9f7a4fe9f89 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Process Reward Models for LLM Agents: Practical Framework and Directions DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.802927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.802927Z digest=sha256:f49defa91414ae1cd6d7ece2d3510f3d4270fea1f2903929fe0a3cde1b97c793

Observation 795341fa-f86d-4957-9fc0-5dba5cfc891b · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor.

Process Reward Models for LLM Agents: Practical Framework and Directions Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.807987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.807987Z digest=sha256:0c7628c64bc5ec825a22210fe02eb170ed2b56368175275f4f9128107b3ee085

Observation 865e7792-c477-415c-bb8c-12224dca57a6 · outbound

This paper cites Let's Verify Step by Step.

Process Reward Models for LLM Agents: Practical Framework and Directions Let's Verify Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.812425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.812425Z digest=sha256:1201401df44eb3a90246e1f77031c8be2330076e4078a9c6063e11f9ae31f095

Observation 46fd1c97-dbad-4f5d-a1fc-f50c27e64761 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Process Reward Models for LLM Agents: Practical Framework and Directions Solving math word problems with process- and outcome-based feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.817543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.817543Z digest=sha256:65ecbf82f23f1f55464c4d7eaeec06ae5cc71e2c4e3d53195f78eebdab9ba05b

Observation 75077441-2f49-42de-9c94-49732407c624 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Process Reward Models for LLM Agents: Practical Framework and Directions Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.823226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.823226Z digest=sha256:333435b95228c039f576c6651b91e741b0d768b316946bd5255d3d87cfe1a5b1

Observation 14972fb4-64bb-42ef-a4bb-5adc0b21d005 · outbound

This paper cites Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D.

Process Reward Models for LLM Agents: Practical Framework and Directions Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.828413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.828413Z digest=sha256:cbdaf037763f97258fba19dca32a7207d39e56706f7c00d94fe916d063b220aa

Observation 9cbdf9f7-fb8d-423a-890a-7227154da4da · outbound

This paper cites Trl: Transformer reinforce- ment learning.

Process Reward Models for LLM Agents: Practical Framework and Directions Trl: Transformer reinforce- ment learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.577236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.832427Z digest=sha256:4c7b4a47cb92e63903d537f3106bb8141cce7fcfca1e301f1f909d40000d6496

Observation 85075978-3419-4847-9b71-1bfabbf77a6b · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

Process Reward Models for LLM Agents: Practical Framework and Directions Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.564620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.837193Z digest=sha256:70b8ef9876e57454df9531ac9c7c0f99b9035569ce2d14c67f5b31918e1d615c

Observation 6d609886-5765-4909-9e77-13fddea298f3 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Process Reward Models for LLM Agents: Practical Framework and Directions Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.841455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.841455Z digest=sha256:e09e8ff115c0522032cc1a8f468ef62ce62fda586f7b1ee0359b2811a92ef224

Observation cb6ae2fc-b4b4-4ef9-9707-b4925880981f · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Process Reward Models for LLM Agents: Practical Framework and Directions Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.846166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.846166Z digest=sha256:7445864b47e683c81f0721eceac3af8b5465b0c9c088461ed90a9089dbdd2de5

Observation 2f8d128d-da9f-4de8-aaaf-7aa5796d1fa1 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Process Reward Models for LLM Agents: Practical Framework and Directions SGLang: Efficient Execution of Structured Language Model Programs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.850883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.850883Z digest=sha256:c5652dec0307598650e10fa18efa25bb29e8fd3e8f383cc9bc75aa165ca36fd3

Observation 480c3fff-be9a-4d01-90fa-ca98236fc61d · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Process Reward Models for LLM Agents: Practical Framework and Directions Gonzalez, Hao Zhang, and Ion Stoica

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.855683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.855683Z digest=sha256:f92aa743b4a9bf16d1bb94939213933fd824c31dc509cf3a6da12af2cb2933c0

Observation fa296111-14c9-4750-8da9-c34cd389ed1e · outbound

This paper cites Proximal Policy Optimization Algorithms.

Process Reward Models for LLM Agents: Practical Framework and Directions Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.859703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.859703Z digest=sha256:1e7e7a4dc050130b89c73ddbd441b26008e44d6d6d3787d572ceac07c5d98d6f

Observation 238780b0-f35b-4771-8497-9b696b970edf · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Process Reward Models for LLM Agents: Practical Framework and Directions Direct Language Model Alignment from Online AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.864500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.864500Z digest=sha256:7e01a8b6c510ef3b203865da3bf0dc7727ff81a1ce73ded82e56f5e0638ba1bf

Observation 940adcd1-0907-4494-98a3-a84a8a47de01 · outbound

This paper cites The Llama 3 Herd of Models.

Process Reward Models for LLM Agents: Practical Framework and Directions The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.868713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.868713Z digest=sha256:ef336c22985972646ce07269c41c53b0ec644a799eab36011d55e41f791e215b

Observation 45ee08a1-44e9-4689-b40c-b0b6cbfba721 · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Process Reward Models for LLM Agents: Practical Framework and Directions Approximately optimal approximate reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.538108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.873203Z digest=sha256:54dba945088c4da931ce067dc508b6ba1756c7de798832d751d037b1f74c7733

Observation 20e198df-fcc8-49c0-a158-c3d59f5a49c0 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Process Reward Models for LLM Agents: Practical Framework and Directions ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.876369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.876369Z digest=sha256:63788501ce3aea283782db3f114e9297329bb9711de16b979faa0e800751781c

Observation 47c2017d-dc43-4906-b04b-f2f1c83551b6 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

Process Reward Models for LLM Agents: Practical Framework and Directions AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.880327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.880327Z digest=sha256:665a2bcca60e6df57ea563b8d2372c49e69b7282204d47081127a10d5da75b5e

Observation 7540d552-8a8b-4cc8-95dc-7404d5504fe3 · outbound

This paper cites Expel: Llm agents are experiential learners.

Process Reward Models for LLM Agents: Practical Framework and Directions Expel: Llm agents are experiential learners

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.883873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.883873Z digest=sha256:eef070dc525339e22c286ac9c24e35fbe585a4bd6d6e22a03c319e0184558c6b

Observation 30be6acc-cfc4-40fe-adc3-c181d7b2dedb · outbound

This paper cites Adaplanner: Adaptive planning from feedback with language models.

Process Reward Models for LLM Agents: Practical Framework and Directions Adaplanner: Adaptive planning from feedback with language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.518028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.888101Z digest=sha256:05d8aa6d8b488b8225239f9264abefc3b043f0a69830dd2c4a9aeca88d271b3b

Observation 7cdd03be-64ae-47f4-9570-d595f98f981d · outbound

This paper cites Specification gaming: the flip side of ai ingenuity.

Process Reward Models for LLM Agents: Practical Framework and Directions Specification gaming: the flip side of ai ingenuity

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.506420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.891380Z digest=sha256:2990b471c2d2888ec5baa758aa2cdab9c48ba6487cbd889a6f9c0496f5ff27d5

Observation acf07846-add6-4c65-8f52-42a23a81bfd8 · outbound

This paper cites Reward hacking.

Process Reward Models for LLM Agents: Practical Framework and Directions Reward hacking

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.494882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.895086Z digest=sha256:e31d257f4a52e1ff19b4b986d29d4d36956458f4cb508f6b71c24c552565383d

Observation e9fadb62-7de0-4a95-bd4c-860a6a6e6b0d · outbound

This paper cites Rank analysis of incomplete block designs: I.

Process Reward Models for LLM Agents: Practical Framework and Directions Rank analysis of incomplete block designs: I

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.898874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.898874Z digest=sha256:f3c7b633b64f0181909956852717e96f07a1ec416ea63e474717be79c33e79d9

Observation a3e78fb8-059f-4345-8f08-a6e238ab33e5 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Process Reward Models for LLM Agents: Practical Framework and Directions The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.902896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.902896Z digest=sha256:2138f51f40bf61a9dd6cc6d03ee0d07f439e5e707de46052405d116a3a1bbc68

Observation 23838738-6982-4d19-ba4c-4bd32154662b · outbound

This paper cites Iq-learn: Inverse soft-q learning for imitation.

Process Reward Models for LLM Agents: Practical Framework and Directions Iq-learn: Inverse soft-q learning for imitation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.907824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.907824Z digest=sha256:fc8491094444f4224b2df07e4b979ba5cfc1aacdf817267961691d5a37bcfc4d

Observation de5e6fa7-7e9d-45d7-9be7-329845718221 · outbound

This paper cites Q* approximation schemes for batch reinforcement learning: A theoretical comparison.

Process Reward Models for LLM Agents: Practical Framework and Directions Q* approximation schemes for batch reinforcement learning: A theoretical comparison

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.469910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.911571Z digest=sha256:4eacb11eefd6e89ae2b64cdc1a2076f80601f941d8cc6191fd7c70d4497d10e3

Observation 1d4189ef-0bfb-401f-9c73-96bf7b58654b · outbound

This paper cites Better than Your Teacher: LLM Agents that learn from Privileged AI Feedback.

Process Reward Models for LLM Agents: Practical Framework and Directions Better than Your Teacher: LLM Agents that learn from Privileged AI Feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.915767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.915767Z digest=sha256:3766f696eff3ae9595de79268268f458f28c4ad677e37d17a6db84207bb191d2

Observation 0ff61f82-f325-4891-85bc-ac7f2d5a3d9f · outbound

This paper cites Inverse reinforcement learning without reinforcement learning.

Process Reward Models for LLM Agents: Practical Framework and Directions Inverse reinforcement learning without reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.456384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.920612Z digest=sha256:20bc9826ace6da1d8eebecb357cfe31c3d908979a250421eac2259cda2bb1145

Observation 536882f9-2500-4b91-ad23-5ae28625e367 · outbound

This paper cites Policy search by dynamic programming.

Process Reward Models for LLM Agents: Practical Framework and Directions Policy search by dynamic programming

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.444147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.924821Z digest=sha256:c62bfacc953c88815fe9114781e4c766ba1f4f0f038e9a5729aa189db1c108eb

Observation c9fa1438-73ad-46d2-a0fe-2ac4cd37db80 · outbound

This paper cites (more) efficient reinforcement learning via posterior sampling.

Process Reward Models for LLM Agents: Practical Framework and Directions (more) efficient reinforcement learning via posterior sampling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.929780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.929780Z digest=sha256:58acd57664b1afbe88dc21f4cf7c2e4f43af3b1d610e4c7f0f885e56b76d07e8

Observation d4af72bf-b343-4303-97d6-e6883edf15fc · outbound

This paper cites Sequence model imitation learning with unobserved contexts.

Process Reward Models for LLM Agents: Practical Framework and Directions Sequence model imitation learning with unobserved contexts

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.424316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.933921Z digest=sha256:91abb39100ef89be5c005b0f111a8846cd78d15ac8cd414c55393c6bf3daa5c9

Observation 4f327c8a-59a6-4dc9-a50e-ddad0f33dad2 · outbound

This paper cites Data-driven planning via imitation learning.

Process Reward Models for LLM Agents: Practical Framework and Directions Data-driven planning via imitation learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.411598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.939188Z digest=sha256:eccc4a64bf3d9174e18bf6f8e9dcf92d2f3abe9e6c63ab98ba644ad3de1b6c60

Observation dc7d1419-d01e-4ced-8706-81e920bfd194 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning.

Process Reward Models for LLM Agents: Practical Framework and Directions A reduction of imitation learning and structured prediction to no-regret online learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.399365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.943371Z digest=sha256:ebc14b47ec3a570d6856b5040cd0ba22ca995b2f24f6189313225bf916dc83a7

Observation 43cafe49-345a-4fb5-833a-27e5436e9c9d · outbound

This paper cites Reinforcement and Imitation Learning via Interactive No-Regret Learning.

Process Reward Models for LLM Agents: Practical Framework and Directions Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.948274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.948274Z digest=sha256:45a7c9f365c635a2ef959047d440b20936c34df7a1dce94c938903d910834099

Observation 69f4abdb-ad3c-4dde-98c1-bb180b8d91bb · outbound

This paper cites Deeply aggrevated: Differentiable imitation learning for sequential prediction.

Process Reward Models for LLM Agents: Practical Framework and Directions Deeply aggrevated: Differentiable imitation learning for sequential prediction

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.386533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.952283Z digest=sha256:f1b6e61b650bf7fca361a477e3314011ac8dcf7b1722a22fc4c13366c421b84b

Observation 99c9dd62-a3a6-4702-aa67-06c87b1fb4fa · outbound

This paper cites An application of reinforcement learning to aerobatic helicopter flight.

Process Reward Models for LLM Agents: Practical Framework and Directions An application of reinforcement learning to aerobatic helicopter flight

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.375311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.956574Z digest=sha256:01fd79a893496e5f60d77185cf3e2e09c97ee143887dc8166a2c122088fbe453

Observation 70f67bdd-8ec0-45ba-92ab-761890ded74f · outbound

This paper cites Learning dexterous in-hand manipulation.

Process Reward Models for LLM Agents: Practical Framework and Directions Learning dexterous in-hand manipulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.961506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.961506Z digest=sha256:c0532eb16cf4303a55680842a081a7e6d996b4a78f62f8a3ab242849bc4ef5bd

Observation af5b97ec-dc3c-4ed0-8bda-5cb18d802ef3 · outbound

This paper cites Model-based reinforcement learning with a generative model is minimax optimal.

Process Reward Models for LLM Agents: Practical Framework and Directions Model-based reinforcement learning with a generative model is minimax optimal

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:39:38.357863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T18:39:37.965190Z digest=sha256:6bd09fc9ee74b0f3d4954ea4b6f4b7f04bb2313df1ee4524241184a7ad18217d

Observation 457d090c-c7d2-4f26-bd28-09183312eb08 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Process Reward Models for LLM Agents: Practical Framework and Directions AgentBench: Evaluating LLMs as Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.969168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.969168Z digest=sha256:0624022e284e2772be746b413d730ca35283b0d0e92594b1524bb5b9e6eda7a4

Observation e42f8356-178c-4eef-ad20-ca3a0fd8591c · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

Process Reward Models for LLM Agents: Practical Framework and Directions Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.974199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.974199Z digest=sha256:42bbd50a101886ae1ba65879660ef0865ab65b8fecc3154647b97a90f3019189

Observation 4981e9e1-0e35-4790-817d-1e1059a38d25 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

Process Reward Models for LLM Agents: Practical Framework and Directions AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.978380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.978380Z digest=sha256:74cd0c8bbcacfe735f91982953f9bb4d74713c27a834661774de622bf66fee4c

Observation 11484c73-dde9-4c93-aa03-456ef3ded975 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

Process Reward Models for LLM Agents: Practical Framework and Directions ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.983121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.983121Z digest=sha256:d0c5d25b9d909525c67094e165ed7824bbdb11bd82ed452e704292642b69d126

Observation 9adc526c-84b4-43c9-8e7d-71b06f5a52f1 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Process Reward Models for LLM Agents: Practical Framework and Directions Training Verifiers to Solve Math Word Problems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.986993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.986993Z digest=sha256:60df26bb167b336001ee952b4856eac8699e32c9d7198e3df9ad7e123b446e81

Observation d7130766-2dc1-4947-8e84-3b1fe1efe616 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Process Reward Models for LLM Agents: Practical Framework and Directions Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.991227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.991227Z digest=sha256:5a471a4b6e80bc08e964dd7764557c9dc7ddc3043f2ba030c54041e8e0db1e65

Observation 8971eb30-f7e4-4df1-bb88-cf78aa2f2fcf · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Process Reward Models for LLM Agents: Practical Framework and Directions DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.994928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.994928Z digest=sha256:e7facc9e3e5e04c9d19a75735645aaf82994ac75fd24fb6f3aa602bed4889ab8

Observation 63542aa5-0688-4083-970c-3d12a35bb618 · outbound

This paper cites Let's reward step by step: Step-Level reward model as the Navigators for Reasoning.

Process Reward Models for LLM Agents: Practical Framework and Directions Let's reward step by step: Step-Level reward model as the Navigators for Reasoning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.998225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.998225Z digest=sha256:f743a4f94ce26aa6506763f29d1356f421eaadfb82e22dad1b7c598bdf24aefb

Observation b525f024-300e-443d-b921-1ffd9a720aab · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Process Reward Models for LLM Agents: Practical Framework and Directions Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:38.001803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:38.001803Z digest=sha256:d21b420f14cd449f6433b7819be7044bec273b6a5a3d4765da5f31cf3c6435c8

Observation c98a33f2-13cf-4e37-974c-51909915dc0d · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Process Reward Models for LLM Agents: Practical Framework and Directions Teaching Large Language Models to Reason with Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:38.006409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:38.006409Z digest=sha256:353599ab86694bfb4cae3f53cd0ae809a95e51e150c91aac3891c599d14bcac6

Pith citing papers

Observation f95ff518-c0a2-4cf2-ba9b-6f309bdde87b · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:42.043925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:b8f68710c3eebe0a1f692e89854365a9422487238c1555366ff528e29dfcac96

Observation 3e4079b9-57fd-495f-91f3-1d34e1693736 · inbound

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution cites this paper.

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:00.891477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:00.891477Z digest=sha256:8a27c1658932da5799df8d02c2dc0554f30ebc148146a727fe750c037f87a845

Observation 46e7171b-ec46-4bee-b76d-f05f0060db78 · inbound

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning cites this paper.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.569399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.569399Z digest=sha256:569b72568faa7254b1cbc06938495716034393bc10f9ffd09c02874f00d9a6be

Observation 7b677e12-6bf5-45d4-a241-5e8c246a3fcc · inbound

A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement cites this paper.

A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:30:46.105396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-21T23:26:38.457193Z digest=sha256:0a15e801eafb6ca9630ff3525f78c091449b7d45da5cdc718b1be69c67de2640

Observation e9162390-ac2b-46fe-8592-48cd973764e3 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 235

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.209112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:25e95c4cd67ff396b5936c061cdaea55aa9ef005dbb7cc0cbbdc405afbfd36b2

Observation 4baf07d6-19cf-4135-a5b3-e240d8b1c703 · inbound

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning cites this paper.

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:14.788155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:44:14.788155Z digest=sha256:ca931fd136c63761762a826532f258552136a164b6512d3b49d3e0302fc9e378

Observation ef0ee11a-e246-440f-8cbe-5c3556f6446b · inbound

Reinforcement Learning for Machine Learning Engineering Agents cites this paper.

Reinforcement Learning for Machine Learning Engineering Agents Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T12:24:02.368310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:24:02.368310Z digest=sha256:8d00047ab08563425c40d76a7c779c27d2ff852078a24b2f27581229658c7e77

Observation c4d0eb90-6505-4219-a8fb-e67540e346a8 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 268

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.656981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:37d71f390476030f8d6deebae90e778b757d603509437361f491499abc8bf5be

Observation 69ce9a63-a132-439c-b4ef-6b5382d7eb0c · inbound

MASPRM: Multi-Agent System Process Reward Model cites this paper.

MASPRM: Multi-Agent System Process Reward Model Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.023170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.023170Z digest=sha256:02debc65649a58e8aa5c1195f93d30c69cdbd4abdb99c69d5fc11e839e6e9e82

Observation 8c44698b-6b8a-4b98-94a4-386534cb4aa6 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.473321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.473321Z digest=sha256:c3fa2640443bcb19976b0be7b972b44ab35e3e9fb62d5b131cf69137581a10a0

Observation 735e5750-b183-46b6-acf9-150fa5b7b843 · inbound

ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuning cites this paper.

ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuning Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:38.272747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T17:29:34.145855Z digest=sha256:2ee94e9c001626b12d7b359454c7dc5f254472bcad0ecc9502522d7dbee5fa1b

Observation 76ecb7bf-6c00-4e45-b94c-db02d10d2c44 · inbound

A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping cites this paper.

A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:56:08.764793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T10:41:46.675257Z digest=sha256:8449df3aa1ac6902c0cdcb11502702a861a3838e19869711157498511863c279

Observation d1017cb9-f713-412d-af1d-b9b8e08afcd1 · inbound

TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents cites this paper.

TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:26.376883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T04:30:06.869916Z digest=sha256:5d5764278c6319c61b53614cf8082614d2a8e4bed2534d695d82b681cd7f2946

Observation ab9ed4e5-0ccb-48b8-8d12-bc2b2b370c58 · inbound

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning cites this paper.

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:22:06.790866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T02:19:27.345348Z digest=sha256:5e4f8522bedde1cf47ed3a623680d20422530eeaea84cb5d90caf9dab25498cd

Observation 64ff4a60-ed75-4796-8d76-ca22865b8c53 · inbound

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents cites this paper.

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:47:09.644058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T22:24:41.353303Z digest=sha256:cdfddf2fbe7b4ed80400c21822ec7247d8f1e46b809d18924d9c02c36b45fb0b

Observation 750b729d-0b41-4b49-85f6-316fd64d974a · inbound

Self-evolving LLM agents with in-distribution Optimization cites this paper.

Self-evolving LLM agents with in-distribution Optimization Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:57:09.563002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T22:18:27.021136Z digest=sha256:dec133befc5b0d8f06e47ae0bf04c603a57ea6d534302c38f5e1b8386fdf462b

Observation ab4bda4e-7106-4530-bc4d-d072e9e97af4 · inbound

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents cites this paper.

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:19:38.638684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T14:33:50.123077Z digest=sha256:267d36af967d82788f250e6bbe17090f365fa2c08fe06260c6d653f9437ba771

Observation 2811bb84-3870-4bee-964b-335b758adde9 · inbound

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents cites this paper.

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T10:43:52.874792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:43:52.874792Z digest=sha256:24ba160027356acd36054fe710e3257edb01622e45ddbebb4540274ec4677b5f

Observation 24e18af3-2760-4054-af69-5b4f913f1edf · inbound

Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention cites this paper.

Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:37.873379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T14:24:38.527103Z digest=sha256:759871185f21d85118ec047e5891b98267257e1753b1850daa82b2a556327882

Observation 1dbb1f91-93d6-4b4a-a8a2-d9717bde3001 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:7efa9f1f7622eede113b64eb75a11ba091ebb4114efd780a83eab7d76323e09b

Observation 63338ea6-52ae-41cc-a02b-f373b9e1161b · inbound

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents cites this paper.

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.668163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-25T20:16:39.347676Z digest=sha256:7c0cb9a58980d283976e14cd374d7833825b52b9b121dc1d9ee0820d87746d3c

Observation 4694c7c6-38ac-4a83-aab6-7911e90eeabb · inbound

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents cites this paper.

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:40:07.853911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-25T19:55:51.114244Z digest=sha256:bc5ca3f58a9a3233a70da41d727f8a02bb2997b40ecbc2d6be363ea91b8e4557

Observation c7939e92-aad6-47c9-bf80-6d2e901c8b1c · inbound

ECHO: Learning Epistemically Adaptive Language Agents with Turn-Level Credit cites this paper.

ECHO: Learning Epistemically Adaptive Language Agents with Turn-Level Credit Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.609798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T04:36:35.213066Z digest=sha256:e3c6f923ee2e05554eed0bf6a4d35a627f734cb64e56ea020a2fb3a8f1202e26

Observation fd198367-8d74-4d80-8bef-178f768b551c · inbound

A Diagnostic Framework for AI Agent Behavior cites this paper.

A Diagnostic Framework for AI Agent Behavior Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:08.965286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:08.965286Z digest=sha256:7ce630bbebf65d8c140331e2aad9712fe6abe3b5e71b810ad03cb995764732d0

Observation f81e7495-b455-4cc6-a571-75adc88caacb · inbound

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents cites this paper.

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-14T04:40:55.391008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:40:55.391008Z digest=sha256:a677c336cdb0f75aedbfd030c43f320204acf549f4390a790ac163f655cee620