Pith. sign in

Paper Citation Record · LEDGER

RRO: LLM Agent Optimization Through Rising Reward Trajectories

As of 9 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2505.20737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20737 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:43.605901Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0faff669-6ec0-4b39-a99e-fe973c712998 · outbound

This paper cites write newline.

RRO: LLM Agent Optimization Through Rising Reward Trajectories write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.015662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.015662Z digest=sha256:fc1808ade548ee9d1b9c845751a2b34bf21c4291d77f288a1959c82171730c49

Observation 4c2f4002-50e6-489d-adf2-d869cb36b850 · outbound

This paper cites Language models are few-shot learners.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.148103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.148103Z digest=sha256:b6e9e7986dea0ba9a18450b52bc2e807d9a2f4eee3a27d277e6efab54fab3322

Observation 34e23431-e069-445d-90ab-5951361985e9 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

RRO: LLM Agent Optimization Through Rising Reward Trajectories FireAct: Toward Language Agent Fine-tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.263415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.263415Z digest=sha256:e621ce3f79f006d5fd5283d0f33d94c05f737810ac568c99d0d80cb386fe7b1f

Observation bd212bbe-d3f0-4958-8578-56ae96444fd3 · outbound

This paper cites an unresolved cited work.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:51:44.799382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:51:41.347777Z digest=sha256:e8c03cd8b44f867a21638f0d1d544d1f4bd429a5a44528e6f05386401b27d927

Observation 0ae589bc-8a04-4b7e-92cb-0bc154a358af · outbound

This paper cites Deep reinforcement learning from human preferences.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.442326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.442326Z digest=sha256:b17d548e2795eaa212001032016ccaa9be278cb8f6f90549a52c332301b4ef74

Observation 9651b20e-3ac5-41c8-a74d-50a085c7d770 · outbound

This paper cites Complexity-based prompting for multi-step reasoning.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Complexity-based prompting for multi-step reasoning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:44.616698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:51:41.525532Z digest=sha256:6a57ccf61b5b881bc4e514da677b7e6167694a490247de0ef7f99bf1c5d2353e

Observation 55751f3d-ce1b-40cb-a013-31ebdee19dab · outbound

This paper cites Large language models are zero-shot reasoners.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Large language models are zero-shot reasoners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.647817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.647817Z digest=sha256:51762b8e2eda746a1b45fda763f9687716c089b3d4d58d61ae425839d823b87e

Observation 0fa6e4be-6a92-465d-9c0b-61f713909791 · outbound

This paper cites Let's verify step by step.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Let's verify step by step

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.770139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.770139Z digest=sha256:1fa54002f9d4ce3ff9f74052ba5fc214935b321e54ead685d6615aa2c0c12a7f

Observation bfbee95e-0654-4dbc-9ef7-d123955fa6bf · outbound

This paper cites Let's verify step by step.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Let's verify step by step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.851235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.851235Z digest=sha256:2388fb3de3bc88d8ff6f36577246bf59898016489730f69d411ef1046a3eaccf

Observation e3ca2ecf-ebce-4e01-8414-30e0e656cd66 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.974215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.974215Z digest=sha256:6ac710cc88741e5824f8a707b22abfb8db9d1044a0a6ea6f7ab2399f6f154ad0

Observation ba0de3c6-a8df-4580-8830-c8cb4ec0ff69 · outbound

This paper cites Training language models to follow instructions with human feedback.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Training language models to follow instructions with human feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.072863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.072863Z digest=sha256:8360cd89714ff8247358917cfd02dc9decdfd3303f3179d2ac8e3e29b2e54403

Observation c7627424-038d-4c6c-b72c-cd0a68fa5f26 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Direct preference optimization: Your language model is secretly a reward model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.132509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.132509Z digest=sha256:f8fabe1c6f5c179cdb9baaffa283ec7cb8286e658a043595d6ad21db6a4cda9e

Observation 0a5508d8-bd6f-4483-a7d7-4feb62be20b3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Proximal Policy Optimization Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.247181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.247181Z digest=sha256:516476e7d3e89694480559e46be5a3dae00af1b373134aa1c16c989361e59d60

Observation ca9dc9fc-7c7f-414a-9001-8d0c75cf5792 · outbound

This paper cites Trial and error: Exploration-based trajectory optimization of LLM agents.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Trial and error: Exploration-based trajectory optimization of LLM agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.374197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.374197Z digest=sha256:4c2a400c92e5ec831961a0770485036ca996268a011ebcf0d612b2f951cb306c

Observation 0b254d60-96f2-4a93-8945-e698f000a5d8 · outbound

This paper cites Math-shepherd: Verify and reinforce LLM s step-by-step without human annotations.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Math-shepherd: Verify and reinforce LLM s step-by-step without human annotations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.452605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.452605Z digest=sha256:acf84dcd7feb9219357c151ce3de436938f36b7d22fecc2836129fec3be513c4

Observation ad41f962-9107-4136-8949-6e47c652b4d7 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:44.417785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:51:42.540673Z digest=sha256:b25d73b7b9345a84308bf8812fe925e87aa8442cc29f921c02384267407cc173

Observation 742ac0a8-38d1-4a25-b960-0002ac15c9f1 · outbound

This paper cites Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.634639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.634639Z digest=sha256:82b285a10f0f6fd5738647a33e4ab3cdc578099d92f89074e337a14f621dbf47

Observation 21b930cb-cdb7-4214-ab74-04d336c0e2a8 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Chain-of-thought prompting elicits reasoning in large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.713789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.713789Z digest=sha256:1a05f1c40ac4faca268734022ca162b005e3986f3ce9cbfc7eb8507963734b90

Observation 9eb5afc0-64c0-48da-b24b-8ccaca91f770 · outbound

This paper cites Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.816272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.816272Z digest=sha256:c51fc1e49ece5851812a4cd971212fd2a4039f276a2b81baaa53637bd8ad4ecf

Observation 168df566-aeae-4b62-bb23-f970c537a8b9 · outbound

This paper cites Intercode: Standardizing and benchmarking interactive coding with execution feedback.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Intercode: Standardizing and benchmarking interactive coding with execution feedback

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:44.176909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:51:42.879826Z digest=sha256:761362d47275447555ca0bfba15a12011056689b291436ba5734f10b68c1767c

Observation 26ca5ac3-7e97-4b7c-8ede-2f9a8aa11425 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.009609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.009609Z digest=sha256:a77ceeb3550d56c328550c63873e473571ef5bb1ae8bd217b3dc09fa7fabe6d7

Observation 8a53a5e1-3352-4d9a-9141-eb749838ead0 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

RRO: LLM Agent Optimization Through Rising Reward Trajectories React: Synergizing reasoning and acting in language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.080279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.080279Z digest=sha256:35470c3461fda8f5b185a8e461945048238a5b2053e7c271b12dec1569187ef7

Observation eb73a048-7729-42ef-8dd7-71ebfd3c7c38 · outbound

This paper cites Debug like a human: A large language model debugger via verifying runtime execution step by step.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Debug like a human: A large language model debugger via verifying runtime execution step by step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.139184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.139184Z digest=sha256:9ba6cc07e4aa6b370b5a3dcb681dd01f5bd4b8c7eee72732a896ea6ccdb8d2d4

Observation c9aaadfe-1183-415f-8367-027decb7b1da · outbound

This paper cites Least-to-most prompting enables complex reasoning in large language models.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Least-to-most prompting enables complex reasoning in large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:43.901711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:51:43.257177Z digest=sha256:2da51fe71b4d4f4a4139228a5df50c988685a9e089fa8c022cd6a2aa2a639b9a

Observation 89da84bb-3e8e-4bba-bb80-1899d29d02a5 · outbound

This paper cites @esa (Ref.

RRO: LLM Agent Optimization Through Rising Reward Trajectories @esa (Ref

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.371191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.371191Z digest=sha256:0bbd6dc4208b93332ab7c4ed39b45477d552d8b30c6427a4c0d37ffe54dc5067

Observation 3a0a0449-7562-4032-8f2a-51fe5889d20f · outbound

This paper cites an unresolved cited work.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.499214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.499214Z digest=sha256:020756b8af4fd7c9833ba9e74453c819668351d2f6e38ea1c6766161c1f55917

Observation 6c2b883c-d0da-4453-b261-85cf64a77b70 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.605901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.605901Z digest=sha256:d3fdd429eefd8b91da815f3fd0306e507ddf717509b71994b993703df2daae1e

Pith citing papers

No inbound Pith citation observations are available.