Pith. sign in

Paper Citation Record · LEDGER

RRO: LLM Agent Optimization Through Rising Reward Trajectories

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2505.20737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20737 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:43.605901Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0faff669-6ec0-4b39-a99e-fe973c712998 · outbound

This paper cites write newline.

RRO: LLM Agent Optimization Through Rising Reward Trajectories write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.015662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.015662Z digest=sha256:676df2125eba5c5937556c42957bf9f879c94228f4a576836ee8de24a6c40d69

Observation 4c2f4002-50e6-489d-adf2-d869cb36b850 · outbound

This paper cites Language models are few-shot learners.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.148103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.148103Z digest=sha256:1117d98c6e8bb60ed3e24dd0acf0871b788c8f3c8a5a98aca744986ccd18d5b6

Observation 34e23431-e069-445d-90ab-5951361985e9 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

RRO: LLM Agent Optimization Through Rising Reward Trajectories FireAct: Toward Language Agent Fine-tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.263415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.263415Z digest=sha256:87056dc985a8d8f4daa550204179d1318db92f86a3c63eb058a30c20d7ed832f

Observation bd212bbe-d3f0-4958-8578-56ae96444fd3 · outbound

This paper cites an unresolved cited work.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:51:44.799382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:51:41.347777Z digest=sha256:9d27062dd9722f4a8584228fed260acde891861c843b05e83d1ed1f5249dcdfc

Observation 0ae589bc-8a04-4b7e-92cb-0bc154a358af · outbound

This paper cites Deep reinforcement learning from human preferences.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.442326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.442326Z digest=sha256:7ac18f5ba71ee131ec933e4bce36126049019cb832088404dc373d1bb414463e

Observation 9651b20e-3ac5-41c8-a74d-50a085c7d770 · outbound

This paper cites Complexity-based prompting for multi-step reasoning.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Complexity-based prompting for multi-step reasoning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:44.616698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:51:41.525532Z digest=sha256:42337e10429da07b6407e80d508063265a7c4e10c35a2a7674b10224dd44c9a3

Observation 55751f3d-ce1b-40cb-a013-31ebdee19dab · outbound

This paper cites Large language models are zero-shot reasoners.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Large language models are zero-shot reasoners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.647817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.647817Z digest=sha256:086b404e200a48d991137aeac5fd6d2f0d0a324d4a0bc060bf43b7b98df0ff6a

Observation 0fa6e4be-6a92-465d-9c0b-61f713909791 · outbound

This paper cites Let's verify step by step.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Let's verify step by step

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.770139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.770139Z digest=sha256:3ab9927735303bca5183e759d767ac2a40362e81ed25915b7065d191e9dd2441

Observation bfbee95e-0654-4dbc-9ef7-d123955fa6bf · outbound

This paper cites Let's verify step by step.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Let's verify step by step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.851235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.851235Z digest=sha256:903f3b63f7fb6e95e0fcca7176c9c75f6652a6d97cceb9dbd3312dd22c45a2fc

Observation e3ca2ecf-ebce-4e01-8414-30e0e656cd66 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.974215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.974215Z digest=sha256:69ac0661889ace6bb17c815cc607392d76dd1e1dc302d9e25ada525c24c6b546

Observation ba0de3c6-a8df-4580-8830-c8cb4ec0ff69 · outbound

This paper cites Training language models to follow instructions with human feedback.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Training language models to follow instructions with human feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.072863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.072863Z digest=sha256:87284851a10eb234df9a55f5d80b1b908b9420db647ce76d6e79af0c89438f3b

Observation c7627424-038d-4c6c-b72c-cd0a68fa5f26 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Direct preference optimization: Your language model is secretly a reward model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.132509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.132509Z digest=sha256:308a0eb2505578eca2872656b7fbbbf85ba0390982ea64ec9a5b784acae44d5e

Observation 0a5508d8-bd6f-4483-a7d7-4feb62be20b3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Proximal Policy Optimization Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.247181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.247181Z digest=sha256:ba62544724463145e82d203f29c143aab15eaf5e8a105b4a2ed1182b84ade0c9

Observation ca9dc9fc-7c7f-414a-9001-8d0c75cf5792 · outbound

This paper cites Trial and error: Exploration-based trajectory optimization of LLM agents.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Trial and error: Exploration-based trajectory optimization of LLM agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.374197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.374197Z digest=sha256:d25e71d8f9d6179e9fbcf5478c246253e5aac3302bca276354adc7c9032014a1

Observation 0b254d60-96f2-4a93-8945-e698f000a5d8 · outbound

This paper cites Math-shepherd: Verify and reinforce LLM s step-by-step without human annotations.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Math-shepherd: Verify and reinforce LLM s step-by-step without human annotations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.452605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.452605Z digest=sha256:6537d89b986ffbae2620eebb4ab9496f0bbfb099079063c30fc93f64cd8f2271

Observation ad41f962-9107-4136-8949-6e47c652b4d7 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:44.417785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:51:42.540673Z digest=sha256:b98265d7e0cda3cd76849349cc43287b824c22c4e84a03595cc3d02c0a20becf

Observation 742ac0a8-38d1-4a25-b960-0002ac15c9f1 · outbound

This paper cites Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.634639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.634639Z digest=sha256:544af7697535c92fd53cb3a87906bb654bbaa9cfaff72f7d259f65df74794927

Observation 21b930cb-cdb7-4214-ab74-04d336c0e2a8 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Chain-of-thought prompting elicits reasoning in large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.713789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.713789Z digest=sha256:d84fbb1cc2774279a38c796a0d058a5bb103eb1b35362c52044ca4eb219b0a20

Observation 9eb5afc0-64c0-48da-b24b-8ccaca91f770 · outbound

This paper cites Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.816272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.816272Z digest=sha256:c2b866d08162d3e03f5e4a65d504a497dc82dde6b52810be25ab19dd12ffee13

Observation 168df566-aeae-4b62-bb23-f970c537a8b9 · outbound

This paper cites Intercode: Standardizing and benchmarking interactive coding with execution feedback.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Intercode: Standardizing and benchmarking interactive coding with execution feedback

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:44.176909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:51:42.879826Z digest=sha256:20a7da84cf63985ef55e9b1e595325d8e26a504879a3e24ea59d544413039024

Observation 26ca5ac3-7e97-4b7c-8ede-2f9a8aa11425 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.009609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.009609Z digest=sha256:b8474d0c3dd0275cfb727978e94573d1122d4e06f653daeffe4f19ccc242b2bf

Observation 8a53a5e1-3352-4d9a-9141-eb749838ead0 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

RRO: LLM Agent Optimization Through Rising Reward Trajectories React: Synergizing reasoning and acting in language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.080279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.080279Z digest=sha256:f2cc09cdf94578bacb21c78a5f0bbe132e3be5de583f95c9c1feb45176d68096

Observation eb73a048-7729-42ef-8dd7-71ebfd3c7c38 · outbound

This paper cites Debug like a human: A large language model debugger via verifying runtime execution step by step.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Debug like a human: A large language model debugger via verifying runtime execution step by step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.139184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.139184Z digest=sha256:c4150f1f50a6f9e21d0a059a1de428a4255089a7d1e09f612133306fb44811a0

Observation c9aaadfe-1183-415f-8367-027decb7b1da · outbound

This paper cites Least-to-most prompting enables complex reasoning in large language models.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Least-to-most prompting enables complex reasoning in large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:43.901711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:51:43.257177Z digest=sha256:76d6585cbd8697128136bfe444845623e32c658e920c7428a95990588973e06c

Observation 89da84bb-3e8e-4bba-bb80-1899d29d02a5 · outbound

This paper cites @esa (Ref.

RRO: LLM Agent Optimization Through Rising Reward Trajectories @esa (Ref

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.371191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.371191Z digest=sha256:2d17e5738b1305284632a483ea18ef52bc71b8d220f4505fad177fe0e3560b7b

Observation 3a0a0449-7562-4032-8f2a-51fe5889d20f · outbound

This paper cites an unresolved cited work.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.499214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.499214Z digest=sha256:11caf88f1d526222228205e1b5e3438a86fd3966394f62ca403b35dd6d1a95f8

Observation 6c2b883c-d0da-4453-b261-85cf64a77b70 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.605901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.605901Z digest=sha256:4b914d34781a974d7379ada9045c1f593a4670e7eb108f0567c5b91d486a5330

Pith citing papers

No inbound Pith citation observations are available.