Pith. sign in

Paper Citation Record · LEDGER

Exploring Expert Failures Improves LLM Agent Tuning

As of 18 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 2 inbound Pith citation observations for arXiv:2504.13145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.13145 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:18:49.217321Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:35:18.372917Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:37:49.973211Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 249b3db7-8aff-4e77-b34a-f77a8c814a26 · outbound

This paper cites GPT-4 Technical Report.

Exploring Expert Failures Improves LLM Agent Tuning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.351840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.351840Z digest=sha256:cad7f40b27249a8b54f1421d39e6a214ed254eeb73486a2199abd118df7e942f

Observation d0d3ecae-0ccb-4ac0-9a07-f0eaa1188143 · outbound

This paper cites Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models.

Exploring Expert Failures Improves LLM Agent Tuning Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.456453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.456453Z digest=sha256:94108be194584926e8c6ab59ec9904fed9034c935710ffda98c18f076cbf8b85

Observation 84cfe1d2-0b4d-4f22-8fa9-e6bc399ca847 · outbound

This paper cites ATLaS: Agent Tuning via Learning Critical Steps.

Exploring Expert Failures Improves LLM Agent Tuning ATLaS: Agent Tuning via Learning Critical Steps

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.461379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.461379Z digest=sha256:a8c382adf9f9b392df8f1f0326a0bed9e07f2dd54c4ec5ba4c7e4739edb8b0fc

Observation 72cc724b-ec18-4e01-8074-87fae4431a62 · outbound

This paper cites Contextual Markov Decision Processes.

Exploring Expert Failures Improves LLM Agent Tuning Contextual Markov Decision Processes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.505040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.505040Z digest=sha256:88b369383fadf8f876d4ab1a1aff4224da229d5960b9cdc271f0b12ee7c54201

Observation cc2a8a1b-8fc2-475e-bf22-77fb5cbf60a5 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

Exploring Expert Failures Improves LLM Agent Tuning Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.636400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.636400Z digest=sha256:ef1039f732d53882f7dcb893f91aa913f014bf63e5ed1eb6d67c2a85c38151cd

Observation 44a0e70f-1438-40bc-8e6c-d88f0b49fba0 · outbound

This paper cites Can Agents Run Relay Race with Strangers? Generalization of RL to Out-of-Distribution Trajectories.

Exploring Expert Failures Improves LLM Agent Tuning Can Agents Run Relay Race with Strangers? Generalization of RL to Out-of-Distribution Trajectories

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:18:49.545440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:18:48.642240Z digest=sha256:382fda0d195927fbe6ce7b587085f4e64538664606592aa145047853970e3f5d

Observation 8a194436-aeff-4e10-bc78-a7631828a1d7 · outbound

This paper cites AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System.

Exploring Expert Failures Improves LLM Agent Tuning AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.647661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.647661Z digest=sha256:7e5d1f987f690bf923f509c05f1bea517b8dd3f696d1e7a289558bea1a78c1ea

Observation b03e05ab-cc8b-4679-b438-1eb760f011d3 · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.

Exploring Expert Failures Improves LLM Agent Tuning AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.652037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.652037Z digest=sha256:09a0151934d0d7148e576f5c899c6f0e0005af9d9cf72b1bf75c8b73224386a7

Observation 8c65e9a6-bc17-4964-b91e-9929624fb5c5 · outbound

This paper cites Let's reward step by step: Step-Level reward model as the Navigators for Reasoning.

Exploring Expert Failures Improves LLM Agent Tuning Let's reward step by step: Step-Level reward model as the Navigators for Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.731690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.731690Z digest=sha256:54c568c40033788c23ada11fcae041d17c364bc8248696afc7bbae83344750e0

Observation 0734153d-3925-46e4-9365-2bd38a29105a · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Exploring Expert Failures Improves LLM Agent Tuning Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.806241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.806241Z digest=sha256:11f4e847026fdf5b8c3b6f05a6b63a4f7ffd0125fb525af36f6e4510e57bb490

Observation 44660778-45d2-4f54-9d0c-b38c82ffa363 · outbound

This paper cites Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Peter J.

Exploring Expert Failures Improves LLM Agent Tuning Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Peter J

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:18:49.744493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:18:48.811365Z digest=sha256:1c42b3a8c138417c2bc9fd990cd3ab354936b132f2b79c5e590a3a3a9747441e

Observation 76c7a648-bc90-4e61-a3c0-e7323f7616ec · outbound

This paper cites AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories.

Exploring Expert Failures Improves LLM Agent Tuning AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.816530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.816530Z digest=sha256:fcfd135aedfe8c3c0e42ff329d90b23b6a48e3c00b13475d917df8d1dc7f9698

Observation d39c7cb4-c684-413d-937d-243cf74acd92 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Exploring Expert Failures Improves LLM Agent Tuning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.822368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.822368Z digest=sha256:423b506e3bb5736ed425ffd9301a01031c3f4965d0c28583be1276d7c78227dc

Observation 28eeec0e-a312-48d7-88fc-48956ab785d8 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Exploring Expert Failures Improves LLM Agent Tuning Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.932683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.932683Z digest=sha256:35d884bb4696de9ffa1500d3704f218b82b8580c3eb1a4d481d58c4822c1b3ae

Observation 592ad90b-a4d2-4c0a-a807-e590420720b9 · outbound

This paper cites Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement.

Exploring Expert Failures Improves LLM Agent Tuning Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:49.043336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:49.043336Z digest=sha256:a70d0f858513b9fa4d494e33c2a7bfe907117f2782a059788d9c31332024a528

Observation c99db4df-5e29-4fb1-8941-6487d8232cc0 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Exploring Expert Failures Improves LLM Agent Tuning Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:49.048092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:49.048092Z digest=sha256:381be830ca23fdb968c9351e3fe029ff9e93b633caaad91f5f82ec4724172bec

Observation 0395c9fb-f92f-411d-9e33-5911a80ef306 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

Exploring Expert Failures Improves LLM Agent Tuning AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:49.053848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:49.053848Z digest=sha256:fb3f3f899f570595a7ef69adf3de94c02f659c3e547b98ef5ca16c62efda22a4

Observation 038236ee-fc72-4ba4-8dc0-6b827899e086 · outbound

This paper cites AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning.

Exploring Expert Failures Improves LLM Agent Tuning AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:49.121774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:49.121774Z digest=sha256:33cb4e58b77ca87412194d4dede4b48fcc07f88b1ea77d554dec4967a693cb9c

Observation 21a63a2a-87fe-4779-8e77-355d2a965055 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Exploring Expert Failures Improves LLM Agent Tuning WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:49.217321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:49.217321Z digest=sha256:6066ae373737bad60a2014ec39b8849faf95c6810df2e20913c7ea2a6243e0d7

Observation ebf6be87-6040-4376-b23e-04bfd0ca8bc8 · outbound

This paper cites From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning.

Exploring Expert Failures Improves LLM Agent Tuning From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.467217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.467217Z digest=sha256:0ac93779da51b059ab8a43c9cb2936bf6c3f8b84f4387ae32a4576b8a25d0de7

Observation 72ea9de0-b53f-46ba-a609-7c688668cfed · outbound

This paper cites GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements.

Exploring Expert Failures Improves LLM Agent Tuning GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.585351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.585351Z digest=sha256:9c4d665a5ac70e014f3867f5b811ec159bea99ccc56989b842ba22827c5d8d5d

Observation 83a868b4-051e-4786-b2b0-982c031b2ea0 · outbound

This paper cites Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision.

Exploring Expert Failures Improves LLM Agent Tuning Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:49.037721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:49.037721Z digest=sha256:96ee0c3b246fa2e404fb5173e5c9047541cd7979bcc84d50e8b62840923ae329

Observation c5dd1a83-41b9-4a2f-97c9-555ac768a72b · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Exploring Expert Failures Improves LLM Agent Tuning FireAct: Toward Language Agent Fine-tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.450541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.450541Z digest=sha256:992b666b8efb60da7ab22417279345082927225e55747f7e440e377123b24fdb

Observation 98cd6fc3-0cdb-4b1e-9952-546dc8930ff9 · outbound

This paper cites ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent.

Exploring Expert Failures Improves LLM Agent Tuning ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.417057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.417057Z digest=sha256:0c9a36c7b6a1294f0d670c2dfd868acff559ad5f59fb0e35b438673321937577

Observation 64679209-0df5-4952-8128-8e06271f7c27 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Exploring Expert Failures Improves LLM Agent Tuning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:49.032082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:49.032082Z digest=sha256:d03db5057f24077bd2e36a95eb310dd0aa9ab6e8cb523e146e3c0652d1f6e55d

Pith citing papers

Observation 58e481f7-a907-4cbb-9051-e260b61b71d4 · inbound

ProgRM: Build Better GUI Agents with Progress Rewards cites this paper.

ProgRM: Build Better GUI Agents with Progress Rewards Exploring Expert Failures Improves LLM Agent Tuning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:37:50.068075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:37:45.644348Z digest=sha256:4c4fc9dc2e67a0528056773154525932628c86b5d4e2d44d16afba2d546be3bc

Observation ac2fa8c3-3207-4f37-aa2d-4f7a74718ad6 · inbound

FailForge: Distilling Procedural Competence from Persistent Failures into Code Agents cites this paper.

FailForge: Distilling Procedural Competence from Persistent Failures into Code Agents Exploring Expert Failures Improves LLM Agent Tuning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:18.372917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:18.372917Z digest=sha256:2510cfcdf89e97aea847bd6f3843a9e35c1e16a7d778ba32759ae0eec46bfc5c