Pith. sign in

Paper Citation Record · LEDGER

Hell or High Water: Evaluating Agentic Recovery from External Failures

As of 16 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 3 inbound Pith citation observations for arXiv:2508.11027.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11027 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:35:04.819123Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:05:47.294068Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T09:55:40.559308Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e7be9ab1-4521-419e-877a-27faf18df70e · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Hell or High Water: Evaluating Agentic Recovery from External Failures Evaluating Large Language Models Trained on Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.570607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.570607Z digest=sha256:0fb58fcbc8dbeb271fd5c5243321dc0e9033059bf4cd3e4b83b3d9b06ede637f

Observation 69b255ff-cd94-48b1-904a-8db69fe30c74 · outbound

This paper cites Teaching Large Language Models to Self-Debug.

Hell or High Water: Evaluating Agentic Recovery from External Failures Teaching Large Language Models to Self-Debug

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.581696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.581696Z digest=sha256:f7de6c6e7ed91c12c675b36604ee8efa1ccf9d0d856a0ce41dc10f5489c34d56

Observation 86078d18-7922-4534-a362-c275eb6ccb9e · outbound

This paper cites ChatCoT: Tool-Augmented Chain-of-Thought Reasoning on Chat-based Large Language Models.

Hell or High Water: Evaluating Agentic Recovery from External Failures ChatCoT: Tool-Augmented Chain-of-Thought Reasoning on Chat-based Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.588917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.588917Z digest=sha256:44f51ea2aa6047ee9b7bba5acab67e88f170ac3adb1dd7ccc5a76f32cbc64e1f

Observation 5b5c3286-c8f7-41fb-a2a6-e062187493a2 · outbound

This paper cites AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls.

Hell or High Water: Evaluating Agentic Recovery from External Failures AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.595115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.595115Z digest=sha256:2d78e06989cf53e1a45a10c2428048868e8a0eea3632f85219a2ef237e1a6653

Observation c7164496-35eb-4b23-bfad-11bc1ce8a902 · outbound

This paper cites The Llama 3 Herd of Models.

Hell or High Water: Evaluating Agentic Recovery from External Failures The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.601215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.601215Z digest=sha256:30bd3a168b87389ce645a13685f500cfc18081dc3664bb4a461e673e95b9d4fe

Observation 26d9cb98-d90e-4767-bd61-7325d99c7da5 · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

Hell or High Water: Evaluating Agentic Recovery from External Failures CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.607118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.607118Z digest=sha256:604a0ec7b10fb42ad15c90ed06cae4d449f231fb506b3615df1049053afbed20

Observation aeacd5d1-46f9-4b70-ada7-6bb6f5f6d577 · outbound

This paper cites ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings.

Hell or High Water: Evaluating Agentic Recovery from External Failures ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.613057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.613057Z digest=sha256:a86f779236d8c4920e7dc6dbd277c9c0d80f79b6e121877daf2bbcfd833dbe0a

Observation 7ee430b8-f4d6-4960-bcf1-b0529eafa17d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Hell or High Water: Evaluating Agentic Recovery from External Failures Measuring Mathematical Problem Solving With the MATH Dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.618128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.618128Z digest=sha256:09a791f2329f8249ca6a44381bb494d5206e67f52ba453a20f7d0507afe965bb

Observation 9158879b-5bb3-4b4e-b9e0-dbe9c8be432d · outbound

This paper cites Not All LLM Reasoners Are Created Equal.

Hell or High Water: Evaluating Agentic Recovery from External Failures Not All LLM Reasoners Are Created Equal

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.623901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.623901Z digest=sha256:76665dd23597ef812bfafab32c21d3acc4a7d6ddd3d782c2ce53c0ea0942ef24

Observation 65f5f28c-1edc-420c-a0fd-2972f5e3973a · outbound

This paper cites Creativity in AI: Progresses and Challenges.

Hell or High Water: Evaluating Agentic Recovery from External Failures Creativity in AI: Progresses and Challenges

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.629737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.629737Z digest=sha256:15b7f25685fa4f5a5ac172f75186f40f12d6bcefceeeeb696776db170fa0d033

Observation 9e97a6f0-fb64-4a58-abf3-4ed741cb1c8c · outbound

This paper cites Feedback friction: Llms struggle to fully incorporate external feedback, 2025.

Hell or High Water: Evaluating Agentic Recovery from External Failures Feedback friction: Llms struggle to fully incorporate external feedback, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.637475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.637475Z digest=sha256:e32e5d0aa236cfe7c6461959efb5a8cd51ac21dcda1e062aec6ef5d66c6ce4a6

Observation 741ba194-7da5-48fd-be8f-1366c203fa9b · outbound

This paper cites Chain of Code: Reasoning with a Language Model-Augmented Code Emulator.

Hell or High Water: Evaluating Agentic Recovery from External Failures Chain of Code: Reasoning with a Language Model-Augmented Code Emulator

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.642526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.642526Z digest=sha256:e6c9231640d934d648a2bb77a35a9f28861ef22278954beebf1fd629a84710ec

Observation 814c3afb-4b6c-4c02-934a-1bed1aaa7ccd · outbound

This paper cites Api-bank: A benchmark for tool-augmented llms.

Hell or High Water: Evaluating Agentic Recovery from External Failures Api-bank: A benchmark for tool-augmented llms

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:35:05.932190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:35:04.648302Z digest=sha256:dbeb75d2a5500e353088cb9930297b34269a0fe90c47ad78cd6c559cdb91750b

Observation c8254a3f-bd8b-4506-bc88-76b597e87b29 · outbound

This paper cites ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities.

Hell or High Water: Evaluating Agentic Recovery from External Failures ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.652799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.652799Z digest=sha256:f7288ea37847ea9649906a0774ea8898caeb646f74276477dc8238c1244862e7

Observation 1bcfe983-083a-40b5-b0cf-ef6fc35cdb95 · outbound

This paper cites GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution.

Hell or High Water: Evaluating Agentic Recovery from External Failures GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.657890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.657890Z digest=sha256:503fd8f32a34bdc4aded110768b936392925bba3f71c95613e32592b50eb5506

Observation 250807b2-2fa2-4e0e-8ca5-cf06a0b6d7d3 · outbound

This paper cites Benchmarking Language Model Creativity: A Case Study on Code Generation.

Hell or High Water: Evaluating Agentic Recovery from External Failures Benchmarking Language Model Creativity: A Case Study on Code Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.662940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.662940Z digest=sha256:02aafa136735602f4d9c52484678818a87a71ac6e175fb6222631265db75a68b

Observation 085a074b-e743-4ed1-ad98-aec6a4b0dba5 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Hell or High Water: Evaluating Agentic Recovery from External Failures Self-refine: Iterative refinement with self-feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:35:05.913259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:35:04.667737Z digest=sha256:e98c84eee6dc8da4bfaf5940c067e2ddc8a2f870a6109196ce83e1a451e3e33f

Observation 103481f0-5e4e-4254-ab21-8a132555ba2b · outbound

This paper cites TALM: Tool Augmented Language Models.

Hell or High Water: Evaluating Agentic Recovery from External Failures TALM: Tool Augmented Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.673499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.673499Z digest=sha256:60c64afa211da83979bed5ed79f348ce48c9380f4dc4b27b7d654d6392818de0

Observation ee7ba190-35e7-452c-8664-caede273682a · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

Hell or High Water: Evaluating Agentic Recovery from External Failures Gorilla: Large Language Model Connected with Massive APIs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.678437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.678437Z digest=sha256:be94fa8a2070fd83e2e1f1a4aa0db633477c8d4d172aaab16dc98074526b7a19

Observation 2708f164-0128-422e-971f-a02351dac6cf · outbound

This paper cites EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents.

Hell or High Water: Evaluating Agentic Recovery from External Failures EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.683908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.683908Z digest=sha256:0db298db035402a92dbe39b5b5e40c5ef6a99ec5049dc6afb69b8763169ccf8a

Observation f7455207-d292-4bce-a197-55269bc3314d · outbound

This paper cites Making Language Models Better Tool Learners with Execution Feedback.

Hell or High Water: Evaluating Agentic Recovery from External Failures Making Language Models Better Tool Learners with Execution Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.689047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.689047Z digest=sha256:241ec60dc064c594ef88ed61e7f498205d2ef3bce735d3b646b7ac46817fa7ba

Observation e5cb56fe-b445-4770-b59f-314d345d9320 · outbound

This paper cites Tool Learning with Foundation Models.

Hell or High Water: Evaluating Agentic Recovery from External Failures Tool Learning with Foundation Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.694571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.694571Z digest=sha256:a6b61739dc51999e53484d6658c4fa550d25e76903cf6c32eee19caca454e3d8

Observation a18fdcfc-245c-47e3-a709-6c0ad197b9b4 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Hell or High Water: Evaluating Agentic Recovery from External Failures ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.700385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.700385Z digest=sha256:c2a6c70077481798f8b68409dc74f6e25c05e68156cbdcd16173613e29acb9f6

Observation 98c7743f-3199-4ff2-a7eb-fafdf924c562 · outbound

This paper cites Qwen2.5 Technical Report.

Hell or High Water: Evaluating Agentic Recovery from External Failures Qwen2.5 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.705490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.705490Z digest=sha256:af89918cf098fbaef88c248a59727f7e36c8d957243333a294333c6e4b574ac1

Observation 022af19a-0b06-47bb-8342-8da9b9302231 · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Hell or High Water: Evaluating Agentic Recovery from External Failures Self-critiquing models for assisting human evaluators

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.711079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.711079Z digest=sha256:fce80afe415bca836372a13289b02acff02a72ab6678d69f18c9697032c1b52d

Observation 6bc8ebe8-cb12-4d80-8080-dd0908b53d68 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

Hell or High Water: Evaluating Agentic Recovery from External Failures Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.718094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.718094Z digest=sha256:58e5bfaf159b47e14c55ff690b0f775902154d5e6278709d8ecbd229462080b8

Observation 84308df0-181e-4cb6-80f7-551a42aec180 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Hell or High Water: Evaluating Agentic Recovery from External Failures Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.724861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.724861Z digest=sha256:3bb6857c029a1f9ec279e3cf0a43ae788a7cb4a4095974eda41eabd45e0cce78

Observation f528b36d-0afc-4384-b1c0-811c75bf4ceb · outbound

This paper cites Tools fail: Detecting silent errors in faulty tools.

Hell or High Water: Evaluating Agentic Recovery from External Failures Tools fail: Detecting silent errors in faulty tools

Reference 28

Resolution
verified exact
doi, observed 2026-08-15T17:35:04.865552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:35:04.730123Z digest=sha256:6812de16e72dc8abbf05338b0e5ff88a3cad801fe944ba7b4385ea27227a3442

Observation 2c27f36c-6b43-464e-93d1-8c4bf7c1001d · outbound

This paper cites Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing.

Hell or High Water: Evaluating Agentic Recovery from External Failures Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.735848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.735848Z digest=sha256:cace4b3dcfde276ab6fee24375becf4bbc2dcdadcce6233dbd3c78e919f3dfc2

Observation 77161aab-ed0a-49f5-b8e1-51fb9c144aba · outbound

This paper cites MacGyver: Are Large Language Models Creative Problem Solvers?.

Hell or High Water: Evaluating Agentic Recovery from External Failures MacGyver: Are Large Language Models Creative Problem Solvers?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.741789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.741789Z digest=sha256:fd91a6bb384d87d2e8a5c6d3edb18d719cd6efabcc7d9fd4ed45079ace200f29

Observation b4a7301e-ab49-4020-b33b-caf333fbfc26 · outbound

This paper cites AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents.

Hell or High Water: Evaluating Agentic Recovery from External Failures AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.748315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.748315Z digest=sha256:a808e16092114c85c6e0d7c93f2d4b74c340ee6e7caea12626939d613094792e

Observation 4c2c4c0f-7dbe-4066-a0c0-40a413774888 · outbound

This paper cites LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error.

Hell or High Water: Evaluating Agentic Recovery from External Failures LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.755171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.755171Z digest=sha256:de4ec805f6a596a49279eafb81943d94a82d8b34bfc1e5e87d548599bd7d6098

Observation af94b5a2-a022-406d-ab01-2f71680bf1b8 · outbound

This paper cites Executable Code Actions Elicit Better LLM Agents.

Hell or High Water: Evaluating Agentic Recovery from External Failures Executable Code Actions Elicit Better LLM Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.761224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.761224Z digest=sha256:0cd63e5fd39a67af2b8c598724205d0bc70f574219d9bd0f27106bd093e4ac3d

Observation fc266d72-cba0-415a-8b7c-d31c58da90c9 · outbound

This paper cites Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus.

Hell or High Water: Evaluating Agentic Recovery from External Failures Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:35:05.896423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:35:04.766498Z digest=sha256:df8196a55fcfcdb1d71a9ed22304fac0df7bfd088df303245b446446cdb7eab6

Observation cd00d879-5d49-41dd-a89f-20c9d2c3469d · outbound

This paper cites TravelPlanner: A Benchmark for Real-World Planning with Language Agents.

Hell or High Water: Evaluating Agentic Recovery from External Failures TravelPlanner: A Benchmark for Real-World Planning with Language Agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.771353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.771353Z digest=sha256:ed865d5a392a1569892feca41e712e3c99a9dbdc8b55aceb517ad40697244081

Observation f02a91bc-171e-46bf-b8d6-73b96e307186 · outbound

This paper cites GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction.

Hell or High Water: Evaluating Agentic Recovery from External Failures GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.776568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.776568Z digest=sha256:be796e581c7fce937346b0adf65bbc5ac2b7ecad1907865b9cec998573f4462a

Observation 325dcc46-6bb2-40a4-9651-aa78fd779d91 · outbound

This paper cites Narasimhan, and Yuan Cao.

Hell or High Water: Evaluating Agentic Recovery from External Failures Narasimhan, and Yuan Cao

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:35:05.879324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:35:04.781665Z digest=sha256:0852d318fde265031c7a631fee20ad167db9f9d879d988de4871dce764b4e7e6

Observation 293a55cf-fdbe-48fc-b32b-0eaef57149e1 · outbound

This paper cites ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use.

Hell or High Water: Evaluating Agentic Recovery from External Failures ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.786251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.786251Z digest=sha256:cb05f72c3da08f2be9eb2d08069ec59d5aaa81d3f2a01306ee9e6af12335d5d3

Observation 10fb510f-3660-4539-9db2-1b87d62a0fa0 · outbound

This paper cites S pider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to- SQL task.

Hell or High Water: Evaluating Agentic Recovery from External Failures S pider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to- SQL task

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.791562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.791562Z digest=sha256:d2b2011803ae9fec9036ae50446760c93ae4e11693b0ab2196f393b6a3ce74e5

Observation 80662e47-dfbc-4851-ab83-11f7d49bc0e2 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Hell or High Water: Evaluating Agentic Recovery from External Failures Instruction-Following Evaluation for Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.797558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.797558Z digest=sha256:642ae227dba32bbbc3a5e0538d7b3f48e99464e2ddd24e624c15b397c64bd243

Observation 50ac5ae7-4539-48fc-b1e4-a38fd99ade02 · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.

Hell or High Water: Evaluating Agentic Recovery from External Failures Toolqa: A dataset for llm question answering with external tools

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:35:05.850841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:35:04.802632Z digest=sha256:c47a5e43a560128970c0f9bafa19c64b59e4b14eed00bf6abe2e75c746c710ca

Observation 37e372b8-237a-4c2d-917b-61c9b90ac51a · outbound

This paper cites @esa (Ref.

Hell or High Water: Evaluating Agentic Recovery from External Failures @esa (Ref

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.807665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.807665Z digest=sha256:9d26994183357a3e8b1e238e61f9071358cadae5fa581494e9c6721fc391729c

Observation 697dbf48-d41e-4789-a822-439c811d3938 · outbound

This paper cites an unresolved cited work.

Hell or High Water: Evaluating Agentic Recovery from External Failures Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.813952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.813952Z digest=sha256:f8eacdced1b34da8d19b55a3f23dcd196026931173deb6e649bb0fa55c334ea3

Observation f6373f42-2c47-40e5-88fd-03d49e3b9023 · outbound

This paper cites an unresolved cited work.

Hell or High Water: Evaluating Agentic Recovery from External Failures Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.819123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.819123Z digest=sha256:528d0d43e2c83b8a7d8b2a1a96bfa0cdba2da8c2592494742d9ff1f95f4d532c

Pith citing papers

Observation 9c3c0eba-9da5-4d37-9924-5c6c4ef896d5 · inbound

Robust Agent Compensation (RAC): Teaching AI Agents to Compensate cites this paper.

Robust Agent Compensation (RAC): Teaching AI Agents to Compensate Hell or High Water: Evaluating Agentic Recovery from External Failures

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:30.457941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T16:54:15.433189Z digest=sha256:10fe56baccf8047e80943f723fd14cc9e1eaa6b33ced012e423c3e0a37ca75af

Observation 0adb1785-cff9-4e83-beb7-24b80ffce843 · inbound

Robust Agent Compensation (RAC): Teaching AI Agents to Compensate cites this paper.

Robust Agent Compensation (RAC): Teaching AI Agents to Compensate Hell or High Water: Evaluating Agentic Recovery from External Failures

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:23:52.457861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T00:22:54.128310Z digest=sha256:36bfc488249124d4a0b3e353f228e8688d638c1023d5de2379f76acf563fe2f1

Observation d82411c5-65f6-4a75-be0b-952622749dbd · inbound

When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue cites this paper.

When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue Hell or High Water: Evaluating Agentic Recovery from External Failures

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:55:40.560856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-01T06:05:47.294068Z digest=sha256:84bc3171302d0b10d640c035b8ad4d794ebef5384a50a01a55bafe8f5ec742dc