Pith. sign in

Paper Citation Record · LEDGER

Hell or High Water: Evaluating Agentic Recovery from External Failures

As of 19 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 3 inbound Pith citation observations for arXiv:2508.11027.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11027 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:35:04.819123Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:05:47.294068Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T09:55:40.559308Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e7be9ab1-4521-419e-877a-27faf18df70e · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Hell or High Water: Evaluating Agentic Recovery from External Failures Evaluating Large Language Models Trained on Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.570607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.570607Z digest=sha256:e282511c93f41d8c262a4d8cc84d61ba13a805927e261f4ee26a794659e52056

Observation 69b255ff-cd94-48b1-904a-8db69fe30c74 · outbound

This paper cites Teaching Large Language Models to Self-Debug.

Hell or High Water: Evaluating Agentic Recovery from External Failures Teaching Large Language Models to Self-Debug

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.581696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.581696Z digest=sha256:5d91c8796c590d780ab4b398c1ef68240898c12a9d24b6d5d45ed50872ac3010

Observation 86078d18-7922-4534-a362-c275eb6ccb9e · outbound

This paper cites ChatCoT: Tool-Augmented Chain-of-Thought Reasoning on Chat-based Large Language Models.

Hell or High Water: Evaluating Agentic Recovery from External Failures ChatCoT: Tool-Augmented Chain-of-Thought Reasoning on Chat-based Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.588917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.588917Z digest=sha256:66b4a8bde35059c3ac5bc5c534915a9cacb9c891e5e3f1b1522640cca642cec7

Observation 5b5c3286-c8f7-41fb-a2a6-e062187493a2 · outbound

This paper cites AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls.

Hell or High Water: Evaluating Agentic Recovery from External Failures AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.595115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.595115Z digest=sha256:baf2819a6acd7251fade3465811290db06dc9d48ec790afc877ff1cd81a5860e

Observation c7164496-35eb-4b23-bfad-11bc1ce8a902 · outbound

This paper cites The Llama 3 Herd of Models.

Hell or High Water: Evaluating Agentic Recovery from External Failures The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.601215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.601215Z digest=sha256:58cfc59070894ff34d7d1017e112dfd627c4d5bbb85cecb105c476180abeda43

Observation 26d9cb98-d90e-4767-bd61-7325d99c7da5 · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

Hell or High Water: Evaluating Agentic Recovery from External Failures CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.607118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.607118Z digest=sha256:30e6774b990098a2b9f03974b5a2659434cc7769aadfda4093b8212ecde6c818

Observation aeacd5d1-46f9-4b70-ada7-6bb6f5f6d577 · outbound

This paper cites ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings.

Hell or High Water: Evaluating Agentic Recovery from External Failures ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.613057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.613057Z digest=sha256:9ffd014a666c75fe89da335e634835f19997042878eff01870be2a5337331439

Observation 7ee430b8-f4d6-4960-bcf1-b0529eafa17d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Hell or High Water: Evaluating Agentic Recovery from External Failures Measuring Mathematical Problem Solving With the MATH Dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.618128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.618128Z digest=sha256:98d388ca94cf33184494da67c1dced26d1250abdad3a6a146583e2e391f8ee01

Observation 9158879b-5bb3-4b4e-b9e0-dbe9c8be432d · outbound

This paper cites Not All LLM Reasoners Are Created Equal.

Hell or High Water: Evaluating Agentic Recovery from External Failures Not All LLM Reasoners Are Created Equal

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.623901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.623901Z digest=sha256:05a01980ca1345a10d1eb497beeb32f9ffdc4a414ed4a42d3dc849e13014a7ee

Observation 65f5f28c-1edc-420c-a0fd-2972f5e3973a · outbound

This paper cites Creativity in AI: Progresses and Challenges.

Hell or High Water: Evaluating Agentic Recovery from External Failures Creativity in AI: Progresses and Challenges

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.629737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.629737Z digest=sha256:557da28bd43b9c5c637952863a898244f6d5b6767d6637b802782a55fd410526

Observation 9e97a6f0-fb64-4a58-abf3-4ed741cb1c8c · outbound

This paper cites Feedback friction: Llms struggle to fully incorporate external feedback, 2025.

Hell or High Water: Evaluating Agentic Recovery from External Failures Feedback friction: Llms struggle to fully incorporate external feedback, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.637475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.637475Z digest=sha256:2f7c4463fdf8c2bce8e516b17b3047cc9a9e315471d5e4214a91d5e603b009b3

Observation 741ba194-7da5-48fd-be8f-1366c203fa9b · outbound

This paper cites Chain of Code: Reasoning with a Language Model-Augmented Code Emulator.

Hell or High Water: Evaluating Agentic Recovery from External Failures Chain of Code: Reasoning with a Language Model-Augmented Code Emulator

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.642526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.642526Z digest=sha256:18e7ca6a52a356bf5c7768cadc67ed9cbc22b64384ddd2dfb311243c734d6513

Observation 814c3afb-4b6c-4c02-934a-1bed1aaa7ccd · outbound

This paper cites Api-bank: A benchmark for tool-augmented llms.

Hell or High Water: Evaluating Agentic Recovery from External Failures Api-bank: A benchmark for tool-augmented llms

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:35:05.932190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:35:04.648302Z digest=sha256:6011d6ed751fabe6844a0853ff8ca0ba83039361d015de1ed67052fe3ebf5201

Observation c8254a3f-bd8b-4506-bc88-76b597e87b29 · outbound

This paper cites ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities.

Hell or High Water: Evaluating Agentic Recovery from External Failures ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.652799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.652799Z digest=sha256:588868d33b98b52f472273f5b60521485339702f0d1188244776fbe6289eb2c9

Observation 1bcfe983-083a-40b5-b0cf-ef6fc35cdb95 · outbound

This paper cites GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution.

Hell or High Water: Evaluating Agentic Recovery from External Failures GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.657890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.657890Z digest=sha256:7cdf955c6fbb4ecce9496d9d3cde37c431d27dc139212f773086bdda4b152297

Observation 250807b2-2fa2-4e0e-8ca5-cf06a0b6d7d3 · outbound

This paper cites Benchmarking Language Model Creativity: A Case Study on Code Generation.

Hell or High Water: Evaluating Agentic Recovery from External Failures Benchmarking Language Model Creativity: A Case Study on Code Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.662940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.662940Z digest=sha256:69166bbbf423fafdf4f66fd9246337c3dde834701aa6f33f24e542c4b81a17ed

Observation 085a074b-e743-4ed1-ad98-aec6a4b0dba5 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Hell or High Water: Evaluating Agentic Recovery from External Failures Self-refine: Iterative refinement with self-feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:35:05.913259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:35:04.667737Z digest=sha256:fa9e55323a10e76cdc2003dd179fc81644f87d71b137fc84b6398e185bb1069a

Observation 103481f0-5e4e-4254-ab21-8a132555ba2b · outbound

This paper cites TALM: Tool Augmented Language Models.

Hell or High Water: Evaluating Agentic Recovery from External Failures TALM: Tool Augmented Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.673499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.673499Z digest=sha256:3d2c7f2957c51ec88a4d151ec6544d928df6a2a3bfe32629a2c92a90a9063229

Observation ee7ba190-35e7-452c-8664-caede273682a · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

Hell or High Water: Evaluating Agentic Recovery from External Failures Gorilla: Large Language Model Connected with Massive APIs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.678437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.678437Z digest=sha256:5de7adc5f88dc272ac412b3f5eedf2fdc433707faa032a6833e673eb6fbdaf34

Observation 2708f164-0128-422e-971f-a02351dac6cf · outbound

This paper cites EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents.

Hell or High Water: Evaluating Agentic Recovery from External Failures EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.683908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.683908Z digest=sha256:96c5b9e485ee84d83eda1bdc66d9dfeab9fd9eed74f5e2c67d5125a9b9792fdf

Observation f7455207-d292-4bce-a197-55269bc3314d · outbound

This paper cites Making Language Models Better Tool Learners with Execution Feedback.

Hell or High Water: Evaluating Agentic Recovery from External Failures Making Language Models Better Tool Learners with Execution Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.689047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.689047Z digest=sha256:043aa04c299bda9539adf0f7b5de9e2d1351d912f46e0d6306b5c2dea726c04c

Observation e5cb56fe-b445-4770-b59f-314d345d9320 · outbound

This paper cites Tool Learning with Foundation Models.

Hell or High Water: Evaluating Agentic Recovery from External Failures Tool Learning with Foundation Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.694571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.694571Z digest=sha256:d6133415bc89e5ae515253e0e3f9bee958b3b562614cdbf95530fefcc1e478dd

Observation a18fdcfc-245c-47e3-a709-6c0ad197b9b4 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Hell or High Water: Evaluating Agentic Recovery from External Failures ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.700385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.700385Z digest=sha256:20dbf0c62e087f87187f56fc3cf7ec227ef70c00e718d2d34ccbb2fd4485bd1e

Observation 98c7743f-3199-4ff2-a7eb-fafdf924c562 · outbound

This paper cites Qwen2.5 Technical Report.

Hell or High Water: Evaluating Agentic Recovery from External Failures Qwen2.5 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.705490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.705490Z digest=sha256:d660cb2545d2fab895cf28ce969dd2bb90b8b80d28d98cdafccc77bfda554fdc

Observation 022af19a-0b06-47bb-8342-8da9b9302231 · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Hell or High Water: Evaluating Agentic Recovery from External Failures Self-critiquing models for assisting human evaluators

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.711079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.711079Z digest=sha256:be0b812ad6b0d72c8c15f3b090c20bbbf36557bd96719406770a0f8956c776f6

Observation 6bc8ebe8-cb12-4d80-8080-dd0908b53d68 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

Hell or High Water: Evaluating Agentic Recovery from External Failures Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.718094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.718094Z digest=sha256:6c4918db44112ad297e94da5e157400ef27910f6634e1aa72523a5a2be4cd1be

Observation 84308df0-181e-4cb6-80f7-551a42aec180 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Hell or High Water: Evaluating Agentic Recovery from External Failures Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.724861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.724861Z digest=sha256:b58bfd9a065b33ee460d73e25d762e8e46021d4401aaa8b4f66e7f4c992febe7

Observation f528b36d-0afc-4384-b1c0-811c75bf4ceb · outbound

This paper cites Tools fail: Detecting silent errors in faulty tools.

Hell or High Water: Evaluating Agentic Recovery from External Failures Tools fail: Detecting silent errors in faulty tools

Reference 28

Resolution
verified exact
doi, observed 2026-08-15T17:35:04.865552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:35:04.730123Z digest=sha256:15dae0d5fc74c862abbf5338c077572c268b1f668e10db27bf7bf6cfbbfe6520

Observation 2c27f36c-6b43-464e-93d1-8c4bf7c1001d · outbound

This paper cites Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing.

Hell or High Water: Evaluating Agentic Recovery from External Failures Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.735848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.735848Z digest=sha256:fcccb089f5e9078d179cefe7b9edfd52c2048f0c660c753227af445ec4fe77db

Observation 77161aab-ed0a-49f5-b8e1-51fb9c144aba · outbound

This paper cites MacGyver: Are Large Language Models Creative Problem Solvers?.

Hell or High Water: Evaluating Agentic Recovery from External Failures MacGyver: Are Large Language Models Creative Problem Solvers?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.741789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.741789Z digest=sha256:807c4d2645d483352b25ac51406fa03c875b60ff8b8d98d80b272151003ef96a

Observation b4a7301e-ab49-4020-b33b-caf333fbfc26 · outbound

This paper cites AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents.

Hell or High Water: Evaluating Agentic Recovery from External Failures AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.748315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.748315Z digest=sha256:a883b5c15610e0b307699aa201f436a6c1589937e8235818d6b5b01055832f9b

Observation 4c2c4c0f-7dbe-4066-a0c0-40a413774888 · outbound

This paper cites LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error.

Hell or High Water: Evaluating Agentic Recovery from External Failures LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.755171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.755171Z digest=sha256:1f83e39daf5594c7be9e7955de4b3621a7689eb375214fe86df6aa1a2faab995

Observation af94b5a2-a022-406d-ab01-2f71680bf1b8 · outbound

This paper cites Executable Code Actions Elicit Better LLM Agents.

Hell or High Water: Evaluating Agentic Recovery from External Failures Executable Code Actions Elicit Better LLM Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.761224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.761224Z digest=sha256:11e945cec76f6b9c3658f2e1e6ea3454d5598603b47da30fc07e6fc69eb37bca

Observation fc266d72-cba0-415a-8b7c-d31c58da90c9 · outbound

This paper cites Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus.

Hell or High Water: Evaluating Agentic Recovery from External Failures Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:35:05.896423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:35:04.766498Z digest=sha256:063934abf6360856fab4f6b82cc292b6e704da864bd59ac38d5b9982ece718fe

Observation cd00d879-5d49-41dd-a89f-20c9d2c3469d · outbound

This paper cites TravelPlanner: A Benchmark for Real-World Planning with Language Agents.

Hell or High Water: Evaluating Agentic Recovery from External Failures TravelPlanner: A Benchmark for Real-World Planning with Language Agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.771353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.771353Z digest=sha256:258fbdaa62aaeca5c5670be704bd945b033be3c1381b9943494d62b0da23436d

Observation f02a91bc-171e-46bf-b8d6-73b96e307186 · outbound

This paper cites GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction.

Hell or High Water: Evaluating Agentic Recovery from External Failures GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.776568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.776568Z digest=sha256:b63a9a9f3a3d6d9d75ad8c15940b1a976a1f7b8e0ee6e98688eefcea3821192e

Observation 325dcc46-6bb2-40a4-9651-aa78fd779d91 · outbound

This paper cites Narasimhan, and Yuan Cao.

Hell or High Water: Evaluating Agentic Recovery from External Failures Narasimhan, and Yuan Cao

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:35:05.879324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:35:04.781665Z digest=sha256:e00bb8a5802ef72f2f3f72559f43f1415910845bb175104ae0399dcd8d3eab55

Observation 293a55cf-fdbe-48fc-b32b-0eaef57149e1 · outbound

This paper cites ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use.

Hell or High Water: Evaluating Agentic Recovery from External Failures ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.786251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.786251Z digest=sha256:16bb1381f87bb88386fb81213864e6c27462172d3a139ad81b50bc066777dded

Observation 10fb510f-3660-4539-9db2-1b87d62a0fa0 · outbound

This paper cites S pider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to- SQL task.

Hell or High Water: Evaluating Agentic Recovery from External Failures S pider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to- SQL task

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.791562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.791562Z digest=sha256:4878d54c701345567054bb60e1b6ed9523fbd14d60b5a05180bec134f552526c

Observation 80662e47-dfbc-4851-ab83-11f7d49bc0e2 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Hell or High Water: Evaluating Agentic Recovery from External Failures Instruction-Following Evaluation for Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.797558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.797558Z digest=sha256:1f9c475a84a3a84b96791271a4d4c41fc6e14a2ed8e458ad96fd23b4100520fc

Observation 50ac5ae7-4539-48fc-b1e4-a38fd99ade02 · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.

Hell or High Water: Evaluating Agentic Recovery from External Failures Toolqa: A dataset for llm question answering with external tools

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:35:05.850841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:35:04.802632Z digest=sha256:af23be73e6da372ba97bd819b1d7ee4181c8a09ced650a19a087fa0d83080a13

Observation 37e372b8-237a-4c2d-917b-61c9b90ac51a · outbound

This paper cites @esa (Ref.

Hell or High Water: Evaluating Agentic Recovery from External Failures @esa (Ref

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.807665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.807665Z digest=sha256:2e2be177cb27786be869c8f220b1e286f08dd1dc0f5d8e8cb11e05ce50ef4e6a

Observation 697dbf48-d41e-4789-a822-439c811d3938 · outbound

This paper cites an unresolved cited work.

Hell or High Water: Evaluating Agentic Recovery from External Failures Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.813952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.813952Z digest=sha256:5da7353b34efc437c7393c7e6d0891a357b9a1a2b3065c49e9781a6321a80d09

Observation f6373f42-2c47-40e5-88fd-03d49e3b9023 · outbound

This paper cites an unresolved cited work.

Hell or High Water: Evaluating Agentic Recovery from External Failures Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.819123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.819123Z digest=sha256:74d382d4def3aacf6233fd5c3bfc6b2acb39d8b310fb83437d50eba04bdc1a6e

Pith citing papers

Observation 9c3c0eba-9da5-4d37-9924-5c6c4ef896d5 · inbound

Robust Agent Compensation (RAC): Teaching AI Agents to Compensate cites this paper.

Robust Agent Compensation (RAC): Teaching AI Agents to Compensate Hell or High Water: Evaluating Agentic Recovery from External Failures

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:30.457941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T16:54:15.433189Z digest=sha256:32abde3eb3a0928388024aa79685a6086e7c71ac1b8c8912e4df3bc1390707b5

Observation 0adb1785-cff9-4e83-beb7-24b80ffce843 · inbound

Robust Agent Compensation (RAC): Teaching AI Agents to Compensate cites this paper.

Robust Agent Compensation (RAC): Teaching AI Agents to Compensate Hell or High Water: Evaluating Agentic Recovery from External Failures

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:23:52.457861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T00:22:54.128310Z digest=sha256:7f5cec1eeb2db8a42b4b659a3600e5b3087330bdb80c7c44ec4c3da5d474c322

Observation d82411c5-65f6-4a75-be0b-952622749dbd · inbound

When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue cites this paper.

When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue Hell or High Water: Evaluating Agentic Recovery from External Failures

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:55:40.560856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:05:47.294068Z digest=sha256:c5f18869b7176753bbf7e2b52d2dc369b6f6595f66247e7adaaa18ba46c4ad9a