Pith. sign in

Paper Citation Record · LEDGER

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks

As of 19 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 7 inbound Pith citation observations for arXiv:2508.13143.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13143 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:19:51.873664Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:20:42.058799Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T04:26:38.042285Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b859add-951e-47b5-b3ed-7250b0ba4d21 · outbound

This paper cites TaskWeaver: A Code-First Agent Framework.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks TaskWeaver: A Code-First Agent Framework

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.742803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.742803Z digest=sha256:cc04aa59cb1e6b86baf161894c33a195d02f68abd78cc475023a1c47429504cd

Observation 132a18b8-c8d5-446d-8c7e-351c836e48fd · outbound

This paper cites Autogen: Enabling next-gen LLM applications via multi- agent conversations,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Autogen: Enabling next-gen LLM applications via multi- agent conversations,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.280593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.747671Z digest=sha256:3bc0b3b665374f78221f5378c97afa8165a764af47f82e1e85d01f7ac7de757c

Observation df860d48-b452-4c3b-b986-be57d8f0f674 · outbound

This paper cites CodeAgent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks CodeAgent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.267893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.751664Z digest=sha256:b8cce6e9a64cfde138d4dedadeb4f11505d51be1596eaff97a6d9308740871bf

Observation 0f70a63c-72a9-4a61-b1e6-faecf997c16d · outbound

This paper cites Executable code actions elicit better llm agents,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Executable code actions elicit better llm agents,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.254063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.755536Z digest=sha256:2eb3a0177fa8a5c74261d21e576b526cdf336ac3de39344e3736882ca072a1b0

Observation afcf02cd-bb57-4233-9f3c-22867bfca568 · outbound

This paper cites Solving challenging math word problems using gpt-4 code interpreter with code-based self-verification,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Solving challenging math word problems using gpt-4 code interpreter with code-based self-verification,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.240211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.759413Z digest=sha256:f2af0bc25b89c464b63e85c563aec8767a4b73873fea52fea5c4e87134834d18

Observation 39966860-d1b3-41e3-8f0c-2b3f2a872f0f · outbound

This paper cites CIBench: Evaluating Your LLMs with a Code Interpreter Plugin.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks CIBench: Evaluating Your LLMs with a Code Interpreter Plugin

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.764153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.764153Z digest=sha256:9c56c20af385c67d847ccf543dc254e02ab82cde6367ddb78d30f76841d5418f

Observation 486df829-a9a7-4ac3-86c5-19c55f4e3f19 · outbound

This paper cites Data Dialogue with ChatGPT: Using Code Interpreter to Simulate and Analyse Experimental Data.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Data Dialogue with ChatGPT: Using Code Interpreter to Simulate and Analyse Experimental Data

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.769010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.769010Z digest=sha256:23b5198ced055bc77087e4aa4c9b3eddc4236370162182f350081411ec0c1085

Observation 4d996104-93ca-42f6-b079-926f71cb59ed · outbound

This paper cites Webpilot: A versatile and autonomous multi-agent system for web task execution with strategic exploration,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Webpilot: A versatile and autonomous multi-agent system for web task execution with strategic exploration,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.226686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.773421Z digest=sha256:3c210ffc412135a4ee5b89ad9ea1be379fe0e083ca69774d2f0e65fcae829691

Observation 14af42c9-f294-4aa1-831f-833ed575c4d7 · outbound

This paper cites Redcode: Risky code execution and generation benchmark for code agents,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Redcode: Risky code execution and generation benchmark for code agents,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.777304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.777304Z digest=sha256:5e810908cd3541d1be2ff4c2df1e99cfe701997c20e50a59200d9e3df48f3287

Observation a87fecc5-a8d4-48ce-9b1d-9ea05a7f7655 · outbound

This paper cites Large Language Model-Based Agents for Software Engineering: A Survey.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Large Language Model-Based Agents for Software Engineering: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.781158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.781158Z digest=sha256:ce2925cae92d1702c5735876b83cf7e8688314a69294ff8c40ee24905fc72f01

Observation 5b941e0d-4dcd-4cab-ade5-e2ec7ef6ca3a · outbound

This paper cites Llm-based multi-agent systems for software engineering: Literature review, vision, and the road ahead,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Llm-based multi-agent systems for software engineering: Literature review, vision, and the road ahead,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.204562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.785301Z digest=sha256:0b0bbc2f0b8d93a185660322146f4fcefbdc3eafe35b4e939c80ce020e254c43

Observation dd9554a4-7dff-4c04-9c3b-8ce2ecf791b5 · outbound

This paper cites Demystifying llm-based software engineering agents,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Demystifying llm-based software engineering agents,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.190809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.789370Z digest=sha256:2a4ab13547b2f8332409e541a475f5f988f1b9a4fb5ee0d51a668ebc18e61d4a

Observation ce0f2418-a4a7-4b4e-82a8-97aa0fd538cc · outbound

This paper cites Self-collaboration code generation via chatgpt,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Self-collaboration code generation via chatgpt,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.793365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.793365Z digest=sha256:0de9bd2387705850fade47d8456ad7b34eff1cd86de624a65b960198a9d4bb9e

Observation 9e1fd98c-f23d-4115-8467-c2e54db0ca76 · outbound

This paper cites MARE: Multi-Agents Collaboration Framework for Requirements Engineering.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks MARE: Multi-Agents Collaboration Framework for Requirements Engineering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.797921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.797921Z digest=sha256:d72d453c9c413037871dd76942e25c7cacb6f3bbf1db6932ce3e2d544eb77dd9

Observation b8e0a342-7230-4f28-8ffd-4977df7b8829 · outbound

This paper cites Requirements are all you need: From requirements to code with llms,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Requirements are all you need: From requirements to code with llms,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.802907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.802907Z digest=sha256:7ded50ac818cd5b593ca8c592e98c7436eccb18e286542cbfcd90b19a7039a2a

Observation 8aa8365b-eb61-46c8-b6ee-a451620c4b9c · outbound

This paper cites MarsCode Agent: AI-native Automated Bug Fixing.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks MarsCode Agent: AI-native Automated Bug Fixing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.806704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.806704Z digest=sha256:124991508347a0d345fdfa34b3b7ef4fad4dab51d936aeb4c2a78d45e2fc4278

Observation a663a997-2139-4c4d-9187-c37c3ddafd7d · outbound

This paper cites Cycle: Learning to self- refine the code generation,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Cycle: Learning to self- refine the code generation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.161275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.810956Z digest=sha256:952c4c3265a14ca44a5fb66786b3b745b363792ca3cf0c1ebc15c5a7ba3d4271

Observation 77cc3a4a-abcc-43ff-81c5-a79ed0233ee2 · outbound

This paper cites Self-planning code generation with large language models,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Self-planning code generation with large language models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.814865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.814865Z digest=sha256:44c5709bb02e44d6f7b02fb65b2701022590b2f14817c93eb94c7e0b509f5ad3

Observation eb548101-a0d1-40d0-ae6f-ac13b1f5dbb3 · outbound

This paper cites Swt-bench: Testing and validating real-world bug-fixes with code agents,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Swt-bench: Testing and validating real-world bug-fixes with code agents,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.139286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.818980Z digest=sha256:f6ffa05ea7209a94d1d68776359f69dbfde1041392c8e372307287f096c834c5

Observation 19c2dbed-31a0-4942-8cb4-d552127c0739 · outbound

This paper cites Make llm a testing expert: Bringing human-like interaction to mobile gui testing via functionality-aware decisions,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Make llm a testing expert: Bringing human-like interaction to mobile gui testing via functionality-aware decisions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.125785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.823042Z digest=sha256:fc0b9797983fda8c3aa5daa5dfc878964ea70af2b75cf3252baab41910090ef7

Observation ee73889a-f54e-4a17-98a3-e39bdf63c076 · outbound

This paper cites Chatunitest: A framework for llm-based test generation,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Chatunitest: A framework for llm-based test generation,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.826907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.826907Z digest=sha256:5ff6e5e9d05cfeba2868c1c7bde3e0642c3f231cd5452cac60228f0b227fe487

Observation 101b8ae6-0b90-46f8-8abb-0bd5abac6f65 · outbound

This paper cites Infiagent-dabench: evaluating agents on data analysis tasks,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Infiagent-dabench: evaluating agents on data analysis tasks,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.105708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.831028Z digest=sha256:e3fa1d2c27078e4dd702ffdc24d9bfcaac8d4ecf433a7a72fc05745471289e0d

Observation 1c8d51f8-b7b9-4cb0-8bcb-21fae86dae5c · outbound

This paper cites Super: Evaluating agents on setting up and executing tasks from research repositories,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Super: Evaluating agents on setting up and executing tasks from research repositories,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.090688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.834901Z digest=sha256:6dfccf29c4eaf49abca9da5c8427aa76101b55462e8e170c8b38242602d46dce

Observation 8035ecc3-05f8-4840-80d8-4e59ce525bda · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.838765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.838765Z digest=sha256:a924df84b3d0362f7cc7765db950d942cf51138ea14ad2cf0a893c8623d72a35

Observation 67b0c347-4043-43bd-93f1-c234ef6e3ed8 · outbound

This paper cites MetaGPT: Meta programming for a multi-agent collaborative framework,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks MetaGPT: Meta programming for a multi-agent collaborative framework,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.076511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.842643Z digest=sha256:efb031e1fd25fd21d6430350fe92a142b103bef63c737728ef5f0577fd69c9df

Observation f1841ab7-e1a5-4eef-847d-9777b13c8713 · outbound

This paper cites an unresolved cited work.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.847048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.847048Z digest=sha256:cb006cb73f7f596b780d6a3c6d5738ff561b173e36603bc4ea60ff943bb04d58

Observation aefe7158-e617-416e-b9d5-0b0c094368f4 · outbound

This paper cites Gpt-4o mini,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Gpt-4o mini,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.054404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.850949Z digest=sha256:2a6d5d5f6722eecf4721949f000a6b8e077d5d270cc933670dc17e8ac1804709

Observation bcbf1272-6d2c-4349-a632-fba873876dbe · outbound

This paper cites Multi- stage large language model pipelines can outperform gpt-4o in relevance assessment,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Multi- stage large language model pipelines can outperform gpt-4o in relevance assessment,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.040811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.855292Z digest=sha256:940718e9c8ed44056e04cf8a63f40093121fd844f0c7ba49e65d24c073c7eacc

Observation c05875ca-3556-4be6-84bc-910169701c21 · outbound

This paper cites React: Synergizing reasoning and acting in language models,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks React: Synergizing reasoning and acting in language models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.858815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.858815Z digest=sha256:bce824726535783df9e7b0a01030f67bd61f22a1be6384dbd7d086624aa83bae

Observation b248f53f-df27-4f4c-ada2-e3c60889f4e2 · outbound

This paper cites Reasoning with language model is planning with world model,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Reasoning with language model is planning with world model,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:52.018133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.862543Z digest=sha256:da0dc44f387d1929b7e0539dcfdf4efc46c320925c4032c1456265b20b542d1f

Observation dcb32af2-a2bf-4939-9385-17f973d52429 · outbound

This paper cites Automated program repair via conversation: Fixing 162 out of 337 bugs for $0.42 each using chatgpt,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Automated program repair via conversation: Fixing 162 out of 337 bugs for $0.42 each using chatgpt,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:51.866296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:19:51.866296Z digest=sha256:fcc4914f19d224a32e9654a79d5e0cd05fbf3cb34c160540b4ce59b057affe29

Observation aa93284e-e4c2-4e8a-af87-22e013d59cee · outbound

This paper cites Perfcodegen: Improving performance of llm generated code with execution feedback,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Perfcodegen: Improving performance of llm generated code with execution feedback,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:51.997652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.870060Z digest=sha256:0bfe24c8f03d8bed1a2ee3d4cd3f92372e4d0c9b0e8dd10db54972eb84fcbb2f

Observation 2a697131-4fe8-467d-8137-8d622e8c72e4 · outbound

This paper cites Self-edit: Fault-aware code editor for code generation,.

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Self-edit: Fault-aware code editor for code generation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:51.984781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:19:51.873664Z digest=sha256:8eb3071f718dca5ea2d7dbfe3348b3816bfd79c145c293a3f183d5569f7747b9

Pith citing papers

Observation d0b511dc-7e95-42fe-896e-5aaf32b7ae16 · inbound

Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes cites this paper.

Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:46:08.247213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T14:42:40.925632Z digest=sha256:dc6bf565612ae08e37848cd36b440ae98a5d9207000d9ea10e81df695d754b25

Observation 01ff125c-8a27-4781-b121-41c694677c4c · inbound

Profile-Then-Reason: Bounded Semantic Complexity for Tool-Augmented Language Agents cites this paper.

Profile-Then-Reason: Bounded Semantic Complexity for Tool-Augmented Language Agents Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:03:01.282736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T17:01:26.747812Z digest=sha256:18a2743175a4a805ff508bbbae613c0c316d2123565b33d4d25fe82f1e40d7cb

Observation add86010-c736-43ab-9f90-8bea8dd87bdc · inbound

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors cites this paper.

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:17.896521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T22:07:58.614654Z digest=sha256:c14f20c83b4b481020e50a7ce0829316eedd0cfeb3a8cdc4c7f01531c98ff005

Observation b9206f08-eb4b-4ef2-aa0c-c5c37f9572a0 · inbound

Inference-Time Budget Control for LLM Search Agents cites this paper.

Inference-Time Budget Control for LLM Search Agents Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:07.777341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T11:51:02.872129Z digest=sha256:b4b22fd8d3566e73ad20620d4eb01211b5ac4e566a92da84c3552d1a1e75d1c2

Observation f9c724bb-3eab-4df4-83d3-c6aa3bf6de69 · inbound

When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems cites this paper.

When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:26:38.045622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-25T04:25:26.710488Z digest=sha256:d858158da1aa620890fffc224b5bffab90f453f2dffd8de008400f5b7237b114

Observation 85e33e3a-a616-403b-80f7-eba747086055 · inbound

Failure as a Process: An Anatomy of CLI Coding Agent Trajectories cites this paper.

Failure as a Process: An Anatomy of CLI Coding Agent Trajectories Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T02:31:04.773248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:31:04.773248Z digest=sha256:5a6da7fe18074ef82ed208ae14f68dd7609db6282cbe99e6b84700236defd8b8

Observation 4a4d8dca-9e04-46be-bc1c-95eef9990c49 · inbound

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution cites this paper.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.058799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.058799Z digest=sha256:07c508d1202609d6b3821ec862ffdbbf439d654a19ffd3a048d235f174b34609