Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:19:51.873664Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 7 inbound Pith citation observations for arXiv:2508.13143.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:19:51.873664Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T10:20:42.058799Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-25T04:26:38.042285Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0b859add-951e-47b5-b3ed-7250b0ba4d21 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks TaskWeaver: A Code-First Agent Framework
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132a18b8-c8d5-446d-8c7e-351c836e48fd · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Autogen: Enabling next-gen LLM applications via multi- agent conversations,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation df860d48-b452-4c3b-b986-be57d8f0f674 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks CodeAgent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0f70a63c-72a9-4a61-b1e6-faecf997c16d · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Executable code actions elicit better llm agents,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation afcf02cd-bb57-4233-9f3c-22867bfca568 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Solving challenging math word problems using gpt-4 code interpreter with code-based self-verification,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 39966860-d1b3-41e3-8f0c-2b3f2a872f0f · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 486df829-a9a7-4ac3-86c5-19c55f4e3f19 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Data Dialogue with ChatGPT: Using Code Interpreter to Simulate and Analyse Experimental Data
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d996104-93ca-42f6-b079-926f71cb59ed · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Webpilot: A versatile and autonomous multi-agent system for web task execution with strategic exploration,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 14af42c9-f294-4aa1-831f-833ed575c4d7 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Redcode: Risky code execution and generation benchmark for code agents,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a87fecc5-a8d4-48ce-9b1d-9ea05a7f7655 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Large Language Model-Based Agents for Software Engineering: A Survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b941e0d-4dcd-4cab-ade5-e2ec7ef6ca3a · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Llm-based multi-agent systems for software engineering: Literature review, vision, and the road ahead,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dd9554a4-7dff-4c04-9c3b-8ce2ecf791b5 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Demystifying llm-based software engineering agents,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ce0f2418-a4a7-4b4e-82a8-97aa0fd538cc · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Self-collaboration code generation via chatgpt,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e1fd98c-f23d-4115-8467-c2e54db0ca76 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks MARE: Multi-Agents Collaboration Framework for Requirements Engineering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e0a342-7230-4f28-8ffd-4977df7b8829 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Requirements are all you need: From requirements to code with llms,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aa8365b-eb61-46c8-b6ee-a451620c4b9c · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks MarsCode Agent: AI-native Automated Bug Fixing
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a663a997-2139-4c4d-9187-c37c3ddafd7d · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Cycle: Learning to self- refine the code generation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 77cc3a4a-abcc-43ff-81c5-a79ed0233ee2 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Self-planning code generation with large language models,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb548101-a0d1-40d0-ae6f-ac13b1f5dbb3 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Swt-bench: Testing and validating real-world bug-fixes with code agents,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 19c2dbed-31a0-4942-8cb4-d552127c0739 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Make llm a testing expert: Bringing human-like interaction to mobile gui testing via functionality-aware decisions,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ee73889a-f54e-4a17-98a3-e39bdf63c076 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Chatunitest: A framework for llm-based test generation,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 101b8ae6-0b90-46f8-8abb-0bd5abac6f65 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Infiagent-dabench: evaluating agents on data analysis tasks,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1c8d51f8-b7b9-4cb0-8bcb-21fae86dae5c · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Super: Evaluating agents on setting up and executing tasks from research repositories,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8035ecc3-05f8-4840-80d8-4e59ce525bda · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b0c347-4043-43bd-93f1-c234ef6e3ed8 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks MetaGPT: Meta programming for a multi-agent collaborative framework,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f1841ab7-e1a5-4eef-847d-9777b13c8713 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aefe7158-e617-416e-b9d5-0b0c094368f4 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Gpt-4o mini,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bcbf1272-6d2c-4349-a632-fba873876dbe · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Multi- stage large language model pipelines can outperform gpt-4o in relevance assessment,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c05875ca-3556-4be6-84bc-910169701c21 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks React: Synergizing reasoning and acting in language models,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b248f53f-df27-4f4c-ada2-e3c60889f4e2 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Reasoning with language model is planning with world model,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dcb32af2-a2bf-4939-9385-17f973d52429 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Automated program repair via conversation: Fixing 162 out of 337 bugs for $0.42 each using chatgpt,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa93284e-e4c2-4e8a-af87-22e013d59cee · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Perfcodegen: Improving performance of llm generated code with execution feedback,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2a697131-4fe8-467d-8137-8d622e8c72e4 · outbound
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks Self-edit: Fault-aware code editor for code generation,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d0b511dc-7e95-42fe-896e-5aaf32b7ae16 · inbound
Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 01ff125c-8a27-4781-b121-41c694677c4c · inbound
Profile-Then-Reason: Bounded Semantic Complexity for Tool-Augmented Language Agents Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation add86010-c736-43ab-9f90-8bea8dd87bdc · inbound
When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b9206f08-eb4b-4ef2-aa0c-c5c37f9572a0 · inbound
Inference-Time Budget Control for LLM Search Agents Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f9c724bb-3eab-4df4-83d3-c6aa3bf6de69 · inbound
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 85e33e3a-a616-403b-80f7-eba747086055 · inbound
Failure as a Process: An Anatomy of CLI Coding Agent Trajectories Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4d8dca-9e04-46be-bc1c-95eef9990c49 · inbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.