Pith. sign in

Paper Citation Record · LEDGER

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

As of 17 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 49 inbound Pith citation observations for arXiv:2505.00212.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00212 v3

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:52:52.515047Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:46:05.668170Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T11:24:54.802233Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eeb18f8a-e0ef-4baa-93f4-2acffdad1482 · outbound

This paper cites write newline.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.315414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.315414Z digest=sha256:113ebc1f180ae0fc126ba22025175eaebc05bde8a4c43a29adaab4e30ca70257

Observation 612b08d9-55f2-45ea-8380-917ed6650f4f · outbound

This paper cites GPT-4 Technical Report.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.321230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.321230Z digest=sha256:e2caca1d224fa0583003eb2c569863d48f6b5f8661dfc622eea8a3ce335df70b

Observation 5bcde0db-295e-4216-8e79-3844f4d43e01 · outbound

This paper cites Chateval: Towards better llm-based evaluators through multi-agent debate.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Chateval: Towards better llm-based evaluators through multi-agent debate

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.263639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.326732Z digest=sha256:7fe9e83d3154e2ba11a2cf0e0c459522ed88ea90a2dd7621b047b9a354b14fd1

Observation 6c4f2676-a02e-4899-897e-ab1efd31f5a3 · outbound

This paper cites an unresolved cited work.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:52:53.248940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.331073Z digest=sha256:77e5a6d52a789ceb004d3f184307bb22951132b6fb6cb88a2719a6cd85a7e160

Observation 4e9c293b-6212-4654-a1c7-ce0214a6744c · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Process Reinforcement through Implicit Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.335539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.335539Z digest=sha256:61886a6627df6716026b776aaaa440b184c09a2425437ebcfd3dd05c9cca05a6

Observation 434bb8f4-00d0-4874-a66a-05dfc1a1468a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.340472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.340472Z digest=sha256:2949a271a9e9fac495caf0c2afd2996b7ad74fd8df4458e3084232a849f37343

Observation 21c9123b-1996-418a-9ba7-9a3a98c2dfd9 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Mind2web: Towards a generalist agent for the web

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.234429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.345024Z digest=sha256:b009b28840d319d1a86bd124b9d66f7fd929a9040f4f0d79282b315bad909a8a

Observation 86d3c182-ed28-472f-a674-da328553c038 · outbound

This paper cites Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.349993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.349993Z digest=sha256:e9362f2238f009a0c6823b21d80f3e11a57b4b288af80fc5391ad81c6df75074

Observation 0578aa97-27e7-4016-87af-e7af5e83d630 · outbound

This paper cites Gptscore: Evaluate as you desire.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Gptscore: Evaluate as you desire

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.219902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.354606Z digest=sha256:8b568a6f3e418717b6995886cd5747255f8546f75aada376b4e3de28b6de9ab6

Observation aaf638b5-586b-4faa-9b55-ed3638351b34 · outbound

This paper cites SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.358888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.358888Z digest=sha256:71066b425626b2cf621df69ed27b7f887ad26e860afb769e0741c845e95e0341

Observation 9bd1be5d-f5ab-40ac-87bf-b0389080e709 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems A Survey on LLM-as-a-Judge

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.363628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.363628Z digest=sha256:fc7cb810a86f98ad59c264c17be4e1b18814d467496d1d6ed37e66df33d13413

Observation 47dbc562-0599-444a-a976-e082db0a0e05 · outbound

This paper cites an unresolved cited work.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:52:53.205473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.368191Z digest=sha256:23dda21742ffe39021cd27bdc7a63989b07a79d90c2b2866e268ce3c59adaec5

Observation 0fc95672-48d2-43f7-8149-805c8ddf91f9 · outbound

This paper cites Language model preference evaluation with multiple weak evaluators.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Language model preference evaluation with multiple weak evaluators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.372548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.372548Z digest=sha256:cd09c8b25ac09250bb0db64cf12a7529f3a4d46c05bbf549c32d4845b7e8018a

Observation 68f24689-dfff-42d7-b02d-14357cd5f136 · outbound

This paper cites and Chang, K.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems and Chang, K

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.191161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.376662Z digest=sha256:00d30786b72e30d81c9e885e9976a917dbaf8b59cbca741ed2636321a1fe2123

Observation 90069297-e888-44b4-83ea-f9370fb869b8 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Large Language Models Cannot Self-Correct Reasoning Yet

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.380916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.380916Z digest=sha256:6942f5a988aec2497f9c2c75e10fd33ba43eba5852921ba688da64d89e194128

Observation e107323d-9ca6-4c56-98d6-3b72754803a5 · outbound

This paper cites E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.176151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.385288Z digest=sha256:225d84fe259ebc9c1fbfb47cb3e3bdd78c155ac037a689d872e6ae0b5ff604cd

Observation 16c63f3b-9e16-4e04-8b96-f2b054407ec8 · outbound

This paper cites Camel: Communicative agents for" mind" exploration of large language model society.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Camel: Communicative agents for" mind" exploration of large language model society

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.161097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.389596Z digest=sha256:b7b1d9a1dd434dd2cf83794bc7a1c37d330892b8f3e6b0021ef7d18ac7087a48

Observation b4323b50-fe6b-4bd4-a884-2a2146dc8a3d · outbound

This paper cites an unresolved cited work.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:52:53.145767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.393725Z digest=sha256:ef9dd1b5e5754a12b5edac4e4380f7033abb703912a4234ec85226053eb848ad

Observation 79adb9b2-9bf1-4d8a-9282-a624da5df62f · outbound

This paper cites Let's verify step by step.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Let's verify step by step

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.131040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.398128Z digest=sha256:56c0eb5e4b674ed96b2f0bae5503290ef4aa233da0c0f31cf6fadd3b59fd74fe

Observation d9f0fce0-4341-4061-a273-2974b25bf652 · outbound

This paper cites G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.402348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.402348Z digest=sha256:2ecf095b1600388c92c55e6cb4e4fa6f99418ebb884aa90d9f89695b12c9ea02

Observation 9a9e113d-62bd-466f-94be-4960f86b3daa · outbound

This paper cites Gaia: a benchmark for general ai assistants.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Gaia: a benchmark for general ai assistants

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.116325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.406591Z digest=sha256:79fdd9119b70895052e4fc3e3354e4cd6bde85ecb6f62ae92c3eb9c3ca0a293c

Observation 5ec88ca0-9bdc-46fe-adda-e41600265fcc · outbound

This paper cites W., and Rainforth, T.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems W., and Rainforth, T

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.099840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.410781Z digest=sha256:adece13de8f10adc6fac815152d883714c70b51d439c9d2b83b681d76dce75d1

Observation 4b4f0114-a585-42ac-a365-4d527526664c · outbound

This paper cites Needle in the Haystack for Memory Based Large Language Models.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Needle in the Haystack for Memory Based Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.415074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.415074Z digest=sha256:acfd8635ae5889a6a654e3785e84ad0b923b4df4dc4a6dfca4c21394df292da2

Observation c173ac21-371d-4361-8f91-533f9933c469 · outbound

This paper cites Evaluation thesaurus.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Evaluation thesaurus

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.085514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.419455Z digest=sha256:df5668fe6899bcb0fa27afc24991fa5dba9d1fc98d6fa788534836cd2d585b58

Observation 03a0d808-5fe6-4f0b-ae90-1e7b8c8cac9d · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Reflexion: Language agents with verbal reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.069810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.423594Z digest=sha256:011344ea1edba1217e79cb9aafedb1202ec86b463c97a26299ebaacbd3e5dcc5

Observation c7fad762-37a0-41ef-87e1-b5d486f71d07 · outbound

This paper cites Adaptive In-conversation Team Building for Language Model Agents.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Adaptive In-conversation Team Building for Language Model Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.427782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.427782Z digest=sha256:97095a329ef99a1fd2b60d10849dea276e78f27220707070d35dccfba4e9abc5

Observation 4f647eab-1785-4289-83e3-ecc0419328eb · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.432244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.432244Z digest=sha256:aec4b66259034417d1b678b28deb9d6413f9fdbc74e91d49714fb02664713c86

Observation 72c61dc7-ef0f-4d48-853d-115de6b50840 · outbound

This paper cites an unresolved cited work.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:52:53.055296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.436745Z digest=sha256:a66c175a8723fab64d052d0a5ab3111e09d77cc2b877aa0f8e3e0f48313e7a26

Observation caafc080-b63f-4d76-854d-89e4b1855026 · outbound

This paper cites A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.441020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.441020Z digest=sha256:0c3f87705d68be297a52b72ede367ae2cc01eba65482eba5dafc854de5054ab3

Observation e717bc1c-aa6e-4dd8-9c4b-4a45ce41771f · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.445451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.445451Z digest=sha256:ee737ebc0c52e41d3104fc8a858d945a7df19c15a23ce0a8bfd3db10a70c3730

Observation 6b2f2f81-e2de-4d29-9ce3-e32b09deab70 · outbound

This paper cites V., Zhou, D., et al.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems V., Zhou, D., et al

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.040009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.449941Z digest=sha256:e001f252f734fb12cf1d58afd4ab35fd089ebd94203e8ddb9889dfa25aac33da

Observation 4a528a6b-da18-4817-998c-0bcfc1971114 · outbound

This paper cites Autogen: Enabling next-gen llm applications via multi-agent conversation framework.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Autogen: Enabling next-gen llm applications via multi-agent conversation framework

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.024879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.454323Z digest=sha256:6815c2cf01216a90ace7373777dd4dfe706656eb7a7acfab1e293de9f7bdd7d3

Observation 4e26073e-1684-4554-962b-a529dc555f27 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.458239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.458239Z digest=sha256:e25a1cf8a68afda914b1d1763e94bd55b0af67a11da9a72f23e7a712fa128dc8

Observation 1827b626-9416-438b-a685-7787c2cd9a49 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems React: Synergizing reasoning and acting in language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:53.009732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.462474Z digest=sha256:e6bf8c9e50a2d2640cfd849bca5e11df7050413b35be5360c72e921da30baa33

Observation 1667d99c-9a08-4a98-8f5c-2efed1ad85aa · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Tree of thoughts: Deliberate problem solving with large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:52.992320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.466791Z digest=sha256:868f4c6c6aba05f7587c663d150a30574ac9bbafc9cc6ae98cc8c2157c24fdc9

Observation 8623e744-fe0f-4e4e-a591-3e35c3c5c8e2 · outbound

This paper cites J., Malaviya, C., Bogin, B., Press, O., and Berant, J.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems J., Malaviya, C., Bogin, B., Press, O., and Berant, J

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:52.977650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.470974Z digest=sha256:3a6a0768714c040d89bf9c0fc7532cedb27e8a38fa2f0a109e6b9c20e1accf3c

Observation 3ecd998e-0387-4c81-88d3-d2505118e300 · outbound

This paper cites Hyper-Parameter Optimization: A Review of Algorithms and Applications.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Hyper-Parameter Optimization: A Review of Algorithms and Applications

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.475316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.475316Z digest=sha256:f42017f16f6ccbb6c5baf23baf87264c57756d6ecec12ea2a4b707f9ddc8a291

Observation e8792fdf-4b1b-4289-8c47-b9d21335793b · outbound

This paper cites EcoAssistant: Using LLM Assistant More Affordably and Accurately.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems EcoAssistant: Using LLM Assistant More Affordably and Accurately

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.479845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.479845Z digest=sha256:ab09bc4ab05dfa333008bb49fd6cd620c30ff8b683ca19d6b105011bfe151005

Observation 16062771-5e72-4f79-8e56-8ec704c3b6d7 · outbound

This paper cites Targeted hyperparameter optimization with lexicographic preferences over multiple objectives.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Targeted hyperparameter optimization with lexicographic preferences over multiple objectives

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:52.963215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.484393Z digest=sha256:af2ff3dac3c9000091584190a86c87bd206264b9641791a2ca04443e13976114

Observation 6b260e53-dede-40c1-bbc6-45eed70d9a6a · outbound

This paper cites EcoAct: Economic Agent Determines When to Register What Action.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems EcoAct: Economic Agent Determines When to Register What Action

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.488577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.488577Z digest=sha256:a51bd4fea3cb599c8cb80c20d3ee3323f2c18f9b8d4a03f107a585fbfcfa62ce

Observation e17d1601-1c5c-48c2-891b-836cb6fe5ba1 · outbound

This paper cites Offline training of language model agents with functions as learnable weights.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Offline training of language model agents with functions as learnable weights

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:52.948781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.492901Z digest=sha256:9542a1fb71d12d0940070341a79439107e2d4177ebfa083b8325673bfe722ee4

Observation 55fba8a4-b624-42d2-84ac-ee7f9f39423a · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.497186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.497186Z digest=sha256:0aaa0b12b1df1ebfe5b9cb7883ea8eeb067581354052d485ce83efd973e16d4d

Observation 7af62ab7-d0ab-40e0-b01d-01ccbb1c483e · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.501689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.501689Z digest=sha256:1dcda8fefde437c94abdcbe9a9365bfba5777a31175cf38f5300575422afa45c

Observation 61ba5cd1-7e64-4f95-9a46-891fcb6b3f80 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:52.933992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:52:52.506300Z digest=sha256:074b3d389f56e60c7c5d753eb08fd12bf1660b787135a2904767d5fb6db63a6a

Observation 91966d66-931c-4012-a1c1-ca91d6e7fac4 · outbound

This paper cites A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.510706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.510706Z digest=sha256:bb2134139323ab3796715cbef917c84aa4bd284da18a3f26e27f41645613db09

Observation 29b727f4-68d3-4b07-8e77-bbe15fe8f17c · outbound

This paper cites Agent-as-a-Judge: Evaluate Agents with Agents.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Agent-as-a-Judge: Evaluate Agents with Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.515047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.515047Z digest=sha256:bb37635d0c6de934ee74e4f0c008ed7d69a075a82c696b5fc7b27f1d01ad6800

Pith citing papers

Observation 3d8cb982-91c4-47b9-a72c-c8dbaaee7738 · inbound

Why Do Multi-Agent LLM Systems Fail? cites this paper.

Why Do Multi-Agent LLM Systems Fail? Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:58.410764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:16d67ac8c5afa9187bffa3b4798fc5d83e09a2ceb3a48beb62c7d5e5b5f70ae4

Observation eb8fb925-58c7-4aed-840a-5c3e0cb7fcf2 · inbound

Divide, Optimize, Merge: Fine-Grained LLM Agent Optimization at Scale cites this paper.

Divide, Optimize, Merge: Fine-Grained LLM Agent Optimization at Scale Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T23:46:05.668170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:46:05.668170Z digest=sha256:fef594443243da0c8b4dedfd7596a90d7746f242a1d8287e11b58af563539dc5

Observation e14e2a30-71dd-469b-ad7a-be5ae7bc46c1 · inbound

Single-agent or Multi-agent Systems? Why Not Both? cites this paper.

Single-agent or Multi-agent Systems? Why Not Both? Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:05.109462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:05.109462Z digest=sha256:f0f8dfd05fcffa3d28acf8aca1e6550df514321ced2e18d2007dcefa8811a5f5

Observation b910eb31-f4f4-4c9a-bbb4-11084b0e3987 · inbound

MermaidFlow: Redefining Agentic Workflow Generation via Safety-Constrained Evolutionary Programming cites this paper.

MermaidFlow: Redefining Agentic Workflow Generation via Safety-Constrained Evolutionary Programming Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:09.849760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:09.849760Z digest=sha256:07c53e6e66de501f2cc94b7dd897ca5d608919c245ad0086fcf2ae14ad4d5959

Observation f065c39b-f19b-499f-91e6-fe61124a1dbd · inbound

Why do AI agents communicate in human language? cites this paper.

Why do AI agents communicate in human language? Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:21:49.544808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:21:49.544808Z digest=sha256:5919d998c46813ac3299ab02b6ad0d233cf12677c5e3f13bbd8ecad5fbaef8e2

Observation bf08a241-c5e3-4f44-959d-52709e8348d3 · inbound

From Virtual Agents to Robot Teams: A Multi-Robot Framework Evaluation in High-Stakes Healthcare Context cites this paper.

From Virtual Agents to Robot Teams: A Multi-Robot Framework Evaluation in High-Stakes Healthcare Context Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.392106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.392106Z digest=sha256:23eeb97676fb75dba65c2162e12b40a84855fe12dbc9a0ba8cd64a62b16c6702

Observation 4f8666fa-1932-439f-8757-23d98e7b827b · inbound

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems cites this paper.

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:36.139157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:42:36.139157Z digest=sha256:4b6f948d0f9417a32f7510635adbee72425afc8cefb21167d357e1367b420a6e

Observation a7f712c4-12aa-415c-b9a9-875897882311 · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:06:17.780496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T11:02:55.529271Z digest=sha256:f87aa2de08573555bd91970a73710ab85a7dfe9266272b663f948572681e1046

Observation 0af9bf48-d95d-45eb-9b86-0e2ea41017a1 · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T12:44:13.174346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:44:13.174346Z digest=sha256:57aa3f6de1b69a6b7afc0be5433c4c6c5ea89778337882b0e8222bf5e92de410

Observation 5c2dbfb6-b98f-4d43-b06a-937a33b0ad92 · inbound

CAM: A Causality-based Analysis Framework for Multi-Agent Code Generation Systems cites this paper.

CAM: A Causality-based Analysis Framework for Multi-Agent Code Generation Systems Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T05:31:41.146745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:31:41.146745Z digest=sha256:0ad02fe078249415eca051801c0bb81b18f1aebe27b319af0bdf2fb11d8bed5a

Observation f9d11f55-f1a9-4122-ae70-370c5ce14490 · inbound

AXE: Grey-Box Exploitability Confirmation for Localized Vulnerability Reports cites this paper.

AXE: Grey-Box Exploitability Confirmation for Localized Vulnerability Reports Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T23:18:01.253850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:18:01.253850Z digest=sha256:def5075a339afff7b39b90128a8c1e1f4e8213e7966a3e64980925954838fb22

Observation e7c5c23e-e7f1-4887-9288-f5711a0115de · inbound

From Spark to Fire: Modeling and Mitigating Error Cascades in LLM-Based Multi-Agent Collaboration cites this paper.

From Spark to Fire: Modeling and Mitigating Error Cascades in LLM-Based Multi-Agent Collaboration Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:40:10.646859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T16:36:43.330447Z digest=sha256:327215f3b4e18ba8553fe743a5c7a6417409c6289f59e005b47db616c0e5711e

Observation 896e7db2-07d7-4352-afcb-4447582ebb29 · inbound

Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes cites this paper.

Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:46:08.225092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T14:42:40.925632Z digest=sha256:88381fba8e317a15a704211a6f86b971a0f020ae69b6e38d594864c46fa41fe4

Observation e8cb84db-0ff8-42c0-9fc8-df116f2c05b8 · inbound

Trace-Level Analysis of Information Contamination in Multi-Agent Systems cites this paper.

Trace-Level Analysis of Information Contamination in Multi-Agent Systems Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:41:27.091767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T09:41:15.185578Z digest=sha256:fd6c3ae30de2f423fa73b23bd80d52f4113544e70d931bb1402f0834a7c43e8b

Observation 25e214e2-46f2-4a84-95f0-a99b7c877398 · inbound

Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge cites this paper.

Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 32

Resolution
malformed identifier
arxiv_id, observed 2026-05-12T07:51:41.638609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T01:47:30.987626Z digest=sha256:904de80bd4411d1bb3dc123d30ea9aaa8f02db75831ac89a9272b3f85b491724

Observation 9a3a72c5-bcde-4688-9590-5a720aaf206a · inbound

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems cites this paper.

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:27.629042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T01:19:49.062330Z digest=sha256:76dacdbed25dcd3536b82ffcd01bfc35fac269751da588ebb27f668eddcf5b7d

Observation 5717c6b2-33fb-4f61-852b-41c93bc36cd3 · inbound

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems cites this paper.

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:25:03.791076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T05:24:54.265411Z digest=sha256:fc8ca753566eb01e3233dcaa3dfcff808336ffb5365c0eb6f72d188e7d7ff50f

Observation e2f4fe93-550e-41bc-a68a-928bbdba5c73 · inbound

PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement cites this paper.

PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:22:06.193495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T02:22:04.669760Z digest=sha256:84bf992b091099731b93968c642a064022cb24aced7980ffd918f33bd90f4103

Observation e069b85e-9e4f-4e97-b71d-99d68955af95 · inbound

Do Coding Agents Understand Least-Privilege Authorization? cites this paper.

Do Coding Agents Understand Least-Privilege Authorization? Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:40.010720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T16:34:14.379419Z digest=sha256:19189638b1697f6aadec0946522e586b2ab146f48068851aae321c97cbb03f7e

Observation 73fb5a1f-9f6e-4b2f-bc9f-3f4736e37d6e · inbound

STAR: A Stage-attributed Triage and Repair framework for RCA Agents in Microservices cites this paper.

STAR: A Stage-attributed Triage and Repair framework for RCA Agents in Microservices Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:23:40.788771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T19:23:19.135749Z digest=sha256:1a2c7e770236c4fe8c63c01821e3fe790835f05fbbc186e19f616791122786c1

Observation 046f5e26-b0ee-420c-8490-095aef43beb1 · inbound

AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question Answering cites this paper.

AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question Answering Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T13:48:19.921035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T13:43:42.097447Z digest=sha256:03fac47e664115cd97e6f627963bfc6068f5a9aff12c1023d52507bdb4c9a3e8

Observation 69456704-5e79-432c-9be2-0dd4d4800e62 · inbound

CASPIAN: Online Detection and Attribution of Cascade Attacks in LLM Multi-Agent Systems via Cross-Channel Causal Monitoring cites this paper.

CASPIAN: Online Detection and Attribution of Cascade Attacks in LLM Multi-Agent Systems via Cross-Channel Causal Monitoring Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T03:13:00.421602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T03:11:36.534055Z digest=sha256:5043f5516f06d99b284ec47c438abaf1fd66f47e7903dcc3489c89691da4fb22

Observation ca20721f-d306-474e-9568-73a7c654c5fd · inbound

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents cites this paper.

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:23:57.550468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T04:20:43.780849Z digest=sha256:0d6f5df568a58c3eafe30b9807caab434ad91fc42ca27223f6281f228b6c2316

Observation 9a0b3bf1-ac63-4c48-9802-8d833da036b0 · inbound

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents cites this paper.

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:51:21.867557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T09:46:26.124683Z digest=sha256:65d1c842b990d1b8a4570e4e6d9e1386bfa513aaa53746ca3b324e4eb3ca3af7

Observation 19200b64-ac8b-4aaa-8dc5-ab7da96f0200 · inbound

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents cites this paper.

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.667709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T17:31:25.538977Z digest=sha256:269e9815eb6b79fa917ae6bd43350d15a592f1b1f518385b41c63610e105671f

Observation 229681f8-3d11-4643-9978-952c05f7d6ac · inbound

When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems cites this paper.

When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 139

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:26:37.914057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-25T04:25:26.710488Z digest=sha256:58fdf531fff6292618ba62ba61435043a509e4a8e04d784d3861579b8a630244

Observation c595fe95-14a9-49cc-a60e-a884ca192363 · inbound

Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems cites this paper.

Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:13:59.460712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T21:13:03.346422Z digest=sha256:0072edbdcdcbd63b8b83549c544da9306ccc058e6bdeaba4834422ed919cf4ec

Observation 3bc51b50-d00e-4825-97d5-c7d9535c231d · inbound

FALAT: Tracing Failures in LLM Agent Trajectories via Dependency-Guided Search cites this paper.

FALAT: Tracing Failures in LLM Agent Trajectories via Dependency-Guided Search Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T18:42:29.585718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T18:38:56.787942Z digest=sha256:8ce12d082a53190adf3cacb33d453d865a63be5be1e1f6c38f7279f78c5f3957

Observation 962780bd-e35e-4e07-b705-746c026de1e3 · inbound

POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems cites this paper.

POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:19.973490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T14:44:21.487169Z digest=sha256:88f7a1e2e1600eb297f0ee8c5577ab95612e87e606d33ea0a47c108c368b5200

Observation 99b43385-5d72-4512-b4b7-b088278826c5 · inbound

StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent Systems cites this paper.

StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent Systems Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:16:33.986357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T10:12:35.956616Z digest=sha256:8f3a6c549bf912e86ccc02170b3950140008aa2da7d32279c29bf995ed7ace69

Observation c0e26903-a502-42c2-ae6c-c7b4c3ce95a0 · inbound

Strained Coherence: A Pre-Failure Signal in Coding Agent Execution Trajectories cites this paper.

Strained Coherence: A Pre-Failure Signal in Coding Agent Execution Trajectories Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:57:09.805473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T22:16:45.910467Z digest=sha256:7b4dda9e13c70e909b03b9d17905bb266daf8967defbdd49fa588eff98c3f452

Observation 8a1eb7ce-983f-44ea-9e18-291f8fce76a7 · inbound

Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures cites this paper.

Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:24.302055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T19:54:26.300961Z digest=sha256:d2a179b7c65428a936b5074b762df7dd4d1273a53d3e27031ef7740fc474e4a4

Observation 67700a4a-a9b7-43bd-adf5-f899e65c50f4 · inbound

GRADE: Graph Representation of LLM Agent Dependency and Execution cites this paper.

GRADE: Graph Representation of LLM Agent Dependency and Execution Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.499687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T09:44:37.914985Z digest=sha256:6d40dafa91a8d7d4a5f6552cdf43fde7c141069a4680d8c86117fc9c5c5e9de4

Observation 85b6f485-41a9-4589-890d-3b317fe7a7b6 · inbound

MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems? cites this paper.

MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems? Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:45.584925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T09:15:50.722199Z digest=sha256:816b2368dcda5e03d4b40c48982e8928927846502755dca11f4867fa83b73aae

Observation 313f4da9-cd96-4e70-90a6-4caa787e36fe · inbound

Delayed Verification Destabilizes Multi-Agent LLM Belief: Instability Thresholds and Optimal Corrector Placement cites this paper.

Delayed Verification Destabilizes Multi-Agent LLM Belief: Instability Thresholds and Optimal Corrector Placement Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:55:58.656482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T01:25:53.612641Z digest=sha256:2dba9d9ac6b2604071e51dd67cdf23107b5be7b04cfe956550ba60275505616d

Observation b8b77c80-58ac-4253-b5c9-004d533d1da4 · inbound

Diagnosis-Driven Automatic Repair for Agentic Workflow via Symbolic Inference cites this paper.

Diagnosis-Driven Automatic Repair for Agentic Workflow via Symbolic Inference Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T06:24:52.086585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:24:52.086585Z digest=sha256:ae4f8e90febc537cdd5ba75b91b1a6b4116f00c6c9ecb69a8a27cc72165f07c6

Observation df89a4a6-4360-450e-9096-25d5a752cd14 · inbound

AgentTether: Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operation cites this paper.

AgentTether: Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operation Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T11:24:54.805914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T11:24:43.002769Z digest=sha256:08b65b90f970743bae6e7657d561a8ace5838694f5cee0a88e6cdad8ee1d4e46

Observation a6b426ec-ee9e-40f7-943e-5c125a4e55af · inbound

Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity cites this paper.

Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T04:32:36.973681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T04:32:36.973681Z digest=sha256:cb785ee91e1ca997389d5c24f444d160a51ebf5dcef8777a306b81ddb52efc10

Observation 01477f78-7530-4dde-b7eb-922974ec845e · inbound

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak cites this paper.

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T06:30:56.999697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:30:56.999697Z digest=sha256:c04cb019b01f9512166ad19a9a5d054233145a67c270914878a173dc21d43ae0

Observation ab51730b-cf8f-44f8-970a-998dc40932a5 · inbound

Fantastic Adaptive Taxonomies and How to Use Them cites this paper.

Fantastic Adaptive Taxonomies and How to Use Them Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T21:14:43.971973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:14:43.971973Z digest=sha256:79dbd8fe0e8c2af6eccac0bdddafb3bf301499fde4ea837ba2312fde9a113174

Observation 67f04fec-d243-4f66-b1aa-d54894df2bb3 · inbound

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures cites this paper.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T00:25:05.394072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:25:05.394072Z digest=sha256:f4deec0cbdc33c64a35596b5630d3bf1bbaec09a5350bc741c9a5fec95eb06d8

Observation fc8b0c28-f92c-4da6-97cf-44accdd097b3 · inbound

AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair cites this paper.

AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:21.493741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:21.493741Z digest=sha256:f89730df6ce2b23efa557c9367b6300af18c7e5936bbe224602662c3ceaa5e86

Observation 66dc7322-7689-4529-a48b-97425fc0ecc7 · inbound

Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucinations cites this paper.

Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucinations Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:23:28.627046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:23:28.627046Z digest=sha256:f13eb89a9ff819434cfcdfbe1c82807445166f8667cd64d0d0060965194d4fd9

Observation ddc7c400-5ab8-470a-b18a-c945d00cd642 · inbound

Real-Time Detection and Repair of LLM Agent Failures cites this paper.

Real-Time Detection and Repair of LLM Agent Failures Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T06:54:58.850261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:54:58.850261Z digest=sha256:961af5ca5196c72452f784fe23386b1337a069276b3ccd1925b49a0580899e19

Observation 3f7cdcfb-68ca-4dbf-9693-bebea510f99b · inbound

OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality cites this paper.

OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:25.086245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:25.086245Z digest=sha256:e13db56978d4ebe2276115e3efd7023d4efc5c96159107ef656dcef677ad9e6c

Observation 4c0b8961-d3e1-400c-97f1-ef60d7ad02dc · inbound

TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure cites this paper.

TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:11.176776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:11.176776Z digest=sha256:924202d01f779f62b668fb7bca7ed63d11775b0bb3363845a7371c7dcb0c0917

Observation 77fceb62-7aef-4cec-8979-e48bd6b371ae · inbound

RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough cites this paper.

RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:37:48.064617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:37:48.064617Z digest=sha256:2cf475c29ca17beb543fd37cfefe426a757d7c261aeace46e3db9d594b149d4d

Observation 0e58611c-65fe-4701-84f4-5c94e6ca5366 · inbound

Tangent: An Empirical Study of Testing Practices for LLM-Based Agent Applications cites this paper.

Tangent: An Empirical Study of Testing Practices for LLM-Based Agent Applications Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-14T04:39:16.661890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:39:16.661890Z digest=sha256:3f032da38e25e7da13934edde2669aa6fded1ac2d9b3a6c6e257d7df7272f729

Observation c22278dd-6d74-44fc-994c-fc5b120df8f0 · inbound

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations cites this paper.

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:01.368996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:01.368996Z digest=sha256:78d94b95f2bd4f2bae573244da117ed44742068b5728e51065d5ee242b49d876