Pith. sign in

Paper Citation Record · LEDGER

Evaluation and Benchmarking of LLM Agents: A Survey

As of 19 August 2026, this Paper Citation Record lists 100 of 143 outbound references and 6 inbound Pith citation observations for arXiv:2507.21504.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21504 v1

Coverage vector

measured 100 of 143 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:44:21.813422Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T01:59:52.218232Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:27:56.022361Z

Reference resolution

100 of 143 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved95
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01cdcaa9-3927-49bf-beb2-97395950eaba · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.482897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.482897Z digest=sha256:440f9a6f3c9a5abf3f1cf1a9138a9ee252393fef2fc70a4c481e8d1adefd9199

Observation 99349306-a237-44b4-b97b-f48a3ee2428a · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.486623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.486623Z digest=sha256:8847833a707c650a84c5d0c56b2205c274beaac6aaf137b668febbc2ad634a26

Observation d6b75862-6f1c-4a82-b403-cf6c41ea3fb6 · outbound

This paper cites 2024.Inspect AI: Framework for Large Language Model Evaluations.

Evaluation and Benchmarking of LLM Agents: A Survey 2024.Inspect AI: Framework for Large Language Model Evaluations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.489909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.489909Z digest=sha256:2b6e9ecf2e5d813c5418270e2df6d8e8303da65db3e6ccaa55a52a0fd44f109c

Observation 3acd0a5f-50e6-42ab-b9c6-5501b546d41d · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.493073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.493073Z digest=sha256:d09063f1b6ae9d8a11072c03aef1cf3ac7a252d441d578537a4818b55165ef40

Observation a0183b9e-8c90-4431-93d2-6f8bc626cc76 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.495973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.495973Z digest=sha256:a2bbdd99703344e180b03dc8a266d167fe4a53d0f6e1a638cf6c2193962d7ef9

Observation 3e084299-b3f2-4353-8702-1ac4c1ff8a31 · outbound

This paper cites Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs.

Evaluation and Benchmarking of LLM Agents: A Survey Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.502718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.502718Z digest=sha256:c29a1fe1da1277d7eade9f0c89c8bc9d0217cec050a0b549a35fe7875735e163

Observation f046b40a-e071-4cab-b409-b5d75cd66385 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Evaluation and Benchmarking of LLM Agents: A Survey ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.506012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.506012Z digest=sha256:2053c678765e684f0e822eb4e6f1e456383e18c3eb89f95c665542e28cdf4ff5

Observation 6b32d450-e913-47cb-8d89-384fcf71762e · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.509324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.509324Z digest=sha256:93ce4b18a2947c044b4b053596b712548b619cec8fad204cb2beba69872410fd

Observation e0ce98c8-6a51-450f-a48f-dfd9539403e7 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.512128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.512128Z digest=sha256:8a6e88d054522e46583d85669546b7cd90ea123400db1f018b5df0a702e589da

Observation a5fc4b80-d188-49ad-920f-98e5b6681c0a · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.517960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.517960Z digest=sha256:6e8dbd9abe61427d346e2f664f27bb1ab3bb4e9787443531b1796fb38404d3e8

Observation 2094d729-a247-4398-969c-1bfd1da0a94a · outbound

This paper cites ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery.

Evaluation and Benchmarking of LLM Agents: A Survey ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.520560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.520560Z digest=sha256:d7b7d68ab9a7ee754a2b6fc0e428fb97eee15dbc97718fac8c3ea93277308bdf

Observation 724229a0-0d2e-477b-bc0a-f9fcc9135313 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.523530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.523530Z digest=sha256:4000a66a2ad26cb5b8780c371ea7cb203d325026a91b412298bb55bb78fd8193

Observation 33451500-2922-4100-8d9f-25563c283b73 · outbound

This paper cites AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases.

Evaluation and Benchmarking of LLM Agents: A Survey AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.530025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.530025Z digest=sha256:5ddf2d5ac15545913f6620155a147ab546b3c9c39ca3305bbbe2d3ff72e588e1

Observation ae999c9c-d449-48de-a343-0099d907a1c3 · outbound

This paper cites The BrowserGym Ecosystem for Web Agent Research.

Evaluation and Benchmarking of LLM Agents: A Survey The BrowserGym Ecosystem for Web Agent Research

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.532807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.532807Z digest=sha256:be322d57aea63a2268f5ddf3ed8655df33dff1bda1ee8ae584ff562c9e03bbb6

Observation af3d1478-45bd-4d65-b7e9-7e7899f8da05 · outbound

This paper cites T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step.

Evaluation and Benchmarking of LLM Agents: A Survey T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.526623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.526623Z digest=sha256:796a4611307f4265cd8bee6330469c8b19821b7e122d7ff56b7975e22725e7f8

Observation e3010720-da73-41a1-9bf6-9021a3d30020 · outbound

This paper cites Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI.

Evaluation and Benchmarking of LLM Agents: A Survey Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.538622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.538622Z digest=sha256:1d86ac8c707e3e34659409da8827f0a69ab0027e5823eb65f52c1a7dd2f24b49

Observation 25d1dbed-283a-4454-84b1-19d51b7eb010 · outbound

This paper cites AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.541517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.541517Z digest=sha256:d0fa457fcd0f4ff469769d57ca6052e688d6d1ab92c7d5bdb9b6bd459e74376d

Observation fd5e71eb-3fd5-4fc8-bee5-f08a2d330abe · outbound

This paper cites GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.535787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.535787Z digest=sha256:9278e5472e2f091dacbb65db4c1e43be11cb6fc418b8728d3da24be0587026da

Observation f830591c-6c6c-4755-b8b5-a1772cfe8ebc · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.547500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.547500Z digest=sha256:2e56a9da8d23105cd62a6b848d3dbe2cfce3c68b1259d83a3d6a4ce5145ae28c

Observation 93374fe1-760a-482a-98e8-9082b5f5cfd0 · outbound

This paper cites Agent AI: Surveying the Horizons of Multimodal Interaction.

Evaluation and Benchmarking of LLM Agents: A Survey Agent AI: Surveying the Horizons of Multimodal Interaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.553434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.553434Z digest=sha256:4616bd5d712ab0525374ea548a8b4bac33087b637ecb21b55bd5abffb0c40cc8

Observation 33514e0b-4b21-431b-b16a-68848f3dad5e · outbound

This paper cites Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents.

Evaluation and Benchmarking of LLM Agents: A Survey Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.544451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.544451Z digest=sha256:0f33305da4a3dab0ea4855f15b98e121f402bbdeb2ef56ff0bdff9f2ca9deea4

Observation 8c2992d1-6232-476b-b113-711a9cbf67ba · outbound

This paper cites LLM Agents can Autonomously Exploit One-day Vulnerabilities.

Evaluation and Benchmarking of LLM Agents: A Survey LLM Agents can Autonomously Exploit One-day Vulnerabilities

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.562169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.562169Z digest=sha256:108030dd681c84af78f437527388996563d26ff5bc9fa0de24db3890a4b6c522

Observation dac87e5d-7d37-4502-a571-c8eec1d99f71 · outbound

This paper cites Re-ReST: Reflection-Reinforced Self-Training for Language Agents.

Evaluation and Benchmarking of LLM Agents: A Survey Re-ReST: Reflection-Reinforced Self-Training for Language Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.550610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.550610Z digest=sha256:1e4800e6427834f2164f883456472884c0f503c3d2a16a64b288e4c9fc207160

Observation 8f11e5aa-5671-45f3-8e40-c48c16c85aec · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.568290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.568290Z digest=sha256:f15a511ba84438a1ef0ff2b445390e1f3b12e62155763d14ddcabc9577312379

Observation 5f71b096-20f3-4f48-95de-22c4c408171b · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.556360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.556360Z digest=sha256:c2f5827f48b44e70e311e017a5c8315543b6606ec226b60ad4b0907c856fe20b

Observation 0dd12506-e3ec-4b8b-bf1a-21442e31dc7d · outbound

This paper cites Ragas: Automated Evaluation of Retrieval Augmented Generation.

Evaluation and Benchmarking of LLM Agents: A Survey Ragas: Automated Evaluation of Retrieval Augmented Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.559149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.559149Z digest=sha256:e33b9612ba4d56b121dfbb071ddb6e96d797433981061ee105e1abece4bfd12e

Observation 98fbabe5-a8ae-4a36-b3a1-d6b09064913c · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Evaluation and Benchmarking of LLM Agents: A Survey RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.577059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.577059Z digest=sha256:88f402207dd5adb524c227e3b0843852204219e9b8008c75ef2592823ad0e4b3

Observation 11fba66a-4d87-4501-bf86-cb4f7ade44bc · outbound

This paper cites Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks.

Evaluation and Benchmarking of LLM Agents: A Survey Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.565142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.565142Z digest=sha256:9cf55f9689322561c1b2657ff2a4228230d8f0695ad522c39493eaca7c4c19b9

Observation f4164373-270c-472f-a5c0-c1411bfdc6db · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.583195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.583195Z digest=sha256:1791608b9897bec35f73f306b391eb096d1cd87ce9037b6deab143c02b4819d3

Observation 45062438-4da3-4cff-adb6-6ce4ec281027 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.570992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.570992Z digest=sha256:6d82944616440708bd5506d5dc25f218f4cb4601a837496dbed401141b76adeb

Observation 03ed874f-470c-43f2-a573-5a733ab8619a · outbound

This paper cites Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents.

Evaluation and Benchmarking of LLM Agents: A Survey Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.574030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.574030Z digest=sha256:1b616671b0c09d731396523d4f22e55cc4944b7ff83c00f41a8d8c64b546b525

Observation 3be0e896-3a10-443f-9f86-725f2db6b00d · outbound

This paper cites Large Language Model based Multi-Agents: A Survey of Progress and Challenges.

Evaluation and Benchmarking of LLM Agents: A Survey Large Language Model based Multi-Agents: A Survey of Progress and Challenges

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.591815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.591815Z digest=sha256:e2abe389490bdf566bb43b3a9dae35608718cad5794a96dfd8baf4db23e0a005

Observation c0640dad-e5f3-42c0-9f55-e0a2bed5c23e · outbound

This paper cites AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.580085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.580085Z digest=sha256:8efa17af8c8daf4063a7801834867a8c171be572950869be4480a4e271049bc0

Observation d720d498-29d4-46c4-8d38-e7aed3ab185d · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.597845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.597845Z digest=sha256:25994e1bd0fba741febc6123a2d14555363cd8d2a0eb181e262aa8253145e8e5

Observation 67888622-73c4-45c2-9b2e-e5bc28d3bc5c · outbound

This paper cites A Survey on LLM-as-a-Judge.

Evaluation and Benchmarking of LLM Agents: A Survey A Survey on LLM-as-a-Judge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.585641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.585641Z digest=sha256:c0020dd65ca3687172e602160ab6bc94c4dfaff195266db7ba4aca9d36c2dc71

Observation 57d265d4-3b1b-4bb5-b9cd-4067450ccb32 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.588784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.588784Z digest=sha256:9f5c7242c2b732f1312adeb9ac533089435eb7e19373c36f9b9336b53a16a664

Observation 03971990-d79d-4f98-931d-b8d22cc98ff8 · outbound

This paper cites Understanding the planning of LLM agents: A survey.

Evaluation and Benchmarking of LLM Agents: A Survey Understanding the planning of LLM agents: A survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.606949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.606949Z digest=sha256:7ae7e8492f246cbaa6f3e9861daee0d5a6fe8f78ebcba782eb85c462524a604e

Observation 1cf4dcc5-30b3-49f2-955f-a82ac905765a · outbound

This paper cites LLM Multi-Agent Systems: Challenges and Open Problems.

Evaluation and Benchmarking of LLM Agents: A Survey LLM Multi-Agent Systems: Challenges and Open Problems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.594720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.594720Z digest=sha256:e086b07c8e81fbfc23eeb5e11734878e130ff0848e2ac95ecb8f15f431803f84

Observation a47616f2-fdfc-43ee-8e73-50016babc9c6 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.612824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.612824Z digest=sha256:d84de6993e689de1fe52f336f0ecbf004eeaf3348ce5c179dde36ac33c2967b2

Observation 7e4ffcf8-ac71-47ad-a8f1-b9f26e1d22c1 · outbound

This paper cites AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation.

Evaluation and Benchmarking of LLM Agents: A Survey AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.600583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.600583Z digest=sha256:d9be2120f183d5c766c76a45cdff392b12b97cd435e5e547f0e26edae78371c5

Observation c0f909b1-34af-4a68-bfb1-59360d7225e9 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Evaluation and Benchmarking of LLM Agents: A Survey Measuring Massive Multitask Language Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.603848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.603848Z digest=sha256:dd0dbef3e80aba66e6262585946ff8a2f8179b6b612851d821b673472c12b86b

Observation 38e49dfc-b559-4e10-8a8e-db2336f3cace · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.624201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.624201Z digest=sha256:3362ceb57ad90ac1965365f99eb721b210d8895d64ebc0dc9d5760d64f20e726

Observation 1576480d-e593-4139-8f1a-353f60d92f52 · outbound

This paper cites MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use.

Evaluation and Benchmarking of LLM Agents: A Survey MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.609937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.609937Z digest=sha256:83e7161c7c1165094f59fa16a318db70b7c100d481a42cf64c02d665d3ce3a12

Observation f5195e60-8380-49c0-be77-aced989195e1 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.633063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.633063Z digest=sha256:c1086a1aa64ccb62f994fc035178530d6b4f6cd2dfd826500f2b2ff1241279ce

Observation 17afe007-792d-4036-81d8-4f6e86d4e29c · outbound

This paper cites LangSuitE: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments.

Evaluation and Benchmarking of LLM Agents: A Survey LangSuitE: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:44:22.818063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:44:21.615493Z digest=sha256:55a293e320be31ac37991f4ff718016c0e8da09c8bc0e77bfabc354e65f3e641

Observation 6fbcd1ef-99ba-44ae-9490-f3f8136fb2d4 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Evaluation and Benchmarking of LLM Agents: A Survey SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.618485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.618485Z digest=sha256:e1147fb9661f9b24b62e1aab7572ee11fdddb2e906536193b455b8cc86437c3a

Observation b75ef5d4-56c0-4b5e-a026-a442f6ea8500 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.621618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.621618Z digest=sha256:2ee676f1eee518191c2115c4c55305aa2d9e79588fc66f7ae8816f0fbe054f85

Observation fae8e12c-84d0-4d3a-8cf8-5b793eede1c1 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.645335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.645335Z digest=sha256:11903d22017a09ef762cd02ece18f30ba0db1c112c583aac8ffa7320da6bcf37

Observation 95dbed4c-c632-46b2-847c-0137bac07bca · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

Evaluation and Benchmarking of LLM Agents: A Survey VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.627127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.627127Z digest=sha256:510076b75d1d1166389262147b32d9ee71652edf416cb5f41a2ced25db92a3e1

Observation 08dd9ff0-baad-4333-b7dc-9e5ef96fc258 · outbound

This paper cites LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization.

Evaluation and Benchmarking of LLM Agents: A Survey LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.630164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.630164Z digest=sha256:74ac634e1c86eeeee227313617ee34741a6464c6362545a2a61f395d71037f37

Observation 988d73f9-f5c6-475f-830e-f6a5bc06d29f · outbound

This paper cites Manning, Christopher Ré, Diana Acosta-Navas, Drew A.

Evaluation and Benchmarking of LLM Agents: A Survey Manning, Christopher Ré, Diana Acosta-Navas, Drew A

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.653693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.653693Z digest=sha256:d4e3c9879e4ed0b56ec2f8902bd21e77355a03ac26e3dca40cfefe8c1e7480a6

Observation a7a1f55d-5608-463f-ba76-40eaeb9259b8 · outbound

This paper cites IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems.

Evaluation and Benchmarking of LLM Agents: A Survey IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.636050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.636050Z digest=sha256:a809af2a2c83a8ce818bd0e7bb06bf23add7a636a689b1f547258bc77721d244

Observation 1ea9c14a-43b7-48a2-9a75-5b196197df94 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Evaluation and Benchmarking of LLM Agents: A Survey LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.639176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.639176Z digest=sha256:a515a815a71f170c03bdb97e07594ccb48005ff4f00740cd2dd80aa092695edc

Observation 4f5134ab-59be-4e29-8c6c-d1a72f2f2df9 · outbound

This paper cites API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs.

Evaluation and Benchmarking of LLM Agents: A Survey API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.642196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.642196Z digest=sha256:18d0bb5307428e53e2c20e6126fa0c450a1ead0ca4489bc6fb5df4b733b1d361

Observation 32246f49-b35c-4ff5-9d98-25867ca6336b · outbound

This paper cites Autonomous Agents for Collaborative Task under Information Asymmetry.

Evaluation and Benchmarking of LLM Agents: A Survey Autonomous Agents for Collaborative Task under Information Asymmetry

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:44:22.763154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:44:21.671369Z digest=sha256:2d87f71a981ec5bba514668f68f469639d1deb0da7195834e56cfee611f0ff2f

Observation a16e74b8-bc20-4e6f-afc8-922a972f43e3 · outbound

This paper cites Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security.

Evaluation and Benchmarking of LLM Agents: A Survey Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.648044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.648044Z digest=sha256:0d73daeb9c5d783b4c64250c29c81efbb7ba52a641720ee8ff20e7d46c3a926f

Observation 87cae1f6-0021-42e7-abc1-8c40a2a3919c · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.651084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.651084Z digest=sha256:476d7c4d728f9bc7d293dcba3e931e8dcaa34479dfcfb6813d15a897088bd83e

Observation 5d0eda4b-5def-4aba-aad3-b03c122d49fe · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Evaluation and Benchmarking of LLM Agents: A Survey AgentBench: Evaluating LLMs as Agents

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.679806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.679806Z digest=sha256:62ddf852834a3744a7f13271056421f4f257e58b2ae917a8172dcc1224af7e19

Observation 4d83681e-34b9-4e2d-a9f5-7ea8018dcf15 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.682381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.682381Z digest=sha256:91831168f01da4b05029cb1d6f2757eb5a73655f69cc995884c63bffb9274fe3

Observation 10e97fee-f5b5-4387-a5b8-1543f55ffa76 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.659528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.659528Z digest=sha256:ee4ca545fecbaeba28507a9eaccf1ae378d670e38feeb3fe65ac4b3702469c10

Observation acb862c6-b9b7-40ed-b0b5-1650e4f96031 · outbound

This paper cites AgentSims: An Open-Source Sandbox for Large Language Model Evaluation.

Evaluation and Benchmarking of LLM Agents: A Survey AgentSims: An Open-Source Sandbox for Large Language Model Evaluation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.662265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.662265Z digest=sha256:f026257890b5f9ac11e131b183b5f8b8b8b22fba508c6046d979c72995f9bf88

Observation a71eb74e-e7d4-4f48-a9fc-14f70eaf0f26 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 62

Resolution
verified exact
doi, observed 2026-08-06T12:44:22.136258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:44:21.665219Z digest=sha256:b3c69f8fcdea70f0bc2c14c98a2cd8cfc387d88096b33148a793bf2b03fedb9f

Observation b4988fd1-00f9-411c-a20b-cd5f66f6298d · outbound

This paper cites Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems.

Evaluation and Benchmarking of LLM Agents: A Survey Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.668218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.668218Z digest=sha256:8ca0a8c6bf204f70e95b704fc622280ad6bba16bc99442bb54774d2a310d0cae

Observation ff4c10d5-7239-4858-9037-39477c88bd91 · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.699335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.699335Z digest=sha256:0ef97ad5e7b3d8afff12e4ee8d29372e0d49a4a1dc93790dc18a12dc49346f4a

Observation 4f007bf9-f524-4fe6-bddb-4a592f38fa06 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.674191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.674191Z digest=sha256:730a003dfef88457a612f8c29e2fe8457e452ea3f7231a1d5eab7e9fbfb53289

Observation 8693fcb6-3f7d-4d62-ba10-250858c969b7 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Evaluation and Benchmarking of LLM Agents: A Survey AgentBench: Evaluating LLMs as Agents

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.676884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.676884Z digest=sha256:0500617fb51fbadf77708810ff61b6026fe424259408bbad2eeb0045e905c48d

Observation 5d6b6591-30e0-40d4-a330-824cbe460198 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.707889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.707889Z digest=sha256:1d6d20501f51373ce30dbd2dcac5f0e20024a4f48e63e291b2303f47fde43423

Observation a3bb5815-e184-4050-875e-aff77b02f5b8 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.710615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.710615Z digest=sha256:b1cc95d40519a1dbd00d1b7f912f50eeed411cc2fec3d6a5c891900be8a0dba2

Observation e4833b80-cd87-46f3-804a-6d73e1bbc477 · outbound

This paper cites Chatting with GPT-3 for Zero-Shot Human-Like Mobile Automated GUI Testing.

Evaluation and Benchmarking of LLM Agents: A Survey Chatting with GPT-3 for Zero-Shot Human-Like Mobile Automated GUI Testing

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.685043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.685043Z digest=sha256:ea48f2bf419f39f7575fb64d6976dd59e4fc32dddd530988bcdb4fbba65e0854

Observation d7f7fd7c-d72d-4031-b562-f6a09f14307b · outbound

This paper cites AAAR-1.0: Assessing AI's Potential to Assist Research.

Evaluation and Benchmarking of LLM Agents: A Survey AAAR-1.0: Assessing AI's Potential to Assist Research

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.687855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.687855Z digest=sha256:fb74b50ad976a0c086eaa5a1e9d35df052f9406cab34ddf651c6c09cedf66fd2

Observation 332611bb-3bbf-433e-ab58-d8aa519a7dda · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.691267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.691267Z digest=sha256:95bb94c4eeb5429c480bc733a4dfc94fc0079ab150e8374e00e09ee204becf9f

Observation 51527181-da25-4aff-a12a-678c0932dc05 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

Evaluation and Benchmarking of LLM Agents: A Survey The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.693952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.693952Z digest=sha256:76ebf6e50a43b6d93c81c5f08cbf6e771104edf7e3c4dc5514c97fe46465ffab

Observation 2793d074-519f-4448-b9f2-cfc2d738c4fa · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.696779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.696779Z digest=sha256:e413fbbe6d5efbb1230b72d1219d739e5579a9df0d957abffeff7a4ce9581a1e

Observation d82f2c72-baa2-4505-9ab0-5d77ac44e58e · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.730067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.730067Z digest=sha256:c6af71a9712dc6cfa2e61ffc027120b52bb90e052ed23b1408d038e65ef4aa83

Observation ea8717b2-0ba8-4a3b-8d82-ba41f72e027e · outbound

This paper cites Evaluating Very Long-Term Conversational Memory of LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey Evaluating Very Long-Term Conversational Memory of LLM Agents

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.702317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.702317Z digest=sha256:cfb481695ba7fb9e342ebc603b6aeaff700efb426944aeb23f7cadc8107cface

Observation 8d34f258-59af-46a8-8de8-a4132c3b0ccb · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.705264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.705264Z digest=sha256:55d1f867fe79eee4a17312236621f9127b764541105d05887f700cea6b6972c0

Observation c47bb156-6098-4617-81f1-d02a36e835ab · outbound

This paper cites Evaluating Cultural and Social Awareness of LLM Web Agents.

Evaluation and Benchmarking of LLM Agents: A Survey Evaluating Cultural and Social Awareness of LLM Web Agents

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.741655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.741655Z digest=sha256:fda5661659f527b64acbef3061d29e06f462a091c83de72a3fbec1ddd1bbc457

Observation f34adac2-6809-4442-b63e-0ee4f1d601d0 · outbound

This paper cites Know What You Don't Know: Unanswerable Questions for SQuAD.

Evaluation and Benchmarking of LLM Agents: A Survey Know What You Don't Know: Unanswerable Questions for SQuAD

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.744477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.744477Z digest=sha256:8e3acdef21c66a66b3a634a34fc660ddd6d3723568b8f35a61b337d41256b35c

Observation a63d1cf0-e7dd-4d36-8075-c300f73d5f27 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.713486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.713486Z digest=sha256:3a8039f1cddb2f49a3fb739ce01ef5f2f5829a7f9c7e0cd407d5e71ae4009e6c

Observation f021281d-6cf7-46de-9b3c-47488b8685eb · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.716073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.716073Z digest=sha256:4b9a8e919adf03c8f0da202eb0ee8c6f17f9e57be4d9137820f0bf60c499c0c4

Observation e4b2c4c2-a9cd-420d-b5c8-9c8bceee7f3f · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.718892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.718892Z digest=sha256:9e47a5e94deedeba5cdb511e59d36d92d722d0022d47f067dc27eaac76c1216b

Observation a32272ee-a3ea-4d59-bd10-7c943f08e3e2 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.721508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.721508Z digest=sha256:e4d38d065e1a0e06ffaa03932acae337051abd02b9b941bab1a8288597e0cb9f

Observation 40861379-0377-47af-9bf4-7b6d5b2f7118 · outbound

This paper cites BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games.

Evaluation and Benchmarking of LLM Agents: A Survey BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.724101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.724101Z digest=sha256:1273479e761ea314f6805a1ab2ea53b3996b87e421e57f7e4636f88ed7deaa42

Observation 56cfd1f3-cdee-4b97-9412-567365435198 · outbound

This paper cites WebCanvas: Benchmarking Web Agents in Online Environments.

Evaluation and Benchmarking of LLM Agents: A Survey WebCanvas: Benchmarking Web Agents in Online Environments

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.727061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.727061Z digest=sha256:2d4963d09cfc2c4d49c821eed82a177c86a8a99d8601dd1b63c82218acf43411

Observation 670d85e0-0722-4c56-8bcb-46e51cbff8b5 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.768017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.768017Z digest=sha256:794baad030c9d82e0d7d7e270e6a7d69a6cd58249cf4833b99e3ba3c0870d736

Observation 67ee1029-8780-44f8-8fee-756590eea613 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

Evaluation and Benchmarking of LLM Agents: A Survey Gorilla: Large Language Model Connected with Massive APIs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.732658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.732658Z digest=sha256:0e47f7426b7b473f9de8061d1e4c8b13e9617999fe051e089c916cf041c89ec2

Observation d30536c9-4172-4044-8786-5528383e4f5b · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.736007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.736007Z digest=sha256:cc65fb28721e13f8ecca4079a1335b9a7fcff769c55bff6a3afd750c3f49afe7

Observation b42bbf4c-8db5-45a2-910d-5e0cbcef3c63 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Evaluation and Benchmarking of LLM Agents: A Survey ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.738675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.738675Z digest=sha256:6870da53afe54c16ef8464443fa90822bae648eb4ea118604a8040c5f795f8e6

Observation 3df9fabe-2935-42bb-bc9e-917387192ed9 · outbound

This paper cites Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.782168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.782168Z digest=sha256:1cb3905f9be296c7e3e77be4cb484d27a7dc12102c9f9ef49f598d90c10caea7

Observation 78653e92-0386-4bef-a808-53abdf6e81e3 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.785167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.785167Z digest=sha256:7932f7a02fa6b25d7c3d967f743ca6c850ba03ea34b8e0fb5baba2cd41fa2151

Observation 9e607f34-0109-49b2-9994-8cd5f0bcc300 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.747431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.747431Z digest=sha256:12373e525c9baa60cd999c639df1b645dc84b82137567fffbbed81ec604edab8

Observation e8dab5f9-2e2a-445e-8ec8-310bff789f07 · outbound

This paper cites MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.793466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.793466Z digest=sha256:fa4939465f51f0ce56bbcfd6ce55c69c63ed3792987545e9bafb3ce5a06943c9

Observation a939c8f0-8b53-4922-aa76-00d617573ea3 · outbound

This paper cites Reimann, Catharine Oertel, Florian A.

Evaluation and Benchmarking of LLM Agents: A Survey Reimann, Catharine Oertel, Florian A

Reference 93

Resolution
verified exact
doi, observed 2026-08-06T12:44:22.059811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:44:21.752880Z digest=sha256:e3564c351b746640bf3f0da4c01a6ddd549dec6e8a727c99c8a798ebd9f1e517

Observation ff967955-d427-497b-a773-035c266ccb71 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.755622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.755622Z digest=sha256:6a676c107c794160e85fcc992bab177c3ee539f64d1f855fc2a01d9b3cc165de

Observation 78cc4a51-1f84-4584-8b9f-0ba4a884a3f5 · outbound

This paper cites Identifying the Risks of LM Agents with an LM-Emulated Sandbox.

Evaluation and Benchmarking of LLM Agents: A Survey Identifying the Risks of LM Agents with an LM-Emulated Sandbox

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.759303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.759303Z digest=sha256:6b25d9f84284f1037d69e96755a2d9596270e42300b15977322b52c828eae87f

Observation 8fd32775-a877-4f2f-8338-4d5d8e800dc2 · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation.

Evaluation and Benchmarking of LLM Agents: A Survey TaskBench: Benchmarking Large Language Models for Task Automation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.762114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.762114Z digest=sha256:64bec449251f4b29aef9b66b577e9353f382417eef641e10415ddc26a6136fe1

Observation ab70d92b-935f-40c1-8002-2c449028ec71 · outbound

This paper cites Enhancing Cluster Resilience: LLM-agent Based Autonomous Intelligent Cluster Diagnosis System and Evaluation Framework.

Evaluation and Benchmarking of LLM Agents: A Survey Enhancing Cluster Resilience: LLM-agent Based Autonomous Intelligent Cluster Diagnosis System and Evaluation Framework

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.765199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.765199Z digest=sha256:5d75d39b3ab220fd9636cae8322740cc4a4d83aa4502df8ed225872e95858d71

Observation 85876bdb-f874-4a30-8cf2-1ece2333c195 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 98

Resolution
verified exact
doi, observed 2026-08-06T12:44:21.999884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:44:21.810777Z digest=sha256:369218db1340cf5387ef7f56c10a0949eb322802355c06896dd5753902f336b1

Observation 474df2fa-d0e8-460b-b8bf-1b720f13b87d · outbound

This paper cites MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration.

Evaluation and Benchmarking of LLM Agents: A Survey MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.813422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.813422Z digest=sha256:dc070af1845ccea3c81723ee518cbbd750868ec0610a1159bb9b4729c3a48ce9

Observation 13f5ad70-d8d5-4c0b-9765-0743eb507494 · outbound

This paper cites CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark.

Evaluation and Benchmarking of LLM Agents: A Survey CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.773326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.773326Z digest=sha256:6c96a070caa4173e78515e49936daf8c8934244ffe4e5c7164f68a68354f6b44

Pith citing papers

Observation 502c01ef-29cc-448f-98a6-0fec0e76d9b2 · inbound

Herding CATs: ALARA for Agent Harness Engineering in Portable Composable Multi-Agent Teams cites this paper.

Herding CATs: ALARA for Agent Harness Engineering in Portable Composable Multi-Agent Teams Evaluation and Benchmarking of LLM Agents: A Survey

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:14:08.422394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T11:12:46.626382Z digest=sha256:18b59eb3cc887e78fcf2063da359a1a59456d53521f692584e1143a9f892c688

Observation b79c7602-7f7c-4d27-a357-45e108c435fa · inbound

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents cites this paper.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Evaluation and Benchmarking of LLM Agents: A Survey

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.680558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:40981af6a59121ac7b279d874cefbe0895360d7bc7d7eb7a3b87d38648902209

Observation 83459fa1-372a-4927-952a-000830629ab0 · inbound

The Scaling Laws of Skills in LLM Agent Systems cites this paper.

The Scaling Laws of Skills in LLM Agent Systems Evaluation and Benchmarking of LLM Agents: A Survey

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:13:37.607891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T18:10:08.737710Z digest=sha256:e197747923a849cbca40b1b7242b33e39c9a1f316d5a770ff8ff6bb25e4d6445

Observation 9e9ad036-2431-4c7d-a36a-e36b80b5c2d7 · inbound

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems cites this paper.

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems Evaluation and Benchmarking of LLM Agents: A Survey

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:47:17.744554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T21:53:37.616447Z digest=sha256:c22339f6da000b428684a908740b28262e674caa87f71087ce683f29eb53b97e

Observation b0fd8271-1371-4f80-9afb-b8ec1bb1dc4e · inbound

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness cites this paper.

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness Evaluation and Benchmarking of LLM Agents: A Survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.024402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:04:29.791558Z digest=sha256:aa9025de1505768a731cad97909cc0f72bf8da490fccb078b47ebf819933ec7e

Observation 6e863949-2b1e-4d52-90bb-120f43c84ea3 · inbound

CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI cites this paper.

CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI Evaluation and Benchmarking of LLM Agents: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T01:59:52.218232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:59:52.218232Z digest=sha256:cb5e394476f77627addfb24abbf4d1e8c655ac55cc4f2ea4eaf8f275d3e8ff3a