Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T06:00:19.491104Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 209 outbound references and 4 inbound Pith citation observations for arXiv:2506.11102.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T06:00:19.491104Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T12:52:20.911788Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T12:53:26.388274Z
100 of 209 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e9f6a353-6bd8-4e4c-ad86-066f7a6b64b0 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Attention Is All You Need
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a874fb-0745-42ba-ae6c-4dbe13914100 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4204d53-69d2-4844-9220-e09b75a3b614 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey LLaMA: Open and Efficient Foundation Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd209cb1-83af-4bed-a32e-42fe2a9e3eef · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Gemini: A Family of Highly Capable Multimodal Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f1aea1c-d210-439c-826e-ad41f4eb41da · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Qwen Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9955fec4-7265-4424-801f-20e4dbd9cb3d · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69fa2201-e8bc-4705-b486-9f157ea47c3b · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce1507e-654c-4c19-9d1f-29057183678a · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Beyond Browsing: API-Based Web Agents
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b486343d-eeb8-4628-8e1f-34b2dbe3db46 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Reflexion: Language Agents with Verbal Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25195080-ca4c-4c89-b675-24aa4716ba9c · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey A Survey on Evaluation of Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67aa1f7b-6ad2-4d05-81d8-0bebbd80e45a · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a852ba-7372-4622-9b0c-2306905712da · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Large Language Models: A Comprehensive Survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e22630de-a20b-4ab0-a2ba-22e7415e6c38 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Large Language Models Trained on Code
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c35e0c2-bf2a-4915-8d74-55c1a2c5b3a0 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8725aa85-b4ef-4110-94c3-4875641fb7c4 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Codereval: A benchmark of pragmatic code generation with generative pre-trained models,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9a44159-a09f-4e11-8a89-a47582ee1cb6 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb73fe5c-1553-4d57-80b8-eec45055522d · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b00341c-253f-4bcc-9dbe-67c6ddb7b03f · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f4399a5-4ada-423b-b1de-9ac0369b8db8 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb373681-2cd1-4688-80b1-e656e8507ee3 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Natural Language to Code Generation in Interactive Data Science Notebooks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38c40404-8e4a-442d-92f3-7c537a4581bd · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a21d675-1c13-4335-81f9-84ac9f795873 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1309315-93d3-4e28-9d01-2ab65f09bbd6 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6f4fa44-f487-4298-89f8-fc866145363d · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac6f0cbc-2bd6-4dc9-8551-cdf4d36fc239 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Swe arena: An open evaluation platform for automated software engineering,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 260f0122-a06f-458b-8013-6a447543a989 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Large Language Model Critics for Execution-Free Evaluation of Code Changes
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 813acb66-9098-400d-ad9c-6e5b2a189a9b · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-Bench+: Enhanced Coding Benchmark for LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e418d1de-4524-4de4-8d16-e35ddd9d35d4 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f139a15e-d5af-4b93-a480-42d0a4391586 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Swt- bench: Testing and validating real-world bug-fixes with JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 15 code agents,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9b9dfe3-3ea2-422c-ac78-ec6c47107748 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ChatDev: Communicative Agents for Software Development
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8794d234-6f6e-4577-af53-7fe360f293b7 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d59ecd38-95a2-4267-810e-79a9d494d2f9 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64669182-ef2a-4c15-8438-b5fbcd23cb2e · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey PyBench: Evaluating LLM Agent on various real-world coding tasks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfa6a7b2-935b-4330-954d-b906b47d2b12 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey World of bits: An open- domain platform for web-based agents,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6a0b09c-4fd8-442b-9c61-e7b4509ca6bb · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28696c09-9fb9-4fee-a302-0c0c984f2934 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a000fe7e-1b6c-4350-857e-19800d4bce97 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mind2Web: Towards a Generalist Agent for the Web
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32195e48-97c1-47c4-91bc-c333b7e5b9a8 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16c6613a-41f3-443b-9ed3-312f57df1290 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Weblinx: Real-world website navigation with multi-turn dialogue,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b841e096-a322-4e05-b1e9-be5b2f421b13 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab2c25f-bea8-497c-9224-96dacde30869 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d16271a0-9a49-4f04-8f55-39f65d233eca · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 630117aa-28ce-4993-80b4-114cec9d210a · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1186c679-70d6-4b64-88d5-97feafe877be · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mmina: Benchmarking multihop multimodal internet agents,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f46d40d4-b103-481d-bb00-172bd36a716b · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1cf617c-6338-4d35-ba44-47b9195a9237 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WebCanvas: Benchmarking Web Agents in Online Environments
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4662cf8d-cc72-4bb3-ad01-cc6b06f34723 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c94bc8b-9eed-45ac-aa11-f94971c1673d · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2078033d-11f9-4c4b-aa9a-50f46a02f291 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Tur[k]ingBench: A Challenge Benchmark for Web Agents
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f07bc4b2-1d36-420b-9040-ebb798af31bc · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey BEARCUBS: A benchmark for computer-using web agents
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e668517-a3fd-4001-90bb-2b2b25d90438 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d4cc33b-2ebb-423f-aa49-f9c9382bc14e · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Waber: Evaluating reliability and efficiency of web agents with existing benchmarks,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b2137e9-47f1-41d3-ba0c-d1117c0951a1 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Visualagentbench: Towards large multimodal models as visual foundation agents,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 462df6dc-603f-4a8e-bcde-f53f8d7eae9f · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey An illusion of progress? assessing the current state of web agents,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f6c04dd-390b-4807-87bb-94cf5d86252d · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Available: https://arxiv.org/abs/2408
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a93fdb7-817a-49d1-b70e-8ece429887cd · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef6a25c2-4649-443a-b98a-a168c8efa59a · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bac919f-dc2c-4aae-ba7e-4e4107315075 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Osworld: Benchmarking multimodal agents for open-ended tasks in real com- puter environments,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f0563c7-aef4-4f8c-a779-69619ed51102 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Chatshop: Interactive information seeking with language agents,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dda98239-a6b9-4277-95b6-c42b36d3fad9 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Available: https://arxiv.org/abs/2404
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19e26d6-6b73-4def-8ebd-53bbd598a4f0 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Omniact: A dataset and benchmark for enabling multimodal generalist JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 16 autonomous agents for desktop and web,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a74440e1-7362-4698-933a-030d5e0788dd · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 926aa596-72e1-4422-8377-9852b44ba4a2 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AgentStudio: A Toolkit for Building General Virtual Agents
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1161f127-6a82-41d7-a01d-7c7dc859b437 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mapping Natural Language Instructions to Mobile UI Action Sequences
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c9775cf-fff2-41b9-8e65-cec65a483072 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c4ce685-e80a-427b-8709-453443680570 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c24c036c-e0ba-44ec-92ac-ba87eb081604 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Androidinthewild: A large-scale dataset for android device control,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c92253-b41c-4aeb-bb37-9ff2de041d01 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey UGIF: UI Grounded Instruction Following
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b16d9ff5-45a8-41e3-84d9-6a62c76dd8ad · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mobile App Tasks with Iterative Feedback (MoTIF): Addressing Task Feasibility in Interactive Visual Environments
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c516cf87-86ac-4af6-ad50-884687aaba0e · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Benchmarking Mobile Device Control Agents across Diverse Configurations
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c35fdbb1-a917-4fae-82f7-bb9d1f4e8f26 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey On the effects of data scale on ui control agents,
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 597d503c-9e1f-4749-b85a-2961a22930c9 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 662a50e3-fee9-4a80-a395-7d688686b468 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey A3: Android agent arena for mobile gui agents,
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d43916d4-6038-4196-9424-1085d5bd84e3 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8946c67a-a648-4d4d-a8b3-ed1c6b20f1df · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b013bde-7dff-41c4-a047-955f874f7890 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mobilesafetybench: Evaluating safety of autonomous agents in mobile device control,
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61b6aa15-c081-4984-9703-32da31ddb09a · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 880efb3e-e422-4412-8589-7df047937b48 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Understanding the weakness of large language model agents within a complex android environment,
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e51160c5-3605-4bfa-a3ea-9c04408c3173 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Llamatouch: A faithful and scalable testbed for mobile ui task automation,
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8277fa0c-d3c7-465f-bcbe-5fb718485389 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b46f2ab7-81a9-45b5-9190-efb3a3381d3d · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Spa-bench: A comprehensive benchmark for smartphone agent evaluation,
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9249483-3387-4fc9-807e-6b48abb01468 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb119f8e-8e28-43f3-8f4a-0b5c711da058 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d98393-a742-4dc2-aff9-a6ba44f84d3d · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Autoeval: A practical framework for autonomous evaluation of mobile agents,
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fdde1d4-9b54-4d4e-b018-1050ccf0519e · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Qasa: advanced question answer- ing on scientific articles,
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 791bbd92-ebc3-44b0-beff-a12e5c7f077b · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e72e4600-702a-421d-b460-bf671fe7bd72 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ac21333-8046-40ab-83e6-38322acfbae1 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MS2: Multi-Document Summarization of Medical Studies
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da3053c0-5c76-460c-9b66-332a6e24514e · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Automated Focused Feedback Generation for Scientific Writing Assistance
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 028d9b46-f4d3-4629-9976-e3ae1e18f5ef · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AAAR-1.0: Assessing AI's Potential to Assist Research
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1ca246e-b03b-465f-9f0b-cfa9a805e722 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03d5c5a1-2f1b-4f20-af18-7af2816ab996 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb95e92e-a272-40c9-b8f7-2e872a71d6f5 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e1a027-28d6-459e-96e5-bb96895989a5 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec61084-7634-407c-97d2-09d132ee4d14 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey PaperQA: Retrieval-Augmented Generative Agent for Scientific Research
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27b08d10-0e8a-46ed-84aa-a9b8314678d3 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ScienceWorld: Is your Agent Smarter than a 5th Grader?
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 043cd23b-d07a-4272-8573-e9748501667f · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mlagent- bench: Evaluating language agents on machine learning experimentation, 2024,
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a271e68-1308-4909-8500-8452e86d43b9 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79bccf71-c36d-4f0c-ba6c-fac4d942ced4 · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b944511-a60a-467d-87f4-4285fa1271aa · outbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18e21d6-4917-4748-9379-93e8566336c1 · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
Reference 154
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3e2c672f-c607-4eae-81a4-f70ce2797404 · inbound
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
Reference 201
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c5748b6-1302-4008-9625-1e82f498a169 · inbound
SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 068b5531-8028-47f4-a264-5b52127cbf69 · inbound
OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.