Pith. sign in

Paper Citation Record · LEDGER

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

As of 11 August 2026, this Paper Citation Record lists 100 of 209 outbound references and 4 inbound Pith citation observations for arXiv:2506.11102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11102 v1

Coverage vector

measured 100 of 209 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:00:19.491104Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T12:52:20.911788Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:53:26.388274Z

Reference resolution

100 of 209 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9f6a353-6bd8-4e4c-ad86-066f7a6b64b0 · outbound

This paper cites Attention Is All You Need.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Attention Is All You Need

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:15.697976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:15.697976Z digest=sha256:d9dc0c1d44c9698806aaf1b0f694b61d0ff6c4e8d9cdea5f6eeb053d1cb4c0ff

Observation 20a874fb-0745-42ba-ae6c-4dbe13914100 · outbound

This paper cites GPT-4 Technical Report.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:15.738449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:15.738449Z digest=sha256:a65ed34c7e1403b963c99002c1b5a6ebb35ef8a575534540580e74746b3d243e

Observation e4204d53-69d2-4844-9220-e09b75a3b614 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:15.823334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:15.823334Z digest=sha256:b71ef82d729ee0d8731e4d19d8860234ba93f9c4bf3e6f4ef80637472f7c010d

Observation bd209cb1-83af-4bed-a32e-42fe2a9e3eef · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Gemini: A Family of Highly Capable Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:15.888246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:15.888246Z digest=sha256:10a3d679607bd1e9d9d9b20faeca59649314899bbcf698aba2d11dd3aea46378

Observation 1f1aea1c-d210-439c-826e-ad41f4eb41da · outbound

This paper cites Qwen Technical Report.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:15.977297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:15.977297Z digest=sha256:cbba90ddba375914a23ff1aa9713976631b67da6ba7ce2754983d0a1e4c4fc03

Observation 9955fec4-7265-4424-801f-20e4dbd9cb3d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.066521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.066521Z digest=sha256:5a814befd5029633a7dd380c24c5af9d33e5c800d7dfb8d084f138679afa5ee3

Observation 69fa2201-e8bc-4705-b486-9f157ea47c3b · outbound

This paper cites Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.156522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.156522Z digest=sha256:0331112f6c767a9cf68a48b58cbf995bc24a9b88623e8c853bf80fb12bcd5fc9

Observation bce1507e-654c-4c19-9d1f-29057183678a · outbound

This paper cites Beyond Browsing: API-Based Web Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Beyond Browsing: API-Based Web Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.253072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.253072Z digest=sha256:10c9c4506945c6eea725e91c8985159a5a78a40d39ff566954051cb22658e121

Observation b486343d-eeb8-4628-8e1f-34b2dbe3db46 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.334469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.334469Z digest=sha256:64f14baa8a9a5d9d4477e4e54217ecdae6e77bc2bde63b06100a3e67e7eed4c5

Observation 25195080-ca4c-4c89-b675-24aa4716ba9c · outbound

This paper cites A Survey on Evaluation of Large Language Models.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey A Survey on Evaluation of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.424040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.424040Z digest=sha256:cb72187970ca9fdfcb3e9af35d1f3b5af62245962b82b7a593e23fb7c29a5d30

Observation 67aa1f7b-6ad2-4d05-81d8-0bebbd80e45a · outbound

This paper cites A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.519703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.519703Z digest=sha256:ac01dbbd6594794c50dacaad336f860b289f84c11a06eee6853d0ae8b29c6968

Observation a6a852ba-7372-4622-9b0c-2306905712da · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Large Language Models: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.603597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.603597Z digest=sha256:8b20b0cd20b0db2917186434b6ce75ae9e354c5038800b5d9b64a3c2b901b093

Observation e22630de-a20b-4ab0-a2ba-22e7415e6c38 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Large Language Models Trained on Code

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.716446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.716446Z digest=sha256:0efa5f30e589a6e9801a172ec1914867e464fa79bae0ca44dec3d261ec015fd6

Observation 6c35e0c2-bf2a-4915-8d74-55c1a2c5b3a0 · outbound

This paper cites ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.785966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.785966Z digest=sha256:8c1f6007126d494a8f4522e0f0e58eb2bb0db9dd0daa6e4a8e24fb30b26b0c31

Observation 8725aa85-b4ef-4110-94c3-4875641fb7c4 · outbound

This paper cites Codereval: A benchmark of pragmatic code generation with generative pre-trained models,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Codereval: A benchmark of pragmatic code generation with generative pre-trained models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.847762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.847762Z digest=sha256:08d105c2c1ce892659ff3708e1f1dc330d2cbd42b14bfad4a56a3590079abf76

Observation a9a44159-a09f-4e11-8a89-a47582ee1cb6 · outbound

This paper cites Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.923044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.923044Z digest=sha256:b1fb1bdc64d4c8e9e6ac737fed2d11287e11908fb5d0fbbb6264593834a546d9

Observation bb73fe5c-1553-4d57-80b8-eec45055522d · outbound

This paper cites CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.038881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.038881Z digest=sha256:025aa034b5a67a86366754256b853d52b1ead5816c89ac44de482ad1385f7cb7

Observation 6b00341c-253f-4bcc-9dbe-67c6ddb7b03f · outbound

This paper cites CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.129393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.129393Z digest=sha256:0e4e5eede55c1e5a24338f2996b2fc6ac3b81b3897da01d6bd1c81eea0bc14e2

Observation 4f4399a5-4ada-423b-b1de-9ac0369b8db8 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.220729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.220729Z digest=sha256:9c9bdec3b34da7f7b54ffddbba3d9dcdadaf9c259d181a20b541c22c91187613

Observation fb373681-2cd1-4688-80b1-e656e8507ee3 · outbound

This paper cites Natural Language to Code Generation in Interactive Data Science Notebooks.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Natural Language to Code Generation in Interactive Data Science Notebooks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.290883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.290883Z digest=sha256:79b35de77d08714f4b115e5dffa58390b1866a09374039fca0d22dab6490e993

Observation 38c40404-8e4a-442d-92f3-7c537a4581bd · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.379245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.379245Z digest=sha256:925747ceffad253d651ecf19dc469941eef915368adf5b655a73121e9018da45

Observation 9a21d675-1c13-4335-81f9-84ac9f795873 · outbound

This paper cites SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.468272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.468272Z digest=sha256:58001c05a2616106d9cb4f38462b0296c4266adb1b18803c77f4fa3f23d12ad6

Observation a1309315-93d3-4e28-9d01-2ab65f09bbd6 · outbound

This paper cites SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.591067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.591067Z digest=sha256:a86b29dfa4213d670c9d0eff83911774d97d54b8448ffbbb70757c16b44b2109

Observation a6f4fa44-f487-4298-89f8-fc866145363d · outbound

This paper cites SWE-bench-java: A GitHub Issue Resolving Benchmark for Java.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-bench-java: A GitHub Issue Resolving Benchmark for Java

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.668524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.668524Z digest=sha256:bf97009dc22779db1c9b7b39ce737e042a55a374797a68c706b3733b15fe5bc0

Observation ac6f0cbc-2bd6-4dc9-8551-cdf4d36fc239 · outbound

This paper cites Swe arena: An open evaluation platform for automated software engineering,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Swe arena: An open evaluation platform for automated software engineering,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.755574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.755574Z digest=sha256:0d3b749bd8f86c56a61a694913408a558ccd335aeef84a0037cc80ebd303f07c

Observation 260f0122-a06f-458b-8013-6a447543a989 · outbound

This paper cites Large Language Model Critics for Execution-Free Evaluation of Code Changes.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Large Language Model Critics for Execution-Free Evaluation of Code Changes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.825915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.825915Z digest=sha256:74bf240f6f2dd966554f9a5cdfa26005c612da5c4a6d2a5ceffab0e11a829716

Observation 813acb66-9098-400d-ad9c-6e5b2a189a9b · outbound

This paper cites SWE-Bench+: Enhanced Coding Benchmark for LLMs.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-Bench+: Enhanced Coding Benchmark for LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.909889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.909889Z digest=sha256:04ac7078f8ea497b4ffe06cbc8f79b65fd2b8244d658ebd97bacac66001041af

Observation e418d1de-4524-4de4-8d16-e35ddd9d35d4 · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.987699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.987699Z digest=sha256:1b75092e6660a2b46cdef1e8651f33c7f9d13743e783707fc71195987ed9c8b9

Observation f139a15e-d5af-4b93-a480-42d0a4391586 · outbound

This paper cites Swt- bench: Testing and validating real-world bug-fixes with JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 15 code agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Swt- bench: Testing and validating real-world bug-fixes with JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 15 code agents,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.114116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.114116Z digest=sha256:44ae8c09598ae7bd721a369b363de6d85b86d744307db0c2ca364ec8a7a24590

Observation c9b9dfe3-3ea2-422c-ac78-ec6c47107748 · outbound

This paper cites ChatDev: Communicative Agents for Software Development.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ChatDev: Communicative Agents for Software Development

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.220143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.220143Z digest=sha256:e17813782d678a43ed5f5b931dcbaa43413ce9f3c3de7f48feb85504162c1d66

Observation 8794d234-6f6e-4577-af53-7fe360f293b7 · outbound

This paper cites Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.380320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.380320Z digest=sha256:d37c8f599f83bf499698b3f91c214b096b05949e2ed603803440a640b282d061

Observation d59ecd38-95a2-4267-810e-79a9d494d2f9 · outbound

This paper cites ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.492754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.492754Z digest=sha256:e6cf57671364695b81e061b5c536648c74ac4825896beaa5bdeb27bf13297a3f

Observation 64669182-ef2a-4c15-8438-b5fbcd23cb2e · outbound

This paper cites PyBench: Evaluating LLM Agent on various real-world coding tasks.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey PyBench: Evaluating LLM Agent on various real-world coding tasks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.652929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.652929Z digest=sha256:3d2e7d01b160889604b0d3fe2cab19ed0a6ec83ce220a64be3ffbf0bc8607bab

Observation cfa6a7b2-935b-4330-954d-b906b47d2b12 · outbound

This paper cites World of bits: An open- domain platform for web-based agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey World of bits: An open- domain platform for web-based agents,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.819234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.819234Z digest=sha256:671896aa47ab690d093f19b1eed19820a844bdfa1097e55f4e9102da34af53c8

Observation d6a0b09c-4fd8-442b-9c61-e7b4509ca6bb · outbound

This paper cites Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.979667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.979667Z digest=sha256:c613d7fd518ceda71b2c36755631b18594f280f3001bb641c2b93b6afc336da0

Observation 28696c09-9fb9-4fee-a302-0c0c984f2934 · outbound

This paper cites WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.147928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.147928Z digest=sha256:0064b91f85e25a204dc1c6a04ec9b0ea0b6377bb972f7db9f561cd4003d16e72

Observation a000fe7e-1b6c-4350-857e-19800d4bce97 · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mind2Web: Towards a Generalist Agent for the Web

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.200192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.200192Z digest=sha256:1f73714a250507b271b2c08284e303f0059e9578ce970432f7cf3c9b72edb658

Observation 32195e48-97c1-47c4-91bc-c333b7e5b9a8 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.208743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.208743Z digest=sha256:0c2d758ae5b270fcac14e8206fee1a30967f781d9cc58a2b2f0f0b946aafe743

Observation 16c6613a-41f3-443b-9ed3-312f57df1290 · outbound

This paper cites Weblinx: Real-world website navigation with multi-turn dialogue,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Weblinx: Real-world website navigation with multi-turn dialogue,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.213819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.213819Z digest=sha256:085e7603e5125c1cbd944defa05ea64dcbc5e5c9fa4b48310c673f93befd9f82

Observation b841e096-a322-4e05-b1e9-be5b2f421b13 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.217913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.217913Z digest=sha256:459c4be7cb91358bb6fed22cb773f70487a9cc83c162a319c9ac411f7fe19f23

Observation 2ab2c25f-bea8-497c-9224-96dacde30869 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.222011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.222011Z digest=sha256:f94201d0ed96602307d118463a853049afa160accfc5361e3bd0c307663c5fbe

Observation d16271a0-9a49-4f04-8f55-39f65d233eca · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.226003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.226003Z digest=sha256:14b94ddc66611f7c8e579be9999e90068ddcc21430504d4c94dbfe4492d76212

Observation 630117aa-28ce-4993-80b4-114cec9d210a · outbound

This paper cites WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.230822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.230822Z digest=sha256:5aa16d89faea1bb6461167e57acf0fb7503572261a158cad5de02d44fd661412

Observation 1186c679-70d6-4b64-88d5-97feafe877be · outbound

This paper cites Mmina: Benchmarking multihop multimodal internet agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mmina: Benchmarking multihop multimodal internet agents,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.235463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.235463Z digest=sha256:dbcc69611d15feb1ec76d401ee2605f82eabc48b62794a799804bd85e23fe82e

Observation f46d40d4-b103-481d-bb00-172bd36a716b · outbound

This paper cites AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.244078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.244078Z digest=sha256:113ffd5124a0fd32574d397c67420e57e5d8a288abf8d59839d68cbbe1a9e294

Observation b1cf617c-6338-4d35-ba44-47b9195a9237 · outbound

This paper cites WebCanvas: Benchmarking Web Agents in Online Environments.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WebCanvas: Benchmarking Web Agents in Online Environments

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.248544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.248544Z digest=sha256:6a03b64655946a8544a2083178bbb64bc593787cc36d2a82a7d42a08f904a25c

Observation 4662cf8d-cc72-4bb3-ad01-cc6b06f34723 · outbound

This paper cites ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.252726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.252726Z digest=sha256:d81bbfc89f652c382b29a2b57c06d45fbf9f7da3ed156d2cfeaa6420d47ca989

Observation 9c94bc8b-9eed-45ac-aa11-f94971c1673d · outbound

This paper cites VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.256618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.256618Z digest=sha256:1f05098c66f921c65c332774949fe4e1b322bb36a0768866cfdf0dcb54d46f8b

Observation 2078033d-11f9-4c4b-aa9a-50f46a02f291 · outbound

This paper cites Tur[k]ingBench: A Challenge Benchmark for Web Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Tur[k]ingBench: A Challenge Benchmark for Web Agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.260739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.260739Z digest=sha256:f690e9d181ae530575952a42e187dd99266089be19e75c4fd52d381fa8b6cf10

Observation f07bc4b2-1d36-420b-9040-ebb798af31bc · outbound

This paper cites BEARCUBS: A benchmark for computer-using web agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey BEARCUBS: A benchmark for computer-using web agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.265264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.265264Z digest=sha256:0d8cee8eb0a2be82cd5bfb3f6030e0ec492a0f113d7089ff4d26a15729fb3f72

Observation 5e668517-a3fd-4001-90bb-2b2b25d90438 · outbound

This paper cites TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.269330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.269330Z digest=sha256:a19852a3e74ab9c0ed11fb0f94f337bf7f07f3617be09bd4ff1dd3aead3c42ef

Observation 2d4cc33b-2ebb-423f-aa49-f9c9382bc14e · outbound

This paper cites Waber: Evaluating reliability and efficiency of web agents with existing benchmarks,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Waber: Evaluating reliability and efficiency of web agents with existing benchmarks,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.273386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.273386Z digest=sha256:d07f380bb45940c94b9063eccd710d68a0bc8668775b4176a696f2bbecae7c78

Observation 4b2137e9-47f1-41d3-ba0c-d1117c0951a1 · outbound

This paper cites Visualagentbench: Towards large multimodal models as visual foundation agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Visualagentbench: Towards large multimodal models as visual foundation agents,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.277558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.277558Z digest=sha256:0a83508fe3616e7b765ce4f2a541dea6560de53021685e7f2ff3c51259e277ef

Observation 462df6dc-603f-4a8e-bcde-f53f8d7eae9f · outbound

This paper cites An illusion of progress? assessing the current state of web agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey An illusion of progress? assessing the current state of web agents,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.286579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.286579Z digest=sha256:bd748604d3bc030c0a65490acc9b0d2270a1f8ff25315c116f85680b57645802

Observation 3f6c04dd-390b-4807-87bb-94cf5d86252d · outbound

This paper cites Available: https://arxiv.org/abs/2408.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Available: https://arxiv.org/abs/2408

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.282162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.282162Z digest=sha256:1ab3a881707b3238044417af12c302be3aa81bddfde74b00849fdb2a120113a0

Observation 3a93fdb7-817a-49d1-b70e-8ece429887cd · outbound

This paper cites OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.295587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.295587Z digest=sha256:1d43466be74e251b2f8c3bff02602a4081de1700576b887fccc9b9480af949ca

Observation ef6a25c2-4649-443a-b98a-a168c8efa59a · outbound

This paper cites REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.291115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.291115Z digest=sha256:870023bc3b0e926f6d6eea8d01f10dbcd8e38e2966b2417e7d070bdd583406c9

Observation 3bac919f-dc2c-4aae-ba7e-4e4107315075 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real com- puter environments,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Osworld: Benchmarking multimodal agents for open-ended tasks in real com- puter environments,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.308312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.308312Z digest=sha256:5a6b9cbca4995eea355540ab9cc4aa7f2ebc3e39ad31d31da7f7298de7d08f6d

Observation 7f0563c7-aef4-4f8c-a779-69619ed51102 · outbound

This paper cites Chatshop: Interactive information seeking with language agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Chatshop: Interactive information seeking with language agents,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.300016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.300016Z digest=sha256:3fea86f39ea3149bdb03f9d0ba3b63c5ef11221f9677df223ab05123d206a0c3

Observation dda98239-a6b9-4277-95b6-c42b36d3fad9 · outbound

This paper cites Available: https://arxiv.org/abs/2404.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Available: https://arxiv.org/abs/2404

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.303915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.303915Z digest=sha256:b194ae2f7bacf72620af4cb6b8c5818d812a04ee569787337385c2f265ecb247

Observation b19e26d6-6b73-4def-8ebd-53bbd598a4f0 · outbound

This paper cites Omniact: A dataset and benchmark for enabling multimodal generalist JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 16 autonomous agents for desktop and web,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Omniact: A dataset and benchmark for enabling multimodal generalist JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 16 autonomous agents for desktop and web,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.320520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.320520Z digest=sha256:a76ae355123eb76480ff5421a91ba8ef2b15979b84ff1aae27075c9d969abfb4

Observation a74440e1-7362-4698-933a-030d5e0788dd · outbound

This paper cites Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.312139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.312139Z digest=sha256:5a2a1c37e9e32ba0cf4b0e2e43e69e7c147c19ca3547b1214607b0cf6fac05b7

Observation 926aa596-72e1-4422-8377-9852b44ba4a2 · outbound

This paper cites AgentStudio: A Toolkit for Building General Virtual Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AgentStudio: A Toolkit for Building General Virtual Agents

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.316341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.316341Z digest=sha256:74e04606788e111138e99587cebf6d1f24a2ba6d7703c737ecaffe42ef16f519

Observation 1161f127-6a82-41d7-a01d-7c7dc859b437 · outbound

This paper cites Mapping Natural Language Instructions to Mobile UI Action Sequences.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mapping Natural Language Instructions to Mobile UI Action Sequences

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.333127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.333127Z digest=sha256:2976bc1245f479d682b3b368f074ff8d85c9e9740d273e74ec72c8d6ece5cde2

Observation 7c9775cf-fff2-41b9-8e65-cec65a483072 · outbound

This paper cites OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.324265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.324265Z digest=sha256:bc1e9df1fe3453dd6e68286a717764de9bb39b8cc6b35e91d5debca85c9c04d8

Observation 6c4ce685-e80a-427b-8709-453443680570 · outbound

This paper cites PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.328784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.328784Z digest=sha256:64731336c307bfa30ebe660bb0776285eae5544f8e209560afaff808abfc4b3f

Observation c24c036c-e0ba-44ec-92ac-ba87eb081604 · outbound

This paper cites Androidinthewild: A large-scale dataset for android device control,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Androidinthewild: A large-scale dataset for android device control,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.346134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.346134Z digest=sha256:7902b98f5da487f7df18837e136b5662089eeb5423b9e89d13af3a00d227370a

Observation 53c92253-b41c-4aeb-bb37-9ff2de041d01 · outbound

This paper cites UGIF: UI Grounded Instruction Following.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey UGIF: UI Grounded Instruction Following

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.337239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.337239Z digest=sha256:e3155dbb92bbe249150f22ea5f5b21a440e21a6fd91a753e156bd4c5e01a75aa

Observation b16d9ff5-45a8-41e3-84d9-6a62c76dd8ad · outbound

This paper cites Mobile App Tasks with Iterative Feedback (MoTIF): Addressing Task Feasibility in Interactive Visual Environments.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mobile App Tasks with Iterative Feedback (MoTIF): Addressing Task Feasibility in Interactive Visual Environments

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.341889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.341889Z digest=sha256:6fad702215da603ff6bd99e3e6700b2ac8d53098c0eecccdc95cb8eaa2a541f4

Observation c516cf87-86ac-4af6-ad50-884687aaba0e · outbound

This paper cites Benchmarking Mobile Device Control Agents across Diverse Configurations.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Benchmarking Mobile Device Control Agents across Diverse Configurations

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.358520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.358520Z digest=sha256:bfce3fc91c4b15dc75f0fc6924b92075a38450dced2d958bf67736461612273b

Observation c35fdbb1-a917-4fae-82f7-bb9d1f4e8f26 · outbound

This paper cites On the effects of data scale on ui control agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey On the effects of data scale on ui control agents,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.349977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.349977Z digest=sha256:9a3b2ccc1bf26d0115840d8187e467065c8c2c72f06aba791c19dbd15233b3e3

Observation 597d503c-9e1f-4749-b85a-2961a22930c9 · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.353842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.353842Z digest=sha256:013e15e6c2fd5013152726c196819a642a3a897aabe7b0435d50c7331892956f

Observation 662a50e3-fee9-4a80-a395-7d688686b468 · outbound

This paper cites A3: Android agent arena for mobile gui agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey A3: Android agent arena for mobile gui agents,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.371285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.371285Z digest=sha256:4d314fa3b8378e6ffcfd8a8b6870ead22119a005c107d042b889830e18a4f396

Observation d43916d4-6038-4196-9424-1085d5bd84e3 · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.362479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.362479Z digest=sha256:ba6c8a9cb74773501358ad1fb93b1fbb55d15d0ba06bbee020f4406a0820517b

Observation 8946c67a-a648-4d4d-a8b3-ed1c6b20f1df · outbound

This paper cites Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.366743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.366743Z digest=sha256:b0010a4321dad5eb59f0e3d1dfaf95d710c1f4651c1cd6843a305cbe86834e1a

Observation 7b013bde-7dff-41c4-a047-955f874f7890 · outbound

This paper cites Mobilesafetybench: Evaluating safety of autonomous agents in mobile device control,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mobilesafetybench: Evaluating safety of autonomous agents in mobile device control,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.384014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.384014Z digest=sha256:65e12293e5284a0e3d9702958f18d1f92a8f70be199f9e5d0068cc4e8cfdeaf0

Observation 61b6aa15-c081-4984-9703-32da31ddb09a · outbound

This paper cites Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.375445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.375445Z digest=sha256:4b02c31e705c7b14d34a0230b4356068562cbecf8f45e841dd166b776ed4169a

Observation 880efb3e-e422-4412-8589-7df047937b48 · outbound

This paper cites Understanding the weakness of large language model agents within a complex android environment,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Understanding the weakness of large language model agents within a complex android environment,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.379903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.379903Z digest=sha256:3f65a967fae30f0a231f935acfe54744dcf1a22c6ce51289492e2a6f123b4d70

Observation e51160c5-3605-4bfa-a3ea-9c04408c3173 · outbound

This paper cites Llamatouch: A faithful and scalable testbed for mobile ui task automation,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Llamatouch: A faithful and scalable testbed for mobile ui task automation,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.396863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.396863Z digest=sha256:761131497697196a4526e2634d10a3cb5064ca537d5e248fa4e18e97dfd4f584

Observation 8277fa0c-d3c7-465f-bcbe-5fb718485389 · outbound

This paper cites MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.388589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.388589Z digest=sha256:69648b8a917c7e25675221b8d9ceac19465f19dc53982bf23ea98d6e7a4bdb1c

Observation b46f2ab7-81a9-45b5-9190-efb3a3381d3d · outbound

This paper cites Spa-bench: A comprehensive benchmark for smartphone agent evaluation,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Spa-bench: A comprehensive benchmark for smartphone agent evaluation,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.392685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.392685Z digest=sha256:0e0db14ae7aacbeee95fb35a9269ef1b82416e0db5d19c234be21cef9a5e127d

Observation b9249483-3387-4fc9-807e-6b48abb01468 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.409750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.409750Z digest=sha256:14d35cce6274a211b8c63e64df980fa35e3ee2814a879ca1a187b508c84f6555

Observation fb119f8e-8e28-43f3-8f4a-0b5c711da058 · outbound

This paper cites AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.401380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.401380Z digest=sha256:8855c5a1b29c0f1464425c755bd5d83a0e34e2d11c25b22b3a43494b260abe31

Observation 15d98393-a742-4dc2-aff9-a6ba44f84d3d · outbound

This paper cites Autoeval: A practical framework for autonomous evaluation of mobile agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Autoeval: A practical framework for autonomous evaluation of mobile agents,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.405415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.405415Z digest=sha256:d9325584199af07bee77e034210eb962480f492381bc5b530cfe3b08dd1d0a4a

Observation 0fdde1d4-9b54-4d4e-b018-1050ccf0519e · outbound

This paper cites Qasa: advanced question answer- ing on scientific articles,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Qasa: advanced question answer- ing on scientific articles,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.422789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.422789Z digest=sha256:6cf3c0ef6243873b90d525bff346a1d9f72e5a7510f49e3dbb66396a47d4c89a

Observation 791bbd92-ebc3-44b0-beff-a12e5c7f077b · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.414057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.414057Z digest=sha256:7d704e43e9a51fb3c0449f26774e32e658b5d29c94f8d295a5c2a178715f6c16

Observation e72e4600-702a-421d-b460-bf671fe7bd72 · outbound

This paper cites A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.418439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.418439Z digest=sha256:6db97b21ff57c6952867e8846f92c3eeb852566185b8127e63446c471647d5e0

Observation 3ac21333-8046-40ab-83e6-38322acfbae1 · outbound

This paper cites MS2: Multi-Document Summarization of Medical Studies.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MS2: Multi-Document Summarization of Medical Studies

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.435832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.435832Z digest=sha256:c00f0d8719c2983c662cb68f15ce05aef0b1f5076c79dcedad1a588717280aa2

Observation da3053c0-5c76-460c-9b66-332a6e24514e · outbound

This paper cites Automated Focused Feedback Generation for Scientific Writing Assistance.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Automated Focused Feedback Generation for Scientific Writing Assistance

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.426701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.426701Z digest=sha256:5cfc36d48815b3f65a2da54123762283f194782020123aec07be8433c3dc8dc7

Observation 028d9b46-f4d3-4629-9976-e3ae1e18f5ef · outbound

This paper cites AAAR-1.0: Assessing AI's Potential to Assist Research.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AAAR-1.0: Assessing AI's Potential to Assist Research

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.431100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.431100Z digest=sha256:5f8fb68cc54eb87d269aa4533168b6491b764644a5921abfa486f637cfb30be5

Observation c1ca246e-b03b-465f-9f0b-cfa9a805e722 · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.449455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.449455Z digest=sha256:71596b3df74e620721518787ea0fdd668974b94c27125dc57013c89107e5cd16

Observation 03d5c5a1-2f1b-4f20-af18-7af2816ab996 · outbound

This paper cites LAB-Bench: Measuring Capabilities of Language Models for Biology Research.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey LAB-Bench: Measuring Capabilities of Language Models for Biology Research

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.439930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.439930Z digest=sha256:07b95703f94f15c5066316bb60c5a993f0023dda2afee3c7289236f533e3dfaa

Observation bb95e92e-a272-40c9-b8f7-2e872a71d6f5 · outbound

This paper cites ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.444970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.444970Z digest=sha256:278bda8bb17691acceaa784697a77db6652bfd310c5407ab8cbb82a8d40880da

Observation 49e1a027-28d6-459e-96e5-bb96895989a5 · outbound

This paper cites DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.463504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.463504Z digest=sha256:9f0d824cc2180634d185f247b3431cfd38c8e0ea005188454a17ea73093faa85

Observation fec61084-7634-407c-97d2-09d132ee4d14 · outbound

This paper cites PaperQA: Retrieval-Augmented Generative Agent for Scientific Research.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey PaperQA: Retrieval-Augmented Generative Agent for Scientific Research

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.453993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.453993Z digest=sha256:d356104d832ec7b38029574d43a9047cccb56c0dfc62724f4da15aa6a1320b92

Observation 27b08d10-0e8a-46ed-84aa-a9b8314678d3 · outbound

This paper cites ScienceWorld: Is your Agent Smarter than a 5th Grader?.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ScienceWorld: Is your Agent Smarter than a 5th Grader?

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.458614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.458614Z digest=sha256:e3b814fb1e183d3a96d1115f3aea9c31896d133f647bd9665420c91c5dac49e6

Observation 043cd23b-d07a-4272-8573-e9748501667f · outbound

This paper cites Mlagent- bench: Evaluating language agents on machine learning experimentation, 2024,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mlagent- bench: Evaluating language agents on machine learning experimentation, 2024,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.476908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.476908Z digest=sha256:17abf0fc7d9394eb3880fa240d5396f2f1c24e779966980dacee9cbb963909a2

Observation 6a271e68-1308-4909-8500-8452e86d43b9 · outbound

This paper cites ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.467715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.467715Z digest=sha256:15a4bd3056c734ae15dad1aa2a5a9383d8878cd6fcff6d237d738e1717a84b27

Observation 79bccf71-c36d-4f0c-ba6c-fac4d942ced4 · outbound

This paper cites MLGym: A New Framework and Benchmark for Advancing AI Research Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.472486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.472486Z digest=sha256:e28f09f863be8a56044790cc07ba81c82419c69a6fc4582947d874d75a2d0d18

Observation 3b944511-a60a-467d-87f4-4285fa1271aa · outbound

This paper cites DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.491104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.491104Z digest=sha256:05d1f06adc3aee17956447d73d198f7686e907b39ff177a5a2a5531a9a5f386e

Pith citing papers

Observation c18e21d6-4917-4748-9379-93e8566336c1 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

Reference 154

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:14.879801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:d5efd5a17c25f975ef06ebec54d00277533e3ddf2f7fea6ca2e42c768fe4871f

Observation 3e2c672f-c607-4eae-81a4-f70ce2797404 · inbound

Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering cites this paper.

Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:21:00.461563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T17:40:14.733882Z digest=sha256:d9b22c6f358a1159e6ec42d08846df3d330c4f1065a940acf236f89f2d4a9f25

Observation 0c5748b6-1302-4008-9625-1e82f498a169 · inbound

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory cites this paper.

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:13:40.511046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T19:11:15.761831Z digest=sha256:1cc7af40eacf5b953ba4a0d66437fd9fca3476b16300f2b73cdace1e17abfd94

Observation 068b5531-8028-47f4-a264-5b52127cbf69 · inbound

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents cites this paper.

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:53:26.389766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T12:52:20.911788Z digest=sha256:39c817bcc590e83e909599a30ad510054e0cfa9512c0d6995e5d9901846e7e7a