Pith. sign in

Paper Citation Record · LEDGER

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

As of 19 August 2026, this Paper Citation Record lists 100 of 209 outbound references and 4 inbound Pith citation observations for arXiv:2506.11102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11102 v1

Coverage vector

measured 100 of 209 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:00:19.491104Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T12:52:20.911788Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:53:26.388274Z

Reference resolution

100 of 209 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9f6a353-6bd8-4e4c-ad86-066f7a6b64b0 · outbound

This paper cites Attention Is All You Need.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Attention Is All You Need

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:15.697976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:15.697976Z digest=sha256:88c36cea6e11a1948dd7d1eaedf0da0e1bda863adbe8a6621bd6bb565f631da7

Observation 20a874fb-0745-42ba-ae6c-4dbe13914100 · outbound

This paper cites GPT-4 Technical Report.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:15.738449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:15.738449Z digest=sha256:34c64578550414d94932d64bf1a4299924410665c09a963fbd085b9ffdb0cefd

Observation e4204d53-69d2-4844-9220-e09b75a3b614 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:15.823334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:15.823334Z digest=sha256:b71ef82d729ee0d8731e4d19d8860234ba93f9c4bf3e6f4ef80637472f7c010d

Observation bd209cb1-83af-4bed-a32e-42fe2a9e3eef · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Gemini: A Family of Highly Capable Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:15.888246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:15.888246Z digest=sha256:10a3d679607bd1e9d9d9b20faeca59649314899bbcf698aba2d11dd3aea46378

Observation 1f1aea1c-d210-439c-826e-ad41f4eb41da · outbound

This paper cites Qwen Technical Report.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:15.977297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:15.977297Z digest=sha256:cbba90ddba375914a23ff1aa9713976631b67da6ba7ce2754983d0a1e4c4fc03

Observation 9955fec4-7265-4424-801f-20e4dbd9cb3d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.066521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.066521Z digest=sha256:a229541de6d5f6724cc9c7f55e5a7d333010a062c61b8505b138c5876c3802d7

Observation 69fa2201-e8bc-4705-b486-9f157ea47c3b · outbound

This paper cites Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.156522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.156522Z digest=sha256:350721454d39978f38656ce9be5f3ab83d2d890368b9aff79cca6fbf94ab4007

Observation bce1507e-654c-4c19-9d1f-29057183678a · outbound

This paper cites Beyond Browsing: API-Based Web Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Beyond Browsing: API-Based Web Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.253072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.253072Z digest=sha256:700086982e7651f0c7c16cf9709ea64c1b0159f328d7c4ceb15a938496a10c8c

Observation b486343d-eeb8-4628-8e1f-34b2dbe3db46 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.334469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.334469Z digest=sha256:5dc40a0aaa71762d6537bd922eb29b5adad5a7f5765a9c0f811b22701b4be179

Observation 25195080-ca4c-4c89-b675-24aa4716ba9c · outbound

This paper cites A Survey on Evaluation of Large Language Models.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey A Survey on Evaluation of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.424040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.424040Z digest=sha256:292f16008c1d90bc30c7ef2f5059d3f3cbfd63a64621baa7376a4615c978ab09

Observation 67aa1f7b-6ad2-4d05-81d8-0bebbd80e45a · outbound

This paper cites A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.519703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.519703Z digest=sha256:4b286326affb4b4023799939bdd84cc970ce8b3d2a98c1272d770f761dcd3608

Observation a6a852ba-7372-4622-9b0c-2306905712da · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Large Language Models: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.603597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.603597Z digest=sha256:95fc214d000a6aac65ef96cfc15f769932e4e09e89e082493f56d39428b8e59d

Observation e22630de-a20b-4ab0-a2ba-22e7415e6c38 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Large Language Models Trained on Code

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.716446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.716446Z digest=sha256:0efa5f30e589a6e9801a172ec1914867e464fa79bae0ca44dec3d261ec015fd6

Observation 6c35e0c2-bf2a-4915-8d74-55c1a2c5b3a0 · outbound

This paper cites ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.785966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.785966Z digest=sha256:747f4dcdbb574b366ca24e84cb9db23fda67c1a15f603fd5c1d52df303ea6251

Observation 8725aa85-b4ef-4110-94c3-4875641fb7c4 · outbound

This paper cites Codereval: A benchmark of pragmatic code generation with generative pre-trained models,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Codereval: A benchmark of pragmatic code generation with generative pre-trained models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.847762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.847762Z digest=sha256:08d105c2c1ce892659ff3708e1f1dc330d2cbd42b14bfad4a56a3590079abf76

Observation a9a44159-a09f-4e11-8a89-a47582ee1cb6 · outbound

This paper cites Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.923044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.923044Z digest=sha256:5ca95097caedf60fd092b157782aeddfd218635f6687d52803fa60b3adb60391

Observation bb73fe5c-1553-4d57-80b8-eec45055522d · outbound

This paper cites CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.038881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.038881Z digest=sha256:947f8de83fb99e002436f5826b6206b36e5fc812eb5d0f1f0476c6ce7cf45c5d

Observation 6b00341c-253f-4bcc-9dbe-67c6ddb7b03f · outbound

This paper cites CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.129393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.129393Z digest=sha256:7ea84e7a36a7a8e369d41f937a967de3d2d7297b8d0ce960047d0bcd631fa07a

Observation 4f4399a5-4ada-423b-b1de-9ac0369b8db8 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.220729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.220729Z digest=sha256:03695d496e797fd1f0dff9f50106f2fe42e4a818d2228eb380aca11f9bbea4d0

Observation fb373681-2cd1-4688-80b1-e656e8507ee3 · outbound

This paper cites Natural Language to Code Generation in Interactive Data Science Notebooks.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Natural Language to Code Generation in Interactive Data Science Notebooks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.290883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.290883Z digest=sha256:8075960c6d7a617db021539b6335d146fbadd348c86d46f20f654a65fd82ad16

Observation 38c40404-8e4a-442d-92f3-7c537a4581bd · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.379245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.379245Z digest=sha256:d6fd1b75c8cd7e42c29f617164feaaa91a2697a3bc6e2de2d4fee89e5c4dbb13

Observation 9a21d675-1c13-4335-81f9-84ac9f795873 · outbound

This paper cites SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.468272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.468272Z digest=sha256:184ab2b5328268639c3eb4b9c7425d5c443731011b20880a98dedb525aed5247

Observation a1309315-93d3-4e28-9d01-2ab65f09bbd6 · outbound

This paper cites SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.591067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.591067Z digest=sha256:26b0bf281cc61f3e54227f199f40eaea05cfd0c34b656de1d879eff9538ca1df

Observation a6f4fa44-f487-4298-89f8-fc866145363d · outbound

This paper cites SWE-bench-java: A GitHub Issue Resolving Benchmark for Java.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-bench-java: A GitHub Issue Resolving Benchmark for Java

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.668524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.668524Z digest=sha256:170c23597e5dd19c900f4cb26a6e98e1b4acd954f9038e79c05531f882a4bd92

Observation ac6f0cbc-2bd6-4dc9-8551-cdf4d36fc239 · outbound

This paper cites Swe arena: An open evaluation platform for automated software engineering,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Swe arena: An open evaluation platform for automated software engineering,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.755574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.755574Z digest=sha256:0d3b749bd8f86c56a61a694913408a558ccd335aeef84a0037cc80ebd303f07c

Observation 260f0122-a06f-458b-8013-6a447543a989 · outbound

This paper cites Large Language Model Critics for Execution-Free Evaluation of Code Changes.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Large Language Model Critics for Execution-Free Evaluation of Code Changes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.825915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.825915Z digest=sha256:c1aabfb208bc02a5515b09dece3e47be1449acc382c1c5ab0a21d67acbe1f247

Observation 813acb66-9098-400d-ad9c-6e5b2a189a9b · outbound

This paper cites SWE-Bench+: Enhanced Coding Benchmark for LLMs.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-Bench+: Enhanced Coding Benchmark for LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.909889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.909889Z digest=sha256:19e27f4479a05ad52f4742bbb2d7723f4d7db4bd4613aefd4b70f56ce0e7e684

Observation e418d1de-4524-4de4-8d16-e35ddd9d35d4 · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.987699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.987699Z digest=sha256:c0e792e3e078d3672c7cd2146932ef781dea652fde7b539940bd8632c072c06e

Observation f139a15e-d5af-4b93-a480-42d0a4391586 · outbound

This paper cites Swt- bench: Testing and validating real-world bug-fixes with JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 15 code agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Swt- bench: Testing and validating real-world bug-fixes with JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 15 code agents,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.114116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.114116Z digest=sha256:44ae8c09598ae7bd721a369b363de6d85b86d744307db0c2ca364ec8a7a24590

Observation c9b9dfe3-3ea2-422c-ac78-ec6c47107748 · outbound

This paper cites ChatDev: Communicative Agents for Software Development.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ChatDev: Communicative Agents for Software Development

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.220143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.220143Z digest=sha256:07b936778085c6b371166f80d52024d108d8cd5ac38b2496fb325078e663dc95

Observation 8794d234-6f6e-4577-af53-7fe360f293b7 · outbound

This paper cites Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.380320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.380320Z digest=sha256:1e4711c2d6aa98adfabf760dc743abbb475a7237b8e75a13d2cb39048596bfb3

Observation d59ecd38-95a2-4267-810e-79a9d494d2f9 · outbound

This paper cites ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.492754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.492754Z digest=sha256:c06f448bc1783c88529d0ac739abc93e9300e8b5b80d08368a740666827c454e

Observation 64669182-ef2a-4c15-8438-b5fbcd23cb2e · outbound

This paper cites PyBench: Evaluating LLM Agent on various real-world coding tasks.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey PyBench: Evaluating LLM Agent on various real-world coding tasks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.652929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.652929Z digest=sha256:41098c6c635181e9fca1fff69863f9e34a73b567617deac1e2365a39b9917022

Observation cfa6a7b2-935b-4330-954d-b906b47d2b12 · outbound

This paper cites World of bits: An open- domain platform for web-based agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey World of bits: An open- domain platform for web-based agents,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.819234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.819234Z digest=sha256:671896aa47ab690d093f19b1eed19820a844bdfa1097e55f4e9102da34af53c8

Observation d6a0b09c-4fd8-442b-9c61-e7b4509ca6bb · outbound

This paper cites Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.979667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.979667Z digest=sha256:0b622291e70468a76e13c2586e4ad21bb1acf4698157b21ce935dbb64a5b270b

Observation 28696c09-9fb9-4fee-a302-0c0c984f2934 · outbound

This paper cites WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.147928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.147928Z digest=sha256:23cb1402d57f4b8a918f7dbe150433689f9f90fb86540613781eeea6932a3526

Observation a000fe7e-1b6c-4350-857e-19800d4bce97 · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mind2Web: Towards a Generalist Agent for the Web

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.200192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.200192Z digest=sha256:b43654c372ec02d0062ae10b3a2dbfebbdcd49e647e0d4692f573b2714df81aa

Observation 32195e48-97c1-47c4-91bc-c333b7e5b9a8 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.208743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.208743Z digest=sha256:c8be7cc8fe3d5b7623ce982fbd0ae9c6b4abf0f55aadbfa3bfe0638e1249c680

Observation 16c6613a-41f3-443b-9ed3-312f57df1290 · outbound

This paper cites Weblinx: Real-world website navigation with multi-turn dialogue,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Weblinx: Real-world website navigation with multi-turn dialogue,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.213819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.213819Z digest=sha256:085e7603e5125c1cbd944defa05ea64dcbc5e5c9fa4b48310c673f93befd9f82

Observation b841e096-a322-4e05-b1e9-be5b2f421b13 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.217913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.217913Z digest=sha256:d79cf355f3bd19310af88a01a26e5fe42be12878cdb7f1808c39f6992d5381b1

Observation 2ab2c25f-bea8-497c-9224-96dacde30869 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.222011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.222011Z digest=sha256:157afb4f431df5b1325591fc03b99a3990e7353e6058785f660f7d45c3b67b0b

Observation d16271a0-9a49-4f04-8f55-39f65d233eca · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.226003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.226003Z digest=sha256:14b94ddc66611f7c8e579be9999e90068ddcc21430504d4c94dbfe4492d76212

Observation 630117aa-28ce-4993-80b4-114cec9d210a · outbound

This paper cites WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.230822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.230822Z digest=sha256:211c492448c96a01119d47ed54cf75efa259d990e5b8bb9ac7d9c7b4406b4e92

Observation 1186c679-70d6-4b64-88d5-97feafe877be · outbound

This paper cites Mmina: Benchmarking multihop multimodal internet agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mmina: Benchmarking multihop multimodal internet agents,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.235463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.235463Z digest=sha256:dbcc69611d15feb1ec76d401ee2605f82eabc48b62794a799804bd85e23fe82e

Observation f46d40d4-b103-481d-bb00-172bd36a716b · outbound

This paper cites AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.244078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.244078Z digest=sha256:d469f73bbd96b41525b4b417827979df5e14a344b7c36968deee4d38eeeda660

Observation b1cf617c-6338-4d35-ba44-47b9195a9237 · outbound

This paper cites WebCanvas: Benchmarking Web Agents in Online Environments.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey WebCanvas: Benchmarking Web Agents in Online Environments

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.248544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.248544Z digest=sha256:6a03b64655946a8544a2083178bbb64bc593787cc36d2a82a7d42a08f904a25c

Observation 4662cf8d-cc72-4bb3-ad01-cc6b06f34723 · outbound

This paper cites ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.252726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.252726Z digest=sha256:6791f715498549a851134700014d0c0d2213cbd68e7d5831f3e19f6ed126d1d0

Observation 9c94bc8b-9eed-45ac-aa11-f94971c1673d · outbound

This paper cites VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.256618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.256618Z digest=sha256:ce7a25c93276848f809bcb23474894acb43fdc065f695325495f8458b0893376

Observation 2078033d-11f9-4c4b-aa9a-50f46a02f291 · outbound

This paper cites Tur[k]ingBench: A Challenge Benchmark for Web Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Tur[k]ingBench: A Challenge Benchmark for Web Agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.260739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.260739Z digest=sha256:52708ae4e628c0007ead278e209beaf12957ce9d7e54e88b0de149b58989e877

Observation f07bc4b2-1d36-420b-9040-ebb798af31bc · outbound

This paper cites BEARCUBS: A benchmark for computer-using web agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey BEARCUBS: A benchmark for computer-using web agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.265264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.265264Z digest=sha256:d825738146f62be23fd903efe8bcffab82b93d2c219e4e8956168c069a0c10d4

Observation 5e668517-a3fd-4001-90bb-2b2b25d90438 · outbound

This paper cites TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.269330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.269330Z digest=sha256:e3d031ea2f37e0c84a44e5cb5496eaebaed6c533f5e8ce6ef2936b4fb9d8c0c4

Observation 2d4cc33b-2ebb-423f-aa49-f9c9382bc14e · outbound

This paper cites Waber: Evaluating reliability and efficiency of web agents with existing benchmarks,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Waber: Evaluating reliability and efficiency of web agents with existing benchmarks,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.273386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.273386Z digest=sha256:d07f380bb45940c94b9063eccd710d68a0bc8668775b4176a696f2bbecae7c78

Observation 4b2137e9-47f1-41d3-ba0c-d1117c0951a1 · outbound

This paper cites Visualagentbench: Towards large multimodal models as visual foundation agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Visualagentbench: Towards large multimodal models as visual foundation agents,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.277558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.277558Z digest=sha256:0a83508fe3616e7b765ce4f2a541dea6560de53021685e7f2ff3c51259e277ef

Observation 462df6dc-603f-4a8e-bcde-f53f8d7eae9f · outbound

This paper cites An illusion of progress? assessing the current state of web agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey An illusion of progress? assessing the current state of web agents,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.286579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.286579Z digest=sha256:bd748604d3bc030c0a65490acc9b0d2270a1f8ff25315c116f85680b57645802

Observation 3f6c04dd-390b-4807-87bb-94cf5d86252d · outbound

This paper cites Available: https://arxiv.org/abs/2408.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Available: https://arxiv.org/abs/2408

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.282162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.282162Z digest=sha256:1ab3a881707b3238044417af12c302be3aa81bddfde74b00849fdb2a120113a0

Observation 3a93fdb7-817a-49d1-b70e-8ece429887cd · outbound

This paper cites OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.295587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.295587Z digest=sha256:65d430a08b589aec0f1e49e43f33dea99233505a3273b4b6384eb811aec192c1

Observation ef6a25c2-4649-443a-b98a-a168c8efa59a · outbound

This paper cites REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.291115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.291115Z digest=sha256:7a1d50fe05f007a21f2d8529d7edfc221a2dfbe336a00c29946c3e75f9185e07

Observation 3bac919f-dc2c-4aae-ba7e-4e4107315075 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real com- puter environments,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Osworld: Benchmarking multimodal agents for open-ended tasks in real com- puter environments,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.308312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.308312Z digest=sha256:5a6b9cbca4995eea355540ab9cc4aa7f2ebc3e39ad31d31da7f7298de7d08f6d

Observation 7f0563c7-aef4-4f8c-a779-69619ed51102 · outbound

This paper cites Chatshop: Interactive information seeking with language agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Chatshop: Interactive information seeking with language agents,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.300016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.300016Z digest=sha256:3fea86f39ea3149bdb03f9d0ba3b63c5ef11221f9677df223ab05123d206a0c3

Observation dda98239-a6b9-4277-95b6-c42b36d3fad9 · outbound

This paper cites Available: https://arxiv.org/abs/2404.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Available: https://arxiv.org/abs/2404

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.303915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.303915Z digest=sha256:b194ae2f7bacf72620af4cb6b8c5818d812a04ee569787337385c2f265ecb247

Observation b19e26d6-6b73-4def-8ebd-53bbd598a4f0 · outbound

This paper cites Omniact: A dataset and benchmark for enabling multimodal generalist JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 16 autonomous agents for desktop and web,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Omniact: A dataset and benchmark for enabling multimodal generalist JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 16 autonomous agents for desktop and web,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.320520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.320520Z digest=sha256:a76ae355123eb76480ff5421a91ba8ef2b15979b84ff1aae27075c9d969abfb4

Observation a74440e1-7362-4698-933a-030d5e0788dd · outbound

This paper cites Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.312139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.312139Z digest=sha256:333b6f0221390c42412d1940b37789252ccd6f0201b705599780b979ab072dab

Observation 926aa596-72e1-4422-8377-9852b44ba4a2 · outbound

This paper cites AgentStudio: A Toolkit for Building General Virtual Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AgentStudio: A Toolkit for Building General Virtual Agents

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.316341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.316341Z digest=sha256:0f3feb7a499f2afb1adfdbe194db723242169008b663aabb1b96e73c49cda4c6

Observation 1161f127-6a82-41d7-a01d-7c7dc859b437 · outbound

This paper cites Mapping Natural Language Instructions to Mobile UI Action Sequences.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mapping Natural Language Instructions to Mobile UI Action Sequences

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.333127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.333127Z digest=sha256:ca172c354dcfd1df4058a2ed3e0183df026c2260e3ff29d9a3419d5aa29eee52

Observation 7c9775cf-fff2-41b9-8e65-cec65a483072 · outbound

This paper cites OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.324265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.324265Z digest=sha256:7a18609c7c09fffd773c18c47b96c969a788424476f80028992f0def3d6302d1

Observation 6c4ce685-e80a-427b-8709-453443680570 · outbound

This paper cites PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.328784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.328784Z digest=sha256:6a5b69a25657b1e4cacb07513f3138a6842e224a62f8916ee4008f7a48969766

Observation c24c036c-e0ba-44ec-92ac-ba87eb081604 · outbound

This paper cites Androidinthewild: A large-scale dataset for android device control,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Androidinthewild: A large-scale dataset for android device control,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.346134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.346134Z digest=sha256:7902b98f5da487f7df18837e136b5662089eeb5423b9e89d13af3a00d227370a

Observation 53c92253-b41c-4aeb-bb37-9ff2de041d01 · outbound

This paper cites UGIF: UI Grounded Instruction Following.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey UGIF: UI Grounded Instruction Following

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.337239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.337239Z digest=sha256:4023f7fefe28d2e1cdb13c2095647ff7020e8dab0a124b15cbb8dd4c6b27c63b

Observation b16d9ff5-45a8-41e3-84d9-6a62c76dd8ad · outbound

This paper cites Mobile App Tasks with Iterative Feedback (MoTIF): Addressing Task Feasibility in Interactive Visual Environments.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mobile App Tasks with Iterative Feedback (MoTIF): Addressing Task Feasibility in Interactive Visual Environments

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.341889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.341889Z digest=sha256:f3b5d7d8ad665b95bdb8b8bd6cde11d7014857ae65a3e6a8a875461c58fd7c8a

Observation c516cf87-86ac-4af6-ad50-884687aaba0e · outbound

This paper cites Benchmarking Mobile Device Control Agents across Diverse Configurations.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Benchmarking Mobile Device Control Agents across Diverse Configurations

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.358520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.358520Z digest=sha256:38daa0cde176fce8074c41e07edd9be0884185458808a743b32b7124cf16db4a

Observation c35fdbb1-a917-4fae-82f7-bb9d1f4e8f26 · outbound

This paper cites On the effects of data scale on ui control agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey On the effects of data scale on ui control agents,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.349977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.349977Z digest=sha256:9a3b2ccc1bf26d0115840d8187e467065c8c2c72f06aba791c19dbd15233b3e3

Observation 597d503c-9e1f-4749-b85a-2961a22930c9 · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.353842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.353842Z digest=sha256:a93124141e5b770bccc3bb4bdd480772b45f1438072fa5221b2ce89fd38c17b2

Observation 662a50e3-fee9-4a80-a395-7d688686b468 · outbound

This paper cites A3: Android agent arena for mobile gui agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey A3: Android agent arena for mobile gui agents,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.371285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.371285Z digest=sha256:4d314fa3b8378e6ffcfd8a8b6870ead22119a005c107d042b889830e18a4f396

Observation d43916d4-6038-4196-9424-1085d5bd84e3 · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.362479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.362479Z digest=sha256:a44dd54753930724fe223c2c54a934a9768caab7bfd2db3fef6bdd4f24c92779

Observation 8946c67a-a648-4d4d-a8b3-ed1c6b20f1df · outbound

This paper cites Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.366743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.366743Z digest=sha256:94ee63b1ea2ff646ce03fff494ee7839775f007b2f383d66206d857c1485ba26

Observation 7b013bde-7dff-41c4-a047-955f874f7890 · outbound

This paper cites Mobilesafetybench: Evaluating safety of autonomous agents in mobile device control,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mobilesafetybench: Evaluating safety of autonomous agents in mobile device control,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.384014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.384014Z digest=sha256:65e12293e5284a0e3d9702958f18d1f92a8f70be199f9e5d0068cc4e8cfdeaf0

Observation 61b6aa15-c081-4984-9703-32da31ddb09a · outbound

This paper cites Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.375445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.375445Z digest=sha256:57128938bc920f9a931751cbb1caa97dadbe04d8f32beaf9a8fa0977d15c78ee

Observation 880efb3e-e422-4412-8589-7df047937b48 · outbound

This paper cites Understanding the weakness of large language model agents within a complex android environment,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Understanding the weakness of large language model agents within a complex android environment,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.379903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.379903Z digest=sha256:3f65a967fae30f0a231f935acfe54744dcf1a22c6ce51289492e2a6f123b4d70

Observation e51160c5-3605-4bfa-a3ea-9c04408c3173 · outbound

This paper cites Llamatouch: A faithful and scalable testbed for mobile ui task automation,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Llamatouch: A faithful and scalable testbed for mobile ui task automation,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.396863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.396863Z digest=sha256:761131497697196a4526e2634d10a3cb5064ca537d5e248fa4e18e97dfd4f584

Observation 8277fa0c-d3c7-465f-bcbe-5fb718485389 · outbound

This paper cites MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.388589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.388589Z digest=sha256:011d638e8af3b5bd15aa64244736a29a19aff87cc7d1405b2a2dd294ad9a1e81

Observation b46f2ab7-81a9-45b5-9190-efb3a3381d3d · outbound

This paper cites Spa-bench: A comprehensive benchmark for smartphone agent evaluation,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Spa-bench: A comprehensive benchmark for smartphone agent evaluation,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.392685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.392685Z digest=sha256:0e0db14ae7aacbeee95fb35a9269ef1b82416e0db5d19c234be21cef9a5e127d

Observation b9249483-3387-4fc9-807e-6b48abb01468 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.409750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.409750Z digest=sha256:5221f211c14a21a96ff3dabce363ec231a3e50df79549ffe1cfc5e017c32fac2

Observation fb119f8e-8e28-43f3-8f4a-0b5c711da058 · outbound

This paper cites AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.401380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.401380Z digest=sha256:aa2dbeb77f314e839c40d89cb48ba06826667aeea825f940862b60f2285ae9af

Observation 15d98393-a742-4dc2-aff9-a6ba44f84d3d · outbound

This paper cites Autoeval: A practical framework for autonomous evaluation of mobile agents,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Autoeval: A practical framework for autonomous evaluation of mobile agents,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.405415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.405415Z digest=sha256:d9325584199af07bee77e034210eb962480f492381bc5b530cfe3b08dd1d0a4a

Observation 0fdde1d4-9b54-4d4e-b018-1050ccf0519e · outbound

This paper cites Qasa: advanced question answer- ing on scientific articles,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Qasa: advanced question answer- ing on scientific articles,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.422789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.422789Z digest=sha256:6cf3c0ef6243873b90d525bff346a1d9f72e5a7510f49e3dbb66396a47d4c89a

Observation 791bbd92-ebc3-44b0-beff-a12e5c7f077b · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.414057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.414057Z digest=sha256:e42e2a2f613346fd7527bbc4211f20e29df2cbf9f186d1211e04ff791ad962e7

Observation e72e4600-702a-421d-b460-bf671fe7bd72 · outbound

This paper cites A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.418439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.418439Z digest=sha256:461d7bd4f2988e7cdd5d4b1d395b27e3ee60f0b71104f84646a7a85440be384f

Observation 3ac21333-8046-40ab-83e6-38322acfbae1 · outbound

This paper cites MS2: Multi-Document Summarization of Medical Studies.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MS2: Multi-Document Summarization of Medical Studies

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.435832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.435832Z digest=sha256:c636d47979e53859622125970f7ff14c1c3f035dc8096775a58a3e20cfc65809

Observation da3053c0-5c76-460c-9b66-332a6e24514e · outbound

This paper cites Automated Focused Feedback Generation for Scientific Writing Assistance.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Automated Focused Feedback Generation for Scientific Writing Assistance

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.426701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.426701Z digest=sha256:29af743dd684ac79035623f5c46c6bb37fc46bfd9b21c1fa4e4c72c07343b023

Observation 028d9b46-f4d3-4629-9976-e3ae1e18f5ef · outbound

This paper cites AAAR-1.0: Assessing AI's Potential to Assist Research.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey AAAR-1.0: Assessing AI's Potential to Assist Research

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.431100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.431100Z digest=sha256:c9df5e38906180e1fba675c15bf2051dc37ac95128a54e482bf7a2e0d79000bf

Observation c1ca246e-b03b-465f-9f0b-cfa9a805e722 · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.449455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.449455Z digest=sha256:bf899ef6badcba6c08baaaa03fb66f61c600d7f31694575fa9e83ddc80d0910b

Observation 03d5c5a1-2f1b-4f20-af18-7af2816ab996 · outbound

This paper cites LAB-Bench: Measuring Capabilities of Language Models for Biology Research.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey LAB-Bench: Measuring Capabilities of Language Models for Biology Research

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.439930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.439930Z digest=sha256:594e291dd23b1af97c13e557043ae2b0de60cb9178e6405644862d56b8e17e87

Observation bb95e92e-a272-40c9-b8f7-2e872a71d6f5 · outbound

This paper cites ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.444970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.444970Z digest=sha256:6386a762e45902cbe86a686153b538f630f2f273a914294ed5ea00998369a813

Observation 49e1a027-28d6-459e-96e5-bb96895989a5 · outbound

This paper cites DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.463504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.463504Z digest=sha256:0460e82b3ed3ce5177fda746a99b8490f7f984323364b74f4ba3aa631168f3d6

Observation fec61084-7634-407c-97d2-09d132ee4d14 · outbound

This paper cites PaperQA: Retrieval-Augmented Generative Agent for Scientific Research.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey PaperQA: Retrieval-Augmented Generative Agent for Scientific Research

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.453993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.453993Z digest=sha256:abcaf9c4a456fa5d732a88a1c66d6cb1f972ed0f0cd6ac19cc7f77bb0ae1a2af

Observation 27b08d10-0e8a-46ed-84aa-a9b8314678d3 · outbound

This paper cites ScienceWorld: Is your Agent Smarter than a 5th Grader?.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ScienceWorld: Is your Agent Smarter than a 5th Grader?

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.458614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.458614Z digest=sha256:f518ebe6448abb3ae72440e4917ee736bd56f8d1115309bea8bb8b5c193b47b1

Observation 043cd23b-d07a-4272-8573-e9748501667f · outbound

This paper cites Mlagent- bench: Evaluating language agents on machine learning experimentation, 2024,.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Mlagent- bench: Evaluating language agents on machine learning experimentation, 2024,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.476908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.476908Z digest=sha256:17abf0fc7d9394eb3880fa240d5396f2f1c24e779966980dacee9cbb963909a2

Observation 6a271e68-1308-4909-8500-8452e86d43b9 · outbound

This paper cites ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.467715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.467715Z digest=sha256:36e173e5a555643f54f82ecf35398262969f09eb1a9863f88c2b6c27d26211a3

Observation 79bccf71-c36d-4f0c-ba6c-fac4d942ced4 · outbound

This paper cites MLGym: A New Framework and Benchmark for Advancing AI Research Agents.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.472486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.472486Z digest=sha256:0086b781c7984317e3cae522228b282a607a8c2e2e4cb37c18d3c720b27fc42d

Observation 3b944511-a60a-467d-87f4-4285fa1271aa · outbound

This paper cites DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.491104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.491104Z digest=sha256:074e75499550b1eb3ca93204b353215712e2bb571f204383ac70781edb26528c

Pith citing papers

Observation c18e21d6-4917-4748-9379-93e8566336c1 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

Reference 154

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:14.879801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:3765a30ee4bdf3cf3c882b87580220ffcc084e3c7436695476dc43e81f920611

Observation 3e2c672f-c607-4eae-81a4-f70ce2797404 · inbound

Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering cites this paper.

Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:21:00.461563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T17:40:14.733882Z digest=sha256:66b079f0d61da73228046c15b959afd3b8838e50d0ed2f73fb2a2c9bc419f9b9

Observation 0c5748b6-1302-4008-9625-1e82f498a169 · inbound

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory cites this paper.

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:13:40.511046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:11:15.761831Z digest=sha256:470b02fe2e0f3762fbc32c4f8ef0abb09f1d65eb077b2050413fb47a4126d51b

Observation 068b5531-8028-47f4-a264-5b52127cbf69 · inbound

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents cites this paper.

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:53:26.389766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T12:52:20.911788Z digest=sha256:d1b895c29c770de933ca80cfc189cabb872e33ba15b8e2511d7063f56c8784ef