Pith. sign in

Paper Citation Record · LEDGER

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

As of 10 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 11 inbound Pith citation observations for arXiv:2506.01616.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01616 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:42:34.334089Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:34:41.371201Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:28:07.459324Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0192ec2-8633-4c0c-a7e7-c28e2de56f8e · outbound

This paper cites GPT-4o System Card.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments GPT-4o System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.631650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.631650Z digest=sha256:5040a7dd79729a2fdeafa29f76be64c6a0a213a7bcca3f24a820ad1abcd4c891

Observation 51fbccb8-2d75-4f0e-bb85-4851de27047a · outbound

This paper cites Gemini 2.0 flash model card,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Gemini 2.0 flash model card,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.656884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:32.651030Z digest=sha256:2cab6fc6f5eeaf0ee415aea582e373dde8384943199af4b78812f26351692e43

Observation 1f42dbaf-c2b6-45b9-a41e-257e98bef35a · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments The claude 3 model family: Opus, sonnet, haiku,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.640303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:32.672386Z digest=sha256:52dd4e3d068557be86a34b6575d6f3c0e0f6eddb430025c0048070bcbbb14136

Observation 90b58b83-1de5-45ec-880c-1581d6b3e514 · outbound

This paper cites Pixtral 12B.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Pixtral 12B

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.693528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.693528Z digest=sha256:056a25cab984c793cb97943849d809abe9b29e5ccc8bf55dad47d9fd0c10a68d

Observation d94bb0c9-0ddc-4240-ae38-b04ffcef4434 · outbound

This paper cites Chat with the Environment: Interactive Multimodal Perception Using Large Language Models.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Chat with the Environment: Interactive Multimodal Perception Using Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.714754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.714754Z digest=sha256:6c4c2ad60856b7d61bf5da4ac4b1b0006973473a5f8d0c1ff8396e84b4a8cce2

Observation d6cd318a-e863-4908-8f01-2975ea30476c · outbound

This paper cites Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.742298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.742298Z digest=sha256:28f46c08f2c015e64bd601c5996c86af6d723e2252e2c818dc79e00974772fef

Observation 59419a14-fc70-47db-ace9-0f0f706761aa · outbound

This paper cites VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.768449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.768449Z digest=sha256:2f2e1c6b168739cc1f334ffe9b400ae1f7d49d2474a9c64e711cd7d915cd0461

Observation 8d38a11e-35f0-43d8-a017-45eca82fbdc1 · outbound

This paper cites Large Multimodal Agents: A Survey.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Large Multimodal Agents: A Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.798379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.798379Z digest=sha256:f839d646358ddca2c423f1686d989919a6b0808a2bce69f3005120a63f561191

Observation 4c05c70e-0a23-4e0f-be7a-9fe7db43e7cc · outbound

This paper cites Agent AI: Surveying the Horizons of Multimodal Interaction.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agent AI: Surveying the Horizons of Multimodal Interaction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.859252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.859252Z digest=sha256:13c0651c7b27efd42332c05a0159b6c262c8ca248b6ff5776b4d7cdd9e35cfb2

Observation 1c0f580c-e3ad-480a-b9ef-86ce9a0428c5 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.881419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.881419Z digest=sha256:ab4d35ec3620994673fdbec6d18c4f5b9b31c6ce4194fd0bdae24f630e20292c

Observation a01ae670-76c0-4411-9e86-1048cc89327c · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AppAgent: Multimodal Agents as Smartphone Users

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.902989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.902989Z digest=sha256:bad568aebdc83ca40137e90baf4b2e7d31ef480873a657e4de1fc2ea3f662a4e

Observation f6c743ab-99ed-49e8-ba22-12b1b4c53dbf · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.924240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.924240Z digest=sha256:33620b607f5707363a9576967623693c59f1950d9dc34570ee637868415d3a30

Observation 149c1d36-29e9-4477-9dd0-f1345bbf87dc · outbound

This paper cites Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.948785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.948785Z digest=sha256:b7071eaf85303b2d1d08d60a1462ade5c416eb03461ba729d0f6ebaec9cb885f

Observation b8b15dd0-fbfb-4c34-b722-8c5182ea0749 · outbound

This paper cites Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.973269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.973269Z digest=sha256:c6d2c634ca67047b56a88ea88e402813d7ee33415197d381d93fbabfce114932

Observation 2ea8d8c1-98c1-41c9-b0e1-8a77d39622d5 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.995013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.995013Z digest=sha256:4809949ec9adee56ef78364bb2182eccd70a3724dcb94ec1dbddfbd8fef4f9b1

Observation 58f43906-2a00-43a5-9296-71b30caad99b · outbound

This paper cites VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.016850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.016850Z digest=sha256:49da331e8294480e6f9ab2e50061ce81162d51e5868c9da33910468dc8a5ab1f

Observation 2b3f4c75-d2dd-4309-a2db-847a2f4f4cf2 · outbound

This paper cites WebCanvas: Benchmarking Web Agents in Online Environments.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments WebCanvas: Benchmarking Web Agents in Online Environments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.038850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.038850Z digest=sha256:f8d9f9ae6f3fd1113fd1e724984ccfe59ad038c07aa68697f2d8f8e4a0329756

Observation b6efd011-5a9b-4c64-9ad2-56c558dde720 · outbound

This paper cites Agentboard: An analytical evaluation board of multi-turn llm agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agentboard: An analytical evaluation board of multi-turn llm agents,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.623971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:33.077219Z digest=sha256:7d6f2997d0ddb6bc37951404d00b7ec04ceb45cb43630132927da11b9d138b6f

Observation 88115221-051d-428c-ad1e-1d8407fa7238 · outbound

This paper cites The Positive-Definite Completion Problem.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments The Positive-Definite Completion Problem

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:42:35.160742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:33.108586Z digest=sha256:d2f5a22505323db1f304fc90d0ab5bea2e653898b2d3332a2a1abd874b2f6e97

Observation 213f8aa6-6738-490e-89d8-79fab063984b · outbound

This paper cites OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.129535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.129535Z digest=sha256:a50e5de585aef7b3686c3de2d75c31afe8fbd3bea35185afcde0ae1d274338b5

Observation 4de0b938-f923-4b0a-8e57-4b32cf438f03 · outbound

This paper cites Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.151323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.151323Z digest=sha256:e74f23e37c9be6ac118426a7945713a057751b86fe74f9258d32581f90164312

Observation 373eeb94-1219-4c1d-9fad-bde91477419e · outbound

This paper cites From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.179393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.179393Z digest=sha256:845a418278b62d435e352faf5722335fce4fc023bcb0e2f1e3d09a693ac8f4ce

Observation 444972f3-190f-4558-b62b-bf7ccd680d6c · outbound

This paper cites Agentdam: Privacy leakage evaluation for autonomous web agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agentdam: Privacy leakage evaluation for autonomous web agents,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.211463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.211463Z digest=sha256:82bd899b0513588608f3585f1b240b89f77c5d1a127e7b93dbbc394a2d3f5b06

Observation 76bd7d73-bebe-403b-81bf-2f66c536a251 · outbound

This paper cites Towards trustworthy gui agents: A survey,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Towards trustworthy gui agents: A survey,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.232181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.232181Z digest=sha256:234f23d441c25574bf19ff7005b89f3929838275cb4332fe43314f40a269331e

Observation 69a3638e-6bc9-45ab-91c3-b20690291bc4 · outbound

This paper cites Ai-powered robots can be tricked into acts of violence,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Ai-powered robots can be tricked into acts of violence,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.608567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:33.254892Z digest=sha256:01dd5e8ad9fd65738d78a409dce1e2d854be9412b7be3c7ba05f8b765eaaa6e6

Observation ed95cab0-819a-4d59-b163-9176c90893d5 · outbound

This paper cites R-Judge: Benchmarking Safety Risk Awareness for LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.276048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.276048Z digest=sha256:cfa693ff45ffe7e81364b57791df65865b3d7de4ae59c6bf0a226c75e9221d6c

Observation 932cf7b5-82e0-4c12-9142-af15aa85b84b · outbound

This paper cites Identifying the Risks of LM Agents with an LM-Emulated Sandbox.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Identifying the Risks of LM Agents with an LM-Emulated Sandbox

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.307592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.307592Z digest=sha256:4775a53baf5c12777d84a87f135cfbc6d4495fea33925320790e1e3c7057b2d0

Observation 70068e19-bbd1-4515-b95c-696492d1afa9 · outbound

This paper cites AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.346782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.346782Z digest=sha256:e707e039b09d6f6a8f5bebdd9b0730fa1a7c3771b2f2ba92d5f8e2dadf4219f7

Observation a6a48b9f-64e0-43b3-b6fc-736c4a032d87 · outbound

This paper cites Aligned llms are not aligned browser agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Aligned llms are not aligned browser agents,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.591773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:33.381673Z digest=sha256:d8cbc60d5f5b7daac38dc8b2376aba12832799a129d80eb6569f3ea99e94dd30

Observation d01aae15-96e5-4d0b-9f1a-a9ea4b2c8db3 · outbound

This paper cites Figstep: Jailbreaking large vision-language models via typographic visual prompts,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Figstep: Jailbreaking large vision-language models via typographic visual prompts,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.401642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.401642Z digest=sha256:d7f7e4546ec5836f562fcbc7dac6b4ccff5a3f9c0101d079d4d06c65c73ddadd

Observation 448ccc0f-a26a-4c91-ab4a-9b4750d221e9 · outbound

This paper cites How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.433505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.433505Z digest=sha256:8a77cf14c7cb5fdb925da33aa0fc5dae1f8fac56973b9e7a3704d82d22ad3879

Observation 9e82d4df-4156-413d-8cb6-c20e668115c6 · outbound

This paper cites Red Teaming Visual Language Models.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Red Teaming Visual Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.455009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.455009Z digest=sha256:9971e8420e0c7111f69e2b1b34083af237e3a228accc4b9d88df721ada325e03

Observation a56f2271-a3df-4f32-87ea-508b0d3ac0af · outbound

This paper cites Multitrust: A comprehensive benchmark towards trustworthy multimodal large language models,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Multitrust: A comprehensive benchmark towards trustworthy multimodal large language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.565998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:33.476297Z digest=sha256:7e3274e8c6ed829cd64998add467f57e666d390024efb88f79b9b9fc72dc22a3

Observation ee8890b0-3b36-492c-8951-bdce4cd1322a · outbound

This paper cites Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.498378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.498378Z digest=sha256:21418f8f77331e3660f18e5f5f30bb7745074af53c4b54051fd3ef7c25ea833c

Observation 9a59d0cc-58b3-45b4-98a5-4eaa03a6d04e · outbound

This paper cites Mobilesafetybench: Evaluating safety of autonomous agents in mobile device control,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Mobilesafetybench: Evaluating safety of autonomous agents in mobile device control,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.519896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.519896Z digest=sha256:8d3fbfd60c915a117e709ced82b9d367516c15e56d047eb4ed4bb46d121f5af2

Observation 223afa44-20b6-4b8e-858b-fb977338d684 · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.548948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.548948Z digest=sha256:234f9f6c9312a1a4c50de5ebf0275e940adfe26b7d23fe29f05a26c7d23ab4b5

Observation 50298a47-6967-4c1b-befb-69fbfec31c04 · outbound

This paper cites Cognitive architec- tures for language agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Cognitive architec- tures for language agents,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.550663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:33.570249Z digest=sha256:6e2602c2c51702ec16587962659443a45b891b252494242e2f48e3f5066d21cc

Observation f64a28bd-0e6c-4b7d-9188-11c97fbbbf33 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Vipergpt: Visual inference via python execution for reasoning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.536128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:33.606591Z digest=sha256:a97551134e47bdfecdd411340381a61e5bd80a228d01f7d46f94850b89acdeed

Observation db6e8ac8-d582-41c4-88cd-114c35d4d976 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Chameleon: Plug-and-play compositional reasoning with large language models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.521029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:33.643057Z digest=sha256:eb6d0615595644e202450afd2f9832dea72249649eb02436391c54d4c78999b2

Observation f8eabefb-c52c-482c-af84-86a1ee746690 · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.672003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.672003Z digest=sha256:0e7fc344c77773ba81e6fa45e321f72b0fb358f15427d1890c72336d41be2e79

Observation bd4011c7-7988-4248-8389-bf726766b610 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.692481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.692481Z digest=sha256:99ecd7e388b5a24d661f08429f06aaf8d3739964a14e6020c14c17df8ba14b79

Observation 6b589c43-ab41-4c9e-844f-8fcf586ed2dd · outbound

This paper cites An introduction to microsoft copilot,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments An introduction to microsoft copilot,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.495761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:33.714006Z digest=sha256:16bce737c66bca53f192d4a5c0f22f05fad2b214804c6e0a011c0c57ec7f2b45

Observation dd503f24-bdc1-4c00-8b14-00d50cca5717 · outbound

This paper cites GPT-4 Technical Report.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments GPT-4 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.735322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.735322Z digest=sha256:341fa2425ebd6d3f830e1245ee582059ef73de79e48bd39f9dc7cdaab2dd9ef6

Observation 78b27603-5ec6-4f11-ba85-22c42275fc17 · outbound

This paper cites Improving image generation with better captions,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Improving image generation with better captions,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.757941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.757941Z digest=sha256:761fc11ff32b3bc9ce7dcc7217a237c79836aed89213ee96e4455a7995b74441

Observation 55dd091a-4a61-48f6-9bcf-a79b6e1f5a87 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.780135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.780135Z digest=sha256:fdd26e731be5aed3d3b37296326fc386ff8815bfb8cf9864dcc40fb2bd5538bd

Observation 718ce5d8-1928-457d-a07a-b6561ba0ee61 · outbound

This paper cites Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.802260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.802260Z digest=sha256:798fa3492a751f419832d345c094e11ca102a17e4820fe479d20a18242e4317b

Observation ea697270-1886-4d70-a867-2a1346ad5472 · outbound

This paper cites Evaluating Cultural and Social Awareness of LLM Web Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Evaluating Cultural and Social Awareness of LLM Web Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.824262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.824262Z digest=sha256:202d6ea35b4fad80dd5c5a450aa8b9b2873eb1e821900ca24f38f4720140f1e5

Observation 69fe411c-866a-4592-9402-9765861f448f · outbound

This paper cites A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.844167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.844167Z digest=sha256:ff4c3ec708835e27682f498c8dee1337819d25cdae745bb70bad6459300c05c0

Observation c1a6ef74-4f30-436a-8be2-6312e39575ca · outbound

This paper cites InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.874483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.874483Z digest=sha256:6b3e82f48c89b90a2dfc5797b0e7dd7e064bdb242379ef9134f7b92fbaf7a74c

Observation 70dac990-ffe5-4634-9672-79ef68675753 · outbound

This paper cites EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.904793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.904793Z digest=sha256:441af71a987120bff7060aa2bf22b3c16a2e7332b35e29df24140c99d84f72a9

Observation 608181d0-4092-4d95-9951-316fc7a8fdc0 · outbound

This paper cites GPT-4V(ision) is a Generalist Web Agent, if Grounded.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments GPT-4V(ision) is a Generalist Web Agent, if Grounded

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.931526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.931526Z digest=sha256:08f38ca8b3a0a8c686f7413c98a82620d8cf3d289de8f925d01b13d4662c1040

Observation 9f7c962f-ba98-40db-a27c-ce9b3da6036c · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.953240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.953240Z digest=sha256:d30abe56ed245281db00656071643adfe9b1981dd0fc4e79440bbacad4724b06

Observation 453ea8eb-5968-4b02-90c7-81632c60e4aa · outbound

This paper cites Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.973959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.973959Z digest=sha256:ddf3dcda7367ce3c46f8e9561ccf660b4d1fdaadb0e3a236d488ebf3b77c1ddb

Observation a5588d20-6ca3-4cb0-8fc6-26247e9c7b34 · outbound

This paper cites ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.995822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.995822Z digest=sha256:9f57139924f1cb1938aefcb7b8b7c06927cc70e75ccd130cb7bb4cd572d6a827

Observation a7f792d5-1235-4c28-9446-690af4e73f9a · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.018149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.018149Z digest=sha256:a9ab857c1a6c1f8f74cfffb3393a4a32c70ae0ff8848861ca1c96c16e893641d

Observation feb6032e-6828-4491-bdee-121c427bae65 · outbound

This paper cites Dissecting adversarial robustness of multimodal lm agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Dissecting adversarial robustness of multimodal lm agents,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.467914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:34.043522Z digest=sha256:d7440630a5b841b5003c093958a1f6ba8cd9784681be224f9a828a2a96dd5801

Observation a1b4f668-1c2c-4a02-8c9c-3535b008bb86 · outbound

This paper cites Exploring the robustness of decision-level through adversarial attacks on llm-based embodied models,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Exploring the robustness of decision-level through adversarial attacks on llm-based embodied models,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.451542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:34.064792Z digest=sha256:d786cc2e6c24f982781032955d562940259f93f3e0e05a1f128d769a43a4f477

Observation 2903d421-1a8e-4a1a-8b07-d67a46d176de · outbound

This paper cites AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.095527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.095527Z digest=sha256:4b0c14a2590fb8c8e648af99d79449694418303c033a13ad7d0dfd24c724d57e

Observation 9bf83d3d-5c2c-4003-a8f0-ffc334934005 · outbound

This paper cites AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.147867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.147867Z digest=sha256:00a574a5ba07e81fcb03b42daea567b1c0b0314b0875284166072db8ded834cd

Observation 96e618ba-0d0f-443e-a9e1-362f9ad95e1f · outbound

This paper cites A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.170693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.170693Z digest=sha256:9cb13faf19c392d96cf4bc5762ad722d6802fd684880c11c83534eb325f413e7

Observation d4f3d604-3347-46d1-9a55-af913c0b0a43 · outbound

This paper cites The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.196721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.196721Z digest=sha256:0c25b8b54eb4b241b30a46141b8acf4d39fa5ff502c582c1fddbd83653a99492

Observation 7d6a254d-c1a3-482e-b5e4-d212330de1f4 · outbound

This paper cites Large multi-modal models for strong performance and efficient deployment,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Large multi-modal models for strong performance and efficient deployment,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.436239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:34.218472Z digest=sha256:895b3185d9db6bc9bab1f62a81f93b9cacc6bf61689e42cf7dce1a8c66bc35dc

Observation ece3bf19-f58c-40d7-ab6a-baad12aec4c0 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.239494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.239494Z digest=sha256:97b74c94b68931cbab6f9a2b70c0370650bc43d8575f55badd5d79191f700253

Observation 35cf182c-a56d-466a-89a3-955c20040914 · outbound

This paper cites Qwen2.5-VL Technical Report.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Qwen2.5-VL Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.261444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.261444Z digest=sha256:f6c46a75b00f778be83d7a912324c911a8e34095c1f24ce353adb7cfacb38df4

Observation c8befabd-e994-4862-a7e2-f7d130544460 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments LLaVA-OneVision: Easy Visual Task Transfer

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.285617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.285617Z digest=sha256:6cef070a0ea72a956e4bb461bbd51590fc7fc763ad9a8353bbe86b61f0ee4c24

Observation e7a17782-87a1-4b0b-a5a5-5416f5b02d60 · outbound

This paper cites Phi-4 Technical Report.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Phi-4 Technical Report

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.312600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.312600Z digest=sha256:64c513444c316b460e037693151dd1f7afb1da4b6586e89528e1c0bd3120028c

Observation a1fe2f56-e41f-4c2b-bc8e-f52ccc6e6bf6 · outbound

This paper cites Elemente der exakten erblichkeitslehre. 1909,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Elemente der exakten erblichkeitslehre. 1909,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.421328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:42:34.334089Z digest=sha256:dc3a88df473754f7781a70a97d4a972604f7731dca2e5ec8c1c3b3340aa18391

Pith citing papers

Observation 0ac778f4-c436-4eab-ab2e-8f7db66007f3 · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:42:13.989933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:8b3ef7c4c38637ecbe4e0e58633cb8243d4f2f7b57976006362e5f5207df71f5

Observation 264fee2f-c180-46c9-9f0c-de7b6b451ea7 · inbound

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents cites this paper.

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:41.371201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:41.371201Z digest=sha256:73476bd9d92ac80bfcdae6f29483a2382c3b3ff1f95a79262af3fb3094e64c71

Observation 523b9fb1-1da7-4def-85b9-ddd66c16c86c · inbound

VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents cites this paper.

VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:06:42.621938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T18:06:12.349285Z digest=sha256:404b73d8428f4d4356801c4f5e615472d0046acb974e135c782e09c158091124

Observation aec30d6b-d3d3-4614-91c4-54beb8bf53b5 · inbound

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation cites this paper.

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:11.014186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:11.014186Z digest=sha256:fe80b6f5a30d845123171fb814b13cebfbc00bd323b9b3b7b0e9bffc0d66546d

Observation df4f8f69-492f-4a02-8859-d6f5769654b6 · inbound

Red Teaming Large Reasoning Models cites this paper.

Red Teaming Large Reasoning Models MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:31:28.581073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T03:29:16.164423Z digest=sha256:953e4a4a9406737bdd5be923bb8fc71035813bcaa55bd8233c0c225a280ed85e

Observation b5a7936f-cb66-4614-83a7-704b221795ca · inbound

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents cites this paper.

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:32:48.311462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T11:31:16.087579Z digest=sha256:9e2776ac21812936d443c56abd95701e1f5d41cf47f04cdea26193d739b8c3ac

Observation fd288ce0-6ab2-4df2-8f77-243cf9cb4db0 · inbound

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents cites this paper.

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:11.643104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:51:54.310805Z digest=sha256:cb871afb33b7d9d33a52d5802ae205c7f17f667c609675bd0ac1b546201bd122

Observation 7b5d2f42-792d-4a94-ab49-0c71b8127d2e · inbound

Governance by Construction for Generalist Agents cites this paper.

Governance by Construction for Generalist Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:58.013167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T04:51:41.739358Z digest=sha256:20d93fd512ee39d05c5d390d5afd353528778210aaffb7876573d24d577109c6

Observation 026806d4-6102-4363-8424-3a12856b2553 · inbound

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security cites this paper.

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:45:01.572187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T19:18:40.244556Z digest=sha256:f95824762cd07972518cac9a7e204eef32b94af66b8552e9555404ee56b1b6f8

Observation 32027b2f-b83e-496b-a0d2-92207bcd67a3 · inbound

CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents cites this paper.

CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:28:07.460850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:00:10.153170Z digest=sha256:d3640b88f5773380ed87887bfa4c13439bb28bd02c24b9ae470aa79f7933dd23

Observation abc1ed42-32f8-4e7b-b357-2d381601b7ff · inbound

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion cites this paper.

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T11:50:48.700362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:50:48.700362Z digest=sha256:aaad51d1b4070f88f9fe94e4bc61eba908fd4223d380eec89a650d12cab55c58