Pith. sign in

Paper Citation Record · LEDGER

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 11 inbound Pith citation observations for arXiv:2506.01616.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01616 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:42:34.334089Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:34:41.371201Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:28:07.459324Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0192ec2-8633-4c0c-a7e7-c28e2de56f8e · outbound

This paper cites GPT-4o System Card.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments GPT-4o System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.631650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.631650Z digest=sha256:faab7de37266fe18615df40eb76e3d7c2ff6abc4f18525973c605f9f8ae13ab6

Observation 51fbccb8-2d75-4f0e-bb85-4851de27047a · outbound

This paper cites Gemini 2.0 flash model card,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Gemini 2.0 flash model card,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.656884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:32.651030Z digest=sha256:536f70cb90aa2d1b07cb04cf60018fa51110863bf36ba808d3619020765ad118

Observation 1f42dbaf-c2b6-45b9-a41e-257e98bef35a · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments The claude 3 model family: Opus, sonnet, haiku,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.640303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:32.672386Z digest=sha256:fdbb6159b1418870aa6eb0970e9ae36a821887cfc4b0e00e9756dfbe73baced4

Observation 90b58b83-1de5-45ec-880c-1581d6b3e514 · outbound

This paper cites Pixtral 12B.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Pixtral 12B

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.693528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.693528Z digest=sha256:e9e07f8f7fb4b0460a9f3078dd354a5110e098cf2c29f8ea35d505c3f8b2dcb6

Observation d94bb0c9-0ddc-4240-ae38-b04ffcef4434 · outbound

This paper cites Chat with the Environment: Interactive Multimodal Perception Using Large Language Models.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Chat with the Environment: Interactive Multimodal Perception Using Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.714754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.714754Z digest=sha256:cf2dc728373fb097932f3a49fede29596153edb68161cbe7f93cdc0fef148d92

Observation d6cd318a-e863-4908-8f01-2975ea30476c · outbound

This paper cites Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.742298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.742298Z digest=sha256:06a214870423b768b82cf123e0b819d03badbd735ee8b4ea278163ca4e33b3e1

Observation 59419a14-fc70-47db-ace9-0f0f706761aa · outbound

This paper cites VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.768449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.768449Z digest=sha256:080620eb8545737284960a89ccb5f2b133d6900687f72dc171f0f99fdb0572b1

Observation 8d38a11e-35f0-43d8-a017-45eca82fbdc1 · outbound

This paper cites Large Multimodal Agents: A Survey.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Large Multimodal Agents: A Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.798379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.798379Z digest=sha256:4478dc5fddb17d8dc14eec552657259bc4d94a62d22510bd71297917ca332a3e

Observation 4c05c70e-0a23-4e0f-be7a-9fe7db43e7cc · outbound

This paper cites Agent AI: Surveying the Horizons of Multimodal Interaction.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agent AI: Surveying the Horizons of Multimodal Interaction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.859252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.859252Z digest=sha256:36c9219230783301d5a10f36ea211e90eb5d98779fb9fc45206fbda10e9925c2

Observation 1c0f580c-e3ad-480a-b9ef-86ce9a0428c5 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.881419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.881419Z digest=sha256:4d85e4ddea8ff032fdaade72bf209b195f2b0914b47d0b45fb7569e155a4c4e8

Observation a01ae670-76c0-4411-9e86-1048cc89327c · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AppAgent: Multimodal Agents as Smartphone Users

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.902989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.902989Z digest=sha256:420de631846322d088f03c3283c828de609291815ab6f8f5ddca048141b79c5b

Observation f6c743ab-99ed-49e8-ba22-12b1b4c53dbf · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.924240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.924240Z digest=sha256:4181bb7739102abde16240ecd24573759a3f3eb8d72c4ad78b9eaba79cc9b973

Observation 149c1d36-29e9-4477-9dd0-f1345bbf87dc · outbound

This paper cites Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.948785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.948785Z digest=sha256:cc6d078568392c41a57812a05dcbdfb236a5ff0a5e9db3a4c60c99b81e00c123

Observation b8b15dd0-fbfb-4c34-b722-8c5182ea0749 · outbound

This paper cites Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.973269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.973269Z digest=sha256:97a0c0a9ad96889f3ac0485db6062468ec79bf3629bf8127b03426a2d143ecf8

Observation 2ea8d8c1-98c1-41c9-b0e1-8a77d39622d5 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.995013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.995013Z digest=sha256:04faef2f30a367434334486d0f0a7bf7e7f64c0edaf945b7d867d64673173fe5

Observation 58f43906-2a00-43a5-9296-71b30caad99b · outbound

This paper cites VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.016850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.016850Z digest=sha256:f5bf5dafa93ec6ced437c49ec079a25e45e7d3f50547db86747e7b8bde67e86e

Observation 2b3f4c75-d2dd-4309-a2db-847a2f4f4cf2 · outbound

This paper cites WebCanvas: Benchmarking Web Agents in Online Environments.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments WebCanvas: Benchmarking Web Agents in Online Environments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.038850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.038850Z digest=sha256:ef51c8322c4b510f5fd28f0e4662e9ff465414aac373b448eddd819a327ac44c

Observation b6efd011-5a9b-4c64-9ad2-56c558dde720 · outbound

This paper cites Agentboard: An analytical evaluation board of multi-turn llm agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agentboard: An analytical evaluation board of multi-turn llm agents,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.623971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.077219Z digest=sha256:c6dfae5c2fa47f43e91786faeaa6cda0f8586098e6b97eb20e716015c2bc5c5a

Observation 88115221-051d-428c-ad1e-1d8407fa7238 · outbound

This paper cites The Positive-Definite Completion Problem.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments The Positive-Definite Completion Problem

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:42:35.160742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.108586Z digest=sha256:ec3edc9ade812982d781be5f16069d3abeeeeed518e49b952f432d7537a4a891

Observation 213f8aa6-6738-490e-89d8-79fab063984b · outbound

This paper cites OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.129535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.129535Z digest=sha256:e1d3661d1fc581f4763738c00c990cbadbb4bce56c34492e5e275c60d5872a22

Observation 4de0b938-f923-4b0a-8e57-4b32cf438f03 · outbound

This paper cites Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.151323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.151323Z digest=sha256:5fceb73a8bd1f43e427ca64fb029e1be311854ac779e23f9bbc53547db4255d4

Observation 373eeb94-1219-4c1d-9fad-bde91477419e · outbound

This paper cites From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.179393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.179393Z digest=sha256:96b38c0947cd0047f3b704c5a42c10b1977edc26b54258973d4d9ef69440093a

Observation 444972f3-190f-4558-b62b-bf7ccd680d6c · outbound

This paper cites Agentdam: Privacy leakage evaluation for autonomous web agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agentdam: Privacy leakage evaluation for autonomous web agents,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.211463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.211463Z digest=sha256:c287f0b1f26d46ceaec638b989b150c254debaff0376c2ecbeb16d1bfe30b369

Observation 76bd7d73-bebe-403b-81bf-2f66c536a251 · outbound

This paper cites Towards trustworthy gui agents: A survey,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Towards trustworthy gui agents: A survey,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.232181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.232181Z digest=sha256:2511b6b5e9554db47e208ef897decf3e7056d7e1f185c9a58d531c4425c79229

Observation 69a3638e-6bc9-45ab-91c3-b20690291bc4 · outbound

This paper cites Ai-powered robots can be tricked into acts of violence,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Ai-powered robots can be tricked into acts of violence,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.608567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.254892Z digest=sha256:91f970ee50a1e0294a7d64c0f85f14008977fcff52ebca621d305a233afc9345

Observation ed95cab0-819a-4d59-b163-9176c90893d5 · outbound

This paper cites R-Judge: Benchmarking Safety Risk Awareness for LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.276048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.276048Z digest=sha256:353467fa34a09a566b68c8cdd399e09783bd239842f2d4b897591b11341c4e88

Observation 932cf7b5-82e0-4c12-9142-af15aa85b84b · outbound

This paper cites Identifying the Risks of LM Agents with an LM-Emulated Sandbox.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Identifying the Risks of LM Agents with an LM-Emulated Sandbox

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.307592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.307592Z digest=sha256:eeb1e718ce9ed06ef9243bb3da6b10dcf8cdcba9a338aa33b73be87485b70699

Observation 70068e19-bbd1-4515-b95c-696492d1afa9 · outbound

This paper cites AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.346782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.346782Z digest=sha256:ede7fddd5b3daa23d732befb30932585aba5904cc38160317421e721eabf7b3a

Observation a6a48b9f-64e0-43b3-b6fc-736c4a032d87 · outbound

This paper cites Aligned llms are not aligned browser agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Aligned llms are not aligned browser agents,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.591773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.381673Z digest=sha256:5ba1097d243a8895d06314ea43f2d966d0e2f826d5708123536d5562710a1c4e

Observation d01aae15-96e5-4d0b-9f1a-a9ea4b2c8db3 · outbound

This paper cites Figstep: Jailbreaking large vision-language models via typographic visual prompts,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Figstep: Jailbreaking large vision-language models via typographic visual prompts,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.401642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.401642Z digest=sha256:f799047ce50d055214cb53672b57d219365d5c4da2d92348c1cdb7f62525295b

Observation 448ccc0f-a26a-4c91-ab4a-9b4750d221e9 · outbound

This paper cites How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.433505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.433505Z digest=sha256:737a330a0a04fc7382b057141ee9c06118bfeb269a30b885d71d5da8e08e0a61

Observation 9e82d4df-4156-413d-8cb6-c20e668115c6 · outbound

This paper cites Red Teaming Visual Language Models.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Red Teaming Visual Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.455009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.455009Z digest=sha256:8bc348762a72a6a6d20bfcd45f42835ab3a3f6c5dac19e5a190166c4909b7e18

Observation a56f2271-a3df-4f32-87ea-508b0d3ac0af · outbound

This paper cites Multitrust: A comprehensive benchmark towards trustworthy multimodal large language models,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Multitrust: A comprehensive benchmark towards trustworthy multimodal large language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.565998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.476297Z digest=sha256:65e284cc5775cdeb92b265752f87afdb933f21fac4ac7f294f4bae88e4202984

Observation ee8890b0-3b36-492c-8951-bdce4cd1322a · outbound

This paper cites Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.498378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.498378Z digest=sha256:2312f5b3c343f7b362548a33bb8ada7846500174264371eece6f3b44ade0b375

Observation 9a59d0cc-58b3-45b4-98a5-4eaa03a6d04e · outbound

This paper cites Mobilesafetybench: Evaluating safety of autonomous agents in mobile device control,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Mobilesafetybench: Evaluating safety of autonomous agents in mobile device control,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.519896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.519896Z digest=sha256:3916b74e047fdf57c951c7f4e5ffae6f50096fc3550d680e313202be942f10e1

Observation 223afa44-20b6-4b8e-858b-fb977338d684 · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.548948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.548948Z digest=sha256:9fe4f282192ded5108aeee89770e9533b6c019234cb7a3a138373d265180a48d

Observation 50298a47-6967-4c1b-befb-69fbfec31c04 · outbound

This paper cites Cognitive architec- tures for language agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Cognitive architec- tures for language agents,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.550663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.570249Z digest=sha256:3a8b4ce004f97677b0b7e072a6cfa4a1dcc23d8638fd8626591dc1902ee03222

Observation f64a28bd-0e6c-4b7d-9188-11c97fbbbf33 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Vipergpt: Visual inference via python execution for reasoning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.536128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.606591Z digest=sha256:519a9aaa4f6d4f8f663930b83b2c175a54f54433fe114d8d9d19c6ccfbce182e

Observation db6e8ac8-d582-41c4-88cd-114c35d4d976 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Chameleon: Plug-and-play compositional reasoning with large language models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.521029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.643057Z digest=sha256:361fdace75adc33246ea1dda9432241cd5d89d9ea580a571d37d01e2772de7dd

Observation f8eabefb-c52c-482c-af84-86a1ee746690 · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.672003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.672003Z digest=sha256:a3e58659467a47646f1a2ebf33a1cf944bed9013d17df2ec3ff11df1919ade78

Observation bd4011c7-7988-4248-8389-bf726766b610 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.692481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.692481Z digest=sha256:9e5c4f1d1c7bd168f3a03e8073208a6625ca48112ccca990ef02c2c60672db7b

Observation 6b589c43-ab41-4c9e-844f-8fcf586ed2dd · outbound

This paper cites An introduction to microsoft copilot,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments An introduction to microsoft copilot,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.495761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.714006Z digest=sha256:7db90b9dd0c383aadb87f4a9cd757a20566584e3662da8a5d11eac373c312fcd

Observation dd503f24-bdc1-4c00-8b14-00d50cca5717 · outbound

This paper cites GPT-4 Technical Report.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments GPT-4 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.735322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.735322Z digest=sha256:8b7dd9126ac0bb98f2217998fce7fad162cf96a252b4ceabb5b390c479651a28

Observation 78b27603-5ec6-4f11-ba85-22c42275fc17 · outbound

This paper cites Improving image generation with better captions,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Improving image generation with better captions,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.757941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.757941Z digest=sha256:6160fce0a12fe575542fe0bfe4e53db0e107e3f2b53d15fbb5fae82a3b4b6329

Observation 55dd091a-4a61-48f6-9bcf-a79b6e1f5a87 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.780135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.780135Z digest=sha256:bbbf6415c01b7b20a89126ae6410dddce13a0ceb602186d365ec51928b8c8e33

Observation 718ce5d8-1928-457d-a07a-b6561ba0ee61 · outbound

This paper cites Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.802260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.802260Z digest=sha256:6d8cb89705ab7bd353d08deb7a9f481557668d181a65d1af6eddea152c539eaa

Observation ea697270-1886-4d70-a867-2a1346ad5472 · outbound

This paper cites Evaluating Cultural and Social Awareness of LLM Web Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Evaluating Cultural and Social Awareness of LLM Web Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.824262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.824262Z digest=sha256:ab28bc7d95e8c4c1c06b0d9b137825fc8220cb52b1777f4b2b0077235ef9f97c

Observation 69fe411c-866a-4592-9402-9765861f448f · outbound

This paper cites A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.844167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.844167Z digest=sha256:13f3e9a5aa89216a09322e77a012f8315bd04f6cd95da1384b62bbb65f142bb1

Observation c1a6ef74-4f30-436a-8be2-6312e39575ca · outbound

This paper cites InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.874483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.874483Z digest=sha256:09d7765cba15b6bb73379bc39d777e373922e2843e9989bb900c71d24a27415b

Observation 70dac990-ffe5-4634-9672-79ef68675753 · outbound

This paper cites EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.904793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.904793Z digest=sha256:4b10cb7dce17d41589b01a8e2f43e42284db8e23105a9c9f04215d10c4f14659

Observation 608181d0-4092-4d95-9951-316fc7a8fdc0 · outbound

This paper cites GPT-4V(ision) is a Generalist Web Agent, if Grounded.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments GPT-4V(ision) is a Generalist Web Agent, if Grounded

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.931526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.931526Z digest=sha256:ceb32bcaccd2bde2bcd6572f16b8f6b5c7b78ab6496757750a58d5f2436a8785

Observation 9f7c962f-ba98-40db-a27c-ce9b3da6036c · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.953240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.953240Z digest=sha256:e45d0f249f079a9c5313d09abea03289e55e88f0d3464454914cbe9a8d3e2417

Observation 453ea8eb-5968-4b02-90c7-81632c60e4aa · outbound

This paper cites Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.973959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.973959Z digest=sha256:e9a0639a98a3a426fcef13e9d2117452cbaac75c272f39b67fdc92f82ddea5d1

Observation a5588d20-6ca3-4cb0-8fc6-26247e9c7b34 · outbound

This paper cites ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.995822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.995822Z digest=sha256:ddbc65302b004ab95af2690fa92830e81efd2f023a3758fb9b8d296d91234bdf

Observation a7f792d5-1235-4c28-9446-690af4e73f9a · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.018149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.018149Z digest=sha256:b9212f1fdfa180aae72d8a459fa17fc89df3ae980d6f8e3ceea5fb7cf0ab0b2d

Observation feb6032e-6828-4491-bdee-121c427bae65 · outbound

This paper cites Dissecting adversarial robustness of multimodal lm agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Dissecting adversarial robustness of multimodal lm agents,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.467914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:34.043522Z digest=sha256:90d1bcf81365b8ca77a75fafb68c99c20e3763884b7d9de864b816630409bcb7

Observation a1b4f668-1c2c-4a02-8c9c-3535b008bb86 · outbound

This paper cites Exploring the robustness of decision-level through adversarial attacks on llm-based embodied models,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Exploring the robustness of decision-level through adversarial attacks on llm-based embodied models,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.451542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:34.064792Z digest=sha256:b95e2f155c7750e29ee0b0a8f728f6a0ed769d88f719445718cca13016ce8a8a

Observation 2903d421-1a8e-4a1a-8b07-d67a46d176de · outbound

This paper cites AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.095527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.095527Z digest=sha256:205495368488ff8a7445ab3de49ecdd9edadd2fe5446fa6016e002f76bf8cf00

Observation 9bf83d3d-5c2c-4003-a8f0-ffc334934005 · outbound

This paper cites AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.147867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.147867Z digest=sha256:d1a6f38b4973cda8ea009ffddb5de6ad2d29df21615f0edc4edd4367d3cdaa06

Observation 96e618ba-0d0f-443e-a9e1-362f9ad95e1f · outbound

This paper cites A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.170693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.170693Z digest=sha256:7bf9fa1abdffc32033edaafaaddc4c1b1c49c0e27602fc762e4b9d289dd7bfae

Observation d4f3d604-3347-46d1-9a55-af913c0b0a43 · outbound

This paper cites The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.196721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.196721Z digest=sha256:199e7cb1e7935d7963f1db1ce4c127691374931033c471f8509f7fe92bcc945a

Observation 7d6a254d-c1a3-482e-b5e4-d212330de1f4 · outbound

This paper cites Large multi-modal models for strong performance and efficient deployment,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Large multi-modal models for strong performance and efficient deployment,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.436239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:34.218472Z digest=sha256:1f959f619005bc835591a6cf2449a1b1ebe0563697701d4b62d8f87db1d930d5

Observation ece3bf19-f58c-40d7-ab6a-baad12aec4c0 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.239494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.239494Z digest=sha256:40c2fcf7ce36338dbd417c4b863dc1358ec03eb8b39a52d91b69d22f77f3b351

Observation 35cf182c-a56d-466a-89a3-955c20040914 · outbound

This paper cites Qwen2.5-VL Technical Report.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Qwen2.5-VL Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.261444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.261444Z digest=sha256:001a6ddea200e5ba88f3b9b76a7fbb52b1a8c1ea9face9359d89fdbd25df8e12

Observation c8befabd-e994-4862-a7e2-f7d130544460 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments LLaVA-OneVision: Easy Visual Task Transfer

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.285617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.285617Z digest=sha256:bb4cfa039aee3e1f5f0daa300549b43808b9d9c3a7f9636571a8bea071342adc

Observation e7a17782-87a1-4b0b-a5a5-5416f5b02d60 · outbound

This paper cites Phi-4 Technical Report.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Phi-4 Technical Report

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.312600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.312600Z digest=sha256:80dfd1dec6006d76351f163d6f22b7ceffd40b5396a1abfbd1f250ad173e1790

Observation a1fe2f56-e41f-4c2b-bc8e-f52ccc6e6bf6 · outbound

This paper cites Elemente der exakten erblichkeitslehre. 1909,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Elemente der exakten erblichkeitslehre. 1909,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.421328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:34.334089Z digest=sha256:075b9fa35da968afabc2a5a04349e504020b0f3922d7f474080fffa724e97a48

Pith citing papers

Observation 0ac778f4-c436-4eab-ab2e-8f7db66007f3 · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:42:13.989933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:07ae360259b91eda19eb875c6629aebc1df4c849366accf10697f4dce209f873

Observation 264fee2f-c180-46c9-9f0c-de7b6b451ea7 · inbound

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents cites this paper.

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:41.371201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:41.371201Z digest=sha256:9738c0c720cb17cc37bedf6275aea16e1f60bce278297c582b9ea534287f082f

Observation 523b9fb1-1da7-4def-85b9-ddd66c16c86c · inbound

VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents cites this paper.

VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:06:42.621938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T18:06:12.349285Z digest=sha256:7f6a9742aa2d37f6c0477dec7852ab16605686862eedff9fa3c407eb94ee97c0

Observation aec30d6b-d3d3-4614-91c4-54beb8bf53b5 · inbound

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation cites this paper.

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:11.014186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:11.014186Z digest=sha256:a9e43dbbd371f0bbc7c81ab828d69a96902e6492d40a5c63845d3f4a417136db

Observation df4f8f69-492f-4a02-8859-d6f5769654b6 · inbound

Red Teaming Large Reasoning Models cites this paper.

Red Teaming Large Reasoning Models MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:31:28.581073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T03:29:16.164423Z digest=sha256:da9d6724fc54cc4fdbdbfc1ecf1fb424d6b1269c841037751211e0a19a786d84

Observation b5a7936f-cb66-4614-83a7-704b221795ca · inbound

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents cites this paper.

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:32:48.311462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T11:31:16.087579Z digest=sha256:f786ebf5a5632c08c827a67a11903c4e140c03ebb2c697b7417fcaa97c9bb031

Observation fd288ce0-6ab2-4df2-8f77-243cf9cb4db0 · inbound

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents cites this paper.

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:11.643104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:51:54.310805Z digest=sha256:3f3ff19bc7b281376748a879747ddd44756062027415eba12ad34b954dc80827

Observation 7b5d2f42-792d-4a94-ab49-0c71b8127d2e · inbound

Governance by Construction for Generalist Agents cites this paper.

Governance by Construction for Generalist Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:58.013167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T04:51:41.739358Z digest=sha256:844facb8549f30f7a3d24348239ff2975a7896649d522efe9a651180771a50ea

Observation 026806d4-6102-4363-8424-3a12856b2553 · inbound

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security cites this paper.

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:45:01.572187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T19:18:40.244556Z digest=sha256:05d62a4bbd677b051c28470106a30014d2bfe95484b06efe7f29650470a384d6

Observation 32027b2f-b83e-496b-a0d2-92207bcd67a3 · inbound

CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents cites this paper.

CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:28:07.460850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:00:10.153170Z digest=sha256:1814b852594022efe5a55bd17427840b4f01e1504be3582b467086b73facdacf

Observation abc1ed42-32f8-4e7b-b357-2d381601b7ff · inbound

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion cites this paper.

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T11:50:48.700362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:50:48.700362Z digest=sha256:4a822d52a5cf64b4f8148076ef452124d0c7b07d0292a209f90cfcc79f712825