Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:38:09.367743Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2412.10840.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:38:09.367743Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-19T11:08:27.472508Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T11:08:27.773962Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d2fb6641-a1f2-4151-9814-6decf45c19e7 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8d589e4d-f355-444a-857e-017b8961dca2 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e68d410-381b-432e-bb3a-b05faf70a060 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a4c49e1-d45c-4af0-b7f5-a36a784cf6ea · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning GUICourse: From General Vision Language Models to Versatile GUI Agents
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0c37753-debe-4836-b087-de80a8bccf36 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab83b3f-e46b-49d0-8d44-304eb99b16cb · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 19466a30-d03d-4bd9-b65f-84aec5094156 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b47c38f7-4a7c-4101-981f-85d4b65e43e7 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 671e7e49-cdcf-4e7c-b1f1-0df4486ed79f · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c4bb37b-564a-411c-aaef-10c9a3e84d39 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning CogAgent: A Visual Language Model for GUI Agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46557417-933e-4753-8290-db01ddba6a50 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f4de35f-9bbb-4fce-8e71-6b714387a071 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning H.; Zheng, B.; Deng, X.; Su, Y.; and Chao, W.-L
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7d576b52-1acb-471d-a394-c2f1b5e7cbc5 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 325726aa-3788-45eb-b388-0fab7463eed8 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning J.-J.; Mitchell, T.; and Myers, B
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 80cd5fbf-6109-4123-bbf0-19010af9b68e · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1e4522ea-6091-48da-8bab-c15f4703a846 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8ff745c3-d66f-4160-9ad0-1b8788f23a05 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66533dcb-d213-4fb2-b63a-1e7b8fa93d16 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a54594b6-7844-4211-9ff0-b1403db9f4c4 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef39dc5e-3abd-4498-99ad-5834fcfa612e · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cefd08b3-a486-4088-a248-fd43a289b032 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning MobileFlow: A Multimodal LLM For Mobile GUI Agent
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19cc0eb6-e08e-4e86-a1af-063276b49ad6 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da58cc1-2bc8-448e-83f4-0338dcb8a2d9 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c99ab3dd-7d00-4d12-a2a3-49896aaa80ec · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 33220915-e667-465f-8ad8-474c1c0b5380 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4e78da62-d166-4fbb-98fc-56d284d05383 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f0f534-51d3-4361-9773-5f519f90758f · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 877ec80c-6122-4bc2-ac7d-1dc5a236e381 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning CogVLM: Visual Expert for Pretrained Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bc417ed-0584-4914-8c79-bd0e4b54967a · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning OS-Copilot: Towards Generalist Computer Agents with Self-Improvement
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f24583d1-8fe3-4e4a-a157-18ab71995775 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff5af437-adc0-4c0f-a333-e2d4959a1f46 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ae2c90-139b-4331-bdb1-c62235a0a44c · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6271e389-7c5e-4eff-bb41-327ea28492b7 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5d676d2-4c45-4be3-86af-5206e5ee07b9 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df741db4-474f-4fa7-a644-5d9ec810bf1c · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning , " * write output.state after.block = add.period write newline
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29e19b77-4c9a-42f5-8996-18d562f27639 · outbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning write newline
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 389aaa17-af3a-47dd-a822-c935abd24dc0 · inbound
Large Language Model-Brained GUI Agents: A Survey Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning
Reference 216
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.