Pith. sign in

Paper Citation Record · LEDGER

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning

As of 11 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2412.10840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10840 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:38:09.367743Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T11:08:27.472508Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T11:08:27.773962Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d2fb6641-a1f2-4151-9814-6decf45c19e7 · outbound

This paper cites an unresolved cited work.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:38:09.941800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:38:09.180381Z digest=sha256:7d628cb692a072bd8e54fc98ca044f81d65b3a321ee6a139349ad0621d35c236

Observation 8d589e4d-f355-444a-857e-017b8961dca2 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.187380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.187380Z digest=sha256:757ba22cc91bb1ebb4f0a67786e6117fffc0dd041612871cebe4e54e83feb457

Observation 6e68d410-381b-432e-bb3a-b05faf70a060 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.194044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.194044Z digest=sha256:390d296b88c348f94c9c1deb740337849aca250f23efa8cfcadb1219169adb83

Observation 7a4c49e1-d45c-4af0-b7f5-a36a784cf6ea · outbound

This paper cites GUICourse: From General Vision Language Models to Versatile GUI Agents.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning GUICourse: From General Vision Language Models to Versatile GUI Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.199600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.199600Z digest=sha256:922d59a6f47fbe17038096beea796aa4669b72d1d26ccb2906736fd70e53b797

Observation b0c37753-debe-4836-b087-de80a8bccf36 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.204777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.204777Z digest=sha256:bcad1a2314c6615e1b28dd0db2aedd64dd960ce7de28c614559a77eb474dd2b4

Observation bab83b3f-e46b-49d0-8d44-304eb99b16cb · outbound

This paper cites an unresolved cited work.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:38:09.927142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:38:09.210192Z digest=sha256:632114317c8b82ea47e424e95a6fe10b5635bac8bf624f8562e7aca6fc0e1995

Observation 19466a30-d03d-4bd9-b65f-84aec5094156 · outbound

This paper cites Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.215802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.215802Z digest=sha256:463ea6a13b6816d3810b5d6826f182efa7015976f2d132e8609d7a11e21941fd

Observation b47c38f7-4a7c-4101-981f-85d4b65e43e7 · outbound

This paper cites an unresolved cited work.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:38:09.910905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:38:09.221584Z digest=sha256:168966396f5cce820ff701b3b64214e39f2bdd518c2805c80cf54363704b4afe

Observation 671e7e49-cdcf-4e7c-b1f1-0df4486ed79f · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.226286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.226286Z digest=sha256:7dfb4032400ac702be2db58d80de8ee59001b7fa726b0f57978341fe2b2a2ab8

Observation 1c4bb37b-564a-411c-aaef-10c9a3e84d39 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning CogAgent: A Visual Language Model for GUI Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.231300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.231300Z digest=sha256:a8c18813384d043e86d40355f92a6edbe59ce46104c819bc2576387c686e0c65

Observation 46557417-933e-4753-8290-db01ddba6a50 · outbound

This paper cites OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.236741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.236741Z digest=sha256:32ab34115133bcb24605f6b5d623ab7fc19c8e53601fa7e151ed486c92b8cfbc

Observation 9f4de35f-9bbb-4fce-8e71-6b714387a071 · outbound

This paper cites H.; Zheng, B.; Deng, X.; Su, Y.; and Chao, W.-L.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning H.; Zheng, B.; Deng, X.; Su, Y.; and Chao, W.-L

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:38:09.893773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:38:09.241809Z digest=sha256:1667ef9fef6cb4147e25655a97cf629a9d054328f401d455ce91115570fdb239

Observation 7d576b52-1acb-471d-a394-c2f1b5e7cbc5 · outbound

This paper cites an unresolved cited work.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:38:09.878425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:38:09.246955Z digest=sha256:bd5256d0fa2e56b79788a970ddb9294225097e4a9d037a5ef850a222812d4192

Observation 325726aa-3788-45eb-b388-0fab7463eed8 · outbound

This paper cites J.-J.; Mitchell, T.; and Myers, B.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning J.-J.; Mitchell, T.; and Myers, B

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:38:09.862422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:38:09.252245Z digest=sha256:bee5775a45e7d12d63b0561ac79021ad921d9f60db9b84d2dfd502df301ce096

Observation 80cd5fbf-6109-4123-bbf0-19010af9b68e · outbound

This paper cites an unresolved cited work.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:38:09.846647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:38:09.257429Z digest=sha256:23f10a3a3be9a8f16e1e458b12b8e4d63f0590eae01a10428c1e38af5851feec

Observation 1e4522ea-6091-48da-8bab-c15f4703a846 · outbound

This paper cites an unresolved cited work.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:38:09.831146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:38:09.263234Z digest=sha256:83c804f1052d718ab83158b7c1f450e60910ec9475f656123cba8297633c7ae5

Observation 8ff745c3-d66f-4160-9ad0-1b8788f23a05 · outbound

This paper cites an unresolved cited work.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.268870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.268870Z digest=sha256:71c1f9bffc7609d9d29ab8f8a4284e872d75451e7779bcd03b1aa87e1647160b

Observation 66533dcb-d213-4fb2-b63a-1e7b8fa93d16 · outbound

This paper cites VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.273738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.273738Z digest=sha256:2c2dfcdc53d0afbfde7edb10e8264ed03f30b8a30050d5fdbec0a3af057635dc

Observation a54594b6-7844-4211-9ff0-b1403db9f4c4 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.279824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.279824Z digest=sha256:6d2903b649060a1bb1dd5145f7e26eee768df52500f27007a875899f204d18b3

Observation ef39dc5e-3abd-4498-99ad-5834fcfa612e · outbound

This paper cites an unresolved cited work.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:38:09.805332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:38:09.284887Z digest=sha256:ce70e5c4fc74593013a09f7a626bc9500397587b07972b4ecd6353f03179d1a2

Observation cefd08b3-a486-4088-a248-fd43a289b032 · outbound

This paper cites MobileFlow: A Multimodal LLM For Mobile GUI Agent.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning MobileFlow: A Multimodal LLM For Mobile GUI Agent

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.289700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.289700Z digest=sha256:d4915252f21c0935e6b665467e174156180e7866cbfe9006c704f78a3dc580fc

Observation 19cc0eb6-e08e-4e86-a1af-063276b49ad6 · outbound

This paper cites an unresolved cited work.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.294858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.294858Z digest=sha256:f305110da0d96504c075c4f5fbf4f47ee8fd3172d3a8d2fa51cc8f529d24bec4

Observation 5da58cc1-2bc8-448e-83f4-0338dcb8a2d9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.299593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.299593Z digest=sha256:0ebcae5093924f288553ee27eb3df633ff42cd61d60dcfc8daf803cf109d72d0

Observation c99ab3dd-7d00-4d12-a2a3-49896aaa80ec · outbound

This paper cites an unresolved cited work.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:38:09.779042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:38:09.304938Z digest=sha256:3f193841c1e6980e7a0f5eba9f046ca552e1700ec53748f03cadcd8cf23c2c01

Observation 33220915-e667-465f-8ad8-474c1c0b5380 · outbound

This paper cites an unresolved cited work.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:38:09.763859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:38:09.310195Z digest=sha256:54eefcc94404b041241f3c4d14691ab100b5896fc68abb9aab4e872022fe8659

Observation 4e78da62-d166-4fbb-98fc-56d284d05383 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.315695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.315695Z digest=sha256:5c677ad0e8348a87b6eab4bebf1a3b7cffee4f979b0f4f389ab7895b3877fe0f

Observation 50f0f534-51d3-4361-9773-5f519f90758f · outbound

This paper cites MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.322828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.322828Z digest=sha256:75517776cad4dc6ce62c6cf9a7179b1b710de1dc46203a3468e35dd2c57d6d80

Observation 877ec80c-6122-4bc2-ac7d-1dc5a236e381 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning CogVLM: Visual Expert for Pretrained Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.327899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.327899Z digest=sha256:414b843d8a80454411a0c27a1df3fea3c4c0f8a41d899c7f19a2fc7a5fb09cef

Observation 5bc417ed-0584-4914-8c79-bd0e4b54967a · outbound

This paper cites OS-Copilot: Towards Generalist Computer Agents with Self-Improvement.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning OS-Copilot: Towards Generalist Computer Agents with Self-Improvement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.332870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.332870Z digest=sha256:8c5ffbcc56975c8c6460c0a4a793a0341129cc795fcc59f269cee0b740c1b207

Observation f24583d1-8fe3-4e4a-a157-18ab71995775 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.338376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.338376Z digest=sha256:6a314eb363b672e9c5baec19acd7f809df3d49ebf965f7f6392439f287171edb

Observation ff5af437-adc0-4c0f-a333-e2d4959a1f46 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.343277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.343277Z digest=sha256:ff80e234cfb32905cd6a5ba323ee98d42aad01e2c66c585ee629901df22f8c7a

Observation 71ae2c90-139b-4331-bdb1-c62235a0a44c · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.348135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.348135Z digest=sha256:4d46dea75563c99fb8ab610f9046f855b1e8d64a15efbf1a4020201bd21ccfcd

Observation 6271e389-7c5e-4eff-bb41-327ea28492b7 · outbound

This paper cites Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.352732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.352732Z digest=sha256:18bba578a1e7ecc2a72c5019f93ff682b0dd7929fefc14ddc0b6e2d68fc4e2bb

Observation f5d676d2-4c45-4be3-86af-5206e5ee07b9 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.357490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.357490Z digest=sha256:3dbaaacb845d4aecad750afbb01d79929f5350cb305788e3f6092f84117250bf

Observation df741db4-474f-4fa7-a644-5d9ec810bf1c · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning , " * write output.state after.block = add.period write newline

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.362608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.362608Z digest=sha256:0edbb07ec7c5fa5c4c779be3c826cb0049fd80eaa152e15eb20b3d957e7b265c

Observation 29e19b77-4c9a-42f5-8996-18d562f27639 · outbound

This paper cites write newline.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning write newline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.367743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.367743Z digest=sha256:23ea6f4d9dabdf1608664a5185566faeaca0b00a69d120b494e03ca7b0e54d56

Pith citing papers

Observation 389aaa17-af3a-47dd-a822-c935abd24dc0 · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning

Reference 216

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:27.775558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:0dfb65a17edb1c9433cd559b9fd2c957eefe6aa64569007371fbbedff372c2bf