Pith. sign in

Paper Citation Record · LEDGER

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding

As of 10 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2506.23491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23491 v3

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:44:51.870568Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 992594d9-e92b-4bc9-97ea-9ebde3e7b141 · outbound

This paper cites Qwen2.5-VL Technical Report.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:50.843200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:50.843200Z digest=sha256:b1f5a8d4135169fc5c536518ff4da4965b0fbe189d91e9a6e1a71f4fee84ce0c

Observation 4324f7a0-8bdd-4710-b96e-c080e1e20730 · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:50.910470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:50.910470Z digest=sha256:4ac6ef5c820b280c1f260f1ec613c32675f70a9b1f7ca860e2d40b365b47c8a1

Observation 6e7e588f-5aa1-4c2b-95f4-f77288f202db · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.037486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.037486Z digest=sha256:2d35d3c69c98fa275078d4d527500b57e4c33f6703fa02dbb04e0e2390d4a209

Observation 10bf4088-d97f-444f-af54-c6f7aa581c79 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.133054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.133054Z digest=sha256:6559993801e037677dd8b52e5f3c0ce07c75fc6bd62413b45d7cb217fc9c67b7

Observation cc4f8ab6-75d4-4a7d-bf8b-aa198f685dbe · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding CogAgent: A Visual Language Model for GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.212777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.212777Z digest=sha256:711feff6b1a6265530fdc896e07795cd42a4823f09fb83b4e2244e367eec317f

Observation 943b2969-9783-4047-8f10-92d7dd04fe24 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding LoRA: Low-Rank Adaptation of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.300123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.300123Z digest=sha256:18920043e166233e068a3ab910abb4fab1fe5d1e418077d473e6f87f485ade34

Observation e3cbe5b5-0cc5-473a-8312-81e95ddd2e71 · outbound

This paper cites ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.404442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.404442Z digest=sha256:35fe15ac0274f67b5bae8fc7b6a8e9c59a3d82a6a71455a38bc2e74cb66e78eb

Observation 44f450de-7495-43d0-9caa-5f74dfb736c0 · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.512452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.512452Z digest=sha256:96fe3a2f613db79da5f91a4f33aad19ce9c3e6c849a5a444ea3b50d8b34e526f

Observation 5db9d50c-133d-4525-9eb1-e66ec95fc4c8 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.637524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.637524Z digest=sha256:2804a9c36ed01775c54127b2023bb4a246667008aba7670a02655bc80015ca49

Observation 1df72db2-919b-4610-876e-e429fe62d5e5 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.722301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.722301Z digest=sha256:9b393e82e81b13cfb871d8b978b11069547b76b8e540f72420db2cc7a91dec89

Observation ffa7ece6-be0e-4642-ae6d-871d955bb696 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.798796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.798796Z digest=sha256:ad8c19493c14d58df829619027ef1c834417c005ea9795b10dcf15a08c71ba48

Observation 30a0a296-4d0d-4cc7-8987-ab2dfd049432 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.870568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.870568Z digest=sha256:0bc748237f6b3dc634098770b6c84f7b6820bc2df239038f679195bcb8bbedf9

Pith citing papers

No inbound Pith citation observations are available.