Pith. sign in

Paper Citation Record · LEDGER

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding

As of 18 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2506.23491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23491 v3

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:44:51.870568Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 992594d9-e92b-4bc9-97ea-9ebde3e7b141 · outbound

This paper cites Qwen2.5-VL Technical Report.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:50.843200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:50.843200Z digest=sha256:061036d0ebe8d36b68a60144f638dc0b7860482343c73498561489fffb9c269c

Observation 4324f7a0-8bdd-4710-b96e-c080e1e20730 · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:50.910470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:50.910470Z digest=sha256:e58a37586b03222488fe3d5dc7309fb040292d9069b84ca860076bc5c0f849c6

Observation 6e7e588f-5aa1-4c2b-95f4-f77288f202db · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.037486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.037486Z digest=sha256:543dfff96f419ac81c046548615fe4c4a6f7bbc9a661ff7db325f351e10c9e11

Observation 10bf4088-d97f-444f-af54-c6f7aa581c79 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.133054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.133054Z digest=sha256:334cece574f59fa5e9c70c1b62ccaed00a34a8ba69fdb66d41a95b75bb903887

Observation cc4f8ab6-75d4-4a7d-bf8b-aa198f685dbe · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding CogAgent: A Visual Language Model for GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.212777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.212777Z digest=sha256:905d7ab8014fdf777f177bd1c9155dd952b0ab89cb493655a25d1afc164330ec

Observation 943b2969-9783-4047-8f10-92d7dd04fe24 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding LoRA: Low-Rank Adaptation of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.300123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.300123Z digest=sha256:64930374a82e54802a0738c9944b969acb83a640cb8016b321faea946eb6575f

Observation e3cbe5b5-0cc5-473a-8312-81e95ddd2e71 · outbound

This paper cites ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.404442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.404442Z digest=sha256:3f70a1dd06c9d4b051ab9ab51dc858a0d0ad76b810d45d9ac9bd1e314653c688

Observation 44f450de-7495-43d0-9caa-5f74dfb736c0 · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.512452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.512452Z digest=sha256:afb2028a806def51e4168d719af5bf9f76278d17c3caf161e997f0af63f4094c

Observation 5db9d50c-133d-4525-9eb1-e66ec95fc4c8 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.637524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.637524Z digest=sha256:eef9b94fc7ee5b253bf2e409ff0da70f01ea0826591c02085370a7f4af4eb0da

Observation 1df72db2-919b-4610-876e-e429fe62d5e5 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.722301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.722301Z digest=sha256:5504bd20960342c75d6892b4f7e4370b09cf81cf7b73e240eac0135f94f6b790

Observation ffa7ece6-be0e-4642-ae6d-871d955bb696 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.798796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.798796Z digest=sha256:df627dedabe987559c35b18930615ebb466978cb03042a9cb61645da35536937

Observation 30a0a296-4d0d-4cc7-8987-ab2dfd049432 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.870568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.870568Z digest=sha256:69ca8aebc14ab0b1308416d35e93cec00f7ab4da4dbb1fb520afc594c2d215ce

Pith citing papers

No inbound Pith citation observations are available.