Pith. sign in

Paper Citation Record · LEDGER

ScreenAI: A Vision-Language Model for UI and Infographics Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2402.04615.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04615 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:27:37.560929Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation efc94861-acd1-47a8-ac82-95857e7e65c7 · inbound

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments cites this paper.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.469196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:919ce17e1fb80d66a7ddb720978141f1af57d509d85f7d8a5b85cb280eabf9e1

Observation d5f7a6e5-b103-4298-b2fa-6913a60e9fd6 · inbound

IDEA: Augmenting Design Intelligence through Design Space Exploration cites this paper.

IDEA: Augmenting Design Intelligence through Design Space Exploration ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:37.560929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:37.560929Z digest=sha256:d880e7b033bbc66eebe94068c3934e8f2bc66200831923f5195f12f135bb6a88

Observation 2ad40dd3-42fb-4acf-ae29-0472767675d6 · inbound

Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation cites this paper.

Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:03.738304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:03.738304Z digest=sha256:f8cf1f5c32f091149fd03d9af3a702ed1ad17fd354b1cd92db2b7f96def3247e

Observation 036f74a9-e666-4130-9f28-13f53eba4876 · inbound

Task Mode: Dynamic Filtering for Task-Specific Web Navigation using LLMs cites this paper.

Task Mode: Dynamic Filtering for Task-Specific Web Navigation using LLMs ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:52:28.554432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:52:28.554432Z digest=sha256:9fdfb554ecd430594e677800f9dc29afb640cdae91a9f7f9d98251389077659a

Observation 88d2594b-f915-4cea-af08-e33d0c724c6f · inbound

MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents cites this paper.

MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:42:48.151905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T18:42:29.744124Z digest=sha256:820cb288e5ec2751c48a634cffa76f4825b7526460175d6d8b1fd950f0e390fc

Observation 76aec0dd-56a3-4f8f-bb24-d6e1b854fda7 · inbound

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding cites this paper.

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:32:36.507239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:30:50.984873Z digest=sha256:e0cb2abff371f0ecc5b3e1240f11391cef423daf27cd998f2a4180eae4b97b40

Observation 47ad18ac-f905-425c-b1e9-3c002ea3a82e · inbound

Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible cites this paper.

Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.662122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:59:32.583411Z digest=sha256:a370226870d0c76117e6f3258d350e196ea5a3033203ac18f2da0f26b6e66003

Observation 4c16561f-46ad-42fa-9a2e-2fa43fc39c93 · inbound

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web cites this paper.

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:40:58.946027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:00:34.401698Z digest=sha256:527ab1cbd7f5ce5fff3f4e8a024f34e135bd26b121e34aa9910b1b8b8af3f006

Observation c2861502-5fa4-4b97-8ed2-a8e5a05b0873 · inbound

VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation cites this paper.

VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T22:34:07.748388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T22:24:45.045405Z digest=sha256:968134ad0a46e634284b4ebb8bde1fa345a35e2b54e4633792e08888e6360294

Observation b3f0cf82-9d98-421a-a7e8-b884b80bf2fc · inbound

PageGuide: Browser extension to assist users in navigating a webpage and locating information cites this paper.

PageGuide: Browser extension to assist users in navigating a webpage and locating information ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:26:13.939618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T05:39:04.895180Z digest=sha256:e65544c569fd1f07ba2656e648840e6d5568052484684255895016208c964c8e

Observation b3790203-6fdc-443a-aea9-0e5b544f7cbf · inbound

PageGuide: Browser extension to assist users in navigating a webpage and locating information cites this paper.

PageGuide: Browser extension to assist users in navigating a webpage and locating information ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:25:40.435071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T09:08:44.608520Z digest=sha256:0dd6aa5a0452bf62e9555a1ef960ebae16ab0903f163a2f73c2b8a2e558a0d48

Observation b9b0b7f7-0e78-4a06-8402-1f8dbb1de7ef · inbound

A Pattern Language for Resilient Visual Agents cites this paper.

A Pattern Language for Resilient Visual Agents ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:11:27.824183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T07:27:05.873162Z digest=sha256:2fec9ca7a819893320db9f5c178a14a9e1de7e53683cb0356437e5bcf0759b62

Observation 6f3a98fe-392d-4559-8b81-c780be028b7e · inbound

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability cites this paper.

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:56.807368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:15:19.239355Z digest=sha256:1675aedf62c59a0677b125d7202fcc445095f06641ec433d8da4814a7ed22bda

Observation 5f1a9cde-5f0d-40d6-ac53-97bcf4a89029 · inbound

MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding cites this paper.

MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:12:50.557450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T22:08:27.511727Z digest=sha256:f8c4f40254687edf9ad452a4738cf313b5b099e4343e19166f6a9b4d6cc237fb

Observation bb58c6cf-a02a-46ba-ba02-78cc3f5ff6b0 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.756019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:31bb13afef1c3292f814833a5bcbfc215d9943711e20538d15ce6c381f26c80f

Observation d9d3ccf4-5e45-46a3-b8cf-8b1a5794e987 · inbound

GUI-AC: Enhancing Continual Learning in GUI Agents cites this paper.

GUI-AC: Enhancing Continual Learning in GUI Agents ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:27:36.628860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T13:56:09.049753Z digest=sha256:c21ca7891209332432a58ac43b8269dfe83171e499432c9c0823a821a0307090

Observation 2af225e4-d973-49a8-9c8b-2681469850b2 · inbound

GUI-AC: Enhancing Continual Learning in GUI Agents cites this paper.

GUI-AC: Enhancing Continual Learning in GUI Agents ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T14:27:05.589465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:27:05.589465Z digest=sha256:6812fd598e67a531857cd5e924c9428f3cf45b784c4ba415c5f59678545ce38f

Observation 99a2a8dc-fd5b-4d17-8eac-36529e26e391 · inbound

AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction cites this paper.

AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:44.891305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:41:39.932822Z digest=sha256:9019e6b384d3c794c7f5bb373c5b2532d733bb55bc50f7c3243380e4e880e789

Observation 6f31c190-bb42-403e-bb4a-c5a56f157b2d · inbound

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation cites this paper.

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:21.918959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:15:47.056229Z digest=sha256:094030d316ae26f75639c9f0736112070ad6ad04a84950c94199c22bb204f8ab