Pith. sign in

REVIEW 4 cited by

ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.04615 v3 pith:H2VE2L3U submitted 2024-02-07 cs.CV cs.AI

classification cs.CVcs.AI
keywords modeldatasetsinfographicsscreenscreenaiannotationdesigndocvqa
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Screen user interfaces (UIs) and infographics, sharing similar visual language and design principles, play important roles in human communication and human-machine interaction. We introduce ScreenAI, a vision-language model that specializes in UI and infographics understanding. Our model improves upon the PaLI architecture with the flexible patching strategy of pix2struct and is trained on a unique mixture of datasets. At the heart of this mixture is a novel screen annotation task in which the model has to identify the type and location of UI elements. We use these text annotations to describe screens to Large Language Models and automatically generate question-answering (QA), UI navigation, and summarization training datasets at scale. We run ablation studies to demonstrate the impact of these design choices. At only 5B parameters, ScreenAI achieves new state-of-the-artresults on UI- and infographics-based tasks (Multi-page DocVQA, WebSRC, MoTIF and Widget Captioning), and new best-in-class performance on others (Chart QA, DocVQA, and InfographicVQA) compared to models of similar size. Finally, we release three new datasets: one focused on the screen annotation task and two others focused on question answering.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GUI-AC: Enhancing Continual Learning in GUI Agents

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    GUI-AC stabilizes RFT for non-stationary GUI data by down-weighting noisy advantages and relaxing clipping bounds via a grounding certainty term.

  2. Task Mode: Dynamic Filtering for Task-Specific Web Navigation using LLMs

    cs.HC 2025-07 conditional novelty 6.0 of 10

    An LLM-powered browser extension that filters webpages to task-relevant content reduced screen reader users' task completion time by about half in a 12-participant study.

  3. Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Explicitly storing short-term and long-term text memories, instead of raw screenshots, improves GUI agent accuracy on cross-app tasks, and a new annotated dataset helps 7B models approach 72B-level memory use.

  4. IDEA: Augmenting Design Intelligence through Design Space Exploration

    cs.HC 2025-06 conditional novelty 5.0 of 10

    IDEA combines LLM-generated constraints with Monte Carlo Tree Search over a formal design space to automate design decision-making in data storytelling and pictorial visualization.

Pith tools