Pith. sign in

REVIEW 7 cited by

Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.03347 v2 pith:GFNRU7NU submitted 2022-10-07 cs.CL cs.CV

classification cs.CLcs.CV
keywords languagepretrainingmodelpix2structpretrainedtasksvisualdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visually-situated language is ubiquitous -- sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domain-specific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image-to-text model for purely visual language understanding, which can be finetuned on tasks containing visually-situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provides a large source of pretraining data well suited to the diversity of downstream tasks. Intuitively, this objective subsumes common pretraining signals such as OCR, language modeling, image captioning. In addition to the novel pretraining strategy, we introduce a variable-resolution input representation and a more flexible integration of language and vision inputs, where language prompts such as questions are rendered directly on top of the input image. For the first time, we show that a single pretrained model can achieve state-of-the-art results in six out of nine tasks across four domains: documents, illustrations, user interfaces, and natural images.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 46 citations worldwide. Full citation record

  1. Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images

    cs.SE 2026-07 conditional novelty 7.0 of 10

    Across 675 paired API calls, image input-token reductions are 86.5% (Anthropic), 80.6% (OpenAI), and 75.8% (Gemini) under token-volume weighting, with Gemini images costing far more than text below 200 lines.

  2. Mixture of Cognitive Experts in Large Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Routing CV experts into atomic evidence then Bloom-staged verbalization improves LVLM benchmarks and yields measurable query-conditioned reasoning traces.

  3. Chart Deception in Vision-Language Models: From Vulnerability to Mitigation

    cs.AI 2026-06 conditional novelty 6.0 of 10

    Misleading chart designs shift vision-language models' answers away from the true data interpretation; a paired benchmark measures this shift, and a model-extracted chart summary reduces it for most models.

  4. Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A 7B VLM trained with a structured chart-specification reward beats larger and commercial models on chart-to-code benchmarks using only 3K-4K training samples.

  5. Reverse Browser: Vector-Image-to-Code Generator

    cs.SE 2025-09 conditional novelty 5.0 of 10

    An open-weights system that turns vector images of web designs into HTML/CSS, with new datasets and a multi-scale pixel metric, though accuracy remains below production quality.

  6. ROSA: Addressing text understanding challenges in photographs via ROtated SAmpling

    cs.CL 2025-06 conditional novelty 5.0 of 10

    ROSA, which combines four image rotations with likelihood-ranked sampling, improves VQA accuracy on misoriented text by up to 11.7 absolute points over greedy decoding.

  7. On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools

    cs.IR 2025-06 conditional novelty 4.0 of 10

    Preprocessing financial PDFs into text, tables, and chart data with existing tools improves LLM question-answering accuracy over direct GPT-4o image input in a small private evaluation.

Pith tools