Pith. sign in

Paper Citation Record · LEDGER

Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2210.03347.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2210.03347 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:02.118521Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

46
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 82f926bb-30c7-4d22-ab49-736293420864 · inbound

GPT-4V(ision) is a Generalist Web Agent, if Grounded cites this paper.

GPT-4V(ision) is a Generalist Web Agent, if Grounded Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:40:13.936441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T19:40:13.816153Z digest=sha256:dd89c4639ccba3aecef47ba853a8f8148be22d53cd6c38680a49ab90b3b0d21f

Observation 20fee837-c177-47bd-9c3c-bce890a6f26e · inbound

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents cites this paper.

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:29:27.731687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T09:29:27.173784Z digest=sha256:10e3cf00e0026fa42943a7554d7dbad27ff167dff541b5a60c82a70381f66e8a

Observation 906b7d0b-3b79-4991-82bc-aee6a1aaf401 · inbound

Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective cites this paper.

Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:14:27.818364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:14:27.818364Z digest=sha256:2f256569057d6d7c9198124ee74be7430bfb78f68dde5af8bb3064b7e52645a8

Observation 77827c3d-3760-42be-b18b-06509bbe4f35 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 219

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:02.118521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:02.118521Z digest=sha256:8dbb34dd8aeb6730dac8eed42bb481dd03994d2c8fecae5809c3e25d768219e3

Observation d343a2fb-2145-4c82-ab37-cec424bc1cec · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.721321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.721321Z digest=sha256:5e8b1692e0935c62392dd0472b1106eb1fb3a1a62dd97280e797aad9304d9622

Observation a97f4161-3543-47fd-b634-f125db8ec85d · inbound

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations cites this paper.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.183976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.183976Z digest=sha256:2ec3aeb6849418bdca8044dd9ffd3e68835772c0d59eaf01d303ebc870efc7d9

Observation 8ab3fd26-ed02-41c2-8760-fa99c7fade3e · inbound

Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning cites this paper.

Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T16:34:40.353475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:34:40.353475Z digest=sha256:e86981685c9ddea0ac8e75f7b0280a99044c4cc01ffaac174b20221734400a97

Observation 3c85f384-4edd-41bf-8534-18187f05fe2a · inbound

ROSA: Addressing text understanding challenges in photographs via ROtated SAmpling cites this paper.

ROSA: Addressing text understanding challenges in photographs via ROtated SAmpling Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:59.387453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:03:59.387453Z digest=sha256:258ab792ee86231506b4cc9266221ab38b95df73b212ccc21f998bb267f0b657

Observation fed8b3d8-09b9-48a8-beeb-74bfe88388d9 · inbound

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools cites this paper.

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:36.762356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:27:36.762356Z digest=sha256:bbd837faf4ba23badc04ee23612a60d5f0fca6da159c1e0cfdf03070690de849

Observation 9882f442-de08-46bf-b6d3-00bd7f5cee3c · inbound

Reverse Browser: Vector-Image-to-Code Generator cites this paper.

Reverse Browser: Vector-Image-to-Code Generator Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T05:46:40.906386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:46:40.906386Z digest=sha256:9a3aa2fce7382d720ec009ad664892831931b6b8d94059c9273e3c50ff3ce5d6

Observation d21f2ba5-2268-4a35-94c3-47002be82ff8 · inbound

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding cites this paper.

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:32:36.520554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T08:30:50.984873Z digest=sha256:21c0101a9dd1c069e25eca13007ce95be532aead06ec07870f6324c51e3ea0d1

Observation c2638527-bb42-4cbc-a2ef-8b36c9c55188 · inbound

Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation cites this paper.

Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T06:07:00.634725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:07:00.634725Z digest=sha256:38afa0f14ba280fc720190ba9201fdcccbb31daa73cc3f9ea5bd46651b4de1fa

Observation 785fbaf9-8a17-4fb1-984c-a15bf4f060b2 · inbound

PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading cites this paper.

PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:27:44.691324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T10:24:58.444855Z digest=sha256:a726a8d3f59ef43b3e81e2ca8b5204caf541fd30ed2bc43b98d7baa8527b9849

Observation d983b81e-7374-4113-ab0c-d9e68e9686e9 · inbound

From Handwriting to Structured Data: Benchmarking AI Digitisation of Handwritten Forms cites this paper.

From Handwriting to Structured Data: Benchmarking AI Digitisation of Handwritten Forms Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:16:02.598794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:08:30.056580Z digest=sha256:da745bb50f06cebeb03678bfb2705fa6dbbce28fda28ec3fc6affff046faeed1

Observation 2def160d-64ec-4783-92bd-705abfe97061 · inbound

MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding cites this paper.

MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:12:50.405139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-19T22:08:27.511727Z digest=sha256:995d00ac9588fce94afd911e7e43d8f9cfff883b62d93ca4cc702da30fb56f7a

Observation 62eb262e-d044-4415-a72b-dd82e034a00c · inbound

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs cites this paper.

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.462736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T11:08:36.451147Z digest=sha256:5a613943f17fc6ef84d5134db6c6ab61e770b80c4fca473c09facb014b894fe9

Observation 58987a42-728b-49c5-91d1-0625abc55cec · inbound

Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts cites this paper.

Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:47:37.746354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T13:42:21.520627Z digest=sha256:a23e9e6e34169fb4c69a05bef71f7ebe534d915887a8a0b27ebbd4e960ed7214

Observation c7788795-ae57-4fb6-aad1-c43ba0ca2b9a · inbound

Mixture of Cognitive Experts in Large Vision-Language Models cites this paper.

Mixture of Cognitive Experts in Large Vision-Language Models Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T09:13:07.507164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:13:07.507164Z digest=sha256:b530ad430746cbcad676c0c51bbeb529477f3fd4969a96803152946f3349017d

Observation 0ab67878-b9e3-4eed-a113-05c852b32d5c · inbound

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images cites this paper.

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T08:35:42.464663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:35:42.464663Z digest=sha256:64d0ff5c4e8f7413af7aae315650751d1ea52cdd7596ad8bc390f7644793988d

Observation 18b0d5b4-db97-414a-b5f6-b98474a7c1c2 · inbound

Chart Deception in Vision-Language Models: From Vulnerability to Mitigation cites this paper.

Chart Deception in Vision-Language Models: From Vulnerability to Mitigation Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T11:31:21.464350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:31:21.464350Z digest=sha256:11107349bfb579bdb1e2a30eb0456c41be8bcb73d84976921ad04b31931e9356