Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2210.03347.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:02.118521Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
46
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 82f926bb-30c7-4d22-ab49-736293420864 · inbound
GPT-4V(ision) is a Generalist Web Agent, if Grounded Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 20fee837-c177-47bd-9c3c-bce890a6f26e · inbound
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 906b7d0b-3b79-4991-82bc-aee6a1aaf401 · inbound
Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77827c3d-3760-42be-b18b-06509bbe4f35 · inbound
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 219
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d343a2fb-2145-4c82-ab37-cec424bc1cec · inbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a97f4161-3543-47fd-b634-f125db8ec85d · inbound
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab3fd26-ed02-41c2-8760-fa99c7fade3e · inbound
Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c85f384-4edd-41bf-8534-18187f05fe2a · inbound
ROSA: Addressing text understanding challenges in photographs via ROtated SAmpling Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fed8b3d8-09b9-48a8-beeb-74bfe88388d9 · inbound
On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9882f442-de08-46bf-b6d3-00bd7f5cee3c · inbound
Reverse Browser: Vector-Image-to-Code Generator Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d21f2ba5-2268-4a35-94c3-47002be82ff8 · inbound
CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c2638527-bb42-4cbc-a2ef-8b36c9c55188 · inbound
Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 785fbaf9-8a17-4fb1-984c-a15bf4f060b2 · inbound
PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d983b81e-7374-4113-ab0c-d9e68e9686e9 · inbound
From Handwriting to Structured Data: Benchmarking AI Digitisation of Handwritten Forms Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2def160d-64ec-4783-92bd-705abfe97061 · inbound
MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 62eb262e-d044-4415-a72b-dd82e034a00c · inbound
Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 58987a42-728b-49c5-91d1-0625abc55cec · inbound
Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c7788795-ae57-4fb6-aad1-c43ba0ca2b9a · inbound
Mixture of Cognitive Experts in Large Vision-Language Models Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ab67878-b9e3-4eed-a113-05c852b32d5c · inbound
Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18b0d5b4-db97-414a-b5f6-b98474a7c1c2 · inbound
Chart Deception in Vision-Language Models: From Vulnerability to Mitigation Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.