Pith. sign in

Paper Citation Record · LEDGER

Scalable Pre-training of Large Autoregressive Image Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2401.08541.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.08541 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:06:18.216606Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cc51c86e-8447-4528-87ac-473afe76094a · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Scalable Pre-training of Large Autoregressive Image Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.398124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:7486d0858229754b00bf785bb134d88fbd498db616f487d118590e16c07024a8

Observation 25897824-a30e-4eba-b7bb-6b61635c7873 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Scalable Pre-training of Large Autoregressive Image Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:22:37.351022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:1ec072759a6e78039946bbba225bd72d8f06acd285caf23603eac27758ee6e6d

Observation cce1f160-3ff1-43ee-b35f-ccb84a9f3b8b · inbound

Visual Pre-Training on Unlabeled Images using Reinforcement Learning cites this paper.

Visual Pre-Training on Unlabeled Images using Reinforcement Learning Scalable Pre-training of Large Autoregressive Image Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:18.216606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:06:18.216606Z digest=sha256:2d59118ed68720976e1155a3f5f8eb1112e80d4592916864d480911f22c5802d

Observation 4d8ec47f-687d-4fea-ab49-a7e1a0a11765 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scalable Pre-training of Large Autoregressive Image Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.825159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.825159Z digest=sha256:67471291f8da172149606a63db00baa89c694f480eb3c0a6775c9886f21d2961

Observation a87c79cf-73ab-4575-9598-d4760d41a977 · inbound

Hita: Holistic Tokenizer for Autoregressive Image Generation cites this paper.

Hita: Holistic Tokenizer for Autoregressive Image Generation Scalable Pre-training of Large Autoregressive Image Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:53.354253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:38:53.354253Z digest=sha256:049be0232710364eb169ca52de1ebca3d63e871e6f8945218cced61fbcd8a8ef

Observation 505638ea-f1f5-4a45-97f1-1463a38f024b · inbound

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning cites this paper.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Scalable Pre-training of Large Autoregressive Image Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:32.780054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:32.780054Z digest=sha256:be6ab785239fd1dba53dae6ee216156e98167d92165d9a9df58968659b69a51e

Observation 028815b8-227f-4d79-ac1f-cbd50fab44f5 · inbound

DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection cites this paper.

DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection Scalable Pre-training of Large Autoregressive Image Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:02.098890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:31:36.180645Z digest=sha256:1735d69ac805b56457161a5100a94f6450c3050da7cd9bbaad4ba57254de41e4

Observation 520346e1-1991-4a4b-a34e-0ef4c86f31bb · inbound

DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection cites this paper.

DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection Scalable Pre-training of Large Autoregressive Image Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:14.411668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:14.411668Z digest=sha256:a9d9bc514355b8c706255e83ece4fdca730d40b2ca7d67e85fffbbdec5ea80dc

Observation 1d1d70fe-05f9-4f77-8685-52af7b180094 · inbound

What Cohort INRs Encode and Where to Freeze Them cites this paper.

What Cohort INRs Encode and Where to Freeze Them Scalable Pre-training of Large Autoregressive Image Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:26.363547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:32:38.878151Z digest=sha256:70a0d0f292540508c74f421a95d2ba2a59980f54dfe8b4a67c0a394648ecc304

Observation bbe2c6da-c266-4685-8d4d-3287e7ca10df · inbound

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice cites this paper.

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice Scalable Pre-training of Large Autoregressive Image Models

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:43:51.072841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T22:41:44.510546Z digest=sha256:31454a871b14488ccac008610fd2568fed2b0d1208acefd7534e83d760414bb3

Observation 3f8d4011-4e1d-4d6f-bcc8-43b69459d3a4 · inbound

Weighted Reverse Convolution for Feature Upsampling cites this paper.

Weighted Reverse Convolution for Feature Upsampling Scalable Pre-training of Large Autoregressive Image Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.683562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T14:24:28.728963Z digest=sha256:890aff163578069a1ba8b20944ce39e553b11cd3d6e7271a6936a3706ec210ad

Observation b111038b-75b4-4db8-82f5-9f4b1e6b2a22 · inbound

Weighted Reverse Convolution for Feature Upsampling cites this paper.

Weighted Reverse Convolution for Feature Upsampling Scalable Pre-training of Large Autoregressive Image Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.817998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T08:15:52.610554Z digest=sha256:20ed4981911d6fc1966610817b2ace13f710a56d1ce4716aa5905dfc1038ddec

Observation 58b74b70-2776-4b13-9694-c0a6fc467f3a · inbound

Uncovering the Latent Potential of Deep Intermediate Representations cites this paper.

Uncovering the Latent Potential of Deep Intermediate Representations Scalable Pre-training of Large Autoregressive Image Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.257486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-25T05:36:24.743558Z digest=sha256:73ab6ec9d12156d17bd513aba320806a5d0ebfe9372482e0012ca7fdcd2fe40b

Observation 2545e4e5-890a-4a15-862c-1485f8a63296 · inbound

A World Model of Radiologist Reading for Medical Image Representation Learning cites this paper.

A World Model of Radiologist Reading for Medical Image Representation Learning Scalable Pre-training of Large Autoregressive Image Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:15:01.230775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:55:50.523140Z digest=sha256:8add40418390255d95ac74bf7d0ee39f0cb2b9bbe23373878a30b1db1eed0db7

Observation 2d749fa7-129a-4666-9008-4868700605f6 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Scalable Pre-training of Large Autoregressive Image Models

Reference 244

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.519117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:1d99ec001908a9ace832079c4378bd05ae94e0e3ee57b33a780d723ac46e6de9

Observation a07b8580-0dfb-4a53-8ea3-122d735ad973 · inbound

The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models cites this paper.

The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models Scalable Pre-training of Large Autoregressive Image Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-25T20:38:19.798073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T20:30:45.814418Z digest=sha256:eb7c4cac0504cb928f933630255ee0cffa7699959f20e65968ec1b92bfac5325

Observation d38ac3e9-eebc-4c52-b111-4f18ac8a2700 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP Scalable Pre-training of Large Autoregressive Image Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.719757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:0ee39439d482f0210133fac761b5a7265573d76663024d2657c24a6878031239

Observation 6bf5396e-4475-40ac-b767-1f148fe582af · inbound

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types cites this paper.

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types Scalable Pre-training of Large Autoregressive Image Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T03:24:13.814557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:24:13.814557Z digest=sha256:09eb16cf862b28b9b30ed5f44abba46308c4b4a61cec29971461cd8f5f2027ab

Observation 9585a693-900a-4e5b-b603-895a16a09f84 · inbound

Foundation Models for Astrophysics cites this paper.

Foundation Models for Astrophysics Scalable Pre-training of Large Autoregressive Image Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:48.707484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:31:48.707484Z digest=sha256:98050a9fc028732078c33b28af5f633be44b1b8b3b9b9d84d328065b3e049951