Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of Autoregressive Pre-training from Videos

As of 11 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 6 inbound Pith citation observations for arXiv:2501.05453.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05453 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:18:08.598227Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:01:59.614460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:36:39.313175Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 125e0653-d6f5-458f-8695-c1b08967dcee · outbound

This paper cites Figure 11 µ-Parameterization Learning Rate: We show that µ-Parameterization (Yang et al., 2022), we can train all width Toto models, with an single optimal learning rate of 2−7.

An Empirical Study of Autoregressive Pre-training from Videos Figure 11 µ-Parameterization Learning Rate: We show that µ-Parameterization (Yang et al., 2022), we can train all width Toto models, with an single optimal learning rate of 2−7

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:09.071666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:08.598227Z digest=sha256:72856eaf2ea072f28ce0953211b7444b21589432fc4db3266111718e5a85abb9

Observation dcb101e3-dee2-4aba-ae17-88454352a2c4 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

An Empirical Study of Autoregressive Pre-training from Videos A Short Note on the Kinetics-700 Human Action Dataset

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.404695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.404695Z digest=sha256:932b4596dc84ac9e38f74ec2ee6f4f6116663d5c19dd6802fd8b9812be8ab533

Observation 125c856e-165d-48cf-bc99-03264d6ea67a · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

An Empirical Study of Autoregressive Pre-training from Videos Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.440032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.440032Z digest=sha256:a10a8a02b0aea2e496edd22c2668c67a6c9162da18b1709cc6135cbf85af6e9d

Observation 28df9a01-00f9-4805-8583-3ea9508be52f · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

An Empirical Study of Autoregressive Pre-training from Videos Scaling Laws for Autoregressive Generative Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.446082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.446082Z digest=sha256:d1258cc4de3f550af2dcb96be8412fc5bef92533ae026517f538a8b2fb6159ae

Observation 173044f7-051f-489e-afce-93349f30517c · outbound

This paper cites Perceptual losses for real-time style transfer and super-resolution.

An Empirical Study of Autoregressive Pre-training from Videos Perceptual losses for real-time style transfer and super-resolution

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:09.111669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:08.454097Z digest=sha256:ce5835bdcfb06bfde7c5e2a452672f9b0aa7efab62325811f35a0556a9244bb5

Observation c7e9951d-d7cc-4d3d-80e4-1c1c40c710b5 · outbound

This paper cites Network In Network.

An Empirical Study of Autoregressive Pre-training from Videos Network In Network

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.466279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.466279Z digest=sha256:e93a4691299cafa81f242117e731161f533588ce1e30c5b2621883e02d19f426

Observation de9ade04-c8b0-444b-bf9b-d83bb87b79ba · outbound

This paper cites In-context Learning and Induction Heads.

An Empirical Study of Autoregressive Pre-training from Videos In-context Learning and Induction Heads

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.480459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.480459Z digest=sha256:4f54eaab110fdf98d998c8640c6a2b68181c85b076f07a4e85b40e4f5837af1e

Observation ba918768-af36-4d98-bca1-acfc8a857c47 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

An Empirical Study of Autoregressive Pre-training from Videos DINOv2: Learning Robust Visual Features without Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.486951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.486951Z digest=sha256:7ccf1c3f044d1f3d916b2006e506bbd529db81b5f401daf935e9d43d0c51d6b7

Observation 903a5970-8225-426e-950a-7dc653e6b249 · outbound

This paper cites Video (language) modeling: a baseline for generative models of natural videos.

An Empirical Study of Autoregressive Pre-training from Videos Video (language) modeling: a baseline for generative models of natural videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.497076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.497076Z digest=sha256:c061a24dae0f2a9a54a5eb3ab0d36f61cd8e42dcf26789498a0946f3a2830e25

Observation 2069eb2d-81af-456d-a1ef-0c1eb791918d · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

An Empirical Study of Autoregressive Pre-training from Videos Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.513202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.513202Z digest=sha256:5fdccb01482412b6b542c992eff2700342c234bfa35f3de6b1c95b77f4aa419b

Observation 5e1b2c61-5c8c-477a-91d1-e1a64246ce42 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

An Empirical Study of Autoregressive Pre-training from Videos LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.518021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.518021Z digest=sha256:2389cd5c2b277808a8846c1d3f18fbde2e293484e8152f7ccc56547d5d7649ca

Observation 28ce68f5-b37b-4853-9ffd-55225236cd83 · outbound

This paper cites Scaling Autoregressive Video Models.

An Empirical Study of Autoregressive Pre-training from Videos Scaling Autoregressive Video Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.532079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.532079Z digest=sha256:ab846c7eea661bae45c0175da4371a99c11def5a80a3f58048015c697c3d39f5

Observation 34ded5a1-c0ae-4b11-b6f1-e0e73f66d51e · outbound

This paper cites Masked Visual Pre-training for Motor Control.

An Empirical Study of Autoregressive Pre-training from Videos Masked Visual Pre-training for Motor Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.542978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.542978Z digest=sha256:ca63bc91a5f6da54f34d7954b21a74320ae4e9469d4798113ff89a3b580b4007

Observation f7754097-acf0-4569-b106-0bc45a05e8a2 · outbound

This paper cites TFCNet: Temporal Fully Connected Networks for Static Unbiased Temporal Reasoning.

An Empirical Study of Autoregressive Pre-training from Videos TFCNet: Temporal Fully Connected Networks for Static Unbiased Temporal Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.557957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.557957Z digest=sha256:65579a5df30e95b4b27e46d1b18858b3597aacb18f4291f9470c9a0e116d70f9

Observation 7994cab7-3320-4d49-b9b2-b2bbd6e767b5 · outbound

This paper cites Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model.

An Empirical Study of Autoregressive Pre-training from Videos Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.581095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.581095Z digest=sha256:d1991911ea767ce1a2f7b186d16e84da4ed7f4771279a11f58cd0dfb581b9923

Observation 3ce465dc-f3f4-477f-8982-9f026c46cf69 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Autoregressive Pre-training from Videos Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:18:09.092643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:08.591288Z digest=sha256:71e796c7028e56f3a1cdfec8946485b7101c200dd9cab8a1fd7e12f0b55b7d73

Observation 19e7aba7-a982-4765-9198-3a9749b7aa97 · outbound

This paper cites GLU Variants Improve Transformer.

An Empirical Study of Autoregressive Pre-training from Videos GLU Variants Improve Transformer

Reference 1951

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.507922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.507922Z digest=sha256:ae087a0bbf7271c20a417341165b268af2cda4e39d0574152a3bd4121ebe0a8a

Observation 5a9326d3-f23a-4ee1-96ff-47a0dd24c251 · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

An Empirical Study of Autoregressive Pre-training from Videos BEiT: BERT Pre-Training of Image Transformers

Reference 1954

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.393198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.393198Z digest=sha256:d77ff3ba6d1af999f65663fbd4e102ed0ea0217093abef8dd6172a0335e48b9f

Observation 8d14d8f7-e564-49d2-ae11-52a2216ad168 · outbound

This paper cites Decoupled Weight Decay Regularization.

An Empirical Study of Autoregressive Pre-training from Videos Decoupled Weight Decay Regularization

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.473512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.473512Z digest=sha256:e916cb364e50330319c162dda41659cdd4ab46d8901999fa9eaa87e819c1bf53

Observation 5edd0f72-243b-45dd-91bc-398f37e9ce29 · outbound

This paper cites Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles.

An Empirical Study of Autoregressive Pre-training from Videos Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.503176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.503176Z digest=sha256:c34397e30a05a5ab0147d1f4dd968cbf1945965eb03a8135291e65dd09130982

Observation 06619f79-214b-48fe-b95c-86df15aa6a53 · outbound

This paper cites The Kinetics Human Action Video Dataset.

An Empirical Study of Autoregressive Pre-training from Videos The Kinetics Human Action Video Dataset

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.460067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.460067Z digest=sha256:803cd29073176cabb0a415aa2f01fbcb65b35f5b6b67eff158acdfa4e8973fdc

Observation c63646bf-f532-4508-ba8d-8831ca267467 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

An Empirical Study of Autoregressive Pre-training from Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.522588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.522588Z digest=sha256:257927c2d2245bd7c59bfb7a5a40b36103b413d62df689be78642bd432111785

Observation 4c7dd642-e5fc-40ea-be9d-2e2d5d59c165 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

An Empirical Study of Autoregressive Pre-training from Videos The 2017 DAVIS Challenge on Video Object Segmentation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.492053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.492053Z digest=sha256:059539e3e00e5d6c6288d14e4a8c031b64905ae4e05e9cf9b298ec6a2b53bc52

Observation dfc0a216-49f9-4de8-a788-9aa47949f5c5 · outbound

This paper cites Scalable Pre-training of Large Autoregressive Image Models.

An Empirical Study of Autoregressive Pre-training from Videos Scalable Pre-training of Large Autoregressive Image Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.416521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.416521Z digest=sha256:1eeb7ac17d115e75aca1d992e4d3c810106296911587e32107ea5ed980973ce4

Observation eee79cca-c31b-4d1e-bb49-349427a8ebb9 · outbound

This paper cites Data Filtering Networks.

An Empirical Study of Autoregressive Pre-training from Videos Data Filtering Networks

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.428054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.428054Z digest=sha256:1c29c36e9c60e49993669d73ea1228123dc23c874fcaf8a7dc77ec5be5508c8b

Observation 8124c0fa-9850-4d79-9fed-f9856dae87d9 · outbound

This paper cites Language Models are Few-Shot Learners.

An Empirical Study of Autoregressive Pre-training from Videos Language Models are Few-Shot Learners

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.398820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.398820Z digest=sha256:ff824f73edf7e088651d394b43ba8af1241b089d8d6120c47c09ad5520adf7c4

Observation ad3ab6d0-712e-41b8-933e-0cba572da75a · outbound

This paper cites CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning.

An Empirical Study of Autoregressive Pre-training from Videos CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.433520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.433520Z digest=sha256:44bf178b83a65d95931a23ba40a07e32b5b74dd77fd4934c2e320a02247e455a

Observation c2568aca-c82d-40ee-89ea-7c92c58075f6 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

An Empirical Study of Autoregressive Pre-training from Videos Imagenet: A large-scale hierarchical image database

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.410401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.410401Z digest=sha256:2d7597b3fb36798fc4712076dfed721bf237046358eb6238dbae817401e75943

Observation cf000250-c70d-4a25-8d36-818ad539de03 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

An Empirical Study of Autoregressive Pre-training from Videos Taming transformers for high-resolution image synthesis

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:09.130230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:08.422538Z digest=sha256:8c1bc1e1c453233cdcdd95399fd009d1fd44fc3e95feb5219362937647252a46

Pith citing papers

Observation 0ab97ceb-e762-42b1-88c2-25fb482fee79 · inbound

Poly-Autoregressive Prediction for Modeling Interactions cites this paper.

Poly-Autoregressive Prediction for Modeling Interactions An Empirical Study of Autoregressive Pre-training from Videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T00:01:59.614460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:01:59.614460Z digest=sha256:390919a8b3954d54f82e6753df6ea32b357010c3998236367cd0922a3e6fbd5b

Observation 50356802-4c7d-4a8a-a832-980328cf2c9b · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network An Empirical Study of Autoregressive Pre-training from Videos

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:15.812012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:006bb0796392374e58ac1240207addf162bceb222ae33fcb2f6de2afdc4545fd

Observation c075b0bc-fdc5-4161-8d3d-a77885c85539 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics An Empirical Study of Autoregressive Pre-training from Videos

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:22:37.506893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:151833b33ffef2ac1dcdbd9da9d69154de0860b12d2b5378771fbb4f82506ff8

Observation 93043d74-8061-46d9-9ec8-2118051a49ef · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning An Empirical Study of Autoregressive Pre-training from Videos

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.155917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:f89b0288d8ba72c2d141e0a18f65dd8fc04a990260ddf072c483c75256a06770

Observation 43be2391-8b97-4bb9-a60f-a74db1b4da80 · inbound

Frozen Forecasting: A Unified Evaluation cites this paper.

Frozen Forecasting: A Unified Evaluation An Empirical Study of Autoregressive Pre-training from Videos

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T03:42:57.283557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T03:42:54.620069Z digest=sha256:008441e78fb18eaf5464141958ed38a2fa403d3a12416ad90fbb805cbc3b988a

Observation 61840861-5c14-40c3-8312-df09a22b62aa · inbound

Uncovering the Latent Potential of Deep Intermediate Representations cites this paper.

Uncovering the Latent Potential of Deep Intermediate Representations An Empirical Study of Autoregressive Pre-training from Videos

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.314533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-25T05:36:24.743558Z digest=sha256:6ae24ee42b2c0fb9081e41d723163457f0136d3efe5db661a2061a3d289fec02