Pith. sign in

Paper Citation Record · LEDGER

VidTok: A Versatile and Open-Source Video Tokenizer

As of 18 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 9 inbound Pith citation observations for arXiv:2412.13061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13061 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:30:23.560193Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:20:14.972158Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:28:55.542980Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77a296c0-a3c4-4f0f-877a-366b7ade12d6 · outbound

This paper cites UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing.

VidTok: A Versatile and Open-Source Video Tokenizer UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.456291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.456291Z digest=sha256:3dc76e9fdfb709a3548f108b994ad723c11798d0773fe01512aefb07d50d0206

Observation 5af6f8cb-127d-4553-9916-701a0d77dd2c · outbound

This paper cites OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model.

VidTok: A Versatile and Open-Source Video Tokenizer OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.476606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.476606Z digest=sha256:3f1dc7ed811bcf8e14c646b409a33d32bdbda3e174f9d6b6f65ba397ad9eebd9

Observation d78e62f2-88b6-4212-affb-e5460ddfd3e7 · outbound

This paper cites Layer Normalization.

VidTok: A Versatile and Open-Source Video Tokenizer Layer Normalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.496120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.496120Z digest=sha256:ccdfe6ab16167119731ceaa80662c836819884b889bb31bee41bd6efd9db3e4c

Observation e239a439-bd92-4ea4-bb36-fb3388c6a8a2 · outbound

This paper cites OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation.

VidTok: A Versatile and Open-Source Video Tokenizer OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.537169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.537169Z digest=sha256:0b42e982b042037a08c7f322907bdeba1f7de01d05e55af18dfd2b7ee243cf72

Observation 97ab6a6c-c23e-4a05-83f9-00e8e492a65d · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

VidTok: A Versatile and Open-Source Video Tokenizer VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.548962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.548962Z digest=sha256:e916ea2c29b086d1a07c48da5bb1f7b10064a96fcdec7d76c2e1a31090a17845

Observation 19f470f6-33c2-4dc8-a6fd-297f4b00601b · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

VidTok: A Versatile and Open-Source Video Tokenizer Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.543239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.543239Z digest=sha256:efe23ac6dc3094f91df800499e432309d88b843248686ded9980fc2dfca99ed9

Observation b007f36e-c663-488b-9de0-2b7091447943 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

VidTok: A Versatile and Open-Source Video Tokenizer Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.521476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.521476Z digest=sha256:7d34e24c12e0c79a588314d174ac8337a3effa73d02af725fd7f8a8af8c6786e

Observation 93ccae42-6308-4ddb-b643-350e976d7735 · outbound

This paper cites Kingma and Max Welling.

VidTok: A Versatile and Open-Source Video Tokenizer Kingma and Max Welling

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:23.900088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:30:23.489151Z digest=sha256:197a01bb7a2520d7056e4c576470cfceab69877ea0b44e2abdd4cebddade6070

Observation f4b6b843-8591-45de-9fc5-9e874022bcf7 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VidTok: A Versatile and Open-Source Video Tokenizer VideoChat: Chat-Centric Video Understanding

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.503825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.503825Z digest=sha256:0139aaf49efb9f6544f91f349b8d8633d58b0ef3c0adbda6796f9de3d6afa29c

Observation 15c788fe-4e63-44af-8e05-70643bec5acb · outbound

This paper cites Mcl-jcv: a jnd-based h.

VidTok: A Versatile and Open-Source Video Tokenizer Mcl-jcv: a jnd-based h

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:23.879226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:30:23.529604Z digest=sha256:bb5b9fd57bb8b1fd241f363d59f0790b572418edea4e037afd619d6523c19a28

Observation 26e6b10e-eba2-4e9f-b75a-467e885db8f3 · outbound

This paper cites Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators.

VidTok: A Versatile and Open-Source Video Tokenizer Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.560193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.560193Z digest=sha256:fb7ff2bccfe61f68123bde8b47e291c801a679210d814594ad1af770c14d3094

Observation 29f3b717-890d-4ee9-83b6-8e1dc5890046 · outbound

This paper cites aMUSEd: An Open MUSE Reproduction.

VidTok: A Versatile and Open-Source Video Tokenizer aMUSEd: An Open MUSE Reproduction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.511895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.511895Z digest=sha256:2a4f0424a12b2cd36476babbb58371a476137c23c0701eeaae7388a53209661e

Observation 8b702859-d4c0-4d7c-a5eb-d9652f86e853 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

VidTok: A Versatile and Open-Source Video Tokenizer Imagen Video: High Definition Video Generation with Diffusion Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.482600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.482600Z digest=sha256:f92312e44ad9ec624542387ebec7d6eceee9f55bf3ccfc8a2867a065ff6dd847

Observation cea75237-4c0a-42a1-85e9-6cc0e579c6f5 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VidTok: A Versatile and Open-Source Video Tokenizer CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.554765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.554765Z digest=sha256:a2cfd82319c89929ee27a8876a858dd8febcdea7fcaac7bed7def3beeac0d59b

Observation c8dc4f06-7ca2-4600-a146-0675eee0062f · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

VidTok: A Versatile and Open-Source Video Tokenizer Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.470488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.470488Z digest=sha256:cedd446768a027940516b1fb69db4323fa847dbd249fb2e38fecaa783a88af88

Observation 7a67941b-4140-4d96-8879-d984c906f2eb · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

VidTok: A Versatile and Open-Source Video Tokenizer Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.464444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.464444Z digest=sha256:eac6834584f8b63c4965a84660caad498c8ac7d7347ab9f5a97aa33828d57a6f

Pith citing papers

Observation 85fa5572-3bdd-41d3-ba03-3cfb56585fa4 · inbound

VidTwin: Video VAE with Decoupled Structure and Dynamics cites this paper.

VidTwin: Video VAE with Decoupled Structure and Dynamics VidTok: A Versatile and Open-Source Video Tokenizer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.972158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.972158Z digest=sha256:6b8eb0c290f5f05628a42f6872bf547fc493676bec6a4308fa56f48a4177cc09

Observation 4427e8f9-248f-42a0-b985-b4244c81698b · inbound

Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces cites this paper.

Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces VidTok: A Versatile and Open-Source Video Tokenizer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:16.114855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:16.114855Z digest=sha256:c6d9a3a0a08642f063d471c08c9377d776a9107e955c52f2bebac92c1b270203

Observation bb307e0d-63c7-494c-a2ec-9efa73f29d14 · inbound

Latent Wavelet Diffusion For Ultra-High-Resolution Image Synthesis cites this paper.

Latent Wavelet Diffusion For Ultra-High-Resolution Image Synthesis VidTok: A Versatile and Open-Source Video Tokenizer

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:02:16.754970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T12:00:37.025335Z digest=sha256:c575fa06fbda9637d292c1081b1fffd9903ee7bc496a48a4d6b584ec8cf63c18

Observation ef69d201-f105-4960-bf9c-554ac4c7e766 · inbound

Playing with Transformer at 30+ FPS via Next-Frame Diffusion cites this paper.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion VidTok: A Versatile and Open-Source Video Tokenizer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.159313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.159313Z digest=sha256:f45535bccc532b39691f046f41fbfb84260823b5d199c97985124990307ecac0

Observation 33048bd2-577b-4e5a-97b2-1f0dc8a1447e · inbound

Video World Models with Long-term Spatial Memory cites this paper.

Video World Models with Long-term Spatial Memory VidTok: A Versatile and Open-Source Video Tokenizer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:40.611770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:40.611770Z digest=sha256:9ef864ed084c41adc9f1244a02a733c06a4e2677a12a395ef8a643f5dd23b38a

Observation 6e61c7d4-8052-44c3-8a74-56aecbe1d57e · inbound

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics cites this paper.

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics VidTok: A Versatile and Open-Source Video Tokenizer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:49:52.048245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:49:52.048245Z digest=sha256:b95c5a630a25a772ff73e2e91ce359634adfe163848bd0ad09d1e0d75567bec7

Observation 05f6dd02-a832-45f5-9bea-877cd866512f · inbound

TivTok: Broadcasting Time-Invariant Tokens for Scalable Video Tokenization cites this paper.

TivTok: Broadcasting Time-Invariant Tokens for Scalable Video Tokenization VidTok: A Versatile and Open-Source Video Tokenizer

Reference 104

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:28:55.544391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T01:20:32.508409Z digest=sha256:252318192aef880a1697f5631984852e07428921f2551c429696335b1a69aaea

Observation c365d70a-3fee-412a-8dc8-0e461a32b98b · inbound

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation cites this paper.

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation VidTok: A Versatile and Open-Source Video Tokenizer

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:39.991649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T06:16:13.420481Z digest=sha256:b3020ea1d93444d2dcd2cac139503705e7ac31ddcbe1493ad6fc44a4b529dcee

Observation 817a5b73-23f6-472f-8b79-7989bc4ca8c3 · inbound

FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers cites this paper.

FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers VidTok: A Versatile and Open-Source Video Tokenizer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T00:50:14.926631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:50:14.926631Z digest=sha256:c7678e53d690a4ca6320860f53adfb79b03f8f12174a951adb9b601cf49989eb