Pith. sign in

Paper Citation Record · LEDGER

OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2406.09399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.09399 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:36:50.110055Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:46:13.353259Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3eb7dec3-098c-44d4-a69b-f183c5561a0d · inbound

LaVin-DiT: Large Vision Diffusion Transformer cites this paper.

LaVin-DiT: Large Vision Diffusion Transformer OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T18:31:28.981397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:31:28.981397Z digest=sha256:840a4b3bff645e12c17f6dd9fb1d1342473618f92327773bef9ec93149c1b84c

Observation 54c9b807-7f0a-4045-99eb-cbbf299f48de · inbound

Factorized Visual Tokenization and Generation cites this paper.

Factorized Visual Tokenization and Generation OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.062730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.062730Z digest=sha256:cc70c506953f79179a38b9923ea48a184ea6920333f0515ac11e7e27a831c181

Observation 144da361-88f1-437e-b749-8d3090456f2b · inbound

Scaling Image Tokenizers with Grouped Spherical Quantization cites this paper.

Scaling Image Tokenizers with Grouped Spherical Quantization OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T23:20:11.095353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:20:11.095353Z digest=sha256:9896168ab5ace31e94a3bb6675ce7b30fc8d73459ce461475655935ad573c818

Observation cb6365eb-de8f-4f66-81c8-db862b3ca7d4 · inbound

Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis cites this paper.

Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:34.598427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:34.598427Z digest=sha256:2d37753ddb066278a25eb928a5b8fd991ee805ddc644ef4c5cc8a64a8702c8f6

Observation e239a439-bd92-4ea4-bb36-fb3388c6a8a2 · inbound

VidTok: A Versatile and Open-Source Video Tokenizer cites this paper.

VidTok: A Versatile and Open-Source Video Tokenizer OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.537169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.537169Z digest=sha256:e59cce67bd7393ef6792d4c1f5f645839b672550220b4815c6e2e305b60ac4c3

Observation 22c37472-041e-435b-ade9-d7188fc97392 · inbound

Parallelized Autoregressive Visual Generation cites this paper.

Parallelized Autoregressive Visual Generation OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T11:42:31.772042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:42:31.772042Z digest=sha256:61b0906fb0ac59a5d186e17f73de30a5eb053ced5fa3d229aebcb10dbd1dc8bc

Observation 23112e8d-a45e-42ed-b0cf-ca7b1847bb1b · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:38:45.973996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:47c5b5ab4040687c569c6cad13f9e494f572d5762831df0f92bb6fbe9920d60e

Observation 0282cf07-166c-499f-8aac-f34ba020e292 · inbound

One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression cites this paper.

One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T19:28:58.215804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:28:58.215804Z digest=sha256:05f08a9b9771b9aef1e1351a7540e0c0b584d7f641e3b51312404e304da3b5d9

Observation df6172df-8e15-4382-aa81-bcc7abf63b89 · inbound

Taming Teacher Forcing for Masked Autoregressive Video Generation cites this paper.

Taming Teacher Forcing for Masked Autoregressive Video Generation OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.599371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.599371Z digest=sha256:87b5bd7ba77138a81cbc371f09c470ff0d1883c790f1690ab9343db923934d50

Observation b6d397e4-2449-4189-a894-d9f5393264fe · inbound

Long-Context Autoregressive Video Modeling with Next-Frame Prediction cites this paper.

Long-Context Autoregressive Video Modeling with Next-Frame Prediction OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:05:17.352046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T23:05:17.201790Z digest=sha256:d5f37ccbd9b5a655b518b4108c6aed6ef8568904cb1d802472ca661a7a289b0e

Observation 73528be7-3bdd-4824-800b-bf0bb0a1f567 · inbound

Position: Foundation Models Need Digital Twin Representations cites this paper.

Position: Foundation Models Need Digital Twin Representations OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T04:36:50.110055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:36:50.110055Z digest=sha256:02b8c6f7614d455a60919e06c6bb0f7318a4b6558c45ef51d5016be79b812e92

Observation 34b876c0-3a6f-441b-a88d-7d2684715349 · inbound

Video-GPT via Next Clip Diffusion cites this paper.

Video-GPT via Next Clip Diffusion OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:30.333818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:30.333818Z digest=sha256:75fa726bd0f15dfef37df08f6cc59b418aeded677f66462b0d370d3078b56cc1

Observation 56516b94-7394-4542-9928-be50ede623da · inbound

MambaVideo for Discrete Video Tokenization with Channel-Split Quantization cites this paper.

MambaVideo for Discrete Video Tokenization with Channel-Split Quantization OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:52:28.111511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:52:28.111511Z digest=sha256:e3055c7d4b4e45e1ec33ea770ed401265e184361aaf47377fa302d56b2619f80

Observation cda5a9f7-cd61-4ce3-9d75-581450429c58 · inbound

OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation cites this paper.

OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:09.045343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T08:52:07.545191Z digest=sha256:855df5dddf6be783f202c58a42abf7c3caacc591d6adee06fae47b5d3f3264fd

Observation fc68574a-08c1-4ff7-8feb-bf567ef7ff6c · inbound

MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging cites this paper.

MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.039598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T22:42:20.280777Z digest=sha256:e738b910064f68aed48293cc767170be34daac40bc00769b7fd002c44352fab7

Observation 1e2435d6-d802-41cb-806a-337761e6e8af · inbound

Wavelet as Tokenizer: Preliminary Results on a Shared Wavelet Token Schema for Natural Signals cites this paper.

Wavelet as Tokenizer: Preliminary Results on a Shared Wavelet Token Schema for Natural Signals OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:13.355163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T18:05:11.392504Z digest=sha256:cc484e694ef0abfed956b5c1d5e097ff7c1a3d1b7b7c41989172bd86730dd640