Pith. sign in

Paper Citation Record · LEDGER

Image and Video Tokenization with Binary Spherical Quantization

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2406.07548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.07548 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:04:02.451468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:30:07.369173Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6a329c8a-3a3e-4058-ae43-f914da5e935c · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI Image and Video Tokenization with Binary Spherical Quantization

Reference 248

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:38:46.176065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:2ed93d1298f8857a8676be2e3bdd2d3ad99f0ad11a2db6847971ce14661a2a71

Observation 0d66fc1e-a3a2-4698-87f2-d800bc7c8802 · inbound

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation cites this paper.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.451468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.451468Z digest=sha256:b54e0a58bacfafad66f00aad99c0d44390833eed77a877c1f135a3b1f68bf149

Observation 53a839f8-3586-4ffc-af44-54354bb9b2c8 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.913350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.913350Z digest=sha256:9aac1a0cebbf5d8bae5c66287ec4d5198f983e74e047bdb6a64d7f0c46316995

Observation 042198ce-a5bd-4b17-8d00-45b9a34297a5 · inbound

TokBench: Evaluating Your Visual Tokenizer before Visual Generation cites this paper.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:44.839423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:44.839423Z digest=sha256:7ec189cb32ba1500a84692d4140cbcfab173f5813763a3e68c3039fdc2e5879d

Observation 38e3a0e4-6030-48eb-9e23-e98ba94771b3 · inbound

multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data cites this paper.

multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data Image and Video Tokenization with Binary Spherical Quantization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:09.611149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:32:09.611149Z digest=sha256:45a5e5df081911dbee3241cfa746c9a4194e8202a14b645603ec22fff9a9963a

Observation 82af5880-746c-42df-aed7-16c7075b59c0 · inbound

PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models cites this paper.

PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models Image and Video Tokenization with Binary Spherical Quantization

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:27:18.938661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T13:26:17.991938Z digest=sha256:71c21e49acb3c8ca8ca70d897549487cf3b1671bc926888b9da831d398581324

Observation 0eab98de-a119-4e18-92bd-603a1c5d7329 · inbound

MambaVideo for Discrete Video Tokenization with Channel-Split Quantization cites this paper.

MambaVideo for Discrete Video Tokenization with Channel-Split Quantization Image and Video Tokenization with Binary Spherical Quantization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:52:28.135246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:52:28.135246Z digest=sha256:28c6b524d14d4735a72f518e83375b1b22bdd494cec7119777df777569400e88

Observation 6be68bb3-d653-4d4c-832b-eaf9de91d993 · inbound

Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive cites this paper.

Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive Image and Video Tokenization with Binary Spherical Quantization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:44.578141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:44.578141Z digest=sha256:b20571a366c91fc9206bdf60a22b566982d444c5248f54a8ae06d456b90fc697

Observation e0c3e197-b638-4c99-8321-2b6944591867 · inbound

WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting cites this paper.

WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting Image and Video Tokenization with Binary Spherical Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T23:57:19.687253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:57:19.687253Z digest=sha256:1436a9d8e3a829d783d3ea3cef9f43eb07ef3b9f180b6bb16ba2c19566ddac30

Observation 4db5a033-e809-4e52-969a-df2aef4bdaf1 · inbound

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body cites this paper.

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body Image and Video Tokenization with Binary Spherical Quantization

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:08:36.263008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:04:07.403410Z digest=sha256:c52117ef67f0955f27f33ce56a83196132cf2a0f22bd0d47a1dc52cc287a15d5

Observation 53889dda-c3f6-4c01-a36d-a41535389e37 · inbound

Wireless TokenCom: RL-Based Tokenizer Agreement for Multi-User Wireless Token Communications cites this paper.

Wireless TokenCom: RL-Based Tokenizer Agreement for Multi-User Wireless Token Communications Image and Video Tokenization with Binary Spherical Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T23:56:42.310530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:56:42.310530Z digest=sha256:3a67e354ae3be6c7cbf942706a0835acb297341029892f3caa54d3ca1aed9e4d

Observation 107f7c47-d0e9-4b38-915b-5cb972d24187 · inbound

LGQ: Learnable Geometric Quantization for Image Tokenization cites this paper.

LGQ: Learnable Geometric Quantization for Image Tokenization Image and Video Tokenization with Binary Spherical Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T22:43:28.615040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:43:28.615040Z digest=sha256:62b09df52e6c52a3652e190ea2ce3333db7c83333a79793730b1b7a31875c8b3

Observation 0740cf01-3056-475e-8ac8-d875c6eb6c39 · inbound

Modality-Aware and Anatomical Vector-Quantized Autoencoding for Multimodal Brain MRI cites this paper.

Modality-Aware and Anatomical Vector-Quantized Autoencoding for Multimodal Brain MRI Image and Video Tokenization with Binary Spherical Quantization

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:49.442294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:43:41.036191Z digest=sha256:9cd1296405d7e2cd9eb748a092db20f901b6b00e9703070c65da8daa6163ffdf

Observation d500312b-a983-406d-8833-9560ec6231cd · inbound

ELT: Elastic Looped Transformers for Visual Generation cites this paper.

ELT: Elastic Looped Transformers for Visual Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:06:00.090311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:19:22.543462Z digest=sha256:c8ab5b7f93158c3e8e2e7b6289760dde7d81753863e35d90113031fdb56205b1

Observation b830cbcd-5aa4-4060-93e3-146c61d73648 · inbound

ELT: Elastic Looped Transformers for Visual Generation cites this paper.

ELT: Elastic Looped Transformers for Visual Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T16:35:03.432151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:35:03.432151Z digest=sha256:7a6fa14f7aacc594e94f038440ed0540b8fb0ec2514d800b40a99006098d6d3f

Observation aff94852-2587-417d-bcd5-818e8d7b99e0 · inbound

Generative Refinement Networks for Visual Synthesis cites this paper.

Generative Refinement Networks for Visual Synthesis Image and Video Tokenization with Binary Spherical Quantization

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:01:03.361529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:41:23.057949Z digest=sha256:af50d6953d79e2f87a4f1438af2e165c8cc53b98a1370e7a2e7efd808d59c884

Observation 917f8982-2f24-4a38-96a7-f7a155f52e53 · inbound

Generative Refinement Networks for Visual Synthesis cites this paper.

Generative Refinement Networks for Visual Synthesis Image and Video Tokenization with Binary Spherical Quantization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-12T21:00:35.123495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:00:35.123495Z digest=sha256:2474a9b0ba0b5f2036e90b25620435037476345142262bfed6400146ec4d758f

Observation 2eb6b7d5-2334-4925-adce-1565de30279d · inbound

Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept Learning cites this paper.

Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept Learning Image and Video Tokenization with Binary Spherical Quantization

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:08.120351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T20:05:21.069489Z digest=sha256:9920d6532aaa46519c39f4c3ceba686569e6eed75247e6b6d5940dbfff916e44

Observation 82c4b2ac-23be-4cbd-bfb7-afe1354ffeb3 · inbound

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion cites this paper.

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion Image and Video Tokenization with Binary Spherical Quantization

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:05:57.135970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:57:24.033068Z digest=sha256:1a1903271398459371a6d1d3473c6120dec6d1ea7ab42ee4b3967b121c1ac91b

Observation f670e45b-f998-4c30-8c9d-d7f4fa84749e · inbound

ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation cites this paper.

ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 45

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T06:27:24.788773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:24:01.820975Z digest=sha256:1939bed1fdbb3f8770f7765b60f8251e1c1b4738ed643c7f57763fed3feb696d

Observation 4d54be22-07df-45b3-85a1-13f8abf699a5 · inbound

ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation cites this paper.

ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 44

Resolution
malformed identifier
no resolver link, observed 2026-08-02T14:19:04.486699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:04.486699Z digest=sha256:fe869e5477e0d17a0f90a9a27b538740d3812c86d85d506f60949fd76a4735b0

Observation b3587e9b-bcc7-417b-b6d6-618fc9e189b8 · inbound

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models cites this paper.

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models Image and Video Tokenization with Binary Spherical Quantization

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:03:06.307706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T07:01:01.427273Z digest=sha256:378e09a165afc86f803cb2d81d9176ccd8333d0e1513c3dc8f606308390842ea

Observation 5195f449-27e6-47ae-aa2a-7076b077e4c8 · inbound

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation cites this paper.

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.008832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:00:16.005187Z digest=sha256:26bee0d4ffebd22226165b38f870dd73c201ad70a57e3c94916906fd8314697b

Observation fc5e4f06-5d5c-43dd-bbfb-e88e0fa74da9 · inbound

ChannelTok: Efficient Flexible-Length Vision Tokenization cites this paper.

ChannelTok: Efficient Flexible-Length Vision Tokenization Image and Video Tokenization with Binary Spherical Quantization

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:44.334531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T07:09:25.049534Z digest=sha256:f8b4a0b7b1cac6bd8cd39e9571551625a9e88c495f945602af2b01a099aead5c

Observation 7d5c36a2-c73d-4a12-ac10-fa0a315ef9de · inbound

Concept Removal for Frontier Image Generative Models cites this paper.

Concept Removal for Frontier Image Generative Models Image and Video Tokenization with Binary Spherical Quantization

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:30:07.370684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T21:18:32.620951Z digest=sha256:e0f6267e21a34a331dccc23be23bf80b51623ed495b9276584ec644e2ca27343

Observation 67e9d513-4836-49da-98ce-eb506e862f5c · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Image and Video Tokenization with Binary Spherical Quantization

Reference 204

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.152440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:1d4c2a8305643a5a93e0c877dab32b16b0c75a55cafe8d0fb36653f35fd436b9

Observation 388dd666-a4a8-4db3-81b3-187d5cb46766 · inbound

DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units cites this paper.

DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units Image and Video Tokenization with Binary Spherical Quantization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T00:58:38.644366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:58:38.644366Z digest=sha256:617bb0d97cca85a42824b7b5cc50bd2a3b7af21a6f0771fbcf0b2d750e326450