Pith. sign in

Paper Citation Record · LEDGER

An Image is Worth 32 Tokens for Reconstruction and Generation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2406.07550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.07550 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:23.840987Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:13:49.010757Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c6567b4a-4243-43cb-a42c-2000efc017e8 · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 240

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:38:46.153147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:e230c8e39b566bb71757883dfd3882e6bc0c3fe573e50504662ea441278e5f05

Observation b2baccb6-563b-4823-bd9b-d3ce09306f82 · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:21:45.136776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:1c8631b017c41a07cb473788377b541667f56d13ac4e63be651e3c39e31f62d4

Observation 1406d1ef-0d27-4c18-a42b-e0873baf7cbf · inbound

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation cites this paper.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.840987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.840987Z digest=sha256:c6ba16fd3fc722e7ab2f84d9e5a040624259427bc5e475b69b17dbda48831dc5

Observation 97c6b236-800c-42a5-8c4d-0f0cb0fda1c2 · inbound

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation cites this paper.

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:04.678813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:04.678813Z digest=sha256:d336a6b178c2a0a89df8bdf035da7085fc9b847ea18d00f4fd7203fb94e96586

Observation e40974e0-86b1-4647-ab9f-ab8079e07911 · inbound

Is Visual in-Context Learning for Compositional Medical Tasks within Reach? cites this paper.

Is Visual in-Context Learning for Compositional Medical Tasks within Reach? An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:06.140146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:06.140146Z digest=sha256:f7e7afe381f92974ea7281ab2b56fc3bd466fec41c719c34237a3ff6f4c092ec

Observation 4b2fd8ff-3c60-4015-bbfc-aa3e9bcfd79f · inbound

Hita: Holistic Tokenizer for Autoregressive Image Generation cites this paper.

Hita: Holistic Tokenizer for Autoregressive Image Generation An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:58.541916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:38:58.541916Z digest=sha256:ea7feb032a5714e883b786abf512a499f65f5521edfb9faf06fa4d59e5ad41e4

Observation 5e6fd36e-aeb7-483f-8b76-514cd63103d4 · inbound

Single-pass Adaptive Image Tokenization for Minimum Program Search cites this paper.

Single-pass Adaptive Image Tokenization for Minimum Program Search An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:32.847351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:35:32.847351Z digest=sha256:88edc1def3a7dbfcbbfac746d624c7c4b23d3b03fbe143166ef8edf5ce75cad2

Observation e043c121-f9b0-4b56-942b-a2b6b4a95145 · inbound

CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio cites this paper.

CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T18:41:39.636626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:41:39.636626Z digest=sha256:8b7455e1cd5d8e012d2d75fe7e10d1661124be1ab1910a6f3b21d4e9be9c60e3

Observation 282f7c6d-80b6-4c6c-9a70-54a0225dd3b0 · inbound

Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization cites this paper.

Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T18:12:48.006661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:12:48.006661Z digest=sha256:f6d786cf80f7e3a8dc860e2942fd0092292abe3b260ad6ecbd09c385880bac33

Observation a0028c5e-e42b-4cdf-ab2c-7b66998eca75 · inbound

ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters cites this paper.

ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:46:17.503383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:06:21.438738Z digest=sha256:793778461e5875fe932cae62922be427ed5dbe0b887e54411a0990bfe8befbfe

Observation 8efbf0fa-5750-4b9e-b108-0f04cf73663c · inbound

Does Engram Do Memory Retrieval in Autoregressive Image Generation? cites this paper.

Does Engram Do Memory Retrieval in Autoregressive Image Generation? An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:39:28.137756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:33:17.978924Z digest=sha256:c1d3fea74dcf98fb7632c8c95bec4309ee4b0cff3a6280ca169f0bfd4c7ad346

Observation 4c9f453f-ea96-4bc9-a3d5-e1797b412881 · inbound

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice cites this paper.

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:43:51.039023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T22:41:44.510546Z digest=sha256:1a5630875933eeaa73b1527ed12b63661a5049a665e31adbb09640b552b47a7c

Observation df817617-febd-48ed-9a38-f9ca51fe0792 · inbound

Vision Foundation Models as Generalist Tokenizers for Image Generation cites this paper.

Vision Foundation Models as Generalist Tokenizers for Image Generation An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:03:13.487184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:01:24.738195Z digest=sha256:57b4c4e08cfcd2d9bd93cf004321a7fb69cde1cafe8b26ca736585238a7ca0be

Observation abde206c-91da-495f-980a-3aa274828bc2 · inbound

Structure over Pixels: Learning Variable-Length Visual Programs cites this paper.

Structure over Pixels: Learning Variable-Length Visual Programs An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:13:49.012240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:06:01.684713Z digest=sha256:1df8c1fce616e1383cc4c4cd62a1ad0ead3000de98061c9d3c8b17580841c00f