Pith. sign in

Paper Citation Record · LEDGER

TokBench: Evaluating Your Visual Tokenizer before Visual Generation

As of 14 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 2 inbound Pith citation observations for arXiv:2505.18142.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18142 v2

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:44.972598Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T01:12:13.426640Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T01:12:13.598677Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d9c3038-8e4e-424a-85b7-f4382bcbeb4b · outbound

This paper cites FlexTok: Resampling images into 1d token sequences of flexible length.arXiv 2025, 2025.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation FlexTok: Resampling images into 1d token sequences of flexible length.arXiv 2025, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.416395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:38.942520Z digest=sha256:de744f5538c904eb26f369d71672e2db5ac2826e6d21869dedda81d07f651523

Observation d1d16717-d3f8-4039-adac-75c7e42e31e9 · outbound

This paper cites Scene text recognition with permuted autoregressive sequence models.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Scene text recognition with permuted autoregressive sequence models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.260272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:38.973992Z digest=sha256:b2d17941583de031b005a98fb0c8ef79a60789070037b0bd040d4154889f3d7d

Observation 566e5880-ecd6-434b-9591-28cdbbf86495 · outbound

This paper cites Total-text: toward orientation robustness in scene text detection.Int.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Total-text: toward orientation robustness in scene text detection.Int

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.139417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:39.037498Z digest=sha256:455413a63e5beca5af98d660dc1ae065b96ac7a38dffd5d86c5c4c69f0af81fb

Observation 0e083a73-7b2d-4ba5-a9b0-112aaa6f68e0 · outbound

This paper cites Diagnosing and Enhancing VAE Models.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Diagnosing and Enhancing VAE Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:39.114745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:39.114745Z digest=sha256:d6744cfa15c07b055acb055cf85249b397d2ce5aaf4dfd30924f5e133331b4ba

Observation 5dfdf3e5-23fc-4ce3-9af1-282badee32c1 · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Imagenet: A large- scale hierarchical image database

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.019893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:39.140584Z digest=sha256:dffd659e5357b06ac2deacd193f599eabf8a0df52871e9d33aba91a98bd99044

Observation 1de34a10-11e6-4bf7-9ad8-697e4f619de5 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:39.146369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:39.146369Z digest=sha256:0cf8d4201bc434ca4483affa9878e9b32c8cba31a1f67a2537fa362f4d4973ee

Observation ecd28181-a582-44b1-a2c2-0224422abf5c · outbound

This paper cites Taming transformers for high-resolution image synthesis.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Taming transformers for high-resolution image synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:39.187332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:39.187332Z digest=sha256:20b5964eb1077662079ab5e89ff28110ab02bf31249657963eadb390bc455289

Observation 067ee8d7-cd52-48a9-a7b7-c98492576d0a · outbound

This paper cites Data Filtering Networks.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Data Filtering Networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:39.229146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:39.229146Z digest=sha256:03580fb0c2ca234363d72d44b4731bdb9b52cb51114073f939c01657a19ee45b

Observation cffcd49e-7dfa-4b18-986c-8ccce2b9230a · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Mmbench-video: A long-form multi-shot benchmark for holistic video understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:39.318032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:39.318032Z digest=sha256:22fee802da68ad27aee4ba5a93e5ca9d29ff0a41c1ac013bcce00124cbe17dfb

Observation e6b1b01c-0271-453d-b844-fb382b5b2d93 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:39.414670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:39.414670Z digest=sha256:af0e0fee002502be1be464e5116e86f85d573258082175e48334dd1a8f468bf7

Observation e6cfa880-ca64-442c-9574-e23513a75cc0 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:39.478304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:39.478304Z digest=sha256:8bc1e8ce147c4c5d5d307ad981133c100ba08d4db714966dc91d0b764f058136

Observation 166c86ec-9691-4b59-9743-30e41df6398f · outbound

This paper cites Denoising diffusion probabilistic models.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Denoising diffusion probabilistic models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:39.541439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:39.541439Z digest=sha256:b651e81992ae34a35cdd4defba559d578f88d9d0902c33ee4a4862f6a29c87b6

Observation 330b65b7-c6e4-4b06-9fba-974dec16a38a · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Cogvideo: Large-scale pretraining for text-to-video generation via transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:49.802193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:39.598864Z digest=sha256:127b6bf35eb2f77eb7af66b0fb12841b6df8c18bc23b0fdfe1fe522d22c96480

Observation d084e74e-f496-4c81-83c6-40ef2e818309 · outbound

This paper cites Labeled faces in the wild: A database forstudying face recognition in unconstrained environments.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Labeled faces in the wild: A database forstudying face recognition in unconstrained environments

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:39.639427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:39.639427Z digest=sha256:ddf1167a45f2764ffc419c4ba3910c33d117c076b3c0f76fd743a02ae49cdf15

Observation c15cf2e1-b155-40a3-bd54-428ab09d12f2 · outbound

This paper cites Icdar2019 competition on scanned receipt ocr and information extraction.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Icdar2019 competition on scanned receipt ocr and information extraction

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:49.648190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:39.699579Z digest=sha256:bd43f68537710eff612997d5bf5eb1211735fda854d3000a9dfeda0e88080152

Observation 9e365d38-7e7a-4264-bafb-1dd3f3f3ec3a · outbound

This paper cites insightface.https://github.com/deepinsight/insightface, 2024.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation insightface.https://github.com/deepinsight/insightface, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:49.524309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:39.808619Z digest=sha256:7c223c4786b337ff763d481d9d566170f404aacc8ace22219baaa7025473b66b

Observation 2c70ae5f-1502-4cb4-a9ad-ce61050dbeec · outbound

This paper cites Icdar 2015 competition on robust reading.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Icdar 2015 competition on robust reading

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:49.415521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:39.869169Z digest=sha256:e1a7cb80b9e15a5d887d38678f354e6577ee5b9d59e8906df4496cc84551a5e3

Observation a046604f-6fd7-4b6a-bbfc-812468bca38a · outbound

This paper cites Icdar 2013 robust reading competition.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Icdar 2013 robust reading competition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:49.266858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:39.960164Z digest=sha256:5e8de52837e22cca45090abfb15d8fe60d33d205d0387c4158e17b280708df7b

Observation e1db9af6-462c-4336-90ae-cfc5dcb334f3 · outbound

This paper cites Kingma and Max Welling.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Kingma and Max Welling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:40.038050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:40.038050Z digest=sha256:f49188cebf20942948569ab7534d690531e77d57941111b6e56e8803bc249a89

Observation faf4cb1e-e130-4e5e-a3e9-aa1d59b4ae96 · outbound

This paper cites Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:49.160054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:40.108398Z digest=sha256:1cdd81fac0d71b0e1e8f59a5f703c42fbbc1eaeb014a048e5973123f72bef7ee

Observation a5f80af4-8c69-4a63-9bd6-bcf1f6b4e518 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:40.286864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:40.286864Z digest=sha256:fbc359d932789cded688dbfca298c65c5facbde58121256313603a3d5ab343f8

Observation 1a696323-83ad-44a4-a2a4-4018e7923248 · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2024.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Flux.https://github.com/black-forest-labs/flux, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:40.397855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:40.397855Z digest=sha256:cfb8a85039adb14bd49eba7fd3b435d72a88ef1064441256b2cbabf06408d0b0

Observation b5664d69-fbbf-43f3-a7f2-f3471e1b30eb · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:40.535136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:40.535136Z digest=sha256:9b0145b3d7fe5d2b9e9a9229122d4bc6413a355bbb953fa8b5a773ebbe4d1137

Observation da86b6ff-b655-4e29-9a7c-dee1e1e2455f · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Open-Sora Plan: Open-Source Large Video Generation Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:40.645016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:40.645016Z digest=sha256:eb013fc6f6561425352241cd234e542d0b74310f91384ea090d6fc5a5791ea97

Observation 9e607ba7-41bb-4003-b4c4-02d081620dde · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:40.720001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:40.720001Z digest=sha256:93404f128e263d25b479c60c1e21bae260309e09a9b72b88a2cf457d6d89eb4b

Observation b70b03da-bf99-490b-b971-e2a87c711fa4 · outbound

This paper cites Abcnet: Real-time scene text spotting with adaptive bezier-curve network.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Abcnet: Real-time scene text spotting with adaptive bezier-curve network

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:48.958934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:40.836873Z digest=sha256:75510b68f3bbab734f196b4e04e29c7142c5a572438364b58b06256ded5c8ce1

Observation 66980f6f-f201-49f4-9efd-6164576050e7 · outbound

This paper cites Curved scene text detection via transverse and longitudinal sequence connection.Pattern Recognition, 90:337–345, 2019.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Curved scene text detection via transverse and longitudinal sequence connection.Pattern Recognition, 90:337–345, 2019

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:48.771284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:40.946666Z digest=sha256:291c7f6006f9382c1f892d3d7aec478250470a92aadafad7649e06a034bd6b8c

Observation 3fecfe9b-a345-42f1-aa6a-339a35991465 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:41.105052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:41.105052Z digest=sha256:34f261ddfc00be12f24bd7162d029eaf1d45387f399e3a488c60f1e20c5e6f8e

Observation e64fd548-340e-4b09-bdd8-d31fc4459e51 · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:41.256151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:41.256151Z digest=sha256:5199992448da190cd72fd30fb44ad655576695493747e0d455f4908fb887d668

Observation 45701260-5e76-4360-a0bc-f62457671fb6 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:41.363946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:41.363946Z digest=sha256:a791e70eab8e2057045095200c1cf619bbe1664ae6f043966fffb622f43ca47d

Observation 835bbf6e-6e1d-451b-9d7b-6ec72f43f148 · outbound

This paper cites Hdr-vdp-2: A calibrated visual metric for visibility and quality predictions in all luminance conditions.ACM Transactions on graphics (TOG), 30(4):1–14, 2011.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Hdr-vdp-2: A calibrated visual metric for visibility and quality predictions in all luminance conditions.ACM Transactions on graphics (TOG), 30(4):1–14, 2011

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:48.528744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:41.456959Z digest=sha256:2fd5e47e7ea28025339d017c066ee6419db091f5b38eaa500a155a50e2b2f233

Observation 1f4aaf38-7b58-4bdf-812d-87b8fca32974 · outbound

This paper cites Infographicvqa.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Infographicvqa

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:41.563845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:41.563845Z digest=sha256:9c21bbd5ab7a5e7a04c47d3607bdcd9776056ee67051529ab36f5860de272a8f

Observation dac05f3d-9fce-41ab-aece-4b30779e592e · outbound

This paper cites Docvqa: A dataset for vqa on document images.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Docvqa: A dataset for vqa on document images

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:48.319540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:41.662054Z digest=sha256:92a03a5dd1c71f3a8a8942d7ba4931cacb8c9c59583196e79435efaa71da601f

Observation 1e369fcf-fb81-4b77-81f8-3333278266aa · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Finite Scalar Quantization: VQ-VAE Made Simple

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:41.765000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:41.765000Z digest=sha256:dd2d68f680322f3a0a209025e80afffac10e81a134b0b4b7e271fcf498a4278b

Observation 87defef3-f231-4be0-9e22-b87e63ef7ece · outbound

This paper cites doctr: Document text recognition.https://github.com/mindee/doctr, 2021.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation doctr: Document text recognition.https://github.com/mindee/doctr, 2021

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:48.117767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:41.871166Z digest=sha256:8a0bf1f7e76965c80e76b83867b2a141dc3d5ba8da20b3699e43dce4239694ac

Observation 93e16940-7226-46cf-ac0e-217340c5ddf8 · outbound

This paper cites Icdar2017 robust reading challenge on multi-lingual scene text detection and script identification-rrc-mlt.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Icdar2017 robust reading challenge on multi-lingual scene text detection and script identification-rrc-mlt

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:47.935892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:41.965442Z digest=sha256:694efcbf584875506083151f11d6a61e51c9991f4ecc9f3ef9f8492010b271f8

Observation b963c43a-d83b-4445-93d3-fc8e1f9df061 · outbound

This paper cites Cosmos-tokenizer.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Cosmos-tokenizer

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:47.699216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:42.066493Z digest=sha256:25ce6b01ae5a3a1bb46d600453cf2deed4d5537fbc2b14856b9034e99e9af986

Observation c87173a3-42e2-40fc-aa28-2dab1e7bd884 · outbound

This paper cites Cord: a consolidated receipt dataset for post-ocr parsing.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Cord: a consolidated receipt dataset for post-ocr parsing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:47.483981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:42.187418Z digest=sha256:208b2e562ab8707d95912e954383041aaff161246a1d5d4c1738563da7080928

Observation cba92258-d239-4595-a30e-82fa2a0c9d73 · outbound

This paper cites Scalable diffusion models with transformers.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Scalable diffusion models with transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:42.269247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:42.269247Z digest=sha256:671271ce1f83a84bcd76e4d71e96cc81f0bc8db5844a4a90eb827ee8a96b05f0

Observation 1753689a-7adf-4057-a67d-3d60ac265f26 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:42.347013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:42.347013Z digest=sha256:4b3943e40353c1cbc43cff6d44a59d89b038c98ef3be7935a9344500f54d82f9

Observation 49711da3-9b0c-44c5-8b6b-ba35915ed162 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:42.504964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:42.504964Z digest=sha256:060c5636e058deaba40114f749318a09d7865c503fb7dbb941bea68398eb68aa

Observation 1b9a78cb-c3b5-46f2-a087-0fc2b6dc8cb1 · outbound

This paper cites Zero-shot text-to-image generation.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Zero-shot text-to-image generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:47.237168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:42.630779Z digest=sha256:4091625f166a9132699465d5cf951f023a13ee615463d0e79e1452cabe447ec1

Observation 0f986f42-1425-4178-ae8d-2c254ac4dc00 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation High- resolution image synthesis with latent diffusion models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:42.730140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:42.730140Z digest=sha256:18fabf271502f82871c95e7f6ab65349767b39bcd56b5a8cb31f54b8eb63bb85

Observation ea06fae6-6a0b-4bf1-8665-eab066feac0c · outbound

This paper cites 300 faces in-the-wild challenge: Database and results.Image and vision computing, 47:3–18, 2016.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation 300 faces in-the-wild challenge: Database and results.Image and vision computing, 47:3–18, 2016

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:47.107375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:42.844285Z digest=sha256:0a16d399b257c77a5a7f7f35b0bae467af258b84b00627e755f067afa6d1530f

Observation 55528374-26e8-4789-934d-b1f687d7e714 · outbound

This paper cites Improved techniques for training gans.Advances in neural information processing systems, 29, 2016.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Improved techniques for training gans.Advances in neural information processing systems, 29, 2016

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:42.948794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:42.948794Z digest=sha256:a3b1c27100b8128f3178ed4d4c8bbf9e56342dac7a9e4b7333e0f12dba57e144

Observation d114b335-4c23-403e-a5fd-36ff222695f9 · outbound

This paper cites Patel, Rama Chellappa, and David W.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Patel, Rama Chellappa, and David W

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:46.921190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:43.087477Z digest=sha256:1b55e09edafb6f4444c9e1a95dcb19043fd72be72e051282b7cf104b244d4617

Observation 030d9d74-52f5-4b67-8688-eb3bea14e91c · outbound

This paper cites Tex- tocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Tex- tocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:46.641129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:43.293550Z digest=sha256:45c39f8f1982f456c8afeaa0dae9d88f9500970dc5cc2dc8c69867453d9e6107

Observation ee319603-866d-47de-8cfd-377860463da3 · outbound

This paper cites Denoising diffusion implicit models.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Denoising diffusion implicit models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:43.405547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:43.405547Z digest=sha256:827cb74d0de2a0e81a2082de4b38ba15cf9b2ddd318c870012c3fc47a4918146

Observation d579e32a-5835-4b76-bfcd-d252459bf7e6 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:43.547722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:43.547722Z digest=sha256:0b346ca7a701f4b4ff727b7ca9931fa81e4ca830f3f28f30f1815e2f7f9a2d71

Observation e7f241bc-dc4e-4325-8d68-cd9951e562e9 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:43.641496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:43.641496Z digest=sha256:1137cab5cf929175071ac58b127196c83d021a983dc6ddd287b1a60776f55baa

Observation ddef94bd-2843-4ee2-9a80-afa8db31c78b · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:43.679253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:43.679253Z digest=sha256:55bfb3e75ac2a0e78e42f3bafb6312a8da17ae5a6622ad8a8a3990a3e12fe2af

Observation abeefc72-8c68-414a-9286-e6152dbc2056 · outbound

This paper cites Neural discrete representation learning.Advances in neural information processing systems, 30, 2017.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Neural discrete representation learning.Advances in neural information processing systems, 30, 2017

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:43.700408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:43.700408Z digest=sha256:868600534933e6c60a106f403a2f7e1ccf2496c15e99201de504bfea92eda3fc

Observation 38445a59-5b55-45c8-93b6-7cff775f2e24 · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:43.735167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:43.735167Z digest=sha256:1e0904fbe1ad36e7f0b32fdba22c536f7c81cdb2f1a35426f38f5cd29da1c0a0

Observation f27b6b7a-b5da-4050-9dd8-70a7abc4efa5 · outbound

This paper cites Omnitok- enizer: A joint image-video tokenizer for visual generation.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Omnitok- enizer: A joint image-video tokenizer for visual generation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:46.430205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:43.769211Z digest=sha256:99a22a95334e01e4dfbf362d4b98c448063b0588f6f7972f3cce3efbb60ab026

Observation 09589777-01c0-41d9-a7fa-a71f727c6021 · outbound

This paper cites Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:43.834494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:43.834494Z digest=sha256:f1410edfdf81b9b563adf536263b64d51a85646f6fcb3d2a73967b3ad1b11e46

Observation 7df56203-8bb8-43a9-9ae0-d2e0e0702ffe · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.IEEE Trans.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Image quality assessment: from error visibility to structural similarity.IEEE Trans

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:43.896565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:43.896565Z digest=sha256:7d4b2977d29d5760e66202cd73e8ab0d6c1676f2c3d5502a9ff7e807227b1b38

Observation a281c6ed-8266-4c27-b56b-e7cc58c55743 · outbound

This paper cites Maskbit: Embedding-free image generation via bit tokens.Transactions on Machine Learning Research, 2024.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Maskbit: Embedding-free image generation via bit tokens.Transactions on Machine Learning Research, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:46.225655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:43.936059Z digest=sha256:270355f91cb692408d3395d1dd5b0216d3409b4d85b37d72f4ead7ea391bb0d6

Observation 547639b4-fe7c-403f-be0c-81d633721b56 · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:43.966306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:43.966306Z digest=sha256:110fba8b1a900a4e79a83b13d8a7289036557d11a83c6f83eadd31794bef451e

Observation e2347e9c-686b-42a8-bf4e-ca9775cb6fe9 · outbound

This paper cites Look at boundary: A boundary-aware face alignment algorithm.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Look at boundary: A boundary-aware face alignment algorithm

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:46.011125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:44.008316Z digest=sha256:d6d0fa6a9a8fc79b047afa1a0ba3f42cde5edaac1bb4b4bcfe05b518070a57c4

Observation b9ffd809-d640-46c0-9b32-5213b6ef5073 · outbound

This paper cites Dstext v2: A comprehensive video text spotting dataset for dense and small text.Pattern Recognition, 149:110177, 2024.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Dstext v2: A comprehensive video text spotting dataset for dense and small text.Pattern Recognition, 149:110177, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:45.860861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:44.055684Z digest=sha256:086522603a0c5cf1a2fe921fa0bb00ee881d9bb61424339f0f68ec9653944134

Observation 726cb930-b5b6-4752-ab8d-79fe5e5f7849 · outbound

This paper cites SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:44.118090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:44.118090Z digest=sha256:9af27339007fe2735d3aea4a27fae822033e58e3375a0684dc219a1c4cba4096

Observation 0565b5fa-2ac5-4384-a627-9af4e521fb34 · outbound

This paper cites Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:44.187102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:44.187102Z digest=sha256:e6d7a4057a1611199b7775c89a923d505c60262f2cc5d1fe119995f3121cdc87

Observation 8518d46d-7093-4fb5-8e16-b98a18de85a5 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:44.252818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:44.252818Z digest=sha256:79e1f4e96f6083505a68d6b1ede227bcf6b61b6d916e3c9fad8f1a71ae074953

Observation 526033e1-549e-487a-86db-47208210f042 · outbound

This paper cites Reconstruction vs.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Reconstruction vs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:44.317326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:44.317326Z digest=sha256:4d379d7241f3c778216644535deebc990c295847f20d29f89caba4183ac0c89a

Observation e13fd772-71d1-45bf-a167-9e91cebc4184 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:44.404670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:44.404670Z digest=sha256:5d4a1b663bb7882a8845efd6c37a69594e4ef6ffec7bd96def54a8e1c8705abd

Observation 0f6d5428-0747-4474-a115-9e5cc92707fc · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.Advances in Neural Information Processing Systems, 37:128940–128966, 2024.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation An image is worth 32 tokens for reconstruction and generation.Advances in Neural Information Processing Systems, 37:128940–128966, 2024

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:44.491066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:44.491066Z digest=sha256:8c09ebac44b3664139e8096dc113ee3d12328749e40a18bacbee49082cf74f3b

Observation 3978aac8-65c6-4e68-98ee-f3a3522bc270 · outbound

This paper cites Representation alignment for generation: Training diffusion transformers is easier than you think.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Representation alignment for generation: Training diffusion transformers is easier than you think

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:45.740170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:44.557257Z digest=sha256:bc8d86410ca0415781307d874f470be74cf3391f73d703a05b6f737cc54e1717

Observation e4ff4ff2-b2a0-4462-8016-e90e675fa5a1 · outbound

This paper cites Fsim: A feature similarity index for image quality assessment.IEEE Trans.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Fsim: A feature similarity index for image quality assessment.IEEE Trans

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:45.595744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:44.647673Z digest=sha256:7a89c401a87e4a7cdf3feb9162fa55949b3f3f36199bb7beb27fa35c7116c98e

Observation 72240bd9-087b-4fd1-869a-5549e04fea94 · outbound

This paper cites The unrea- sonable effectiveness of deep features as a perceptual metric.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation The unrea- sonable effectiveness of deep features as a perceptual metric

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:44.720800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:44.720800Z digest=sha256:15830f5d5d7a1d1936d91d73f5aa4a333d1c3638396de16940f7ead9d5103d22

Observation b7c8214f-27be-44a3-bd25-f57e2114444a · outbound

This paper cites Icdar 2019 robust reading challenge on reading chinese text on signboard.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Icdar 2019 robust reading challenge on reading chinese text on signboard

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:45.417894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:44.764238Z digest=sha256:03946dd86c25ca5597e5b933283a311d024f3d5025f1815def4d6dc6cc2a9744

Observation 042198ce-a5bd-4b17-8d00-45b9a34297a5 · outbound

This paper cites Image and Video Tokenization with Binary Spherical Quantization.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:44.839423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:44.839423Z digest=sha256:cc8cff4c3713e69d5fdbc78b05fdd1ef5ff60e66f97119f294d142a5396bb0c9

Observation 495aa698-b650-4e0a-92ee-96a86e7b73dd · outbound

This paper cites Cross-Age LFW: A Database for Studying Cross-Age Face Recognition in Unconstrained Environments.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Cross-Age LFW: A Database for Studying Cross-Age Face Recognition in Unconstrained Environments

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:44.899699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:44.899699Z digest=sha256:faaec111fd091a444fc779bfc653f667d7e3e626bd81e76c7b62c461e7a31090

Observation c59e8b27-7396-4544-93f3-795942258c45 · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Open-Sora: Democratizing Efficient Video Production for All

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:44.972598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:44.972598Z digest=sha256:a1d86d481de80112c8c21b1c8bda94fd0f267378e5224d1e71e98a8cf1e13797

Pith citing papers

Observation e7d36cc0-c561-47c3-9812-501d86938b6c · inbound

Emu3.5: Native Multimodal Models are World Learners cites this paper.

Emu3.5: Native Multimodal Models are World Learners TokBench: Evaluating Your Visual Tokenizer before Visual Generation

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.600645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:6072c7a15af59aa31576923fd8026d0f34efae08152b738919cfa68feaddbe1d

Observation e1dc9a3d-6ca7-47f4-8ae3-15984e8730ab · inbound

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation cites this paper.

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation TokBench: Evaluating Your Visual Tokenizer before Visual Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:13:30.541166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T02:10:18.004724Z digest=sha256:241d66ddc0d131dfeeb4a305978f8fe84c272fd41b66a52e7492639dd2aa2598