Pith. sign in

Paper Citation Record · LEDGER

Factorized Visual Tokenization and Generation

As of 14 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 7 inbound Pith citation observations for arXiv:2411.16681.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16681 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:53:03.179680Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:27:00.494570Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:52:16.755017Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact2
  • verified fuzzy17
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c614e85-7187-4a9c-b35c-ab282131b842 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Factorized Visual Tokenization and Generation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:02.852919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:02.852919Z digest=sha256:b5b9de59467ee4099f3c546c71d2b13f3281fc16295e031ee0a81bdebeded02b

Observation d563a9a7-a414-4d2a-a217-4dee3f80cdd1 · outbound

This paper cites Learning by Reconstruction Produces Uninformative Features For Perception.

Factorized Visual Tokenization and Generation Learning by Reconstruction Produces Uninformative Features For Perception

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:02.859806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:02.859806Z digest=sha256:515e84f759281ebf27666b0348907a318801671fe6c9a71bc8671f6f5330986e

Observation 420267c7-9683-4767-9cfe-ef004302bfac · outbound

This paper cites Language Models are Few-Shot Learners.

Factorized Visual Tokenization and Generation Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:02.866325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:02.866325Z digest=sha256:ec03b80e0c25df2dd7e845cdeb1aca8847610b4f3d2323185ae990027d70dd4f

Observation 985800ee-8c61-4e77-9424-57526d357f3c · outbound

This paper cites an unresolved cited work.

Factorized Visual Tokenization and Generation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:53:04.168996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.872362Z digest=sha256:d54738cdacfa1d2393e46034599507831caaa29f3de9a0e94b1d8c8afd027c76

Observation a08600e4-ae59-4df4-b1f9-65244636bf60 · outbound

This paper cites ImageNet: A large-scale hierarchical image database.

Factorized Visual Tokenization and Generation ImageNet: A large-scale hierarchical image database

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:02.877774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:02.877774Z digest=sha256:04ae092152a8ec2a1c0cabd136b010491630aed0a6a61998a82cd131aa86c05d

Observation 15c5e754-daab-4098-87b9-b8952f39196e · outbound

This paper cites Diffusion mod- els beat gans on image synthesis.

Factorized Visual Tokenization and Generation Diffusion mod- els beat gans on image synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:02.884461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:02.884461Z digest=sha256:7cbf1aee80064171d1e385359631439a4a88fc55ace672a823b21f4ad5a5b4ae

Observation 9b5b1889-8f6a-4a22-af5d-2c649ea8f6d6 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Factorized Visual Tokenization and Generation Taming transformers for high-resolution image synthesis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:04.117916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.890010Z digest=sha256:d0bf2fa82d75481cf5ea68eedf8508b934db548764c6536baca8392434c91287

Observation 46f9553e-9ca2-472c-9c09-7858840829d0 · outbound

This paper cites Rethinking the objectives of vector- quantized tokenizers for image synthesis.

Factorized Visual Tokenization and Generation Rethinking the objectives of vector- quantized tokenizers for image synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:02.896487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:02.896487Z digest=sha256:91f1ea62bb94a865236c9d10d6729dda74e4b7b24aefeb4515066ac6fcc285ee

Observation f8f2ffa2-32ac-40e6-8521-cb0f55b10fa8 · outbound

This paper cites GANs trained by a two time-scale update rule converge to a local nash equi- librium.

Factorized Visual Tokenization and Generation GANs trained by a two time-scale update rule converge to a local nash equi- librium

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:02.902507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:02.902507Z digest=sha256:497d8a1a2124b2208c1d5112a0a016231139a2aca5915ad5b5809ba019cf0f62

Observation 42bde6c4-f2c6-4618-a66a-b25e3ef633fb · outbound

This paper cites Classifier-Free Diffusion Guidance.

Factorized Visual Tokenization and Generation Classifier-Free Diffusion Guidance

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:02.908008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:02.908008Z digest=sha256:d90fa94f666da11c0b29c6d7eb9a8ab34113b2dc3aa0c53724609419cf34c6d5

Observation fec3fcdc-3103-4989-9fac-6861e890df94 · outbound

This paper cites Cascaded diffusion models for high fidelity image generation.

Factorized Visual Tokenization and Generation Cascaded diffusion models for high fidelity image generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:02.915934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:02.915934Z digest=sha256:95a1c301e1cd3c55edb15b796d4f68a82d1b19e824ef692cf9e489bfb0da5788

Observation 94a6bd33-1c1c-47bd-9aa4-b0c49ce7c386 · outbound

This paper cites an unresolved cited work.

Factorized Visual Tokenization and Generation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:53:04.051446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.921660Z digest=sha256:80bd52b8735253a41fa72229d0554d131323a7d80aa2e07e18f805b8e915dc06

Observation 98d4afe9-9a43-477b-a93b-9e5201a94b68 · outbound

This paper cites Autoregressive image generation using residual quantization.

Factorized Visual Tokenization and Generation Autoregressive image generation using residual quantization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:04.024159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.926538Z digest=sha256:47a867ad204504b1062ca7c3ded2147eb20cd9363ef3348556465c2085ef7149

Observation b6b95790-9ace-4aa4-9786-37619b21df0d · outbound

This paper cites Imagefolder: Autoregressive im- age generation with folded tokens, 2024.

Factorized Visual Tokenization and Generation Imagefolder: Autoregressive im- age generation with folded tokens, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:04.001875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.931707Z digest=sha256:f6eebb4fd5f6e2ce924a23dd76aa450064e8cc7725ed7b6637b494e15f9c9a33

Observation b1005603-44f6-4ea2-9ea0-7cb2ebbdc52a · outbound

This paper cites LG-VQ: Language-Guided Codebook Learning.

Factorized Visual Tokenization and Generation LG-VQ: Language-Guided Codebook Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:53:03.555014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.937030Z digest=sha256:b140452dcd1676d18f5d6d40787014674eab0fee5c9cfdea6532c400c9ca1474

Observation 69fd7233-8676-4c4a-b0c8-64fcadb822b1 · outbound

This paper cites Visual instruction tuning.

Factorized Visual Tokenization and Generation Visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.979654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.942467Z digest=sha256:d37d0fa4e09e236484b3f61e7f1171abc97a2e8382bba1756f91e9b5293d80f3

Observation 166bed69-7ad7-4bc3-9344-3235abcffc82 · outbound

This paper cites World model on million-length video and language with ringattention.

Factorized Visual Tokenization and Generation World model on million-length video and language with ringattention

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.957036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.947528Z digest=sha256:19fa9049bcd5db666f5e2557c56e4f271cd952f8f0da8969e6c9ca27ad3d1923

Observation 7655a363-4b84-4ba8-bee4-269bc7234248 · outbound

This paper cites Open-magvit2: An open-source project toward democratizing auto-regressive visual gener- ation, 2024.

Factorized Visual Tokenization and Generation Open-magvit2: An open-source project toward democratizing auto-regressive visual gener- ation, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.932744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.953893Z digest=sha256:3179402fd64c497e6721245bfd245a7515690e7ca43ac73a3b1bd890da92ae4f

Observation 9dfb43ce-bdf4-4104-8d60-d00329ccf7ee · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

Factorized Visual Tokenization and Generation Finite Scalar Quantization: VQ-VAE Made Simple

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:02.960030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:02.960030Z digest=sha256:7bc883ae812bc4af775c2c342797f89a1ef42bafa6cd9d6f5385a7093fc2a6a6

Observation 6a334c9e-31ce-44e2-847d-0ed8a42c5d23 · outbound

This paper cites Dinov2: Learning robust visual features with- out supervision, 2024.

Factorized Visual Tokenization and Generation Dinov2: Learning robust visual features with- out supervision, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.913304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.966277Z digest=sha256:43bb2746b2efc5c920d715db23db6e8db6268692846744dc2a17a742bec7bfdf

Observation de87c126-9f51-41a9-af96-da28e59b0504 · outbound

This paper cites Scalable diffusion models with transformers.

Factorized Visual Tokenization and Generation Scalable diffusion models with transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.895857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.972589Z digest=sha256:eb14f1807aa75e27d3c79ad9478bd1578a0ee2dfd3f6017df4b8beed52313b01

Observation 44d0fbfe-c124-44a5-8d0d-98dae2dc2559 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Factorized Visual Tokenization and Generation Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.876799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.977553Z digest=sha256:78c4220ac4e6ff05eef392a43607e7f4008a4d431827a8267d66acb635e6eb20

Observation ec72ff23-121d-4f49-b8b5-c86498be19bc · outbound

This paper cites SDXL: improving latent diffusion models for high-resolution image synthesis.

Factorized Visual Tokenization and Generation SDXL: improving latent diffusion models for high-resolution image synthesis

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.855310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.982535Z digest=sha256:3a08ff0b7e3dd1647369c84c9e99b528e8fed4ba673e1641fde64433cf055ab4

Observation fc42fd21-809a-4836-b113-3954aa12741d · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

Factorized Visual Tokenization and Generation Learning transferable visual models from natural language supervision, 2021

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:02.987598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:02.987598Z digest=sha256:9864d886da5541612d62d838feae9ae82e559fcde03a9e40bac6277e9a3317bf

Observation 2ceff4d6-08fb-4091-b52b-e5a42cd4b47d · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Factorized Visual Tokenization and Generation Learning transferable visual models from natural language supervi- sion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:02.992267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:02.992267Z digest=sha256:36a95a0139604da39f1724b6db09e119e8f729bf0b0db0b0af8f2695cf867783

Observation c62e8cb0-31a8-445c-8ce2-1835ef851b1b · outbound

This paper cites Zero-shot text-to-image generation.

Factorized Visual Tokenization and Generation Zero-shot text-to-image generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.815524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:02.999013Z digest=sha256:8db15304e716b1d4cbab00b677ea4cc219212580781e0d1f4ce54bdcbe7ce90b

Observation 9de17180-ba7e-4589-81cc-227a255ce5d8 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Factorized Visual Tokenization and Generation High-resolution image syn- thesis with latent diffusion models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.005647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.005647Z digest=sha256:f9e4d2c3cbf6d1f56fee9123fec6199a7cefdec14bf57df0286a2f1645566da1

Observation 79103d46-9670-4c4c-ac6e-913ede894fed · outbound

This paper cites Improved techniques for training gans.

Factorized Visual Tokenization and Generation Improved techniques for training gans

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.013352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.013352Z digest=sha256:5e1d374212250079e45332212f55b32fecaabac8172c9c77c31d417c806a3eaa

Observation 25956553-f118-435e-a25d-22df65b92829 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Factorized Visual Tokenization and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.025896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.025896Z digest=sha256:334f4e62af52f119c03baa17c7ee801f9932430fb957fb300086ccd714670083

Observation c450fb9d-5454-44a6-a87b-ce100a26cf4b · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

Factorized Visual Tokenization and Generation Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.035718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.035718Z digest=sha256:0cc08a37e1ea693830e1171dd23771c53d9676d1ecd5cd5bf3a9db873d6a2278

Observation 90ceecc4-f347-4460-83b7-db98ab3c6a9a · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Factorized Visual Tokenization and Generation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.041801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.041801Z digest=sha256:2aa8a3761eb60380af6d71051204f544742f1a1eb3c9f06ee1ec304b463c2b6d

Observation ee738ead-374f-433f-b6c9-38dd59697846 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Factorized Visual Tokenization and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.048156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.048156Z digest=sha256:86c006700a5559231f271377eabb3a37d8988e86f1330fd25868e40eab2e9fe2

Observation 54c9b807-7f0a-4045-99eb-cbbf299f48de · outbound

This paper cites OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation.

Factorized Visual Tokenization and Generation OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.062730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.062730Z digest=sha256:15c27d7365f98da4d91ab77c6aa0118b6e4089f8782610ea5a6b26031e685c29

Observation 4099974c-d713-4a7b-b95e-18a720bdccda · outbound

This paper cites Image Understanding Makes for A Good Tokenizer for Image Generation.

Factorized Visual Tokenization and Generation Image Understanding Makes for A Good Tokenizer for Image Generation

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:53:03.374057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:03.068683Z digest=sha256:83fd666d2b1c1e0c468f22844fa2be7897c7c4350a8b043d12ebded392754bc9

Observation 8bd9d7e4-f40a-4550-b6b1-287cb3269d92 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Factorized Visual Tokenization and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.074800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.074800Z digest=sha256:3c63b53bf0ac27fce25ff93f8ad56b3b258426089a56e1ad816d010747ed4200

Observation 844579fd-6c73-45cb-b8a6-0a32a85d6b2b · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Factorized Visual Tokenization and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.080519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.080519Z digest=sha256:6b44607b0d2bddecb4b531fccff5dda12c3e1b17aa494d004241e313e5d860ba

Observation 4b998e01-bb12-4cc3-a970-1ffd3712dfa9 · outbound

This paper cites Locally hierarchical auto-regressive model- ing for image generation.

Factorized Visual Tokenization and Generation Locally hierarchical auto-regressive model- ing for image generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.093747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.093747Z digest=sha256:06d5bd6cc0160f7d3d0e40caf6cdb7c6bbbe7578335d1ea0d3819533fdc0d635

Observation c559c6a0-9d71-484f-bb87-2234a9b52235 · outbound

This paper cites Vector-quantized image modeling with improved VQGAN.

Factorized Visual Tokenization and Generation Vector-quantized image modeling with improved VQGAN

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.760620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:03.104147Z digest=sha256:0a787f697429376e25ecb0792711392efa28df08ad8974b41128f7ce33ac3511

Observation 78a318d2-e4c3-4c7c-9dd0-45141a44a15a · outbound

This paper cites Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms.

Factorized Visual Tokenization and Generation Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.740781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:03.113736Z digest=sha256:5d8b9248f1ddc0b78cdcca8e69f4d59adb8169cba4d7c2a113e7e41294fd7f11

Observation 28410548-6a14-4c93-a97e-44d8be7b9f99 · outbound

This paper cites Language model beats diffusion - tokenizer is key to visual generation.

Factorized Visual Tokenization and Generation Language model beats diffusion - tokenizer is key to visual generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.719473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:03.121625Z digest=sha256:b54d76f131f05652e5427821ead6b41efc9ffb222cd4fb5dc9c1e53b3bd975fe

Observation c459ae0a-da6f-4aaf-a8a9-04f49e0688b9 · outbound

This paper cites An Image is Worth 32 Tokens for Reconstruction and Generation.

Factorized Visual Tokenization and Generation An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.133992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.133992Z digest=sha256:089ad522b53a662d13359df6e5aeddf759aa3152879521ac97826d27047de6fa

Observation 37a553b8-3e1c-48b8-9dcf-87197ff5ad94 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

Factorized Visual Tokenization and Generation Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.147900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.147900Z digest=sha256:777f5e906509b7c1428c751b625879c3f7925284b6c197f0f11b61878396c7bb

Observation 7b7e3c6b-a3ab-4102-b0a5-d5cda0cb1924 · outbound

This paper cites Efros, Eli Shecht- man, and Oliver Wang.

Factorized Visual Tokenization and Generation Efros, Eli Shecht- man, and Oliver Wang

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.700269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:03.161975Z digest=sha256:9c27846a59ad3d9c1320e04844085f61eee549a8875d67f62ac5c76babdbd74a

Observation 70454539-6648-401d-bd20-873e75ff6f09 · outbound

This paper cites Movq: Modulating quantized vectors for high- fidelity image generation.

Factorized Visual Tokenization and Generation Movq: Modulating quantized vectors for high- fidelity image generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.682207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:03.169317Z digest=sha256:0fb30a65927417038bfc0f39ce1bc25dca53c5acf29e8a4673be7e2bf029f50e

Observation 0b9c0168-1447-40fb-875a-bb033ec8e620 · outbound

This paper cites Beyond text: Frozen large language models in visual signal comprehension.

Factorized Visual Tokenization and Generation Beyond text: Frozen large language models in visual signal comprehension

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:53:03.660665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:53:03.174730Z digest=sha256:60fa71aecc881b6a8e5881f04c1b1cfb50036a9cff9639c6fb35643489b0ca9d

Observation 93290248-c703-4097-aa4b-77e74cd4525f · outbound

This paper cites Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%.

Factorized Visual Tokenization and Generation Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.179680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.179680Z digest=sha256:31656a702c4f86608f892822c64ae0a0cb464ff6a14a048a484e10b53dcd80a6

Pith citing papers

Observation b5636fde-9d7a-46a8-9872-69808b8f4915 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey Factorized Visual Tokenization and Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:32.893497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:4d9c9d522f42ff02a74a36d5d229b06aa0793081f5154105b337c907bfe2028e

Observation 592587ef-d990-4de0-82ee-026b8526f89b · inbound

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies cites this paper.

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies Factorized Visual Tokenization and Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:52:16.759235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T23:51:43.934329Z digest=sha256:0a7572af44f7caff2a7bf09584bc7fd76f7886233558c0224e0c4349e9db7f17

Observation c6c07830-669c-4a0c-a60d-4e6f2a51b4ff · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Factorized Visual Tokenization and Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.571001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.571001Z digest=sha256:80d4b1cfc880fcca43adfb709aee575bca81e4dc44644a16fe5963ac8842d6de

Observation 1db8c4ea-d040-4ad5-b68b-0601ba76c886 · inbound

MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization cites this paper.

MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization Factorized Visual Tokenization and Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:54.664956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:54.664956Z digest=sha256:dbb2772959353fcc31f812ea2bfcccadcdad8787245ffbb84e091c08d7e0dc7c

Observation 2400ffc2-4e56-4b21-8087-6ff3ed36557e · inbound

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction cites this paper.

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction Factorized Visual Tokenization and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T11:07:59.517713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:07:59.517713Z digest=sha256:1f15c551fa207ed081691ccab3ddec091635a22c129a8e80fa5c565bab5994c1

Observation 53fb9e37-5405-4292-ae70-7f2e9e8d30c8 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens Factorized Visual Tokenization and Generation

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:08:15.849861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:ccd065f0bab16ba881e9299ccc404d1920c13c4d32095be946f86d86599670a9

Observation d66f4a04-3b30-47f3-8d59-c55036c45aed · inbound

Tokenizer Generator Coupling in Medical Image Generation cites this paper.

Tokenizer Generator Coupling in Medical Image Generation Factorized Visual Tokenization and Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:27:00.494570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:27:00.494570Z digest=sha256:e866bb961e07319710a302534afcc52f4748501edabeb922f0bf7d156c4aa65f