Pith. sign in

Paper Citation Record · LEDGER

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

As of 9 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 13 inbound Pith citation observations for arXiv:2502.05178.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05178 v1

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:04:02.460265Z

measured 113 of 113 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:56.983795Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:28:55.819781Z

Reference resolution

100 of 103 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8bd6c170-fe81-48bd-a650-84f219af87a7 · outbound

This paper cites GPT-4 Technical Report.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.010635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.010635Z digest=sha256:06d8d61fb97a5632f4e812d85fd538b1f2be95456d78b014f34324638b235eda

Observation 6a4c86db-f4e0-42f8-8f55-699fa5ff417e · outbound

This paper cites Soft-to-hard vector quantization for end-to-end learn- ing compressible representations.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Soft-to-hard vector quantization for end-to-end learn- ing compressible representations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.016595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.016595Z digest=sha256:be55efd0cd02b5543fde9914260ae1fe2ca8ef638663e3617214eea83fe40a97

Observation 70005a9b-b630-48a4-b360-ed96d1e4ba27 · outbound

This paper cites Beit: Bert pre-training of image transformers.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Beit: Bert pre-training of image transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.021553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.021553Z digest=sha256:a321eabef39ecd696e289ae8f8ef4b8814df2de98f8f967d4b3326ab277d7fed

Observation a550fff0-826a-4228-8e97-170b988d81e5 · outbound

This paper cites Fuyu-8b: A multimodal architecture for ai agents, 2023.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Fuyu-8b: A multimodal architecture for ai agents, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.025789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.025789Z digest=sha256:812e4872566d561adce28c8531a2023709ba9bb310c26640d1ba8a519e619700

Observation fd6f13ba-d675-417d-8964-02ab70d2b707 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.030015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.030015Z digest=sha256:f6a9c1fa63d328da2d4118d472a38adcff875c85b7fa972fbd3869ae5f8d4c9b

Observation 7d77e1e7-729b-43e5-a76e-dca36ec3a450 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation PaliGemma: A versatile 3B VLM for transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.035083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.035083Z digest=sha256:bb270aafa1af5dada3e8e85a6046f6dcc3e2b5308ddcfba3a503732045a7868b

Observation b1754986-a731-4cee-8330-f1097faf8ba9 · outbound

This paper cites Piqa: Reasoning about physical commonsense in nat- ural language.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Piqa: Reasoning about physical commonsense in nat- ural language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.040713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.040713Z digest=sha256:4cdbd3a2ef93cfbb63d56564fd119a4c28f1717ce13f4e8f017d2358352b307c

Observation 454e073c-559f-424a-9afa-6a2da8bf9107 · outbound

This paper cites Maskgit: Masked generative image transformer.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Maskgit: Masked generative image transformer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.045345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.045345Z digest=sha256:6989cbd6aeeb6ded0a18359bd11faee876f609c5c35c1a8d4f0cd34774787f48

Observation 2d31b2e6-2585-473e-bac5-f793d67e00a4 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.050027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.050027Z digest=sha256:6d1a0089293e4c4ccf341d12bde9c2f71ddc137334ac024b02f97f01b7064fb1

Observation cf430773-b946-4491-9606-1a949a34a93e · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Training Deep Nets with Sublinear Memory Cost

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.054853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.054853Z digest=sha256:3e51f3e888ff96cd600d590b5363078b13586eda0f7c54cc5ae33eef1cf449e4

Observation 2e71b303-585c-422d-aefa-e640109b3891 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.059660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.059660Z digest=sha256:8a12eb928a1315202fb4c74ddc2727757b1c04a46fcf8301deb7deda40d4232a

Observation 9d0ea132-5568-4f63-817c-4766f2b741b3 · outbound

This paper cites Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.063916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.063916Z digest=sha256:dd11d77c5b81bd754fb4f0f1bdd0d5919db390b3bf6deb7ab678f6c1440a9a9a

Observation a05f44f9-6282-4efa-863f-973f8f67adcc · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Reproducible scaling laws for contrastive language-image learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.067982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.067982Z digest=sha256:abc38ff094a0318d8dd591a53b07b6f062bc6b4d6ea070e4fa7807dd22f4a93c

Observation f59c9454-ea65-42bb-a951-49cd6d2377b2 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gonzalez, Ion Stoica, and Eric P

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.072545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.072545Z digest=sha256:80214c904174d188cbec99b3e746dae7cc8ad01a304732e114b3616d918a1ad8

Observation 4298893b-9776-4bd9-9e09-8931c84393dc · outbound

This paper cites Scaling instruction- finetuned language models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Scaling instruction- finetuned language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.076761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.076761Z digest=sha256:895fb7f5f62ebbe46c94333b6e274e7724ff8ef3cf6501c55464cae2eb4b9776

Observation 303ad003-cea0-494f-ba3e-2810c68f09b8 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.081513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.081513Z digest=sha256:0e73003522a3f343b8c55aafe3ebac13f7d00c2f2d1d3ec7a2f04addc4eee795

Observation ab67c416-a7ab-4037-99ad-b0399c4fac87 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.087307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.087307Z digest=sha256:ebdd43586c51785216925a6f69d07efdd6e2d34b1879bf77e700c906cf35dbe4

Observation 68008d26-52c3-4387-8c0d-5669bcfcc032 · outbound

This paper cites Vision transformers need registers.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Vision transformers need registers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.091863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.091863Z digest=sha256:0f1ba8379ee51ef15ebeaf26f966a75a61bf744aaf74a888bfa8768750f66d3d

Observation 474d4554-538f-4321-886f-968502382f03 · outbound

This paper cites Scaling vision transformers to 22 billion pa- rameters.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Scaling vision transformers to 22 billion pa- rameters

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.097385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.097385Z digest=sha256:6e98d8d49c7ec9baa384aa6ef548808e17b63758a1e236322b0347d92f5b04d5

Observation cf644a5d-2fe0-4b4f-abb8-bf8c8b01521f · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Imagenet: A large-scale hierarchical im- age database

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.101775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.101775Z digest=sha256:28b8a8e0184d142b32f38da90a90d0ac9591ab66f7d0d26bba101a9ab614930f

Observation aebfa514-5cbb-4c2f-99a4-f716ca22b36a · outbound

This paper cites Unveiling encoder-free vision-language models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Unveiling encoder-free vision-language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.105628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.105628Z digest=sha256:ca17dd10a5229bcb6f03f1e1a44b8226d8b270feb84d3310f8882e1fc120b621

Observation fefafe59-b9a5-4b4a-995f-44bb4b20097e · outbound

This paper cites The Llama 3 Herd of Models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.109316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.109316Z digest=sha256:4ccad4eced1cbb066b25f53df0cebf3e1450241fc2760e833dc69e57bde35255

Observation e96fd0f2-3d88-4bb5-8d7f-860cde74123a · outbound

This paper cites Scal- able pre-training of large autoregressive image models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Scal- able pre-training of large autoregressive image models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.113967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.113967Z digest=sha256:7d499d84b8198bcbcb0427568439b4e095e04b9705950ac6693f314af0b83293

Observation d5922cbe-cd1d-4439-8f5c-3877db231ce5 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Taming transformers for high-resolution image synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.118734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.118734Z digest=sha256:b09faeb6feee9c5fc08364f2231da3570ff3a7fa40bbf1987417e9a34f67be5f

Observation e2648a11-c0fd-4ed2-9ba1-92e466db8701 · outbound

This paper cites Eva-02: A visual representa- tion for neon genesis.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Eva-02: A visual representa- tion for neon genesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.123315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.123315Z digest=sha256:c3bfe73880ce15f0e5efad1061168d7e8a1f5a4e90ded6d2e5b23c25fc59cfd9

Observation 8aae7597-f829-435a-8876-3f2c0b6a72e1 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.128254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.128254Z digest=sha256:d8df533a9dee9320af01d2e9b9d39808e4c4fca37694fe8e9be6118d7feba4d2

Observation 68643c2f-f220-4327-9c53-59953585761a · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Dat- acomp: In search of the next generation of multimodal datasets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.132772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.132772Z digest=sha256:d29e7957bb5ebfeaee8bf717cb32541c21a128160e9734968abaeda498bc7cc3

Observation ac7db086-f8f4-453a-a351-6da544f5c8b1 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.136804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.136804Z digest=sha256:0e3a0028bb9865c6e5efbbf381f4e55296e9098a69ec1bac8b6dbfbe71c8d282

Observation c849945b-773c-451c-b9e5-9512d611fba2 · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Geneval: An object-focused framework for evaluating text- to-image alignment

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.141742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.141742Z digest=sha256:74f4198be2680046b02939aa280b2907410832f28e74026347aa6ee29e7f2d5c

Observation b1d319c2-3025-4ddf-a27e-47606c75b218 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.146184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.146184Z digest=sha256:212b8b650ffc56ec01e9cbd0b627f9bccb35697f4ca25a68dbf35c47a5ce1fd1

Observation c8701ceb-b64e-4d31-9c9a-1f3dea395f7e · outbound

This paper cites Making the v in vqa matter: El- evating the role of image understanding in visual question answering.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Making the v in vqa matter: El- evating the role of image understanding in visual question answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.150055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.150055Z digest=sha256:8dcd325d263d7ec1cfc20f8defa37e267159a989f0f937e7444276c4759b1e3d

Observation d270c97a-4c6c-4981-98bd-4a958a0aa107 · outbound

This paper cites Noise-contrastive estimation: A new estimation principle for unnormalized statistical models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Noise-contrastive estimation: A new estimation principle for unnormalized statistical models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.153789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.153789Z digest=sha256:b12b5c8e03956b65f7f6a79b953087b75ebba9a1af337ffefc529d0539683a32

Observation f9794115-36d8-4401-896c-bc4332b5b033 · outbound

This paper cites Masked autoencoders are scal- able vision learners.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Masked autoencoders are scal- able vision learners

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.157535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.157535Z digest=sha256:8f3163bc3611f77e06a7c7a7e7921b1d45cfa696ad2ce990c009f0fbfca396ea

Observation 69f1c8aa-6680-435a-b8b6-7313cc4c3e6c · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.162100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.162100Z digest=sha256:464870bd7af9f6455ae04dc6dea89ecfec0d7361949ea8a40c65b6925851e7d7

Observation 2d757e65-fedb-44df-bad2-8b04f82e4657 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equi- librium.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gans trained by a two time-scale update rule converge to a local nash equi- librium

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.630265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.166745Z digest=sha256:ff231b788107d47631126431d53111aa99671ebb519a03412fa295a3fb6239d6

Observation 734b2f13-c2ce-4f58-b9ff-f29389b70955 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.171586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.171586Z digest=sha256:34352bf747a9df17160aad801782b20d8c39c23d706a699f57287c002b8ac38c

Observation 3492adc1-c3f7-4d19-9c2f-c4d58fa7e5fd · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.616127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.176426Z digest=sha256:c4d2cd49c02741660ad92bc8ae8af0fbd312f02e3428e5662f50041d6d67d285

Observation 5f215137-e47e-4192-8f11-6eec88dd893b · outbound

This paper cites Coincidence, categorization, and consolidation: Learning to recognize sounds with minimal supervision.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Coincidence, categorization, and consolidation: Learning to recognize sounds with minimal supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.602250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.180404Z digest=sha256:1cdf09192788b0abe205e700a7d50ff118727bc281e6c25b4b53bb2c38d86ad9

Observation 14eb0cb4-5d4f-4b17-94a7-bc8f0199e4a7 · outbound

This paper cites Unified language-vision pretraining with dy- namic discrete visual tokenization.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Unified language-vision pretraining with dy- namic discrete visual tokenization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.587534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.185578Z digest=sha256:40c22d21f59a1ca0040f8811a8546608928bebaee940ddf8ea58ed1985a9a6f7

Observation 79a1eb1c-072e-424c-90bb-900fcc3bbcb1 · outbound

This paper cites Analyzing and improving the image quality of stylegan.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Analyzing and improving the image quality of stylegan

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.573043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.190018Z digest=sha256:e97d2cadb17dc14c66965976441992a7bd6a07535f4a0593aea6925362315570

Observation d6e18f26-832d-464c-9ee6-9e7bbca7cb06 · outbound

This paper cites Segment anything.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Segment anything

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.558335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.194325Z digest=sha256:3bc7b496fa59b26ca11aebb8eea60dbbc6de09f13be63b04e3000a79d318e3b4

Observation d3bd40fc-aa3a-4bb7-9e9d-1789d70b079d · outbound

This paper cites Sentencepiece: A sim- ple and language independent subword tokenizer and deto- kenizer for neural text processing.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Sentencepiece: A sim- ple and language independent subword tokenizer and deto- kenizer for neural text processing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.540802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.198689Z digest=sha256:0e99037066d8b420512eacc8233f3f3de41eabd6287c1fddcb56c29bcd92c7fe

Observation a4077dbf-4bc7-4ede-b7ec-9613a7a755ea · outbound

This paper cites What matters when building vision-language models?.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation What matters when building vision-language models?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.202500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.202500Z digest=sha256:208eb1283e603dfbebbdf3a6c0ab547b0797085cd9f1e6e938f25baaa9bd319d

Observation 2214d419-98f5-4334-8895-303afed07f3a · outbound

This paper cites Autoregressive image generation using residual quantization.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Autoregressive image generation using residual quantization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.524016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.206911Z digest=sha256:714d01858391ff9d6d1b046ee7bd261b58cc011fdb412913cd3c8e7ad59f64b0

Observation 4a9100d2-60fc-4203-a762-1182eb0bc4dc · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation DataComp-LM: In search of the next generation of training sets for language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.211697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.211697Z digest=sha256:426808b8c76baaa3fb1799c8c3ef0cd2af2d9c24996dc1915c8641995e2c9ce3

Observation 8eb81f88-1952-4d37-8eb9-4025db60a17f · outbound

This paper cites Mage: Masked generative encoder to unify representation learning and image synthe- sis.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Mage: Masked generative encoder to unify representation learning and image synthe- sis

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.509256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.216578Z digest=sha256:7a7a9bd590894768abfb96c8910ebb3c3b1a95305ad730ed7bb62ec46f7b8919

Observation 20778747-fe02-45bb-9223-4d91232e0709 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Evaluating object hallucination in large vision-language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.494385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.220361Z digest=sha256:56a461685a6db544431e344e3b468d373a25eedabc8a51ad5c004b3d148c0314

Observation 189cdc51-2d4a-4c52-9488-4c49d257c126 · outbound

This paper cites Microsoft coco: Common objects in context.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Microsoft coco: Common objects in context

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.480719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.224653Z digest=sha256:909510b3a919dbd5611876a14eac6b5e21a678e2cdc76b35915621fc23a52c64

Observation 0952fad6-6b39-4b4f-9d53-f48ea7c7fc11 · outbound

This paper cites Visual instruction tuning.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Visual instruction tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.228660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.228660Z digest=sha256:650616e1c363739b5cd79e9629f6b20a0c03b9b729d04dd9863e88f68f423788

Observation db0fce50-4a05-4b4c-bb37-dfcabc1c592f · outbound

This paper cites Language quantized autoencoders: Towards unsupervised text-image alignment.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Language quantized autoencoders: Towards unsupervised text-image alignment

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.457846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.232849Z digest=sha256:587840cd914749891d93a7b25ad36df0e5c08797244aa4ac55b3d32a491c2186

Observation 5595dd76-13a6-46ac-8d75-33a5649c7c66 · outbound

This paper cites Improved baselines with visual instruction tuning.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Improved baselines with visual instruction tuning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.442955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.237764Z digest=sha256:928a655050c0707d0d567e2214e97adde1c9fa149391bb9d418b9ed9b2f78ef7

Observation ab91549c-e08a-40e2-984e-525243af81ec · outbound

This paper cites Unified-io: A uni- fied model for vision, language, and multi-modal tasks.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Unified-io: A uni- fied model for vision, language, and multi-modal tasks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.427785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.242074Z digest=sha256:6962e9289de9a392e32c09b7b794912333570d7b65e3a6cfcfbabaa5d4073742

Observation fec7a35a-afad-492b-8035-5bd024506d7b · outbound

This paper cites Unified-io 2: Scaling autoregressive mul- timodal models with vision language audio and action.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Unified-io 2: Scaling autoregressive mul- timodal models with vision language audio and action

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.414279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.246605Z digest=sha256:5c7c365ab5bfc00af1db675b13ee43bab3c2b5328f01a9e04457ae5789183459

Observation 7e77d0d2-b0b7-446e-9b91-8b20da7b51cb · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.250674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.250674Z digest=sha256:c60b7c0b8131f0340e177955b9db4017f9b8c02eb51bc3eb07d908a1cf837edf

Observation 252c49b0-6672-4642-99e3-a5b1ec6bec32 · outbound

This paper cites Finite scalar quantization: Vq-vae made simple.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Finite scalar quantization: Vq-vae made simple

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.400187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.255181Z digest=sha256:b02cf785224783b51c0bee192c71639217a11b637ec415ba3fa5525a28b1b813

Observation 4a13c14d-5261-4a09-a6af-fd51658da579 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Representation Learning with Contrastive Predictive Coding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.260427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.260427Z digest=sha256:4077849b77d6071553644dbb2cfaff17190bfdcbf9e0601a62d4e53e08ee3101

Observation 957365a5-9e76-4c88-85eb-300f411f3848 · outbound

This paper cites BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.265135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.265135Z digest=sha256:1f7e91eb60af75a103149c8065c039ae1a84df05aca289bcd3c4524561772f38

Observation 579222c2-f70e-45c7-b5fe-90f53bb5ea27 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.270583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.270583Z digest=sha256:3b77c8174a57c874311c80d6d360076be24e73554aa2cd0ecc63ee97a03ebc5c

Observation ef1bc0ab-2f53-4a0c-8a65-83d6094d99fc · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Learn- ing transferable visual models from natural language super- vision

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.386055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.274919Z digest=sha256:6e7ea4b198eea952ef9d0261c43461ced5dae905148fdce953b288e4a8137bef

Observation c9d1c10e-48f9-4711-9b7f-197257a99594 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.371521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.278871Z digest=sha256:2c8247b1717c2703d564e8d76841dbf7e5a64efc80c9ef50eff92b35592c1416

Observation 70c991cc-d2ef-4978-a734-efc4e8326a6f · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Zero: Memory optimizations toward training trillion parameter models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.282886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.282886Z digest=sha256:721983dad30f8d63b3c71f780721b9132c6ef77efd0bf0b2b54b465dbe95c64f

Observation cf5a9212-2786-4c1b-92f0-4d52f8935b82 · outbound

This paper cites Zero-shot text-to-image generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Zero-shot text-to-image generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.347768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.287680Z digest=sha256:4187c9351bd26586c227a71248359c98cf63ba147169133368cb29ba3a2aa9cc

Observation d7a5a524-7b93-477f-a79e-5bf44a845d72 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.332837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.292062Z digest=sha256:3834b7f0caa1ab723059f1db6ea8f7ebb91ee7e3a002100f703b98c29d575e07

Observation d8bea48c-add4-462d-a853-fe13f497a8bb · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Winogrande: An adversarial winograd schema challenge at scale

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.297126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.297126Z digest=sha256:853f267dce601749d9a79fb694fabc232f2cd8f50b266a3bb347ab06a93e154b

Observation 8f40caa5-dcec-44ef-96dd-85f5620d0a7d · outbound

This paper cites SocialIQA: Commonsense reasoning about social interactions.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation SocialIQA: Commonsense reasoning about social interactions

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.310139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.302005Z digest=sha256:0dfb1dacc1c3c42e112646470ead31e84e9bec2cfbdb8c6cf3bea05c0da604be

Observation 3ac211ea-e85a-4996-bc70-9f34a5f8aedd · outbound

This paper cites SBER-MoVQGAN or a new effective image encoder for generative models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation SBER-MoVQGAN or a new effective image encoder for generative models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.296907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.306771Z digest=sha256:6ffaf21f933e50ef0888c07347ba82430282f506d6a6156ab0a810321047c12e

Observation a7e14180-f7fe-4cfc-8cdd-9b55d184716d · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b-en.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Laion coco: 600m synthetic captions from laion2b-en

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.270760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.315752Z digest=sha256:bda130db238d01c301f38d96ab37072ad19e02048ea6712e30c139b4ba40ded2

Observation 6348c906-7c74-42d8-b82f-6f110828263c · outbound

This paper cites Japanese and Ko- rean voice search.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Japanese and Ko- rean voice search

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.256999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.320176Z digest=sha256:5cf7abc56cb75ce1564bcf98b04f7cd441bbec9495ea5dbda548b0991ff96747

Observation a9cb6e2f-c90b-4a91-8bfb-2fda825d10cf · outbound

This paper cites Multi-task learning as multi-objective optimization.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Multi-task learning as multi-objective optimization

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.243667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.324692Z digest=sha256:7c940081cd1bbbf145b3475f61cd53f2424ef069f269a5069d55f79290ada646

Observation 5b528a2d-eeca-40fc-a17b-4f30eeaa8852 · outbound

This paper cites Neu- ral machine translation of rare words with subword units.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Neu- ral machine translation of rare words with subword units

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.230277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.328915Z digest=sha256:1c6f599a93abb90f2eaf3612dfad6f3cc324fafc52d6bf19e7cf44a26a223c5b

Observation 763974f5-2b29-4718-9563-692d710ec268 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.333188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.333188Z digest=sha256:7a19bb15e70522744af82fbb915817de995fcc54ddfdc2acac9d19e0dfe29754

Observation a79b9ac9-7ebd-480d-99b5-c170ebc32383 · outbound

This paper cites Very deep con- volutional networks for large-scale image recognition.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Very deep con- volutional networks for large-scale image recognition

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.216760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.338045Z digest=sha256:c8175a2605f68341af54b6e2aad69de58a334d33f7d900cb03d89f48e9342f99

Observation 3ed30314-8840-45f0-a088-3f3659833ef6 · outbound

This paper cites Towards vqa models that can read.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Towards vqa models that can read

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.202775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.342447Z digest=sha256:fefc4c011920ed2d37de95b8601111c9efd30f5c2062181c2fed70bacc0828e5

Observation 1f0bda81-6ac2-4fb6-86f0-acb4f34fcbff · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.346378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.346378Z digest=sha256:8ae58f58a691f2665800f0c4a497992bfd406c3fddf8e3bb055668a3cccff2d1

Observation f9e9d59d-a110-4b33-a961-24993950a4dd · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.350771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.350771Z digest=sha256:0f63b215924aabbbe91f74463bb92ba04a41ff7cf07804d935c67b4731a24b6f

Observation 6c28f588-8868-4027-bf90-ff5b54a965e5 · outbound

This paper cites Rethinking the inception ar- chitecture for computer vision.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Rethinking the inception ar- chitecture for computer vision

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.188991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.355141Z digest=sha256:8e58a0ed16e49f23ee1ea3e4923479d4da73a276ee7b80a7a7d6978a354e12e9

Observation a51555ff-f1f5-4c9c-84e1-ec308258bc67 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.359195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.359195Z digest=sha256:085798c0e4552cb593d686034f12f6792ac77feaba63d134622d614386c8fa3b

Observation cf41c74a-f4c7-4a83-92e8-5d0bf36ab6e2 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gemini: A Family of Highly Capable Multimodal Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.363640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.363640Z digest=sha256:49a9c5d3cba7b285438a587ca87bea6e7af1bf41e0b10bd97bfaf4b30cb2193f

Observation 5b9250a5-9f06-4cca-ac87-fe9296e98b1f · outbound

This paper cites Zerocap: Zero-shot image-to-text generation for visual- semantic arithmetic.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Zerocap: Zero-shot image-to-text generation for visual- semantic arithmetic

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.175294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.368347Z digest=sha256:9a3fefb2f318947e54c65168cd9d27e7cfa70df633315bce4a89023e932ea55c

Observation 7d387ad0-554e-4504-b412-870f38858b29 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.161572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.372743Z digest=sha256:bac7c82bd0bdc4bd571e1ffbafa17a7fda5e76ab1fc9088b7cab31df70ea6f75

Observation 5120663f-9535-485f-887a-87118f195826 · outbound

This paper cites Cambrian- 1: A fully open, vision-centric exploration of multimodal llms.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Cambrian- 1: A fully open, vision-centric exploration of multimodal llms

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.146572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.377487Z digest=sha256:d4c64f40cdfaca6610cedbf42f16059fc76a03972050a9fd2a620f0727c30678

Observation 4ef7273a-8b91-487c-913e-eeea0dfea38d · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.381765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.381765Z digest=sha256:0b76f15117816b56aa95d486d431e72258fde9909c24e11550761ccdc1b9d5d0

Observation 54ea1290-4ad9-4b80-a076-ad1d10fac364 · outbound

This paper cites Neural discrete representation learning.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Neural discrete representation learning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.133397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.386558Z digest=sha256:b871e9144f5eab3441949d69a691e307ceb96274afe41171b4792aeab9be4578

Observation 3c79048f-2f57-4847-90c7-e066fc8b41dc · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.390941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.390941Z digest=sha256:3f31b5eb80bc28f90aa71994881d563c46a812a81f2e5adab5502a4825e59203

Observation fe84ed4d-84af-4a46-995e-7f2fdec335e1 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.395739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.395739Z digest=sha256:ce372e24b0257d70a92f63a5a0872cc3e6b5006b176d807e008c25b8bf3ef3ca

Observation afac164b-4635-4703-9a45-4c183a40e4aa · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Image quality assessment: from error visibility to structural similarity

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.120627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.399952Z digest=sha256:1bc7271b416576309856ee9743e0edb972e05918de8e0fcdb1b039dda174454b

Observation 5dbfe271-8bd9-4e92-94f5-23d6550ba9b5 · outbound

This paper cites Diffusion models as masked autoencoders.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Diffusion models as masked autoencoders

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.106327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.404211Z digest=sha256:756f5a8053aa94978e5a9c53dbc658b4d7af4b5aa6f2b8bfdfb6bc7d9e72dca4

Observation 6ccbb9ee-5de6-4445-b8b7-efc5d7d6f211 · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Next-gpt: Any-to-any multimodal llm

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.091696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.408117Z digest=sha256:24133c16edec1c16f40dc9c296de3bbff07a4120e16e42bf47f9144a6660eaf5

Observation 21678f6a-67f3-4732-80f8-d69a1c6719e0 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.412305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.412305Z digest=sha256:c94d0622cc2dc8771ab784f424fc28300ce3f3b76e4149c519ccc9612db6fd52

Observation 4ca01e07-7e29-478c-84da-50fb65cac641 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.416333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.416333Z digest=sha256:0973bbfab2ecc3b9306d8ac56138af17331d9c10d30ebb9a8a53bfa9857adec6

Observation 11781f39-f765-4cfe-ba2a-9bcd7bb75604 · outbound

This paper cites Vector-quantized image modeling with improved vqgan.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Vector-quantized image modeling with improved vqgan

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.075755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.420675Z digest=sha256:22243f14a1b9f2753312a8ebc51943659dfb69086dd7cd54a248668775c80fc2

Observation 1b4ddafe-09af-4ef2-94c8-7fe345d7b8b8 · outbound

This paper cites Spae: Seman- tic pyramid autoencoder for multimodal generation with frozen llms.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Spae: Seman- tic pyramid autoencoder for multimodal generation with frozen llms

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.061179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.424933Z digest=sha256:67c433ad7b713aca6562c9d14a8d2ea56ade8a702ffc65bf2e87cc15c61be56a

Observation 64b8f8ad-79f2-4ea7-9c33-1fd36c74b33b · outbound

This paper cites Lan- guage model beats diffusion–tokenizer is key to visual gen- eration.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Lan- guage model beats diffusion–tokenizer is key to visual gen- eration

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.046218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.429265Z digest=sha256:7fce03def10718d8e72f6f9f161f8aed6a3f6fc5aad2651fdb77dc2eb58ffb0c

Observation bdece9a2-041f-4fa8-ab57-b22e7422889d · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.433718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.433718Z digest=sha256:e3e3e5d97f948610fc6f3bb13533cddb38f354ef5753b7c024f29679712fa489

Observation bf3a4357-bba0-42f1-b08f-455a2bc8821b · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In ACL, 2019.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Hellaswag: Can a machine really finish your sentence? In ACL, 2019

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.031418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.438286Z digest=sha256:b498ffe87246b59b37e856012ec47c8ea0b73bf5f8b64297f9926f9cc5159de8

Observation 394930f6-e2d5-40eb-9008-0dfd22aa5589 · outbound

This paper cites Sigmoid loss for language image pre-training.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Sigmoid loss for language image pre-training

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.016591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.442574Z digest=sha256:a933b75326d5ea4562b45a0f1bc6b6de7f8edd934ed6d0e06d8e1b6888cdfd58

Observation 30127f51-8631-4c74-b980-39de953a66f7 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation The unreasonable effectiveness of deep features as a perceptual metric

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.447133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.447133Z digest=sha256:31b467d0e8272acc9f89081942bee99c83d9a7834483d322a523605a0a1783e5

Observation 0d66fc1e-a3a2-4698-87f2-d800bc7c8802 · outbound

This paper cites Image and Video Tokenization with Binary Spherical Quantization.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.451468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.451468Z digest=sha256:b54e0a58bacfafad66f00aad99c0d44390833eed77a877c1f135a3b1f68bf149

Observation 54aca9c2-4c8d-4dc7-8e30-9885c2e3eafe · outbound

This paper cites Online clustered codebook.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Online clustered codebook

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:02.992318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.456061Z digest=sha256:4116c5c6a48618276ad5a67a8b3222af4f09c76ee934314829284c20d0df53e4

Observation 89097363-e6d8-42b1-b2ac-dea405150346 · outbound

This paper cites Movq: Modulating quantized vectors for high- fidelity image generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Movq: Modulating quantized vectors for high- fidelity image generation

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:02.977744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:04:02.460265Z digest=sha256:09ac8f82f2fee533b817046314acd6e3ef22d6d9e9b4b215bfecfe7498e08aa2

Pith citing papers

Observation cf687e1a-b5f5-4563-a753-5427cbecb4a8 · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.928203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.928203Z digest=sha256:6b5920c090597588598d0c4c92d72ddeeee9d6545b3d4d2c9599615870f40a9e

Observation e389c8ad-e413-4311-96c6-dedb6e571120 · inbound

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings cites this paper.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:26.204540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:26.204540Z digest=sha256:3878e2867fe0ea76a1a9145ba1ba583dabcfdffdb3de451ec8912539a0874d0e

Observation 78ce552d-ceb8-4c56-a03b-f1a590da903a · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:17:18.666961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:b55e44a5a4184054c3956c7d36cd7a1f121a003b2b52d98059b9260e24d05beb

Observation f6deb1b8-5c81-4bfc-b684-300bf9f9134d · inbound

Unified Pix Token And Word Token Generative Language Model cites this paper.

Unified Pix Token And Word Token Generative Language Model QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.591361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:30:45.355012Z digest=sha256:2a9830387757e90660e720304de93f44a1abb1dfb2058f847ae1a9b9260e6af0

Observation 231eadc4-01d1-4ae1-ba6e-263234ac0b94 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.768573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:40374e3ac29194c4c837f8718b677b3c8e92880c7f92b534cd07c069c561fd6b

Observation 0d553d7d-d8bb-4acb-bb92-f667a5f324c6 · inbound

Diffusing in the Right Space: A Systematic Study of Latent Diffusability cites this paper.

Diffusing in the Right Space: A Systematic Study of Latent Diffusability QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.631392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T10:44:24.318786Z digest=sha256:8dbc7039a094ffdf1f8454647aa8c5c1e787e8b3c48e391dcd31adbd83bb4a06

Observation b7cfaf30-7abf-4a31-8663-d096ee5faa72 · inbound

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization cites this paper.

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:27:39.702114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T13:18:47.472178Z digest=sha256:7c310544f5cbaec4f2a75017b8b6185255d3b22e2f1308e1406b207229fbf6b9

Observation 6d93c88c-7d14-4a5d-ae65-d42304fbf072 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.819010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:82af023959d69be5757c8da7529338a9fcb2ca05c93a39d8ff2377c4bf416091

Observation 3d54901a-9046-4a6d-9181-a448ae5835b3 · inbound

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification cites this paper.

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:28:55.821216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:18:03.846908Z digest=sha256:21c9d923cbdfde20cf4a77248abe226ce1786a38e2ab86f6326ecf6136142807

Observation aff2e8d0-fbac-4880-9063-09d70c87d728 · inbound

dRAE: Representation Autoencoder with Hyper-Spherical Codes cites this paper.

dRAE: Representation Autoencoder with Hyper-Spherical Codes QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T05:45:53.491983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:45:53.491983Z digest=sha256:d6006a0c8efe034fae5f3ed8c413c2ac0e649d9148075f5efc8784b1021d222d

Observation 63171561-62da-4f9d-9d63-c7eec9039483 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:50.752053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:50.752053Z digest=sha256:5e16df10dafba9dfd9c0850af5245784ce11704f038e6bf8a0b14352e47fcb0d

Observation dc275fa0-7300-4446-9dc3-cc9c0a1d4d35 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:28.219048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:28.219048Z digest=sha256:bd04b0a3fb140eb297361685bb6a30d562cc21d56a80d31d829d353c50d45e2c

Observation 59f237bf-22fe-4d5c-8e3e-c99693e2e1b9 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:56.983795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:56.983795Z digest=sha256:de22ab4bd961aa51809e0c09e4413039e05c057cb09f57bc5d37ab031a6f7c02