Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T20:04:02.460265Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 13 inbound Pith citation observations for arXiv:2502.05178.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T20:04:02.460265Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:56.983795Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:28:55.819781Z
100 of 103 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8bd6c170-fe81-48bd-a650-84f219af87a7 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a4c86db-f4e0-42f8-8f55-699fa5ff417e · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Soft-to-hard vector quantization for end-to-end learn- ing compressible representations
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70005a9b-b630-48a4-b360-ed96d1e4ba27 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Beit: Bert pre-training of image transformers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a550fff0-826a-4228-8e97-170b988d81e5 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Fuyu-8b: A multimodal architecture for ai agents, 2023
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd6f13ba-d675-417d-8964-02ab70d2b707 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d77e1e7-729b-43e5-a76e-dca36ec3a450 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation PaliGemma: A versatile 3B VLM for transfer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1754986-a731-4cee-8330-f1097faf8ba9 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Piqa: Reasoning about physical commonsense in nat- ural language
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454e073c-559f-424a-9afa-6a2da8bf9107 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Maskgit: Masked generative image transformer
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d31b2e6-2585-473e-bac5-f793d67e00a4 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf430773-b946-4491-9606-1a949a34a93e · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Training Deep Nets with Sublinear Memory Cost
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e71b303-585c-422d-aefa-e640109b3891 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d0ea132-5568-4f63-817c-4766f2b741b3 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a05f44f9-6282-4efa-863f-973f8f67adcc · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Reproducible scaling laws for contrastive language-image learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f59c9454-ea65-42bb-a951-49cd6d2377b2 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gonzalez, Ion Stoica, and Eric P
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4298893b-9776-4bd9-9e09-8931c84393dc · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Scaling instruction- finetuned language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 303ad003-cea0-494f-ba3e-2810c68f09b8 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab67c416-a7ab-4037-99ad-b0399c4fac87 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Flashattention: Fast and memory-efficient exact attention with io-awareness
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68008d26-52c3-4387-8c0d-5669bcfcc032 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Vision transformers need registers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 474d4554-538f-4321-886f-968502382f03 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Scaling vision transformers to 22 billion pa- rameters
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf644a5d-2fe0-4b4f-abb8-bf8c8b01521f · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Imagenet: A large-scale hierarchical im- age database
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aebfa514-5cbb-4c2f-99a4-f716ca22b36a · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Unveiling encoder-free vision-language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fefafe59-b9a5-4b4a-995f-44bb4b20097e · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation The Llama 3 Herd of Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e96fd0f2-3d88-4bb5-8d7f-860cde74123a · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Scal- able pre-training of large autoregressive image models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5922cbe-cd1d-4439-8f5c-3877db231ce5 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Taming transformers for high-resolution image synthesis
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2648a11-c0fd-4ed2-9ba1-92e466db8701 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Eva-02: A visual representa- tion for neon genesis
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aae7597-f829-435a-8876-3f2c0b6a72e1 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68643c2f-f220-4327-9c53-59953585761a · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Dat- acomp: In search of the next generation of multimodal datasets
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac7db086-f8f4-453a-a351-6da544f5c8b1 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c849945b-773c-451c-b9e5-9512d611fba2 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Geneval: An object-focused framework for evaluating text- to-image alignment
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d319c2-3025-4ddf-a27e-47606c75b218 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8701ceb-b64e-4d31-9c9a-1f3dea395f7e · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Making the v in vqa matter: El- evating the role of image understanding in visual question answering
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d270c97a-4c6c-4981-98bd-4a958a0aa107 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9794115-36d8-4401-896c-bc4332b5b033 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Masked autoencoders are scal- able vision learners
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69f1c8aa-6680-435a-b8b6-7313cc4c3e6c · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d757e65-fedb-44df-bad2-8b04f82e4657 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gans trained by a two time-scale update rule converge to a local nash equi- librium
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 734b2f13-c2ce-4f58-b9ff-f29389b70955 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3492adc1-c3f7-4d19-9c2f-c4d58fa7e5fd · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5f215137-e47e-4192-8f11-6eec88dd893b · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Coincidence, categorization, and consolidation: Learning to recognize sounds with minimal supervision
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 14eb0cb4-5d4f-4b17-94a7-bc8f0199e4a7 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Unified language-vision pretraining with dy- namic discrete visual tokenization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 79a1eb1c-072e-424c-90bb-900fcc3bbcb1 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Analyzing and improving the image quality of stylegan
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d6e18f26-832d-464c-9ee6-9e7bbca7cb06 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Segment anything
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d3bd40fc-aa3a-4bb7-9e9d-1789d70b079d · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Sentencepiece: A sim- ple and language independent subword tokenizer and deto- kenizer for neural text processing
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a4077dbf-4bc7-4ede-b7ec-9613a7a755ea · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation What matters when building vision-language models?
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2214d419-98f5-4334-8895-303afed07f3a · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Autoregressive image generation using residual quantization
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4a9100d2-60fc-4203-a762-1182eb0bc4dc · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation DataComp-LM: In search of the next generation of training sets for language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eb81f88-1952-4d37-8eb9-4025db60a17f · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Mage: Masked generative encoder to unify representation learning and image synthe- sis
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 20778747-fe02-45bb-9223-4d91232e0709 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Evaluating object hallucination in large vision-language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 189cdc51-2d4a-4c52-9488-4c49d257c126 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Microsoft coco: Common objects in context
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0952fad6-6b39-4b4f-9d53-f48ea7c7fc11 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Visual instruction tuning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db0fce50-4a05-4b4c-bb37-dfcabc1c592f · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Language quantized autoencoders: Towards unsupervised text-image alignment
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5595dd76-13a6-46ac-8d75-33a5649c7c66 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Improved baselines with visual instruction tuning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ab91549c-e08a-40e2-984e-525243af81ec · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Unified-io: A uni- fied model for vision, language, and multi-modal tasks
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fec7a35a-afad-492b-8035-5bd024506d7b · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Unified-io 2: Scaling autoregressive mul- timodal models with vision language audio and action
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7e77d0d2-b0b7-446e-9b91-8b20da7b51cb · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 252c49b0-6672-4642-99e3-a5b1ec6bec32 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Finite scalar quantization: Vq-vae made simple
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4a13c14d-5261-4a09-a6af-fd51658da579 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Representation Learning with Contrastive Predictive Coding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 957365a5-9e76-4c88-85eb-300f411f3848 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 579222c2-f70e-45c7-b5fe-90f53bb5ea27 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef1bc0ab-2f53-4a0c-8a65-83d6094d99fc · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Learn- ing transferable visual models from natural language super- vision
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c9d1c10e-48f9-4711-9b7f-197257a99594 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 70c991cc-d2ef-4978-a734-efc4e8326a6f · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Zero: Memory optimizations toward training trillion parameter models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf5a9212-2786-4c1b-92f0-4d52f8935b82 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Zero-shot text-to-image generation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d7a5a524-7b93-477f-a79e-5bf44a845d72 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d8bea48c-add4-462d-a853-fe13f497a8bb · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Winogrande: An adversarial winograd schema challenge at scale
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f40caa5-dcec-44ef-96dd-85f5620d0a7d · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation SocialIQA: Commonsense reasoning about social interactions
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3ac211ea-e85a-4996-bc70-9f34a5f8aedd · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation SBER-MoVQGAN or a new effective image encoder for generative models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a7e14180-f7fe-4cfc-8cdd-9b55d184716d · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Laion coco: 600m synthetic captions from laion2b-en
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6348c906-7c74-42d8-b82f-6f110828263c · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Japanese and Ko- rean voice search
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a9cb6e2f-c90b-4a91-8bfb-2fda825d10cf · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Multi-task learning as multi-objective optimization
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5b528a2d-eeca-40fc-a17b-4f30eeaa8852 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Neu- ral machine translation of rare words with subword units
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 763974f5-2b29-4718-9563-692d710ec268 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a79b9ac9-7ebd-480d-99b5-c170ebc32383 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Very deep con- volutional networks for large-scale image recognition
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3ed30314-8840-45f0-a088-3f3659833ef6 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Towards vqa models that can read
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1f0bda81-6ac2-4fb6-86f0-acb4f34fcbff · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9e9d59d-a110-4b33-a961-24993950a4dd · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c28f588-8868-4027-bf90-ff5b54a965e5 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Rethinking the inception ar- chitecture for computer vision
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a51555ff-f1f5-4c9c-84e1-ec308258bc67 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf41c74a-f4c7-4a83-92e8-5d0bf36ab6e2 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gemini: A Family of Highly Capable Multimodal Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9250a5-9f06-4cca-ac87-fe9296e98b1f · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Zerocap: Zero-shot image-to-text generation for visual- semantic arithmetic
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7d387ad0-554e-4504-b412-870f38858b29 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Visual autoregressive modeling: Scalable image generation via next-scale prediction
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5120663f-9535-485f-887a-87118f195826 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Cambrian- 1: A fully open, vision-centric exploration of multimodal llms
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4ef7273a-8b91-487c-913e-eeea0dfea38d · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54ea1290-4ad9-4b80-a076-ad1d10fac364 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Neural discrete representation learning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3c79048f-2f57-4847-90c7-e066fc8b41dc · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe84ed4d-84af-4a46-995e-7f2fdec335e1 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afac164b-4635-4703-9a45-4c183a40e4aa · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Image quality assessment: from error visibility to structural similarity
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5dbfe271-8bd9-4e92-94f5-23d6550ba9b5 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Diffusion models as masked autoencoders
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6ccbb9ee-5de6-4445-b8b7-efc5d7d6f211 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Next-gpt: Any-to-any multimodal llm
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 21678f6a-67f3-4732-80f8-d69a1c6719e0 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ca01e07-7e29-478c-84da-50fb65cac641 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11781f39-f765-4cfe-ba2a-9bcd7bb75604 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Vector-quantized image modeling with improved vqgan
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1b4ddafe-09af-4ef2-94c8-7fe345d7b8b8 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Spae: Seman- tic pyramid autoencoder for multimodal generation with frozen llms
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 64b8f8ad-79f2-4ea7-9c33-1fd36c74b33b · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Lan- guage model beats diffusion–tokenizer is key to visual gen- eration
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bdece9a2-041f-4fa8-ab57-b22e7422889d · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3a4357-bba0-42f1-b08f-455a2bc8821b · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Hellaswag: Can a machine really finish your sentence? In ACL, 2019
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 394930f6-e2d5-40eb-9008-0dfd22aa5589 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Sigmoid loss for language image pre-training
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 30127f51-8631-4c74-b980-39de953a66f7 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation The unreasonable effectiveness of deep features as a perceptual metric
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d66fc1e-a3a2-4698-87f2-d800bc7c8802 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Image and Video Tokenization with Binary Spherical Quantization
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54aca9c2-4c8d-4dc7-8e30-9885c2e3eafe · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Online clustered codebook
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 89097363-e6d8-42b1-b2ac-dea405150346 · outbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Movq: Modulating quantized vectors for high- fidelity image generation
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cf687e1a-b5f5-4563-a753-5427cbecb4a8 · inbound
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e389c8ad-e413-4311-96c6-dedb6e571120 · inbound
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78ce552d-ceb8-4c56-a03b-f1a590da903a · inbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f6deb1b8-5c81-4bfc-b684-300bf9f9134d · inbound
Unified Pix Token And Word Token Generative Language Model QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 231eadc4-01d1-4ae1-ba6e-263234ac0b94 · inbound
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0d553d7d-d8bb-4acb-bb92-f667a5f324c6 · inbound
Diffusing in the Right Space: A Systematic Study of Latent Diffusability QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b7cfaf30-7abf-4a31-8663-d096ee5faa72 · inbound
NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6d93c88c-7d14-4a5d-ae65-d42304fbf072 · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3d54901a-9046-4a6d-9181-a448ae5835b3 · inbound
Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation aff2e8d0-fbac-4880-9063-09d70c87d728 · inbound
dRAE: Representation Autoencoder with Hyper-Spherical Codes QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63171561-62da-4f9d-9d63-c7eec9039483 · inbound
Twins: Learn to Predict Unified Representations with Focal Loss QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc275fa0-7300-4446-9dc3-cc9c0a1d4d35 · inbound
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 157
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59f237bf-22fe-4d5c-8e3e-c99693e2e1b9 · inbound
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Reference 157
Source-reported events for the cited work
Unavailable: canonical work link unavailable.