Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:46:10.909729Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 20 inbound Pith citation observations for arXiv:2506.18898.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:46:10.909729Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:10.728331Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:09:44.750475Z
82 of 82 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2a8b05d9-0d0d-41bb-9ff3-f48526ce05d5 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02165106-c791-407a-b12a-b70780927dae · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c8d748f-0b3c-4ae8-bf3e-0f73b54fe326 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Language models are few-shot learners
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb8d5979-65cd-4d60-95b7-aec8796723a0 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Emerging properties in self-supervised vision transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71185742-bf08-4c1f-a7a9-c8471820dc50 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Maskgit: Masked generative image transformer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be16321-95f5-49d7-adf4-1ec6a9cf714a · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations GenQA: Generating Millions of Instructions from a Handful of Prompts
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f1e4ed2-298c-4259-af56-2fc0ba7ce8e0 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 264035c2-5899-44a2-8b50-c298623e026b · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35677777-b3ea-408b-8cda-0c553aad3ac4 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83dcc0b7-572b-450d-ad69-e218367cb587 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 586b8fb9-59c0-4512-a81c-c5ffd362f748 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Instructblip: Towards general-purpose vision-language models with instruction tuning.Advances in Neural Information Processing Systems, 36, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c084934-6642-4f48-b59c-de8ac2b5b868 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Imagenet: A large-scale hierarchical image database
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52380c48-aa05-434b-bf48-3e5538d1be05 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations DreamLLM: Synergistic Multimodal Comprehension and Creation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191ebea1-7a28-4030-8e60-8a92a5e099e1 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Megalith-10m dataset
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8705d8b0-659a-4b07-8915-93ccf4520060 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Scaling rectified flow transformers for high-resolution image synthesis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f457bce-ce7e-4ee5-8457-347a3cba3215 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Taming transformers for high-resolution image synthesis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aea6666-678d-4dcf-bd4f-d7c7a8949f0c · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4874f2c0-9cf4-41ce-a4e5-11c8d530fa6b · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87f0b4e0-468c-4b1e-8e6f-67262c9afe7f · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Making LLaMA SEE and Draw with SEED Tokenizer
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d983f45c-e6ed-4c69-8873-9923b81d15bf · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e261fb-ce51-4f62-9ed8-b9e2d8eee0dd · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132–52152, 2023
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 262c4abc-65ea-4d13-b945-b4a70d471728 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Onellm: One framework to align all modalities with language
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 058f2010-1150-409f-a561-cfc6a77e9b8b · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Ella: Equip diffusion models with llm for enhanced semantic alignment, 2024
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35e73870-ee4e-44c6-893e-e322254b277a · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eff4a3a0-2253-4095-bb86-979f4d9b1d63 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Auto-encoding variational bayes, 2013
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e185875a-2a49-489e-9dbe-2df7ba051078 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Flux.https://github.com/black-forest-labs/flux, 2024
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89806cc4-cb36-44aa-a95f-949234a05ed9 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 034077b3-d060-4fc2-b60a-5b31755ef0b5 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04102618-35a2-4196-a98a-36007ae4b805 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32e8331b-4d92-4f87-b964-32d1cc340a8d · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations What If We Recaption Billions of Web Images with LLaMA-3?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5127e357-63a3-4562-a4f0-820cdfbf8c02 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0a97ace-21d2-4c51-8669-107465655fd6 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Evaluating Object Hallucination in Large Vision-Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca526121-fa56-4e51-8f05-80d22f967955 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e48c617d-5eca-4dcb-a79f-6b98bf31e043 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Dual Diffusion for Unified Image Generation and Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 949556e6-02a4-4437-8bf9-e4597303a660 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dab2b7f-5826-4ba5-b64d-cb722cf0bc39 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51854d55-f876-4fb6-9fe7-8888a965879d · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Improved baselines with visual instruction tuning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 111f8ab0-005d-4d09-a49c-8a1b442bb9a7 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c185b33-e862-4188-a2b8-9f3421f20bb6 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Visual instruction tuning.Advances in neural information processing systems, 36, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation beb8cb3a-de23-49cd-a9d5-665bf4da196a · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations MMBench: Is Your Multi-modal Model an All-around Player?
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb0e38c1-b263-48e0-b1bd-370b48c9022f · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b7fb901-5b98-4279-a5b6-6d26589a0e41 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d6a1d00-a672-4445-9d96-00aea3c77fb2 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 447f5730-6a88-4060-9784-bbe3b73f246b · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Infinity-instruct
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 850d2f04-1961-4efe-8ec3-28cb99b61174 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations GPT-4 Technical Report
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eccfe5c2-712d-4d3b-8492-6d65f8e19728 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db15cea2-5312-4c43-b77a-d8d651e13b71 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Lumina-image 2.0: A unified and efficient image generative framework, 2025
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe7283d-90d2-46e4-bac0-275ef5b740c1 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Learning transferable visual models from natural language supervision
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7494cf0d-ee65-4710-ad51-3942f9862232 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations ImageNet-21K Pretraining for the Masses
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 478593cf-04bf-4d5e-b884-8d251ac2791b · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations High-resolution image synthesis with latent diffusion models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 850c2977-1277-4897-9d68-10022a75ca4f · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df4be647-4ed0-4ea7-befc-28f8ab83e272 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51c25fbd-1190-46cf-a5c3-ef52bee963ab · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Journeydb: A benchmark for generative image understanding.Advances in neural information processing systems, 36:49659–49678, 2023
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f01d1d-bf66-45ab-a3ef-f3bf08263331 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 609880a2-ff8b-4a09-89d4-2fb5a508ede6 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Generative Multimodal Models are In-Context Learners
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4096da8-359d-4d52-b450-5364c277bddf · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Chameleon: Mixed-modal early-fusion foundation models, 2024
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64b9b31e-343c-4c98-95b6-d0ab1e9ef80d · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Openhermes
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e6cf47de-4fd9-45a6-86ee-e79075a5c43d · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e4717bc0-23c2-40d2-a2b5-1b3d7f39b768 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b770942-13e8-46bb-8a85-ff2cc28e737b · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f0c2827-b3d0-473f-a63b-cd0b13d54d5d · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations LLaMA: Open and Efficient Foundation Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c0c010-b114-4517-880b-2a7dc6e5015b · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fc56a80-b82a-425f-ace0-cb0e0ee69953 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Neural discrete representation learning.Advances in neural information processing systems, 30, 2017
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 054dcf71-15e0-4c1a-ba51-c45bbe14a5bf · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Midjourney prompts dataset
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a9151732-a7c7-44e7-977e-c9cdc38729a6 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbb3b56-c0b8-4258-bde4-a770d0ebed82 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7e61fb5-ecb8-4963-b57c-f6efd4d29193 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Emu3: Next-Token Prediction is All You Need
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb1d9097-5cd8-4c12-bf78-a638a394fb32 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73cd0b52-2d57-4ce9-a299-617838d48ec8 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Liquid: Language Models are Scalable and Unified Multi-modal Generators
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55b97299-d22c-420c-80a1-368e288590f4 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Harmonizing visual representations for unified multimodal understanding and genera- tion, 2025
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1c258c4c-97cf-4c25-b18b-d3f84113457f · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8a5eba-3b99-497d-9515-b583a5ba4319 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Sana: Efficient high-resolution image synthesis with linear diffusion transformer, 2024
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a4af2f2-1b3d-4a9e-8d63-a1cfd1055744 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Sana 1.5: Efficient scaling of training-time and inference-time compute in linear diffusion transformer, 2025
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f33a1d38-0725-4056-a86f-58132c7e20e1 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4c6467-4811-4456-874d-2ce06c622404 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d3a35a9-6b92-4d64-9843-22e1a2998e46 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Qwen2.5 Technical Report
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83e78517-1552-4e32-b0f2-0d627876aaa2 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10abd553-7c79-4bc2-9d7b-0d0318d22831 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Mammoth2: Scaling instructions from the web
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6b36831-97a9-4d23-97d7-6613773bbfb6 · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Sigmoid loss for language image pre-training
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2e259ad-b68b-43cb-b963-6eee0b3bba2c · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3a98e70-0a54-4373-87a0-d90f8e6ce32b · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Qlip: Text-aligned visual tokenization unifies auto-regressive multimodal understanding and generation.arXiv preprint arXiv:2502.yyyyy, 2025
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 78a3921d-5d42-4206-bd62-cfd68a069f4b · outbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcc49f89-bb9c-4d1e-8fde-d8d2797fe141 · inbound
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3a0dc13-a72d-4be2-8da8-a5714a44191b · inbound
TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8379ad9-609e-4ff2-9c11-4af3053796be · inbound
Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16d3aa9f-6f7b-4885-a065-99e9b5a2a79f · inbound
InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d6f549d6-59d9-4b18-bb3e-de99e1335bb5 · inbound
Generative Refinement Networks for Visual Synthesis Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bcb7a2aa-fdb1-4bb5-a13c-454254f918d4 · inbound
Generative Refinement Networks for Visual Synthesis Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f09bd54-8d22-4e48-9d5f-b6c7e89f3d47 · inbound
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f9d5a11f-5ae7-40a6-b4ca-ea818e4428df · inbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d48b7176-f7d3-438e-aec4-f7d2d4da5b79 · inbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 15009bbc-4c51-493f-ad6c-cdb223faa208 · inbound
Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8c3754db-dd38-441f-b252-095f165ad71f · inbound
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d5a69bff-ca17-448e-9ae3-f6bc8539c31e · inbound
MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3687ca55-3ff9-4e3f-a0bc-71af85e59c14 · inbound
Residual Decoder Adapter: ID-Preserving Tokenizer Adaption for Autoregressive Text Rendering Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 941b87cd-4ce3-4a21-ae07-4551c6228197 · inbound
ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e09ca6db-7932-424b-9fbb-11e66550a125 · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 208
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3af3830b-45cc-44c3-9d49-aa06204b0ded · inbound
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 68c228fb-b342-4ecd-a447-7130a6b32ea9 · inbound
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6c72ffd8-1e11-45b2-ab65-c76324a081ef · inbound
Bridging Video Understanding and Generation in a Unified Framework Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 06e61899-6925-4113-a9e1-dfd30d2ed4e2 · inbound
dRAE: Representation Autoencoder with Hyper-Spherical Codes Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5376db81-4679-4749-aa0f-5ca47ef27479 · inbound
Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.