Pith. sign in

Paper Citation Record · LEDGER

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

As of 18 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 20 inbound Pith citation observations for arXiv:2506.18898.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18898 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:46:10.909729Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:10.728331Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:44.750475Z

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved72
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2a8b05d9-0d0d-41bb-9ff3-f48526ce05d5 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.581583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.581583Z digest=sha256:c49933b4a7f49059e73237ef41824edaf1f6fa8b7a1e3c66c2095fd89041ab1e

Observation 02165106-c791-407a-b12a-b70780927dae · outbound

This paper cites Qwen2.5-VL Technical Report.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.585974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.585974Z digest=sha256:93f22a51f260bcbd66135765b5f273274d9f75e7cfd72ebfd4d82e7c24209c00

Observation 6c8d748f-0b3c-4ae8-bf3e-0f73b54fe326 · outbound

This paper cites Language models are few-shot learners.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Language models are few-shot learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.590287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.590287Z digest=sha256:e863dd3c5c14edc29ac2466b3643cbe6fe7f5be142613bb9ee51e3d4b1c415cb

Observation bb8d5979-65cd-4d60-95b7-aec8796723a0 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Emerging properties in self-supervised vision transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.595050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.595050Z digest=sha256:db7b6956c8bffd201ce35b077223886c6792cd85f3e54be300622727851481e8

Observation 71185742-bf08-4c1f-a7a9-c8471820dc50 · outbound

This paper cites Maskgit: Masked generative image transformer.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Maskgit: Masked generative image transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.599529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.599529Z digest=sha256:454c66609cabcbe88ed8c0447b23b7b563356a402f9f6f494adc54ecf7e13570

Observation 7be16321-95f5-49d7-adf4-1ec6a9cf714a · outbound

This paper cites GenQA: Generating Millions of Instructions from a Handful of Prompts.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations GenQA: Generating Millions of Instructions from a Handful of Prompts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.604023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.604023Z digest=sha256:68f2bda9161437cf4821da49a143525d3171fab4750e99c69b47a65bfe34db21

Observation 5f1e4ed2-298c-4259-af56-2fc0ba7ce8e0 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.608151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.608151Z digest=sha256:5760a7f375e826b82a53a99e0d4aeb88be2583a6fabf1b685637d04993ab53a7

Observation 264035c2-5899-44a2-8b50-c298623e026b · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.612181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.612181Z digest=sha256:fb56067bd5484f8097f0c34db45d8feafe21af0aa30eea362cb7f03a23957871

Observation 35677777-b3ea-408b-8cda-0c553aad3ac4 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.616383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.616383Z digest=sha256:aa70cc0ffdacd4a0725af07f486e2e5254c10c8040a5fccfa37c10bf07c6667c

Observation 83dcc0b7-572b-450d-ad69-e218367cb587 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.620504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.620504Z digest=sha256:bdb3f2dfff482f8f3201a962ef248771856e44e3102da3ea5cc15a895cba9fa9

Observation 586b8fb9-59c0-4512-a81c-c5ffd362f748 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.Advances in Neural Information Processing Systems, 36, 2024.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Instructblip: Towards general-purpose vision-language models with instruction tuning.Advances in Neural Information Processing Systems, 36, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.624611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.624611Z digest=sha256:29a89fe35c4328be15bb0edd338c267316c4703ed4bc1109229376948364e141

Observation 6c084934-6642-4f48-b59c-de8ac2b5b868 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.627952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.627952Z digest=sha256:4145ecaea1c07bca6f52bda760ef6b639ab635e68b0cfb2a8b00a90ff3e52704

Observation 52380c48-aa05-434b-bf48-3e5538d1be05 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.631263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.631263Z digest=sha256:bb838c6657efa5bf4d17f275ea2eb84fd30969002382e5f1c9acb53b907c3b72

Observation 191ebea1-7a28-4030-8e60-8a92a5e099e1 · outbound

This paper cites Megalith-10m dataset.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Megalith-10m dataset

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:11.924067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:46:10.635382Z digest=sha256:7177ea1c21f0f9cd6d6e42df466a7c3ae0990cc7d8013261d07e81c650a0f080

Observation 8705d8b0-659a-4b07-8915-93ccf4520060 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Scaling rectified flow transformers for high-resolution image synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.639752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.639752Z digest=sha256:c54527ac2cdbf93eba851ece66e4e3f11a5f13d0840aeb48a2018d4321b9ec06

Observation 7f457bce-ce7e-4ee5-8457-347a3cba3215 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Taming transformers for high-resolution image synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.643619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.643619Z digest=sha256:a922a998529ef5aca2d8310937d996dc4119d4c385a8bba660d425f9fda5d124

Observation 0aea6666-678d-4dcf-bd4f-d7c7a8949f0c · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.647842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.647842Z digest=sha256:2d99cf32c9448fbc764e9dea56892adf6954f87565f685251b7fbd3528f41f4c

Observation 4874f2c0-9cf4-41ce-a4e5-11c8d530fa6b · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.651887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.651887Z digest=sha256:b50136972839d0febcae85e1cce48853c7a096c92bb5170885236993af4cb428

Observation 87f0b4e0-468c-4b1e-8e6f-67262c9afe7f · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Making LLaMA SEE and Draw with SEED Tokenizer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.655820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.655820Z digest=sha256:528a5f7cd0232edd2d2e9e5121416a34ddf85b89cc8b33d54a35b9a84509becb

Observation d983f45c-e6ed-4c69-8873-9923b81d15bf · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.660038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.660038Z digest=sha256:840225f40ba8b9d649b9b5d379be0c17be75ad5743bf7f7acd6fbb203319ab57

Observation 63e261fb-ce51-4f62-9ed8-b9e2d8eee0dd · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132–52152, 2023.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132–52152, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.664116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.664116Z digest=sha256:dd57e1588eae3b63a5cfc026cc2dde3ee2ce5f4ca0757bf9365989bd47f7440b

Observation 262c4abc-65ea-4d13-b945-b4a70d471728 · outbound

This paper cites Onellm: One framework to align all modalities with language.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Onellm: One framework to align all modalities with language

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:11.877157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:46:10.667851Z digest=sha256:ba055a35d36e7a7965f137dde2e100b45c808769d8b4cea28f37b2151dd0e264

Observation 058f2010-1150-409f-a561-cfc6a77e9b8b · outbound

This paper cites Ella: Equip diffusion models with llm for enhanced semantic alignment, 2024.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Ella: Equip diffusion models with llm for enhanced semantic alignment, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.671870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.671870Z digest=sha256:0ee7bef2f6ded60b5847c0aae8dcc2b0ece8b5ad83dbc389d2ee34f5ee432d77

Observation 35e73870-ee4e-44c6-893e-e322254b277a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.675754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.675754Z digest=sha256:9ab3eeb82b86ee07d72f4e36df7c049c971665b214ede0098f22bac055371cc3

Observation eff4a3a0-2253-4095-bb86-979f4d9b1d63 · outbound

This paper cites Auto-encoding variational bayes, 2013.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Auto-encoding variational bayes, 2013

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.679995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.679995Z digest=sha256:951bf7fe86387fa731344649cd87de849f3f57d023193183eee7c9fcbfc797f5

Observation e185875a-2a49-489e-9dbe-2df7ba051078 · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2024.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Flux.https://github.com/black-forest-labs/flux, 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.683749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.683749Z digest=sha256:572157b487f8d0484f70864a6b093256b6d364ef72e39db1353e4684adf40664

Observation 89806cc4-cb36-44aa-a95f-949234a05ed9 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.687747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.687747Z digest=sha256:7d23ed6b100a47c2fbe341daf31ac6656bb4d261e448eb95de1f267a61261e36

Observation 034077b3-d060-4fc2-b60a-5b31755ef0b5 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.691852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.691852Z digest=sha256:ab8dc6187e53f6e7e9389e2b3032662923e4374477f397cc79a8f81be752675f

Observation 04102618-35a2-4196-a98a-36007ae4b805 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.696268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.696268Z digest=sha256:7553492497756dfa21fcf81ab4a99e3f193289479d64003f02f8bd475ef885ec

Observation 32e8331b-4d92-4f87-b964-32d1cc340a8d · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations What If We Recaption Billions of Web Images with LLaMA-3?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.699688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.699688Z digest=sha256:8a3831ef7bdfda4d5d831a383a84c4016d250366a32c1ed007783851c288d2bb

Observation 5127e357-63a3-4562-a4f0-820cdfbf8c02 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.703247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.703247Z digest=sha256:3e4f4f9ee8f009530cf47b4dd32f08a66770576d46db56c85191b0046df4575b

Observation e0a97ace-21d2-4c51-8669-107465655fd6 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Evaluating Object Hallucination in Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.706804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.706804Z digest=sha256:c6090c548621c6cf0d3a107628b9b81b70cad0048a522e88b9ea0fc87830b8fc

Observation ca526121-fa56-4e51-8f05-80d22f967955 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.710650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.710650Z digest=sha256:4ac5a2fdeb1b7608d2b00d5e5a73b8f7c208adf6ba7f39e45a64dd3004ff1b08

Observation e48c617d-5eca-4dcb-a79f-6b98bf31e043 · outbound

This paper cites Dual Diffusion for Unified Image Generation and Understanding.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Dual Diffusion for Unified Image Generation and Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.714522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.714522Z digest=sha256:228f2d6fb92876e4c055f2e652d8fc68d65b19ceba6ceaec507732812eb9a563

Observation 949556e6-02a4-4437-8bf9-e4597303a660 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.718598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.718598Z digest=sha256:dde28a95e9d42a6b3ef086f1bc43faf6b51286406108e639743767b0ec241cdb

Observation 1dab2b7f-5826-4ba5-b64d-cb722cf0bc39 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.722651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.722651Z digest=sha256:a621173d1d61414e66455254a344b82951ecbf9595dbae8a7cbf735ad2d6ee3f

Observation 51854d55-f876-4fb6-9fe7-8888a965879d · outbound

This paper cites Improved baselines with visual instruction tuning.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Improved baselines with visual instruction tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.726905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.726905Z digest=sha256:f9478ff5993142f77d234459c941654dd5b09f4a3e93d184b2ec004c4edf4ae9

Observation 111f8ab0-005d-4d09-a49c-8a1b442bb9a7 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.730828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.730828Z digest=sha256:d9a575803d0a507e77e7ca3dee049097fab4954c823afdbf795ea4484b3c4df7

Observation 3c185b33-e862-4188-a2b8-9f3421f20bb6 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:11.803315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:46:10.734817Z digest=sha256:1c492ec8964c9742d7b72570c89c4f44f01679cb6dd8b50e160107e607857f29

Observation beb8cb3a-de23-49cd-a9d5-665bf4da196a · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations MMBench: Is Your Multi-modal Model an All-around Player?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.738694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.738694Z digest=sha256:4d10039f647bc2a52c3ed7f23e181a5ad3ce0298f2423bfc6046f20cff8e9e9e

Observation cb0e38c1-b263-48e0-b1bd-370b48c9022f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.743116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.743116Z digest=sha256:57d452a39d787cccd3c8df01ef8adf0a9b245ff0f68ab391f87818b922aec7ca

Observation 0b7fb901-5b98-4279-a5b6-6d26589a0e41 · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.747456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.747456Z digest=sha256:6be1e7c5a677dc6c4633f4896da2c32e54042cf36e1120f58de54694020d8799

Observation 2d6a1d00-a672-4445-9d96-00aea3c77fb2 · outbound

This paper cites Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.751576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.751576Z digest=sha256:adde2f8373800ba7aa55632bfebd039125e848a16e432ca3d88f2e440d4de19d

Observation 447f5730-6a88-4060-9784-bbe3b73f246b · outbound

This paper cites Infinity-instruct.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Infinity-instruct

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:11.789599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:46:10.755806Z digest=sha256:60da414ce7093a35dcba2d4c22313cb5c3838318868c7bdf0f016ab56bbf3558

Observation 850d2f04-1961-4efe-8ec3-28cb99b61174 · outbound

This paper cites GPT-4 Technical Report.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations GPT-4 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.759795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.759795Z digest=sha256:7d5e846312de48fe83a24cb1be8980aeef8e4f2f0c54ba9b9e64931e20eb0a76

Observation eccfe5c2-712d-4d3b-8492-6d65f8e19728 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.763653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.763653Z digest=sha256:80d55aa2a07d1090f075f030fe4425cf98083834bae9109fb15ebeb74efc3de2

Observation db15cea2-5312-4c43-b77a-d8d651e13b71 · outbound

This paper cites Lumina-image 2.0: A unified and efficient image generative framework, 2025.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Lumina-image 2.0: A unified and efficient image generative framework, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.767577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.767577Z digest=sha256:958828051b974eb0bad5404c2dac1a7f3e9649cfdbfb7ed773485a69b0a6be6a

Observation fbe7283d-90d2-46e4-bac0-275ef5b740c1 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Learning transferable visual models from natural language supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.771378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.771378Z digest=sha256:3ee3dde23534205f9ecfdb72cbc67633a6b9379deccff504793237200b58132d

Observation 7494cf0d-ee65-4710-ad51-3942f9862232 · outbound

This paper cites ImageNet-21K Pretraining for the Masses.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations ImageNet-21K Pretraining for the Masses

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.775601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.775601Z digest=sha256:a6d3d321ef13c34d4e088649ba8c8a444e9c13501aa1d190374dcf876fc35878

Observation 478593cf-04bf-4d5e-b884-8d251ac2791b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations High-resolution image synthesis with latent diffusion models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.779777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.779777Z digest=sha256:530b0570c4904b36377a2efb8d4639c9b75f018f29f72b5bd51586707575792b

Observation 850c2977-1277-4897-9d68-10022a75ca4f · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.783929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.783929Z digest=sha256:f92dc323d39994d0b7cea81226d4d52941d319012c49f9119ef9b8f6df002807

Observation df4be647-4ed0-4ea7-befc-28f8ab83e272 · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.788057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.788057Z digest=sha256:9dc6bf1c28c1b6763e7f19afab74daca95e29e01abdf1db17f2b4595f135ba52

Observation 51c25fbd-1190-46cf-a5c3-ef52bee963ab · outbound

This paper cites Journeydb: A benchmark for generative image understanding.Advances in neural information processing systems, 36:49659–49678, 2023.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Journeydb: A benchmark for generative image understanding.Advances in neural information processing systems, 36:49659–49678, 2023

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.792301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.792301Z digest=sha256:148d6e9abde8eb2bf3b79b4c78db60ea2e0c00690abc169394cf6a0fb24d2f84

Observation f7f01d1d-bf66-45ab-a3ef-f3bf08263331 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.797021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.797021Z digest=sha256:e4594860e53807972f138a9930fe6968eb65d3b5e5d08827f07e603d2a0ea490

Observation 609880a2-ff8b-4a09-89d4-2fb5a508ede6 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Generative Multimodal Models are In-Context Learners

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.801367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.801367Z digest=sha256:fc73ec2fba9310903347d9d52d444708f9954dfcbfac0525a26a02df147ead29

Observation f4096da8-359d-4d52-b450-5364c277bddf · outbound

This paper cites Chameleon: Mixed-modal early-fusion foundation models, 2024.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Chameleon: Mixed-modal early-fusion foundation models, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.805167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.805167Z digest=sha256:a9b274ad583e9427146d6238b5a4c2ab55aa01b093e0cca7926fee9546aa3527

Observation 64b9b31e-343c-4c98-95b6-d0ab1e9ef80d · outbound

This paper cites Openhermes.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Openhermes

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:11.725866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:46:10.808900Z digest=sha256:e6a8fd02cde6fa6858f5b026a90dad476186781dccbfb572c9ec8e5262b270b6

Observation e6cf47de-4fd9-45a6-86ee-e79075a5c43d · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:11.712568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:46:10.812302Z digest=sha256:1e987bdaec14e8d3117f2873077d5e9db54f48213ed3906f7e936fcafe78ff5f

Observation e4717bc0-23c2-40d2-a2b5-1b3d7f39b768 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.815851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.815851Z digest=sha256:e08cb7c1bc2ae2b7f26a0e3272a5bf87eaa15d68e92b7ea99c0ace33db5988d3

Observation 3b770942-13e8-46bb-8a85-ff2cc28e737b · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.819298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.819298Z digest=sha256:3f242ef198729e290ad0397c7db5e046421a789c02c4d003117fd6c73c46fb84

Observation 8f0c2827-b3d0-473f-a63b-cd0b13d54d5d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations LLaMA: Open and Efficient Foundation Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.823831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.823831Z digest=sha256:bfe9fcac668d3986c4dfc15dd2d67efe710e9c078c1225837dd018d36182e0e5

Observation f7c0c010-b114-4517-880b-2a7dc6e5015b · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.828120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.828120Z digest=sha256:12ae695af6c72ee27eacfea5f62fd90d5e931924d9b8c2d026e30be7280946a4

Observation 6fc56a80-b82a-425f-ace0-cb0e0ee69953 · outbound

This paper cites Neural discrete representation learning.Advances in neural information processing systems, 30, 2017.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Neural discrete representation learning.Advances in neural information processing systems, 30, 2017

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.832580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.832580Z digest=sha256:13da079686f2f3a470412bf270058bebebcc8848d865ecf904e5fd005914dfb5

Observation 054dcf71-15e0-4c1a-ba51-c45bbe14a5bf · outbound

This paper cites Midjourney prompts dataset.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Midjourney prompts dataset

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:11.681332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:46:10.836554Z digest=sha256:9ab2464fc20bccb9b687cc3804b0d8de706de1a6831b3811ea7efe45ec67aaa5

Observation a9151732-a7c7-44e7-977e-c9cdc38729a6 · outbound

This paper cites ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.840578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.840578Z digest=sha256:6bdda95baae36e24f293bf93bb2aab4e62cd6fbb976a2d09a54f9c79741a7966

Observation 4bbb3b56-c0b8-4258-bde4-a770d0ebed82 · outbound

This paper cites LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.844791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.844791Z digest=sha256:6058798fd09abe92bbfc8f428229fe12f70177961357b08bab8ddc245ff0f912

Observation c7e61fb5-ecb8-4963-b57c-f6efd4d29193 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Emu3: Next-Token Prediction is All You Need

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.848988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.848988Z digest=sha256:1696a125d25fafe4bd81435ce080c7e2f564c1029d3d8d33858c6feb5518456d

Observation cb1d9097-5cd8-4c12-bf78-a638a394fb32 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.853249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.853249Z digest=sha256:aa3b29c42e47daf73cfe9028e1d7e67df494e0368e0d85e61d3c5c88b1f0d1ac

Observation 73cd0b52-2d57-4ce9-a299-617838d48ec8 · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.857784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.857784Z digest=sha256:d5012ccece3d526756becf536546b74ec29dd2b077e2b8f08578bbe893936f7f

Observation 55b97299-d22c-420c-80a1-368e288590f4 · outbound

This paper cites Harmonizing visual representations for unified multimodal understanding and genera- tion, 2025.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Harmonizing visual representations for unified multimodal understanding and genera- tion, 2025

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:11.666763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:46:10.862139Z digest=sha256:1f1a8dd42024836da12df98459c0454c7009156ab38a1c76e5650b068b58408e

Observation 1c258c4c-97cf-4c25-b18b-d3f84113457f · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.866078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.866078Z digest=sha256:41c79565244580b891a7c06df182bacad8bb4b29d1961b6bfd58846bb1b340c7

Observation 8d8a5eba-3b99-497d-9515-b583a5ba4319 · outbound

This paper cites Sana: Efficient high-resolution image synthesis with linear diffusion transformer, 2024.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Sana: Efficient high-resolution image synthesis with linear diffusion transformer, 2024

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.870227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.870227Z digest=sha256:87a072f7d1aeffc0ff24197095efbb82057a8d8737e8d41e2403ff1545777095

Observation 8a4af2f2-1b3d-4a9e-8d63-a1cfd1055744 · outbound

This paper cites Sana 1.5: Efficient scaling of training-time and inference-time compute in linear diffusion transformer, 2025.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Sana 1.5: Efficient scaling of training-time and inference-time compute in linear diffusion transformer, 2025

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:11.643261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:46:10.874364Z digest=sha256:1038ee898392049cffd9a41c6b490c4c6840a9ba5157e856479aaf61a044ab90

Observation f33a1d38-0725-4056-a86f-58132c7e20e1 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.878499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.878499Z digest=sha256:6e4e03ea72928c339fb8a0072920e8d7ab8985d960f799b4f6cc0c2075259c1d

Observation aa4c6467-4811-4456-874d-2ce06c622404 · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.882716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.882716Z digest=sha256:9262d8c6eb4323503a85b2d1ab1e50cb884029d0c482498e6fd43c1a35514633

Observation 5d3a35a9-6b92-4d64-9843-22e1a2998e46 · outbound

This paper cites Qwen2.5 Technical Report.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Qwen2.5 Technical Report

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.886901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.886901Z digest=sha256:f09faaa43f48553c13059468de17f1659ce5ad6f05cf1aa762adb74cfb75c4fe

Observation 83e78517-1552-4e32-b0f2-0d627876aaa2 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.890846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.890846Z digest=sha256:3ad7188c33cb127dfae1a38265aeec5c8c2640e3a0fc9132d1e80e1adcfdfd75

Observation 10abd553-7c79-4bc2-9d7b-0d0318d22831 · outbound

This paper cites Mammoth2: Scaling instructions from the web.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Mammoth2: Scaling instructions from the web

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.894859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.894859Z digest=sha256:796eedcff4366febda4ad2e03d57df453177783a7d1fa00a57bbcb7cd066cada

Observation c6b36831-97a9-4d23-97d7-6613773bbfb6 · outbound

This paper cites Sigmoid loss for language image pre-training.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Sigmoid loss for language image pre-training

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.898956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.898956Z digest=sha256:84a2f4ce60cf9a275ce479cee8263dfa1dafe6347f9900cd6d591e20ad050cf4

Observation e2e259ad-b68b-43cb-b963-6eee0b3bba2c · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.902395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.902395Z digest=sha256:b1365f4b7d19e36822f6cf914c29b6b8f93a91b22f88112076d3fa81606da018

Observation f3a98e70-0a54-4373-87a0-d90f8e6ce32b · outbound

This paper cites Qlip: Text-aligned visual tokenization unifies auto-regressive multimodal understanding and generation.arXiv preprint arXiv:2502.yyyyy, 2025.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Qlip: Text-aligned visual tokenization unifies auto-regressive multimodal understanding and generation.arXiv preprint arXiv:2502.yyyyy, 2025

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:11.602537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:46:10.906310Z digest=sha256:b5037b1e668faf9850275318dd802cfadc61305ca30382cb4b7716e92200c143

Observation 78a3921d-5d42-4206-bd62-cfd68a069f4b · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.909729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.909729Z digest=sha256:a68eb82249674f4703f24079acfdbbdf0e56307c61a773fa2034f699c7c4adb1

Pith citing papers

Observation bcc49f89-bb9c-4d1e-8fde-d8d2797fe141 · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:10.728331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:10.728331Z digest=sha256:69a31d229eac9c1c9d5d8c35f61a9a1d98b0e53ddb7bf2eca8edb48e53bf6f2b

Observation c3a0dc13-a72d-4be2-8da8-a5714a44191b · inbound

TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning cites this paper.

TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:41:11.990825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:41:11.990825Z digest=sha256:b617c93fe600e92c3574e22d4a3bdba89ed7f2d566e226eba13f984a3c265328

Observation d8379ad9-609e-4ff2-9c11-4af3053796be · inbound

Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration cites this paper.

Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T07:38:46.910012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:38:46.910012Z digest=sha256:f60e3412a527b33f35a7cf99c4e10c0712d4be8dc8f170040ff6d1585b055936

Observation 16d3aa9f-6f7b-4885-a065-99e9b5a2a79f · inbound

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs cites this paper.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.419989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:204e49f8d14f462082646f0d0a7a34eb613e74d9c7d847101a5ec17d8518801c

Observation d6f549d6-59d9-4b18-bb3e-de99e1335bb5 · inbound

Generative Refinement Networks for Visual Synthesis cites this paper.

Generative Refinement Networks for Visual Synthesis Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:01:03.247443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:41:23.057949Z digest=sha256:49e9edfc52c25ce17c75cbd9c9532931488bcf4cabd4ec6ef8149af42407658e

Observation bcb7a2aa-fdb1-4bb5-a13c-454254f918d4 · inbound

Generative Refinement Networks for Visual Synthesis cites this paper.

Generative Refinement Networks for Visual Synthesis Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T21:00:35.123495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:00:35.123495Z digest=sha256:a63683a37c5d0e8c20ec262241082d0c4c47744a3f8cae45202c09b5842b93d7

Observation 3f09bd54-8d22-4e48-9d5f-b6c7e89f3d47 · inbound

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation cites this paper.

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.564612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T05:15:22.907880Z digest=sha256:e86973221ac6cf28f99f9efd1dc475d7bd8c66426e16f60383af5655488538b0

Observation f9d5a11f-5ae7-40a6-b4ca-ea818e4428df · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:18.793166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:dc2d63e3811b34d8d4220bc6c19f3eda6fe7f48f2b7df3206f1c60a419517999

Observation d48b7176-f7d3-438e-aec4-f7d2d4da5b79 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.091129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:da20451433b0cc43a86526bb3a670ce9e6a821f3cdcdc5fadabe93b6b69c3459

Observation 15009bbc-4c51-493f-ad6c-cdb223faa208 · inbound

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models cites this paper.

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:18.931338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T04:08:41.452018Z digest=sha256:fd731245f3a2331257818a1c32cea9fe161d9496a82edcfa207636ecb5cad88c

Observation 8c3754db-dd38-441f-b252-095f165ad71f · inbound

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling cites this paper.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.986490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:28b53b7bca3afa479469753dd54457118755267574a1bcf6263e570fc38c6f3f

Observation d5a69bff-ca17-448e-9ae3-f6bc8539c31e · inbound

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality cites this paper.

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:07.935003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T15:04:41.518195Z digest=sha256:f77d64f436dee94174a0c5cf52357872c5649951b93e2c1b6df455a61b6c778d

Observation 3687ca55-3ff9-4e3f-a0bc-71af85e59c14 · inbound

Residual Decoder Adapter: ID-Preserving Tokenizer Adaption for Autoregressive Text Rendering cites this paper.

Residual Decoder Adapter: ID-Preserving Tokenizer Adaption for Autoregressive Text Rendering Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:15.745397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T15:39:15.744954Z digest=sha256:dff0f376f818e1dc82dfb9f013d0c2f8bbc283d5bc74ac04d89c8a08e0fd7676

Observation 941b87cd-4ce3-4a21-ae07-4551c6228197 · inbound

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations cites this paper.

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.533553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T13:29:11.526106Z digest=sha256:835dcb3c5ba754cb22aa3268be67bd08598493da0cc22c6e3972312176d36745

Observation e09ca6db-7932-424b-9fbb-11e66550a125 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 208

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.862033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:599f14a940795b12bd401244d995c550d36fedec5f09384bda18a8cf0a573ad7

Observation 3af3830b-45cc-44c3-9d49-aa06204b0ded · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.751913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T09:08:25.661515Z digest=sha256:75c5182b3a9f2e430ff013e7525635650a1161c26b281ebff268456e3bbb4a46

Observation 68c228fb-b342-4ecd-a447-7130a6b32ea9 · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.609070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T23:15:09.253879Z digest=sha256:615a6c52daee950bd99a49f40f8b036020e54f14f240e7237aa2f8450a52fe21

Observation 6c72ffd8-1e11-45b2-ab65-c76324a081ef · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:40.345861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:9c3f7419ee3406b09e08cd78101fe224fbb2cf69f18e1c4d255c7810fd6c5d60

Observation 06e61899-6925-4113-a9e1-dfd30d2ed4e2 · inbound

dRAE: Representation Autoencoder with Hyper-Spherical Codes cites this paper.

dRAE: Representation Autoencoder with Hyper-Spherical Codes Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T05:45:47.428321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:45:47.428321Z digest=sha256:04c270bae716ed0a0340370c63dc492b8b7e20ede069444911c8b719c123d940

Observation 5376db81-4679-4749-aa0f-5ca47ef27479 · inbound

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation cites this paper.

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T02:14:09.690130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:14:09.690130Z digest=sha256:dacd80ef043295c180439318b418914f1f0bcb39c0df37779b254b2b25d96714