Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:25:56.014437Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 8 inbound Pith citation observations for arXiv:2502.06788.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:25:56.014437Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:47.528488Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T18:51:16.046798Z
100 of 104 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dca388f7-13b1-43f6-bfb8-943b93ac96ec · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ade161-fe1e-4ff1-8266-153d126342bc · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models The claude 3 model family: Opus, sonnet, haiku
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad2f3fe8-dd74-4490-89b1-f0c28c67c4e4 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b58348f8-41d7-4aef-afbd-a9c67865d6bc · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce6090de-bcaf-41a7-b85b-0c90eb67a69b · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Vlmo: Unified vision- language pre-training with mixture-of-modality-experts
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c762701b-d418-48a4-bcfe-4c86aab26ea1 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Introducing our multimodal models, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 221102a1-281e-48c6-bda6-87361a5eb6d7 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models PaliGemma: A versatile 3B VLM for transfer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5432cfb1-ca45-4f38-9fb5-61c215c0d961 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ecb0bf6-542c-4f7b-a0ad-00c64a286e54 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models InternLM2 Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80884fde-78f9-4ef3-86fa-a29d904c4544 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Emerg- ing properties in self-supervised vision transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eafa05a-2662-4281-8b72-ade0241cb59e · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e6df3dd-dabc-4144-b251-8dbbfc64d53f · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d783e8fa-cc1c-4ef8-b4d6-8b90a045d9f8 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c10a6ea-6332-4215-ab54-882ef11255a6 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bde31ad-577c-4d48-bd9e-090266c3d76a · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48b8cc72-30ba-4ce2-9fe5-9d10bffefb69 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Gonzalez, Ion Stoica, and Eric P
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629ec753-3b7b-4d9f-9569-a92d02f31f2c · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d190b8-39bb-458c-b805-d15a854defa4 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae6bffc1-ee39-4e33-abcb-314b493cca86 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unveiling Encoder-Free Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b42a40c-3e5f-4474-b53c-e0a398311143 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unipt: Universal parallel tuning for transfer learning with efficient parameter and memory
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c141db6-6e05-4743-b29e-5de3f755d9b9 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Sherl: Synthesizing high accuracy and efficient memory for resource-limited transfer learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8df1a290-7bc9-4313-b547-3fd70e2a5dcd · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f920d46-2e08-450f-9e60-25503d29d0d0 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Taming transformers for high-resolution image synthesis
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73857e29-a43c-47da-9f67-1ce4521106e6 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models EV A: exploring the limits of masked visual representation learning at scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb0a214-e505-4402-b316-4916b8ea9293 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d493f05-95dd-42c0-bdb9-5841d61728f6 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Dat- acomp: In search of the next generation of multimodal datasets
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a0e6a2d-030f-41ad-bafb-7a18b6d46781 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95f8c69b-3394-4235-866e-88c90fa637d6 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Making LLaMA SEE and draw with SEED tokenizer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51840bca-4d2a-4d65-ae23-6428312a90f4 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecb5431c-b222-4c9f-85f8-60a5fe5f06a1 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models CogAgent: A Visual Language Model for GUI Agents
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ee0b32f-9466-44de-912c-0d4252ff1857 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7b681be-1ef4-4d70-9983-78e3b19a678b · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Hudson and Christopher D
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46e6663a-f404-4645-a07e-9ba0fa342654 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Introducing idefics: An open reproduction of state-of-the-art visual language model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05533886-e4eb-4811-bf0a-ed1473626e14 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models A diagram is worth a dozen images
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9e3624d-5fa5-40f2-8600-4f2a71cbff7c · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Kingma and Jimmy Ba
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9c99db6-6e45-492e-93ce-245dc3ed24cb · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Segment Anything
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37d774e5-ce92-4d38-b783-0bad9ad04198 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78ef67f-3e12-4473-8176-d60d129373ad · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Building and better understanding vision-language models: insights and future directions
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c345cac1-3c26-484d-8599-56a3de933b85 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68e8ec1-3498-43aa-bcbb-d638cf8dffe7 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8659e668-aea5-46f6-af9e-87a2a11fa040 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 434c5d3b-7449-4547-85b9-5bdcaf8734f3 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04183272-f961-4012-b78e-14344ab02465 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfa8f363-afbf-4b8f-b7c7-016e390096fb · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mc-beit: Multi-choice discretization for image bert pre-training
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30286c70-f0d9-41bc-bf97-62ed05f26b43 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models ImageFolder: Autoregressive Image Generation with Folded Tokens
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5b52559-e7e7-42ee-8c3b-0be5be797755 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ea5c4b-be71-47f1-a41a-abb5dfe8ed82 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Evaluating object hallucination in large vision-language models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31365a7f-c765-4263-a232-b30f7ed0088d · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c745885-3cd2-4023-9e08-f5804c1b595c · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f36955-c614-419b-981f-7ea9a73f4a56 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Improved Baselines with Visual Instruction Tuning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd7ef618-850f-410d-8e22-7d8e3e76e45b · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Visual instruction tuning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f033e7b3-34cb-4d64-a059-1ab62a84e988 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c808b132-449d-4dfc-9c58-196374f34cc6 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MMBench: Is Your Multi-modal Model an All-around Player?
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a26125b8-633d-4986-9a9e-0a49ba8c0aed · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8044e42-29ea-4c47-8829-d402045b0052 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57650a85-2453-464f-8f88-97e868b7b9cb · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6ce369cc-c52d-4505-888e-59dff9b73e0e · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5785fe5-8416-4bfb-ae01-70935eb28920 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5e87e700-855e-4fdb-b467-e8950e3bfd15 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models GPT-4 Technical Report
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca3d66c-37d2-482a-8e41-0ba9f8c5a2c1 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f225d90c-9381-4d87-afa2-f12751545a72 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Learning transferable visual models from natural language supervision
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 75ff3c6e-5086-40d9-8139-5aa60808a861 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e63c32b6-ad8c-4af5-996e-cc77c8173dbb · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Towards VQA models that can read
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c5760b1-99d1-4b68-8f66-7d34184d09b8 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93fee7ad-4953-479f-a040-2b380c6b26f9 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Generative Multimodal Models are In-Context Learners
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 291b4ece-9106-4d66-b709-a4847c6ec119 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 896d60e1-29b6-4d45-a5fe-313b2a6f62cd · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Emu: Generative Pretraining in Multimodality
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20477537-a323-4c31-956e-b3057dc27188 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fae62d9-4f4b-4d6d-b363-2743ff019593 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5032d9f8-67c8-4a72-b8b8-0bd64789f976 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a34eda86-776c-44fa-868b-a9e1fb2913bd · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Internlm: A multilingual language model with progressively enhanced capabilities
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe8f12eb-f70a-43da-986b-e18ff6f57368 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9332fce2-4319-4647-8547-2703bc9313cd · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen2.5: A party of foundation models, 2024
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ca957625-1e55-4d79-9348-2ac58e780253 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f37dae-d47b-4967-bf9c-903653c1beda · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3f85f9b3-f2d7-426a-b4d0-b354b2599d2c · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02f7baf5-9b60-4052-bdd7-c7ededa5d4f7 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66893499-3a31-4a5b-a2ee-e32626443428 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Neural discrete representation learning
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc369af7-71b0-4de6-a412-14689454093b · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c6280041-f4b1-4ef6-ae4a-877331f0c2b9 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 929ff2b0-581b-4122-91c6-67f012e7033a · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdbfd165-2d57-4b32-a820-fefd637ea767 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 706d3749-3922-494f-bafe-2a68d02fd51d · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Emu3: Next-Token Prediction is All You Need
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5380f8f9-1095-4541-8124-9df08d2f2b60 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Mio: A foundation model on multimodal tokens
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df984a12-468b-48e4-a4dd-cb0e61fe49b1 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f53753-39f3-43f5-8c53-629bcfa2b0ee · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 065aea9b-d63f-44d1-9a1c-4ab918e95172 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Grok-1.5 vision preview, 2024
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 81412cb5-a096-4187-92c2-ff89caf42563 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23031d0-36ec-4096-8cf9-467741caeef5 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c73ca218-7d4b-46d1-83c3-fb6010d0fa22 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 261b581e-efb5-4d7c-b5d9-71b806c7170c · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models xgen-mm (blip-3): A family of open large multimodal models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d45ce1-dca6-48c6-ace0-54ccde05b893 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen2 Technical Report
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cee62dc-3aea-45e0-9475-11a758954187 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09ba62fb-7ffe-4938-9208-8263578eccf7 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f231966d-8bb1-44e8-a38e-3b12ef6530d4 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33569c35-748e-4081-9cc9-a6372427a5fe · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b9462a-35a9-4365-96b6-38dbc6de95b8 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44cdbe71-d721-4ea0-a560-daa005694dc8 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Sigmoid loss for language image pre-training
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4c7259e3-331b-4aaf-866c-d932f3a74305 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dde0cbd-6c00-4a84-a005-0ae1a1f197c0 · outbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 738bf878-8fe3-40a2-9a84-d4a0079b3f59 · inbound
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation debf2a9a-6d76-491b-8c8b-dd7dab277001 · inbound
Dense360: Dense Understanding from Omnidirectional Panoramas EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 672d1c36-d965-4122-88c2-c3a0e1057c6a · inbound
Show-o2: Improved Native Unified Multimodal Models EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d54a2bed-4601-40be-852d-3d6a8fecd470 · inbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b134bd5a-39fd-43cc-9b02-961aba8666fb · inbound
NeoBabel: A Multilingual Open Tower for Visual Generation EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee704c9-80bf-4ae4-aace-ac300f0c8736 · inbound
Regularizing Subspace Redundancy of Low-Rank Adaptation EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6313292f-a74b-4422-952c-043a5e25126a · inbound
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c7b937c-c260-4d00-b0f2-9190e4cf32e9 · inbound
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.