Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2406.11832.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:10:38.185094Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation b1462b5b-21a8-4213-ba81-c3c614305d62 · inbound
PaliGemma: A versatile 3B VLM for transfer Unveiling Encoder-Free Vision-Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 66838fd0-dab7-475e-b7dd-0ad5dc33381b · inbound
Emu3: Next-Token Prediction is All You Need Unveiling Encoder-Free Vision-Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 81d5e71b-6d5f-4eef-9a8a-2c6473d63b89 · inbound
Interleaved-Modal Chain-of-Thought Unveiling Encoder-Free Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ee3840-c0b6-4cfc-b6cd-a81b11080e4d · inbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Unveiling Encoder-Free Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de7930ef-f9ba-4cdf-8795-e67183fa8da3 · inbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9efcc2f-8978-4d23-97a8-97fad21a7f40 · inbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b77644af-4a32-4f32-9ec4-3fc6f4dbf302 · inbound
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering Unveiling Encoder-Free Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2eaaa8b-b60f-4334-adcf-1a09a07d5bea · inbound
FastVLM: Efficient Vision Encoding for Vision Language Models Unveiling Encoder-Free Vision-Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d57e557-7d78-405d-89fe-b6177df03839 · inbound
Autoregressive Video Generation without Vector Quantization Unveiling Encoder-Free Vision-Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7e1bbfec-005f-4607-b11a-17934885a95c · inbound
ECG-Byte: A Tokenizer for End-to-End Generative Electrocardiogram Language Modeling Unveiling Encoder-Free Vision-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09609b9d-e3e4-4981-bc3b-3ab8bd6810fa · inbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Unveiling Encoder-Free Vision-Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45336d13-7523-4ecb-a115-66b85f1dd4c4 · inbound
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unveiling Encoder-Free Vision-Language Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 353cc42d-074c-4cc4-a546-351ab013e0d6 · inbound
KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification Unveiling Encoder-Free Vision-Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae6bffc1-ee39-4e33-abcb-314b493cca86 · inbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unveiling Encoder-Free Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c81fc57-ac1e-4e04-81bb-c0024add02cb · inbound
Seed1.5-VL Technical Report Unveiling Encoder-Free Vision-Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7778795f-6ccf-4f52-8b8e-cc97d9ed29f7 · inbound
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Unveiling Encoder-Free Vision-Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9713abf4-f732-4f4c-9f1b-cdedf42b3fbf · inbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unveiling Encoder-Free Vision-Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a10af03-a77f-46dc-9977-4ebbbfb433b1 · inbound
Dense360: Dense Understanding from Omnidirectional Panoramas Unveiling Encoder-Free Vision-Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 077632fb-0668-46c5-a881-19e25638e747 · inbound
Show-o2: Improved Native Unified Multimodal Models Unveiling Encoder-Free Vision-Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e7305b86-daf9-40fe-b4aa-958b5de44d65 · inbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unveiling Encoder-Free Vision-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9946d252-2a16-4bd1-a9e0-50e93c74d681 · inbound
NeoBabel: A Multilingual Open Tower for Visual Generation Unveiling Encoder-Free Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e4b70e-3842-4761-9ffd-58b30ea63fba · inbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Unveiling Encoder-Free Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3412c648-d966-4653-9424-1378125f9dd4 · inbound
Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation Unveiling Encoder-Free Vision-Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6820083f-b286-49e2-9af7-96923cc49188 · inbound
MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning Unveiling Encoder-Free Vision-Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3ff67812-984b-4a7f-8fc2-d12182f1d980 · inbound
From Pixels to Words -- Towards Native One-Vision Models at Scale Unveiling Encoder-Free Vision-Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ba1b7639-e74c-4049-9c4a-0414db2ff596 · inbound
Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Unveiling Encoder-Free Vision-Language Models
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cda596e6-7081-46cb-ba46-9774f1433403 · inbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Unveiling Encoder-Free Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.