Pith. sign in

Paper Citation Record · LEDGER

Unveiling Encoder-Free Vision-Language Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2406.11832.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11832 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:10:38.185094Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b1462b5b-21a8-4213-ba81-c3c614305d62 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Unveiling Encoder-Free Vision-Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:20.790119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:321f0be1c225bac3bef73ba11074d9c79d3476cc817a733d4a393083c99e8028

Observation 66838fd0-dab7-475e-b7dd-0ad5dc33381b · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:56:07.501418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:f315f958a5b4a26afaa229abd423375a58e7e33f972fa3a83909eb964aa6257d

Observation 81d5e71b-6d5f-4eef-9a8a-2c6473d63b89 · inbound

Interleaved-Modal Chain-of-Thought cites this paper.

Interleaved-Modal Chain-of-Thought Unveiling Encoder-Free Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:38.185094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:38.185094Z digest=sha256:8abab79668cd7ce6f1ee16d6139c000019751b78e93fb9556d09ebf823b30e6f

Observation 16ee3840-c0b6-4cfc-b6cd-a81b11080e4d · inbound

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding cites this paper.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.661560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.661560Z digest=sha256:e129dd7df05425dd629a4890748292c8339daa9137e29081581a2dbb843c512c

Observation de7930ef-f9ba-4cdf-8795-e67183fa8da3 · inbound

Optimizing Vision-Language Interactions Through Decoder-Only Models cites this paper.

Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.405342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.405342Z digest=sha256:cabfe618310a6b74e6c9e0b04dd41767539ba5405b106d39dee7ceec53fe3e20

Observation d9efcc2f-8978-4d23-97a8-97fad21a7f40 · inbound

Optimizing Vision-Language Interactions Through Decoder-Only Models cites this paper.

Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.412311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.412311Z digest=sha256:198246d7fd4ba9935fd969df9787c1a5b6bfdcf77ad51e6c395e06d9fb827da4

Observation b77644af-4a32-4f32-9ec4-3fc6f4dbf302 · inbound

LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering cites this paper.

LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering Unveiling Encoder-Free Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:13:24.448179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:13:24.448179Z digest=sha256:76904ecb985fe64802fbba9f3d0e4c488e611d3c64761705ddb50451d860b6ee

Observation f2eaaa8b-b60f-4334-adcf-1a09a07d5bea · inbound

FastVLM: Efficient Vision Encoding for Vision Language Models cites this paper.

FastVLM: Efficient Vision Encoding for Vision Language Models Unveiling Encoder-Free Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:19:23.160965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:19:23.160965Z digest=sha256:39430d88bdc16e0d4cdde38dcba4a26fcfae5f996649ec5eb01c45b97e6685e4

Observation 7d57e557-7d78-405d-89fe-b6177df03839 · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization Unveiling Encoder-Free Vision-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.775119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:4e7a00d9d8ce0595d856754b09c70e83e846250aff511e4bb3dfb8783c626d08

Observation 7e1bbfec-005f-4607-b11a-17934885a95c · inbound

ECG-Byte: A Tokenizer for End-to-End Generative Electrocardiogram Language Modeling cites this paper.

ECG-Byte: A Tokenizer for End-to-End Generative Electrocardiogram Language Modeling Unveiling Encoder-Free Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:20:25.637117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:20:25.637117Z digest=sha256:9943db2d0e1fd891b72e3d4049674aa27ef92865feda60826ee7fd3eb5122999

Observation 09609b9d-e3e4-4981-bc3b-3ab8bd6810fa · inbound

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding cites this paper.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Unveiling Encoder-Free Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.009009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.009009Z digest=sha256:a46e39815c547c8e5d92846e27c6d71478c4e6c5a0b4814dce83c2aa0c3bdb70

Observation 45336d13-7523-4ecb-a115-66b85f1dd4c4 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unveiling Encoder-Free Vision-Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.509004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.509004Z digest=sha256:966633524230f3ce254dd1df20e16cf5905cea7643382790e4254834a172bc0d

Observation 353cc42d-074c-4cc4-a546-351ab013e0d6 · inbound

KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification cites this paper.

KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification Unveiling Encoder-Free Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:30.501449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:30.501449Z digest=sha256:41dc1026d688c72e0935bdd56e3fd19ffdc05a5f72315d670a757eff38114803

Observation ae6bffc1-ee39-4e33-abcb-314b493cca86 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.616853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.616853Z digest=sha256:ef8d9923e317d3c40757b5953b795339c6fd481f17c2f5d6e8561f701905e747

Observation 5c81fc57-ac1e-4e04-81bb-c0024add02cb · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Unveiling Encoder-Free Vision-Language Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:05.462877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:3a8f140b0738e01345141686efaead31d09d81965393350a6cfac1a8c9512699

Observation 7778795f-6ccf-4f52-8b8e-cc97d9ed29f7 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Unveiling Encoder-Free Vision-Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:22:37.309970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:ff217cc4f1ffdd3aa1384af76f264b7a07adf3bdde3f32611d260b7cd62166cb

Observation 9713abf4-f732-4f4c-9f1b-cdedf42b3fbf · inbound

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation cites this paper.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unveiling Encoder-Free Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:15.700721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:15.700721Z digest=sha256:6de589b2c0a3a66bcf1b32748f2a2b8a6acdaaf6e876dfa0662a5bb373fabc32

Observation 6a10af03-a77f-46dc-9977-4ebbbfb433b1 · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas Unveiling Encoder-Free Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:47.368477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:47.368477Z digest=sha256:e9cbf39a137905303a5ffb43c5e8f451f662138486fa880122f15f3a2d33fed9

Observation 077632fb-0668-46c5-a881-19e25638e747 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Unveiling Encoder-Free Vision-Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:51:16.044495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:1b87dc569d7f8821af9889f19cdcc1e86f7d022964a0bc3ffcec54a5f2f639ed

Observation e7305b86-daf9-40fe-b4aa-958b5de44d65 · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unveiling Encoder-Free Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.483409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.483409Z digest=sha256:37a0c71f69c6e2e76bd9888ac14f1c4eb1227b8a72bde424d4dd3849f1a47844

Observation 9946d252-2a16-4bd1-a9e0-50e93c74d681 · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation Unveiling Encoder-Free Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.468458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.468458Z digest=sha256:dc6ab2cc459c4b44898ec2db177f2c1383f814a92a1e83d39f70b37cc71b801d

Observation 31e4b70e-3842-4761-9ffd-58b30ea63fba · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Unveiling Encoder-Free Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.218152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.218152Z digest=sha256:17939a2dbf41f9650128b101e61d973013b2d2ebc8acbb41fe86a0add5fa8440

Observation 3412c648-d966-4653-9424-1378125f9dd4 · inbound

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation cites this paper.

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:15:58.922895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:45:36.306400Z digest=sha256:8a2203dad347aae082e5a8dd417810a50ff0cfff0caaa60a3d7d6a0adde6a40f

Observation 6820083f-b286-49e2-9af7-96923cc49188 · inbound

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning cites this paper.

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning Unveiling Encoder-Free Vision-Language Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:13.486294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T17:19:59.247074Z digest=sha256:48d9b6432d21d1959847abfce56baf091eb8961f039b23646ded8b764aeed40f

Observation 3ff67812-984b-4a7f-8fc2-d12182f1d980 · inbound

From Pixels to Words -- Towards Native One-Vision Models at Scale cites this paper.

From Pixels to Words -- Towards Native One-Vision Models at Scale Unveiling Encoder-Free Vision-Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:28.011099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T13:29:53.006064Z digest=sha256:f722491cc83739879c5351ada0f3d88825d3cea5df0d7e4965a86ded18bff75f

Observation ba1b7639-e74c-4049-9c4a-0414db2ff596 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Unveiling Encoder-Free Vision-Language Models

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:32:46.758323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:a9b727f6f2a798a6a33ab1aab35f414fe34c3b207524c590f8684cfec5e6ede7

Observation cda596e6-7081-46cb-ba46-9774f1433403 · inbound

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning cites this paper.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Unveiling Encoder-Free Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:9aa03eb2d3994229a085fbc6579999138fcc0a32078eaf2cf3971e03f9ccc744