Pith. sign in

Paper Citation Record · LEDGER

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

As of 13 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 8 inbound Pith citation observations for arXiv:2502.06788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06788 v2

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:25:56.014437Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:47.528488Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T18:51:16.046798Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved90
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dca388f7-13b1-43f6-bfb8-943b93ac96ec · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.538636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.538636Z digest=sha256:3bb5437d28a78fbc0c736edccf16762be4857a136f4e99ef5232df0071d6d169

Observation 67ade161-fe1e-4ff1-8266-153d126342bc · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models The claude 3 model family: Opus, sonnet, haiku

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.543179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.543179Z digest=sha256:20cfa992f8400707bcce07a2e1105cdd8952cbe7503172e9a76a4b21d069903c

Observation ad2f3fe8-dd74-4490-89b1-f0c28c67c4e4 · outbound

This paper cites Qwen Technical Report.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.547154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.547154Z digest=sha256:64afd40ab7114aa1e7ca2a8aeb9d7f5df3bbc0e403ac5c9ac6cdd61eb13ab2fe

Observation b58348f8-41d7-4aef-afbd-a9c67865d6bc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.551753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.551753Z digest=sha256:cd5dbf9925fb77765e9763b2fc69152487aa5e1605729ce6681d7d6fa1f3c096

Observation ce6090de-bcaf-41a7-b85b-0c90eb67a69b · outbound

This paper cites Vlmo: Unified vision- language pre-training with mixture-of-modality-experts.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Vlmo: Unified vision- language pre-training with mixture-of-modality-experts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.555696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.555696Z digest=sha256:1730446c52de287a64c0c2f4061d22123c2683a63d24f5176e6976152681ce0a

Observation c762701b-d418-48a4-bcfe-4c86aab26ea1 · outbound

This paper cites Introducing our multimodal models, 2023.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Introducing our multimodal models, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.559567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.559567Z digest=sha256:bbdbd856d417498d2cefddb7a2bf1d242c78f74d0b89ff6da15c557075e2b188

Observation 221102a1-281e-48c6-bda6-87361a5eb6d7 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models PaliGemma: A versatile 3B VLM for transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.563893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.563893Z digest=sha256:879b660248263812e4b37ab81f17687562af89d663e6472c2dad78ccf8a0dd23

Observation 5432cfb1-ca45-4f38-9fb5-61c215c0d961 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.568135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.568135Z digest=sha256:ae9604c9b94511b5149209d4d5956e977ec01a76aed06771e92549862d688c12

Observation 8ecb0bf6-542c-4f7b-a0ad-00c64a286e54 · outbound

This paper cites InternLM2 Technical Report.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models InternLM2 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.572341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.572341Z digest=sha256:ccaa03ae8cdb720588b9da78da82563ce375b82055ac8d00e4bf536119034988

Observation 80884fde-78f9-4ef3-86fa-a29d904c4544 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Emerg- ing properties in self-supervised vision transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.577160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.577160Z digest=sha256:2a29d4306b8320a895666cc72134610e2c2f696594c163c3f7d480685b25e0fc

Observation 7eafa05a-2662-4281-8b72-ade0241cb59e · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.581404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.581404Z digest=sha256:c48d98837cf9900f461c15a6fa33328b0faf497ffb9bc69d7bd31bdb0944f03a

Observation 8e6df3dd-dabc-4144-b251-8dbbfc64d53f · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.586228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.586228Z digest=sha256:735280a9ff00c1ddf9213696e035b1ceeab6b13e1beeeb47e634103f7f823cd2

Observation d783e8fa-cc1c-4ef8-b4d6-8b90a045d9f8 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.590338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.590338Z digest=sha256:53462cd4c6aabce16e917adca855d4baa99757d7464371012bb13e35d71a07cf

Observation 6c10a6ea-6332-4215-ab54-882ef11255a6 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.594875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.594875Z digest=sha256:48d3a7e530576becb8a29a15bd23002d21c0d779aac90fcad0d2525c6e1249e7

Observation 7bde31ad-577c-4d48-bd9e-090266c3d76a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.599372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.599372Z digest=sha256:9f5e6ddabe3366e7ed6134c1526f4d256d53b683f3e3aa594ace663c91a29a90

Observation 48b8cc72-30ba-4ce2-9fe5-9d10bffefb69 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Gonzalez, Ion Stoica, and Eric P

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.604025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.604025Z digest=sha256:aef0d84817b8c31a31ca04591da3247f59da706cc8cec5848f28a7bfa32ef08e

Observation 629ec753-3b7b-4d9f-9569-a92d02f31f2c · outbound

This paper cites an unresolved cited work.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.608163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.608163Z digest=sha256:17c7af9f6a07a970c08f72a133326688d338733d595d1244e2b40d77911e0c93

Observation 88d190b8-39bb-458c-b805-d15a854defa4 · outbound

This paper cites an unresolved cited work.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.612376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.612376Z digest=sha256:b1935b6650c3f86ab816dd7a6590ca0350ea5f5aac87f68fe446295e9b6546ef

Observation ae6bffc1-ee39-4e33-abcb-314b493cca86 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.616853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.616853Z digest=sha256:14416b649922e1bad81acba9e87a3fd19cc8ce9fbcc93625d3c0e6a2372753dc

Observation 4b42a40c-3e5f-4474-b53c-e0a398311143 · outbound

This paper cites Unipt: Universal parallel tuning for transfer learning with efficient parameter and memory.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unipt: Universal parallel tuning for transfer learning with efficient parameter and memory

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.626452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.626452Z digest=sha256:9c00b6c5817ee8500f47c369f4965a1579b8ae144490d1e96f570a36d3a29b6a

Observation 8c141db6-6e05-4743-b29e-5de3f755d9b9 · outbound

This paper cites Sherl: Synthesizing high accuracy and efficient memory for resource-limited transfer learning.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Sherl: Synthesizing high accuracy and efficient memory for resource-limited transfer learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.631027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.631027Z digest=sha256:0772449691feb502a2a123989b530a3212bfd6b17debabda7e49332342babad3

Observation 8df1a290-7bc9-4313-b547-3fd70e2a5dcd · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.635179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.635179Z digest=sha256:f04b081889512b5dfac9094b2539f1b978d9f5eebff1670bea940cc4935eae96

Observation 8f920d46-2e08-450f-9e60-25503d29d0d0 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Taming transformers for high-resolution image synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.639618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.639618Z digest=sha256:acd16d1da5c1453e46b706634487f4c53e5066dc2e6ff9cce5fa237bdd963c0c

Observation 73857e29-a43c-47da-9f67-1ce4521106e6 · outbound

This paper cites EV A: exploring the limits of masked visual representation learning at scale.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models EV A: exploring the limits of masked visual representation learning at scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.643640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.643640Z digest=sha256:f4ae1fa70c48aa1f48d40f3a60ba228a809a895f49e6a76c3a3998ee6dd9f282

Observation ffb0a214-e505-4402-b316-4916b8ea9293 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.648170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.648170Z digest=sha256:41dfccdb318c28c74f5e84058710b905ecfe5e809c9dfcde6836df8594880c2b

Observation 1d493f05-95dd-42c0-bdb9-5841d61728f6 · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Dat- acomp: In search of the next generation of multimodal datasets

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.652587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.652587Z digest=sha256:345083112f2f0fe5dc6495cd76171df5fb0481cbb5afb4c86e30a30b03e5f693

Observation 1a0e6a2d-030f-41ad-bafb-7a18b6d46781 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.656727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.656727Z digest=sha256:29e9fa663ba26a6269fafb0a4d23db21fa14aab65e49b6e372bb665965ce15d2

Observation 95f8c69b-3394-4235-866e-88c90fa637d6 · outbound

This paper cites Making LLaMA SEE and draw with SEED tokenizer.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Making LLaMA SEE and draw with SEED tokenizer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.661080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.661080Z digest=sha256:e61a99512bd702b52a0cadfca7a9f489eeeecebd8b67e34adf50718d895cd712

Observation 51840bca-4d2a-4d65-ae23-6428312a90f4 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.665218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.665218Z digest=sha256:710c18574c3860f1493fade3f8dd803abe78c4b83889c287887cf6697e828ee3

Observation ecb5431c-b222-4c9f-85f8-60a5fe5f06a1 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models CogAgent: A Visual Language Model for GUI Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.669886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.669886Z digest=sha256:19fc122f7bae6b216549c9f0cae9f11af4f5b8b956d156711757c7f7b4587443

Observation 4ee0b32f-9466-44de-912c-0d4252ff1857 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.674294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.674294Z digest=sha256:51882eabdc227c691c014e8915a327d29ba1bbee9cac03d1cd04e9b60ceddeea

Observation c7b681be-1ef4-4d70-9983-78e3b19a678b · outbound

This paper cites Hudson and Christopher D.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Hudson and Christopher D

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.678998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.678998Z digest=sha256:e10c27896d79cc3a5cedcb67d72a9414fced026e944b396f2f67b7b9f873e6bf

Observation 46e6663a-f404-4645-a07e-9ba0fa342654 · outbound

This paper cites Introducing idefics: An open reproduction of state-of-the-art visual language model.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Introducing idefics: An open reproduction of state-of-the-art visual language model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.683411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.683411Z digest=sha256:7bff309df74e623c03d51367896f53bdc24f7e32c710a5f61157a16f22769eb0

Observation 05533886-e4eb-4811-bf0a-ed1473626e14 · outbound

This paper cites A diagram is worth a dozen images.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models A diagram is worth a dozen images

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.688325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.688325Z digest=sha256:cea4df81d8ebb37530a3277bc48ff9bc4b6de55d5e8aa3589f500c28cc63e2ee

Observation e9e3624d-5fa5-40f2-8600-4f2a71cbff7c · outbound

This paper cites Kingma and Jimmy Ba.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Kingma and Jimmy Ba

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.692617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.692617Z digest=sha256:2c116f21dccfeec28352f9aa87f87fd0bd2b61e17f701347883c1c333cf3642c

Observation e9c99db6-6e45-492e-93ce-245dc3ed24cb · outbound

This paper cites Segment Anything.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Segment Anything

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.696864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.696864Z digest=sha256:d381a7aabc9413a7970f618f57adf300ce73238315d9bee5cf05b6ea7a41ca16

Observation 37d774e5-ce92-4d38-b783-0bad9ad04198 · outbound

This paper cites The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.702211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.702211Z digest=sha256:501dfac17604fc7705b62b0fed94d99f0e2dd2003d2e08cb9f26c0f1610f9436

Observation f78ef67f-3e12-4473-8176-d60d129373ad · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Building and better understanding vision-language models: insights and future directions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.706584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.706584Z digest=sha256:4230c0348a3d3587a5ba44698fe0917a8c0d82df5cc7a6a3c442029f66251e1c

Observation c345cac1-3c26-484d-8599-56a3de933b85 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.711133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.711133Z digest=sha256:973c61b5c0c40c6d9506e5bccaf04b9b2f4c34401a5df3c3876b5fa7a2672cb6

Observation d68e8ec1-3498-43aa-bcbb-d638cf8dffe7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.715605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.715605Z digest=sha256:76c2b4df868047962f0e769feb7f2312d2560a647badd57162050aeb8f17db8f

Observation 8659e668-aea5-46f6-af9e-87a2a11fa040 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.720298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.720298Z digest=sha256:16741234f7a2d60f886093a14bac1c0aa7b64407c4ca60f9fde6196f287cd1f6

Observation 434c5d3b-7449-4547-85b9-5bdcaf8734f3 · outbound

This paper cites an unresolved cited work.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.725842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.725842Z digest=sha256:2a8f256755d4912ea1dd2421cbfcc67a7491b6dd844c09c7950f7f11d2687dfb

Observation 04183272-f961-4012-b78e-14344ab02465 · outbound

This paper cites an unresolved cited work.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.730026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.730026Z digest=sha256:49215f83d613c8bd4973e17a3a22e5e30df55bf0fc25299bd37b4acf12c95264

Observation cfa8f363-afbf-4b8f-b7c7-016e390096fb · outbound

This paper cites mc-beit: Multi-choice discretization for image bert pre-training.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mc-beit: Multi-choice discretization for image bert pre-training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.734219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.734219Z digest=sha256:f7a3b1c3ecfd91aca69f4d0ebab02013473754154b16951b17bc7a1b06898499

Observation 30286c70-f0d9-41bc-bf97-62ed05f26b43 · outbound

This paper cites ImageFolder: Autoregressive Image Generation with Folded Tokens.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models ImageFolder: Autoregressive Image Generation with Folded Tokens

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.738517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.738517Z digest=sha256:ea714670d48c0c200d10dee9fbb398707f28d07bf65b71a03bc55401e3ab24fd

Observation e5b52559-e7e7-42ee-8c3b-0be5be797755 · outbound

This paper cites DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.743104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.743104Z digest=sha256:cd1c31efe82472265f5c77f5805fd25f58e959746d09787df984a51029c79b4b

Observation 99ea5c4b-be71-47f1-a41a-abb5dfe8ed82 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Evaluating object hallucination in large vision-language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.748054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.748054Z digest=sha256:0d0d2d0033432755891d95bcbeb9cb4b66d1491268c28ba7cb482e9563d29134

Observation 31365a7f-c765-4263-a232-b30f7ed0088d · outbound

This paper cites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.752574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.752574Z digest=sha256:9ebc4f1dc5a9f41f0b251dfc14351f0f9671ea7ee987a39643f0f6ecab8134d6

Observation 9c745885-3cd2-4023-9e08-f5804c1b595c · outbound

This paper cites MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.756888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.756888Z digest=sha256:55b391899c51d1af6baae399391bed54601ef77952ebb1a84ac971e04d8d4b2b

Observation c0f36955-c614-419b-981f-7ea9a73f4a56 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Improved Baselines with Visual Instruction Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.761283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.761283Z digest=sha256:e4f9ff60650f8706a95a255decb1394fafc8363b0e75dc4256e2150a76b2c560

Observation fd7ef618-850f-410d-8e22-7d8e3e76e45b · outbound

This paper cites Visual instruction tuning.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Visual instruction tuning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.765855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.765855Z digest=sha256:9130fcb076413a53a5675c08caeab3607676689b60c4057026096534d5144c7d

Observation f033e7b3-34cb-4d64-a059-1ab62a84e988 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.770240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.770240Z digest=sha256:4db3160dcefa06d95305ed7afa9ab9f0cea20485dced6971528487ba84b8b99a

Observation c808b132-449d-4dfc-9c58-196374f34cc6 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MMBench: Is Your Multi-modal Model an All-around Player?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.774516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.774516Z digest=sha256:7f9aa18ccca8939430f900ba93ad98308963219357ed6e0a38a82036b3e6990e

Observation a26125b8-633d-4986-9a9e-0a49ba8c0aed · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.779074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.779074Z digest=sha256:3691bbc831cd7bc7f38122360d63d2722ce27b27f556fab224301c0d4f17acbf

Observation c8044e42-29ea-4c47-8829-d402045b0052 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.783785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.783785Z digest=sha256:6b706115d1f417ea4f7f26905972bb685ab1cce6163020c0d8a388e86aaa4991

Observation 57650a85-2453-464f-8f88-97e868b7b9cb · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.351165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T14:25:55.788669Z digest=sha256:9f3a978af72dbd52b88556919332195d315c5a22f8680baad539f0980573634f

Observation 6ce369cc-c52d-4505-888e-59dff9b73e0e · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.793608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.793608Z digest=sha256:8942e1b032cdf2f864ddc7a8140502c77497844d7015edad690b510d39b6c40e

Observation c5785fe5-8416-4bfb-ae01-70935eb28920 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.335376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T14:25:55.798660Z digest=sha256:3973afd9dabf36394ef6d311d4a52c7531ed906de5def2c97908607a2018d85c

Observation 5e87e700-855e-4fdb-b467-e8950e3bfd15 · outbound

This paper cites GPT-4 Technical Report.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models GPT-4 Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.803714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.803714Z digest=sha256:2c237a74be3fc5bb30aa580888584af5f5ed7201197f396396548a3d7c843f39

Observation cca3d66c-37d2-482a-8e41-0ba9f8c5a2c1 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.809314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.809314Z digest=sha256:b78327917df441a226647fed28cab0af2c649e7a5418778b0b449e3851bba8b7

Observation f225d90c-9381-4d87-afa2-f12751545a72 · outbound

This paper cites Learning transferable visual models from natural language supervision.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Learning transferable visual models from natural language supervision

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.319691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T14:25:55.814565Z digest=sha256:96da1d8fbf2b0b150e3b1c5fd33a9b6e5ece2986c0b33d0b00f0cc85f683002f

Observation 75ff3c6e-5086-40d9-8139-5aa60808a861 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.303982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T14:25:55.819490Z digest=sha256:6d7f910bab7a60084e0f3706844b632efc4eaf4c9096155e146ff9f99da08771

Observation e63c32b6-ad8c-4af5-996e-cc77c8173dbb · outbound

This paper cites Towards VQA models that can read.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Towards VQA models that can read

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.824630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.824630Z digest=sha256:893e480d364050d11d252e62911850a6eb76f8e9e46065ffb4c8c9056ed75081

Observation 4c5760b1-99d1-4b68-8f66-7d34184d09b8 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.829308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.829308Z digest=sha256:266832b5a640e40fd169ae75efc0134804205099a00388017f456210d8d7b16f

Observation 93fee7ad-4953-479f-a040-2b380c6b26f9 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Generative Multimodal Models are In-Context Learners

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.834228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.834228Z digest=sha256:f903d77a9361e5c59c6115c869902729165895b01ee870955c56f1fe3245b5db

Observation 291b4ece-9106-4d66-b709-a4847c6ec119 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.839354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.839354Z digest=sha256:08e6d58e01f8758df815f3f7b0acb6410771ef567724d5695d039cc56f557804

Observation 896d60e1-29b6-4d45-a5fe-313b2a6f62cd · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Emu: Generative Pretraining in Multimodality

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.844863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.844863Z digest=sha256:5fb407af149d67e241606c0b51a6219a11e06a16d2667e67c355f91330fc4bdc

Observation 20477537-a323-4c31-956e-b3057dc27188 · outbound

This paper cites EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.850157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.850157Z digest=sha256:c9e9a52dac0ca0d9ca42a181e39c65852f4208ab1debd4a01fdabcbd3c592b03

Observation 6fae62d9-4f4b-4d6d-b363-2743ff019593 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.855575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.855575Z digest=sha256:e4a10a9eb66e6a694f2c30769d03e562070be20c6dbda306f67db48765a96835

Observation 5032d9f8-67c8-4a72-b8b8-0bd64789f976 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.861005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.861005Z digest=sha256:5e69eed8c999e944d0ce77f24d2b1e353a95b7451e1214bffdcc0b06df405f53

Observation a34eda86-776c-44fa-868b-a9e1fb2913bd · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Internlm: A multilingual language model with progressively enhanced capabilities

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.866397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.866397Z digest=sha256:5e29d4bcd15b7e8c9421240eff9d0efee24e90196cff8ab102a972350eb665c0

Observation fe8f12eb-f70a-43da-986b-e18ff6f57368 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.268277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T14:25:55.871639Z digest=sha256:af971c57cf66c4022ce2e1de7fd86bc4c1ea5b3d06df2c07cf6f0d7fbb4d18a7

Observation 9332fce2-4319-4647-8547-2703bc9313cd · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen2.5: A party of foundation models, 2024

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.250633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T14:25:55.876978Z digest=sha256:b4deba2c220059d688ac807f1c70c0eb93d2c27e8315c61282b7b85105604f56

Observation ca957625-1e55-4d79-9348-2ac58e780253 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.882208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.882208Z digest=sha256:8595136406ed9c1638612b8b7023d06082656211a94dc3fb44ba73b942898300

Observation d6f37dae-d47b-4967-bf9c-903653c1beda · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.233562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T14:25:55.886692Z digest=sha256:a4e886ae2d519443c0fcdbea068c1dcd51058c6a134cde7facc3524e705532a1

Observation 3f85f9b3-f2d7-426a-b4d0-b354b2599d2c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.891210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.891210Z digest=sha256:75b52de5a34a1413aa320f30769b6a88dce16cfa68d0fa245835a35d0ecb74ce

Observation 02f7baf5-9b60-4052-bdd7-c7ededa5d4f7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.895582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.895582Z digest=sha256:e021b272b9d53af02efa1d4a6e8401f480e7c7294e26cf91e4055712a1488a34

Observation 66893499-3a31-4a5b-a2ee-e32626443428 · outbound

This paper cites Neural discrete representation learning.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Neural discrete representation learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.900288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.900288Z digest=sha256:8d11718f9c0c2a04c4898fdefd98ad02163a38e134aaa8b01084c0b46a3c4414

Observation cc369af7-71b0-4de6-a412-14689454093b · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.205036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T14:25:55.905656Z digest=sha256:22825bf9047784add2db580af4da1c5991d36fe465020747154c8d63f7cde01e

Observation c6280041-f4b1-4ef6-ae4a-877331f0c2b9 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.910395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.910395Z digest=sha256:90b826fe5c0680744b4818b2446e6a7ba35ad9c1ab40a50ed397a92a819e16d1

Observation 929ff2b0-581b-4122-91c6-67f012e7033a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.915568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.915568Z digest=sha256:77f8cf077fda95365dbaeed62ca6a144b63edc4187dbd2972208aa5b5d1b82b1

Observation fdbfd165-2d57-4b32-a820-fefd637ea767 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.920632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.920632Z digest=sha256:d65bba0b082a0f49867feab2e480746ff519dc72de9b69f6aca70d60b773c590

Observation 706d3749-3922-494f-bafe-2a68d02fd51d · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Emu3: Next-Token Prediction is All You Need

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.925834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.925834Z digest=sha256:811675adc7f52cd665861a58776b481597104f325a9143733d3afa1e3792e546

Observation 5380f8f9-1095-4541-8124-9df08d2f2b60 · outbound

This paper cites Mio: A foundation model on multimodal tokens.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Mio: A foundation model on multimodal tokens

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.931088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.931088Z digest=sha256:230ba819e9e4ac47d917f6b1d0c558fe668f162fc97cfc4b7a8540e13f7c2755

Observation df984a12-468b-48e4-a4dd-cb0e61fe49b1 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.935935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.935935Z digest=sha256:53764b63bcc054cd530ab2fcc9af6a122dd4135ad6b0c125570ba016ef4d7703

Observation a1f53753-39f3-43f5-8c53-629bcfa2b0ee · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.941194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.941194Z digest=sha256:074042c676300ef25364566a1e6dca3eb032bc37540f201b5d1c18ae84d8df7e

Observation 065aea9b-d63f-44d1-9a1c-4ab918e95172 · outbound

This paper cites Grok-1.5 vision preview, 2024.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Grok-1.5 vision preview, 2024

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.188693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T14:25:55.946792Z digest=sha256:35da882e93b318789d8864ae69bc0927d179e15fef8395a78f56b0227d254763

Observation 81412cb5-a096-4187-92c2-ff89caf42563 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.951584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.951584Z digest=sha256:f4985365db0d274fd455e3e70c27d55d3947ecb26e4915ebef6bcf231e8c25d0

Observation c23031d0-36ec-4096-8cf9-467741caeef5 · outbound

This paper cites MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.957533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.957533Z digest=sha256:4fd9071e43a43ca7c9e3fa44360231fe9703b9ee4f43a197db796c93ff33f26d

Observation c73ca218-7d4b-46d1-83c3-fb6010d0fa22 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.962910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.962910Z digest=sha256:ee01337c56e06feb226df71e970788d06093d2630e6492d52d5d8af1b6bbd296

Observation 261b581e-efb5-4d7c-b5d9-71b806c7170c · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models xgen-mm (blip-3): A family of open large multimodal models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.968319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.968319Z digest=sha256:7e85fc175d9e4de299180176bfe0b552dbbd16ddebe1b8aa70c03f7eeac2644c

Observation 52d45ce1-dca6-48c6-ace0-54ccde05b893 · outbound

This paper cites Qwen2 Technical Report.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen2 Technical Report

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.973291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.973291Z digest=sha256:df4ff43949238d925f451d7644f48b780b994eac3cd6becdb2ebe5cce589944f

Observation 4cee62dc-3aea-45e0-9475-11a758954187 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.978546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.978546Z digest=sha256:923ec801b13505e76d67cca56dede2f9c54579e0bee4d9eb5bfdbd1f08720f85

Observation 09ba62fb-7ffe-4938-9208-8263578eccf7 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.983562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.983562Z digest=sha256:92cc089fe92981c339142b40d0f6d371ef51755776332adc9e81f5d13485254b

Observation f231966d-8bb1-44e8-a38e-3b12ef6530d4 · outbound

This paper cites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.988746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.988746Z digest=sha256:f5710dcce16549c439aabcb65bd0c614712e59c28eafc5ab9e8850ecfb2f8f5a

Observation 33569c35-748e-4081-9cc9-a6372427a5fe · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.993984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.993984Z digest=sha256:273b54e7bbf80292c4073f4b57adfb46a96fb5925421b71d05c6c843c96ff801

Observation 69b9462a-35a9-4365-96b6-38dbc6de95b8 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.999125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.999125Z digest=sha256:18e721f4262e735c19d234b3e4d34ad75b381cc180433d3bb746bb95e9a6df01

Observation 44cdbe71-d721-4ea0-a560-daa005694dc8 · outbound

This paper cites Sigmoid loss for language image pre-training.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Sigmoid loss for language image pre-training

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.171985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T14:25:56.004175Z digest=sha256:7cbf9d21547594f0d2e1b9d47bdaddbf9f9d40c9d57069eb9b5515c11cf2ec7b

Observation 4c7259e3-331b-4aaf-866c-d932f3a74305 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:56.009237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:56.009237Z digest=sha256:6643898f427abd889c647107072e48522e7208d7ea858040a14f900edb1fdea5

Observation 6dde0cbd-6c00-4a84-a005-0ae1a1f197c0 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:56.014437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:56.014437Z digest=sha256:141fc03d1f5e86fe150ce3a09879899828d443ea85a16f8c78ca649638fef264

Pith citing papers

Observation 738bf878-8fe3-40a2-9a84-d4a0079b3f59 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:22:37.318223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:552672a6be976f82045d43a71503955b2f2c93961b3107d8923cef0cd545805b

Observation debf2a9a-6d76-491b-8c8b-dd7dab277001 · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:47.528488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:47.528488Z digest=sha256:0062cf34da6ab7c2d5db82d8c998d3299817a063b22c7752943d2d13b8a3d83d

Observation 672d1c36-d965-4122-88c2-c3a0e1057c6a · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:16.049795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:c18b9369f32935fd6a0d85f268f406038a6f57641fe39612388947851c073fdf

Observation d54a2bed-4601-40be-852d-3d6a8fecd470 · inbound

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs cites this paper.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.482217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.482217Z digest=sha256:7d9e1bc51fe57dd230eceb5aeaf1ff9599ee57e41cd1ee12cdf66e28cb24a985

Observation b134bd5a-39fd-43cc-9b02-961aba8666fb · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.683136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.683136Z digest=sha256:d7dc7f90c65cc074ecdc86d1b8fdc979579f7f5145856b256b75130a0704daf8

Observation fee704c9-80bf-4ae4-aace-ac300f0c8736 · inbound

Regularizing Subspace Redundancy of Low-Rank Adaptation cites this paper.

Regularizing Subspace Redundancy of Low-Rank Adaptation EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:33.793178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:23:33.793178Z digest=sha256:0272479df1199407fd9a52c19264c749b74ad035758b4670a0abd095c1291cc9

Observation 6313292f-a74b-4422-952c-043a5e25126a · inbound

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts cites this paper.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.342135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.342135Z digest=sha256:b7ec6ebc752c14c50916b5b19536969d65c02c442dedd382d1c15585ca4c849b

Observation 3c7b937c-c260-4d00-b0f2-9190e4cf32e9 · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:21.474439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:21.474439Z digest=sha256:abb4b84111b7ab94bbd826f673964bd09bdfffa2cac0db2d626e4ca0d9ae1800