Pith. sign in

Paper Citation Record · LEDGER

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

As of 14 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 10 inbound Pith citation observations for arXiv:2501.04322.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04322 v2

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:42:12.664774Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:20:32.873352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:37.114920Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved76
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d451084-42cc-4037-9506-b6de5e913866 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.382246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.382246Z digest=sha256:1f7ecadf774735bdad2e14ba588fc6468b89680550fc2a7d413bbd335888da87

Observation bf5c3248-f3e2-4ea4-af2b-df9bd936d500 · outbound

This paper cites write newline.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.387018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.387018Z digest=sha256:9ffa3d230b1cadc54bf1eb851a6158c795474f39b460eb5762275d0769f0ecb5

Observation 4bf881bc-0e99-4caf-bb61-9adde70df0d8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.391244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.391244Z digest=sha256:6482d930b10c41d6cad7e4a7d11df81614eeaf43978855371a8f39944c564d8f

Observation fc827713-8028-4d87-be55-0c7700e615c6 · outbound

This paper cites VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.395402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.395402Z digest=sha256:ed507dbcb9462fa6807f6041eb777e66fadc999283af8bef4667cd4b273c3553

Observation a7c42f80-daee-48be-a582-c3d128abda0a · outbound

This paper cites Stable LM 2 1.6B Technical Report.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Stable LM 2 1.6B Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.399199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.399199Z digest=sha256:0b72e14881caf6c180b228feb65d4aff22448d6baabf8a3f72c2df98e4fbbeaf

Observation 528ffb5e-b3de-46df-bebd-62c3cbae8d91 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.402955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.402955Z digest=sha256:8df4845fee9585462c7c096488f06d95bfbbf0367d9671ffc00f20fb27c740f9

Observation a7885e44-5241-40fd-943e-981a12235f99 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.354011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.406540Z digest=sha256:8f1f320e9d217c22887aef435db9464f623cfd8b47f878bcc6de754cec57c203

Observation 9ba9f3fd-6113-4c60-8bf7-7f067e8d82c6 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.409904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.409904Z digest=sha256:167203d00948ab20c36f34c0e2fe473c07bd48a488b36dd8f0767014e463c99e

Observation 997d9897-a14b-4941-969f-f27c0cfdddba · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.413888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.413888Z digest=sha256:f239327cbab9af9b185d91f2f16c2ca85631f0399ef11f7ed5c026ed20701ccf

Observation 9530c575-dd45-4a9b-ab18-c2c1da59390f · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.417609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.417609Z digest=sha256:c9d5cf95b777b5f371b1a33b596651d570985fa9ce57c6c193e44547bbde99ca

Observation a901c482-064d-4f32-9757-53c9a8d624b0 · outbound

This paper cites Improved Baselines with Momentum Contrastive Learning.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Improved Baselines with Momentum Contrastive Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.421641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.421641Z digest=sha256:25b29b814f82d242bb23800fa8e8af0ec5962db2fcddca445557278a2f9786f4

Observation 4470e77a-389d-44db-b612-8055e91ebebc · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.425313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.425313Z digest=sha256:a248bd9d2665b058f2c681ad129cabc95d75f61fde9812cf7043d046d2a28773

Observation 86c3582a-f8f6-4936-9dc4-3d30d16c36f5 · outbound

This paper cites E.; et al.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts E.; et al

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.429121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.429121Z digest=sha256:9ceadf2259ca715ca7e0c4f7008540c972f7f1e19bcc79379c6ecaac1ec204cf

Observation a5ae5a6c-25ed-4757-b29b-542d4ef1dd2b · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.432690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.432690Z digest=sha256:3483004aa704fcb37b4e74a5d90baf935faceac45a559c9d478e68026974a4fd

Observation b88f286f-57d3-40a1-8fdd-ad67ddade48e · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.436308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.436308Z digest=sha256:fd073fcf49aaec6bb99d9be304719c5174c7160fb87f37b562028d9aa10b9f46

Observation caf30f46-7b69-4c04-a4b5-4d0d4cc46265 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.439984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.439984Z digest=sha256:2198e9a8be06013b4cdf081e4acbf397743d7c8a32d9196807b8c8d297d4ac2b

Observation 700667b4-e726-457f-a950-2d447c765204 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.443838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.443838Z digest=sha256:d264efe709f69154dee9aea7ad0d0f152930460caa79164dfbd4a3be51b85415

Observation 0441bc01-05c1-4d5a-994f-afa660cf4869 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.447197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.447197Z digest=sha256:ef4cec09f6f78fa81e2bea4deaa637b38db126bc8edac5c7e4e203a631cf0cae

Observation 76c5eea4-2ad9-40eb-b291-df0bfe4efdc6 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.450585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.450585Z digest=sha256:8656fe85746511d0a72a88bb836a4bf6026e55ae517d73725ac1b3b2643e443a

Observation 2cacafbb-0278-4263-a8ef-92ebf4c6dd73 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.454034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.454034Z digest=sha256:44e6a1002452872f3588c2e9950b375c76a92aa5c7f13c4b45796416d3e16e8e

Observation 37d4c636-ccab-4782-a2c4-7578d70dd9cc · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.457573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.457573Z digest=sha256:444c2a95a8824bc28817c18c4b7d18ed983bb09c710358040bd9a48363149680

Observation b2f8134f-0f81-4bca-88ab-a1eb0cb86a07 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.461362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.461362Z digest=sha256:75cb89d4e70add8039613a57835a39277efea59753673002f96cff47a4dd5d88

Observation abbbb94c-7da2-4203-939c-82dd5abcbb8c · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Measuring Massive Multitask Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.465022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.465022Z digest=sha256:0080d7f31f12b8cc8f5c7e14e6c2012742bb215d56a479e9c6df12a38450db6c

Observation ad1b4f3d-cac2-4f07-a91c-34249e7e90b2 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.468709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.468709Z digest=sha256:d6334a6effa2151d1456b6d7c4f7654b14e72e86462e0003207fc6d7674eab0b

Observation a4e5d319-2795-4d7e-9714-730c70a1af34 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.318039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.472621Z digest=sha256:ea6fb4605b2ae77e94561c95cff7c61b68f8885b596ce35ba7e0f1e2d8477c7d

Observation f9c7fd53-f26e-4f6c-9628-16ce04ee39a6 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.307006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.475913Z digest=sha256:304023fc276552ee71e9a893141a24875cc621b40f897fb84c7dfe3bc5792597

Observation 7adc87b0-0f60-48d8-888b-83f38c188b61 · outbound

This paper cites A.; and Manning, C.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts A.; and Manning, C

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.479477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.479477Z digest=sha256:b31fd76bde2842c7e5934797079ad92820563c67c82c429676a2ca64804895d2

Observation f8d0eac5-fa87-4c8c-9a9a-e77baf23af2d · outbound

This paper cites A.; Jordan, M.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts A.; Jordan, M

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:13.290654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.482825Z digest=sha256:4270eb4acce3a7f39f7a53df69692257b08e879e0ae3e471e173aa50ef7c9925

Observation 5bd11329-9335-4a22-b030-4b83e987769e · outbound

This paper cites Mixtral of Experts.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Mixtral of Experts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.486505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.486505Z digest=sha256:6f55f1da4eb398a29379880506abddfc2b71d9080e1c7d6cdc6317a5c45dfa44

Observation 19e2f290-ecaa-496b-8802-6554e1b91693 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.280387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.490016Z digest=sha256:8e0d27d245ed574cc7c5d8d04fd5cb7301ed9ee554bedff53594eb598f788e4d

Observation 7e0c649d-f098-4509-a60b-38a6120a5777 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.269815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.493534Z digest=sha256:6ca19500b7a1852804d36f38ff23ecc4946acd106d5d8fb32d53866014fc4d6b

Observation 7976cf8b-2fc9-45a8-bc16-fe823f9812cc · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.258486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.497177Z digest=sha256:9b14df9a2086368a23fceab1ae69f6f2c3d97d0b98c863268c8f63bf1975c574

Observation 885830d2-d653-44b0-a230-d18c1a682fba · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts CMMLU: Measuring massive multitask language understanding in Chinese

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.500975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.500975Z digest=sha256:8d7e3fd9ff7a9979019661b97659fbf248f4b583ac2d5502da81cc3f0450c234

Observation 3ed641db-35ae-4c7d-9311-d916a6dde100 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.504611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.504611Z digest=sha256:b6da3c5b91bad0817b3788e513e4c8ad677f265eeaf2d4be58bee15245629c62

Observation bd1231c0-5bab-4062-a9a6-4704c070bda8 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.508338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.508338Z digest=sha256:0780edd5c03f90fa9f21b9ab6f5c709118d2f27b9b5609c09f4cb908037982ef

Observation 07f54a43-9a51-491b-9b86-305f3d3c0dd2 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Evaluating Object Hallucination in Large Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.512239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.512239Z digest=sha256:1838179931e186d53896ea4bdce085a0f56e2ef6c51b9166975232c6e35f280d

Observation 1cfb2fd5-93dd-4eff-8f0a-6d59db422156 · outbound

This paper cites PaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional Experts.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts PaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional Experts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.516256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.516256Z digest=sha256:a9c8cc71359623fab64cb5b6d5a9cb90df90bbe30f98513af3b7232d66ecb4dc

Observation 0fcfa30e-461a-456f-80cb-a501101b0056 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.520221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.520221Z digest=sha256:b85316c98ea834afefd2e2288c5190a8ad8e6ac5c4ecdc62b76dd68e1023b2c9

Observation 659cb86b-cd8d-4705-93a9-be3602a48f78 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.524131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.524131Z digest=sha256:4d6b6039e955d11a5cb7870f70b4376233e856ccadc79d5e1fd3aceea77fb279

Observation 877c007d-f680-4b9a-a2be-3279482ffce6 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.527975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.527975Z digest=sha256:d66e51e70041b0e936366293cad784653255d387e3647632a75c0714e65711c9

Observation d78fc038-9e6d-4f6c-ae73-d089b9a8eebb · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Improved Baselines with Visual Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.531526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.531526Z digest=sha256:726f16d47bc6d9925925cb859a2a47d3b53e136b2f449bec8a96fde65a9a25a6

Observation 478b0675-ae22-4741-8e1c-dbf4ba481a6e · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.241708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.534957Z digest=sha256:41a38c9bf0f0818c2b806af211271d90bd90ad875636750e20a7af8618437511

Observation ae1873cc-32ff-43e8-b359-f5317a84774f · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.231186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.538458Z digest=sha256:c0fec27c19f3ad1fa765aefd92bdde3dea9728bb25c7d9ad072b4b53beda3b45

Observation e243fd5d-125b-4c12-a862-c734d38cadd9 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.220323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.542000Z digest=sha256:3e07f2e2f88b839e765d5603afdaa22bd8d9224462f9473250b2198e64a055a5

Observation 863ae401-a94c-4bbb-a37e-8cc954b8a695 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts MMBench: Is Your Multi-modal Model an All-around Player?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.545211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.545211Z digest=sha256:4caee38ddf082ea73c884bb6205797fb3cc652225771fa68c820642171b3651f

Observation 405653a7-573a-4a1d-b89b-fab4481abda2 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.209257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.548742Z digest=sha256:6545a203bc565263139b0da0116551530365d73dd175daa0543f49f07eed6da6

Observation 0a5a2dd0-3111-45e4-bb5d-bff582f7af49 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.552134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.552134Z digest=sha256:2fd3e1b3f1cafd0d5c0cbf3aafad4d48c69233525cecbedb7b0b61c6260a13a4

Observation c73def13-7609-4afc-955e-af8eadc2a16c · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.555718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.555718Z digest=sha256:2ca340250dd12a25d5afe3189d492475b856b983a6656a010e282c05ef41da10

Observation cdb5bf3e-7504-4e92-a1d0-eb1054360fc5 · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.559437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.559437Z digest=sha256:f05bc5430404795f85a7519e775cf5e45bd60a3cdc024979b72802c669fd6e57

Observation af157eb4-a97c-49ed-b00a-ced4d7fd18f6 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.563030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.563030Z digest=sha256:8f169cb97e42b568ec1c054f4972a0907739f725befc60d9c56c47e1ad9c3287

Observation 28e832ca-d096-41d9-99da-680809a141b6 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.566224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.566224Z digest=sha256:eb594743397232dc1e9147d851dff94324051bff9ac4a80f3ccb0934b8ac9667

Observation e9325067-30ad-4a22-9d42-bbda19a4be4d · outbound

This paper cites Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.569949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.569949Z digest=sha256:8d575042435ceb146cc1ee55b631bb8b897cae0f287c30e58f89a63bd9998d97

Observation 9325ae1c-917e-4426-a23c-fe7931bb8bfe · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.573669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.573669Z digest=sha256:d997996b327f47caf70e6a971778376d9290cb141227f01e341bb9f54eb44d33

Observation a3b5a508-74b2-47e8-a999-8b76bb15d4f6 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Learning Transferable Visual Models From Natural Language Supervision

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.577043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.577043Z digest=sha256:32783828797a4dfbc12e0f2e6f71aa1dacf6def8bc5a0af2ff860e4d65781a9c

Observation dd055cde-c53e-4c06-a375-28b75faa5448 · outbound

This paper cites ImageNet-21K Pretraining for the Masses.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts ImageNet-21K Pretraining for the Masses

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.580841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.580841Z digest=sha256:d40ef12d317272f24247f7e07878c565a73559326f35f24a9dc80b1c04766880

Observation 6223f476-e130-40ef-91c3-5d91b9e0a66e · outbound

This paper cites Scaling Vision with Sparse Mixture of Experts.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Scaling Vision with Sparse Mixture of Experts

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.584558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.584558Z digest=sha256:71671e5ec9fe5def8947d95228048dd42294ff00d400efb15025234d7014c6e4

Observation 019d9510-af28-4f28-a226-627b9d1d6097 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.588460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.588460Z digest=sha256:933a7635af6ce8c36796cab2660d31c7ae5967aaab963750eedd43a308433ca4

Observation dea6e367-469f-419d-a28a-4e382e86d29c · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.592020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.592020Z digest=sha256:e90abd183ab58ba4c3ff34240d0a634b7f88c7843494cede6f8f2e7295f7edc5

Observation 42a7f07d-0e9c-47e6-9eb2-27e3fb6c1b78 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.595562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.595562Z digest=sha256:cc2d9954e7b8978a7bd13c978fa07aa1e97eaead837fcda5bb1fcf1e6e8e493f

Observation 85e20666-c225-4740-8fe5-149ab2ebc6c8 · outbound

This paper cites Measuring Vision-Language STEM Skills of Neural Models.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Measuring Vision-Language STEM Skills of Neural Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.598921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.598921Z digest=sha256:4c59f64cc056bf99ae04d8dc50ce9a4d80d0eb4d4c2d5cb29552b392e64b0311

Observation 783102d4-6863-417c-8daf-09a3d388a22d · outbound

This paper cites Scaling Vision-Language Models with Sparse Mixture of Experts.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.602415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.602415Z digest=sha256:26f1b19e5044820df718871880dc3083a6debd0bd1a6902eece23fa3bd1d9478

Observation 8a9f6e9c-2583-42b5-b382-e50da34bf5c9 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.168228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.606068Z digest=sha256:e83ff2680dd50d2ed8edfb41b2dd20a1ea5ee593246e4c8ab1b0a257ccad4713

Observation 7755448d-70d6-43a7-9964-dc949058f22b · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.609286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.609286Z digest=sha256:c119752c495bc25fda53963bcb2cdf98d36adb48fe2a88b2bb51fed7a942cdd6

Observation 54600042-a95f-4061-b0ba-8bf77a3f1ccf · outbound

This paper cites PanGu-$\pi$ Pro:Rethinking Optimization and Architecture for Tiny Language Models.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts PanGu-$\pi$ Pro:Rethinking Optimization and Architecture for Tiny Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.612551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.612551Z digest=sha256:b119a1edba13eaf15c4b61f710755c389c47da1819847d76d5445f5c02849d55

Observation 894fbb6d-f8f7-41e5-80bc-d53241a1c6e1 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Gemini: A Family of Highly Capable Multimodal Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.616503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.616503Z digest=sha256:16de157ccd944ca6c32764a53e7629b603ee2d5c6623f407c9e02439dd611597

Observation 53edb665-371c-4ef7-9f56-ed1a9a3a3097 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.152004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.620355Z digest=sha256:c540cfb093e9e9b2017555ebf60ba119d546de87f78e4be97de471cbd1d3be3a

Observation 68af8662-981f-4c6b-84f2-bfef7d9c9d43 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.623890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.623890Z digest=sha256:7a23390e5d4a9534c21e0cd6b403359681dfdc4ca0b501ff57b2d73529bb4a49

Observation 675b24f2-a236-4e59-a8b9-61864c16132e · outbound

This paper cites PanGu-$\pi$: Enhancing Language Model Architectures via Nonlinearity Compensation.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts PanGu-$\pi$: Enhancing Language Model Architectures via Nonlinearity Compensation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.627697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.627697Z digest=sha256:2e7fca676d0079b8ad3403fb021682e20da90909c1fe01fc4edb2923d2dc3363

Observation eef4fe56-a567-401d-bc58-8049c12cbdf8 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.631162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.631162Z digest=sha256:d377b3463f9bc3c5c4e382a482c15ad147daea05cfc78c734824d247a75ec4f2

Observation 049351a9-5f5d-4fd3-94a3-e7519a1f4bd9 · outbound

This paper cites FewCLUE: A Chinese Few-shot Learning Evaluation Benchmark.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts FewCLUE: A Chinese Few-shot Learning Evaluation Benchmark

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.634761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.634761Z digest=sha256:1cd3f658f2d8346e2155eca3839771294d315c237d0110c4bb062d7538e91c07

Observation e5e26d06-e49c-4bed-9787-f7b0f29af39f · outbound

This paper cites Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.638612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.638612Z digest=sha256:0a3b364024f713a7ab78fe8b85b1de917e50e251ff8ae1d3cd5054143b29a53f

Observation 6fed9b6c-9c23-4ad5-b1ee-45303176c4db · outbound

This paper cites TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.645932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.645932Z digest=sha256:b6839921cb491cd1385c0f31286ba4c07e71d43aefd59a4d29063cea49960a75

Observation dd87f1b0-543f-4db9-acde-2b3d6eb22148 · outbound

This paper cites an unresolved cited work.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:42:13.135226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T21:42:12.649426Z digest=sha256:f2186b2db1dc6d3d03fa1d143fbc5ed9af220c10322a6d0e46b61b8621570855

Observation 649883b4-5106-42e3-96fd-ab2226a01f29 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.653313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.653313Z digest=sha256:156403309faa64be0324309cf527cbc78022928f51c000646905f64b27fcdfe7

Observation bbe55a0e-8a98-4d18-b655-69014a31488a · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts SVIT: Scaling up Visual Instruction Tuning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.656981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.656981Z digest=sha256:a2c7ce857ac99ab12b93480e729141ddbed20d574ba63ecd6b0bd00a00b140fd

Observation b8e64968-c8ed-4736-8c20-d820a87b9e5a · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.661051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.661051Z digest=sha256:8921abdc4560613446ce032cb7db9ba7351c58da1c43dde3c2f91b101409339a

Observation b47da43e-68bb-4219-90ff-0a9e8e9856fe · outbound

This paper cites LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.664774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.664774Z digest=sha256:e7b0b249911689b221af52fb4d2b8ea5d3851e276713c499d57529f8067ca2db

Pith citing papers

Observation 32dfa201-5a44-4629-935a-341775b2f928 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.527639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:f739bd0f1082186005bf1349e1fcef2748c67c6f1ae2b190d94abf75b61701d1

Observation 4a5a7488-3a23-49d7-8382-328a111d4aa5 · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.710173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:50320503ca7b3df4946a5107ea232a76173011c392eba6c93a6260a3dd796914

Observation cf10a8ec-9ea4-4941-9797-345210f4ae62 · inbound

EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models cites this paper.

EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:32.873352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:32.873352Z digest=sha256:de528b9124bf92748ca48e0a806e4bc6e76991c98d0d5b85c78b5b6455841d3e

Observation 879f13ae-d0b9-4b5c-a844-38f843aaf72a · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.199482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.199482Z digest=sha256:b7c9efbade37dc5a046b8774f95be08f3272c460a6e180b36483bfb0d8b65910

Observation b049d93d-4725-4e24-8888-2c87fa20c840 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.205836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.205836Z digest=sha256:ffa07d63bc15602035fe31f5acf2ebb99fec7d6e2204b8816f9db6020b2bcfc9

Observation 16915555-55e8-4406-84dc-2fad6b5f7017 · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:53.436448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:53.436448Z digest=sha256:9e269121bae19f7d9a3a68fe4af5a13ee4e97b2db07a2d9e904e63edaf247d01

Observation 7305e160-f785-41c6-8060-27508a49775c · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.633432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.633432Z digest=sha256:ff91e18cb16cd06d977bc1d85afc02c6c24778348f34ecba026bec7ac615b518

Observation 2ba8ed5f-3bb5-4fc0-a0b3-d68040702d1a · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.886712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.886712Z digest=sha256:9211447f4b08b1372ae4071ae532c271f93b19c011db09922dea9e3756e0e31a

Observation 37685f8e-c4ee-4ad8-a34f-d2baf8251d9c · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.166141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:7041f79e3124bc15fb2a0fc464429e6d0195136c26e29b10d37fa4bc8897bf33

Observation dc1f5f79-6025-4147-93c1-4f3c8fc73760 · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:37.116550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:e24e19dc1d86052919eee30e400c0429d0c11cfd343d1da8afc679f662a21600