Pith. sign in

Paper Citation Record · LEDGER

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

As of 8 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 2 inbound Pith citation observations for arXiv:2506.16691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16691 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:42:13.979447Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T18:08:56.044278Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:31.831734Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9aee4c1c-04f1-4f91-87fa-21060a6a7576 · outbound

This paper cites GPT-4 Technical Report.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:11.467453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:11.467453Z digest=sha256:a997230938a002f8a87594d2217e595b6184a78bb0219ae5af4649b06cb84400

Observation 5ddc843a-52dd-4c47-bc13-381f8f0a0edb · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:11.568556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:11.568556Z digest=sha256:c99191d053fc73ac5ca798610029eacc84fb57a5f34af19932e8857ffd115568

Observation 02645f41-f1cb-41ec-b2a7-d7dea623420d · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Bottom-up and top-down attention for image captioning and visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:11.685462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:11.685462Z digest=sha256:aa674829adc53e3d50684f7587e3f13b64d7111546d58445c22dcc5febd64c0b

Observation 9dc02497-c659-4e8c-8485-0fe1dc005226 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.Claude-3 Model Card, 2024.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation The claude 3 model family: Opus, sonnet, haiku.Claude-3 Model Card, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:11.820401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:11.820401Z digest=sha256:a17204ffc1d058897a326614e665d4b2d779be26aaab719d510d22472039ba97

Observation a6e64210-2b0a-4bfd-980a-64efcfea572c · outbound

This paper cites Vqa: Visual question answering.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Vqa: Visual question answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:11.964755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:11.964755Z digest=sha256:dc8a4d5d4fd8c2c5bad156d255b66ad60c22067e01c18290eea3855475118cc4

Observation 2c3e6796-1489-41cb-9281-7828da2465a2 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.116980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.116980Z digest=sha256:88dff162dc8eda2a13dee6301fe205b9aba1ae6b992e09f5ab117c8f8745e3a6

Observation 88557c66-ee15-4964-a694-9fe9e21332f9 · outbound

This paper cites Layer Normalization.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Layer Normalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.195256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.195256Z digest=sha256:58956962e2810295423eb0acf13dac412c473a5dc13d3163e5d3e1ce4344abc3

Observation b14f4bd6-4b83-419e-a955-1f6fc8c0609e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.353821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.353821Z digest=sha256:ab77c379f8e4fb0d6942f7ddf628dfdd07e1a0c5c35f55f8fa55f8114857a90a

Observation 20014592-ddc7-4c9d-ac48-4e5e6e5bce6b · outbound

This paper cites Nltk: the natural language toolkit.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Nltk: the natural language toolkit

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.761772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:12.478544Z digest=sha256:da67541e5d2fbca4aa271281eb446ff9b2a2c862da3221965ce3dc50b9cb0f7e

Observation 82b5761f-a244-4e85-a8c3-af36bc21f00c · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.608375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.608375Z digest=sha256:1f789a4d1a7d23dba9da602448892d8e0fe38505b64dd71c6994a19229eec793

Observation 6c0cd4ff-faab-4b10-8712-113e8b02756c · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.745946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.745946Z digest=sha256:d3335547665c6322cbf4c80ac77cd70d78d342f13e34e2382307c3b875e2d2ac

Observation 7381ecb4-c8b7-408c-aeab-a652d13bdaa3 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.861061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.861061Z digest=sha256:dcdafb7bdb7e0461cebd867eac938fd02455e26985548b87c9925e5a98d05c86

Observation 86ed57a2-71bd-40c1-853d-80c33517ed53 · outbound

This paper cites Uniter: Universal image-text representation learning.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Uniter: Universal image-text representation learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.735153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:12.987283Z digest=sha256:6b420b5a9662ad652600f812832e19d5016fc71ab917846430bd3c514a9b8718

Observation 5bee69ee-7818-46a8-9df8-f73dd6b562f5 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101, 2024.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.104331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.104331Z digest=sha256:58089ced4887f74ea76c42436154dfb88d2828ede4f690e6bcb72f8b8feba787

Observation 462315c5-df46-4cd5-902b-3e7e4a06829e · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.115317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.115317Z digest=sha256:06f9e19e362fcb0f7e9fd94ce042daf632cf56ca516d1bc85046658c478ae704

Observation df06f042-e510-4e5b-9b2e-e0738b1bc496 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.221205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.221205Z digest=sha256:0f9589d9cb847de5496620a16281254cccaa6f25398836ea0a0e02e56a118db7

Observation 79fbf471-1e39-4852-ba00-21bc3a6d6044 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Gonzalez, Ion Stoica, and Eric P

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.312034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.312034Z digest=sha256:99e57d86eb9d0eb9a30f437c72ad6f5d6034d16922fd8341f55689c5602f69cf

Observation 0d6832b0-410d-4b1f-92d6-4cc10e3b1c2e · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.454102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.454102Z digest=sha256:dd70723b2d3da71263809bed33da1e4c81c61de60adff4fa2989743c7cd9ae3a

Observation 449027db-9c5c-4f6d-8032-641bcdf7cb38 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.573381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.573381Z digest=sha256:543fed65469bc93807d8225ace05c2724be3f7933ea1422098020ccde81eb787

Observation 13947f52-783b-454a-91fa-8130a6778368 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.643721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.643721Z digest=sha256:2139caa2d229d67cd08804825d12e32d3aadc36f60c9e21d47a74192cd053c58

Observation 1b2efe2f-aa52-4eb1-9131-c6e27d88536d · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.647604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.647604Z digest=sha256:29d7c1f6624ec44d8f110e46560659969b73fef084d7a9f791540d7ab12a5433

Observation 95f8fe7e-d105-4ad8-81f1-b0a2a69cf37a · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.651057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.651057Z digest=sha256:b6ca1859bda9eaf3a38d8bca02c2ed246f0b990731c5972b1abd832dc209a43a

Observation 7e0b6d1c-7c87-40ec-9329-9ffcb2d1820f · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Vizwiz grand challenge: Answering visual questions from blind people

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.659269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.659269Z digest=sha256:de648525564acc9eba323aa0505df372e35009c25ea01144b939cd7aa85e869c

Observation 78709f2e-cfab-4f01-bf8a-e7c4225d944f · outbound

This paper cites Measuring Massive Multitask Language Understanding.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Measuring Massive Multitask Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.664483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.664483Z digest=sha256:607340ebe2709b8cee8aaa476df03e8f24cccc7a0efeb97350e212ba5b6456fe

Observation 48cc5443-fa87-494f-b4f2-7827e6dc9681 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.668729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.668729Z digest=sha256:2a2a46fc9de626acec9cdccbf7adfed3038d99c5f11f679550c42b0872bfcd83

Observation ae2c4f91-7e96-4078-b5a6-3d1d269e129a · outbound

This paper cites Revisiting visual question answering baselines.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Revisiting visual question answering baselines

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.677213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:13.673567Z digest=sha256:cbab62a20b89ab561f2026569a2fec50630b893e5c16752d238d4e6743a823b3

Observation 12d634bf-4016-487f-8859-b9e3e94c7490 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.682523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.682523Z digest=sha256:823db28d69acf7888310ee3622ded850d837246d67a0d7bb46e1bc7413b670bb

Observation ce93ee01-fe91-472f-bcf2-f0a7db191003 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Seed-bench: Benchmarking multimodal large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.686298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.686298Z digest=sha256:3756862b6b63b13b2f5bb1aaef17518876d1e5af62efc86c29ea21041ec8e1dc

Observation 723cb2b3-701b-41ba-b3d1-7188ab3ff7da · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.691664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.691664Z digest=sha256:d506ced5f691dc9865ab57a87111dc2339fe54bdeeb826e5cc709bd3ee1ba5d8

Observation 30b5d714-c6a1-48e9-9f0c-1694aef87db8 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.704172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.704172Z digest=sha256:509419072720a79bf0ba210480d075fcbc63a6822b20e11ed2754e5cb2095ce4

Observation d4b2980e-3966-49a0-8db7-1f6893e8df60 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation.Advances in neural information processing systems, 34:9694–9705, 2021.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Align before fuse: Vision and language representation learning with momentum distillation.Advances in neural information processing systems, 34:9694–9705, 2021

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.707741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.707741Z digest=sha256:6de5c88f361d26c887ec8fffb247c970ca4d5298e9fb5df366cf7c1ec882913b

Observation 97785b96-6907-414c-b42a-18ef2769ef47 · outbound

This paper cites Videochat: Chat-centric video understanding, 2023.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Videochat: Chat-centric video understanding, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.611889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:13.711251Z digest=sha256:a4f884d22d9bb70d5a1ca123d90b13069617af0b6dea0cdd48829b6fce1a9361

Observation e2db0624-1235-4266-9b1c-cdaaacf9840a · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.717277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.717277Z digest=sha256:08bdbf1842d4543d91fbd2f94723f0c7b02f18ce7086d551ca5050125b2f502c

Observation b47862b3-ec9f-4319-b291-e06a96870b66 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llama-vid: An image is worth 2 tokens in large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.725584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.725584Z digest=sha256:e79bf8109f3cf48c040f5b19c7e75ed195050f39435308e05ef58189051422f0

Observation 8622997f-673b-4e3f-a8c5-87bf0c700d2f · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Evaluating Object Hallucination in Large Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.730319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.730319Z digest=sha256:bbecafd2e1288bf4331882edf791f402781ad907f6d7b08e05e9a03ea1fb28d6

Observation a231d075-a609-4905-b05a-9bab02935191 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.735467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.735467Z digest=sha256:01c3b3b743596be2a7500444f7971ee1f7d17dbc6b8d5f75e931c487932c91cd

Observation 5b8cf93e-0509-4b35-8c94-f8154c1d8458 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.740088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.740088Z digest=sha256:d6937fe0517367d4cd14834648fc098546e3c3fad5b25e4eb93f064ee218ece2

Observation ce0f5d7b-3421-45af-b0a9-5c6c2810d65c · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.746024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.746024Z digest=sha256:a443484df8250589748aa87b9e3095a64d8cf5a17be75a3a4924d0f2f642ad8b

Observation 019916fc-4086-4e9e-ba18-06309506835d · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Improved baselines with visual instruction tuning, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.754518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.754518Z digest=sha256:868f72d4cc2afeeca4a6f580fc4bde3acf8f8490ba25cae0d3640d8a76115093

Observation 403461a0-177e-4689-a44c-3525b9891f31 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.759294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.759294Z digest=sha256:aa04b5200616ee12d476568994161d4ba25e9654cfd4d6c9cceff4d5f5183220

Observation 297559bd-906f-46be-b760-ff7b5aa00b18 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.762862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.762862Z digest=sha256:68d6b9495bbc68b4fedf185d04dda1575d53b563d1138fe91db9ff92c8bcdf4d

Observation cdd77114-445b-4867-9373-2a5056bc7b38 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.770024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.770024Z digest=sha256:9b538e52f2b6a37ce9bf4955fd4e4997807f58d5ad0696f7c14fa6bb114d144b

Observation 8519c36c-eef2-429e-aa49-9d5a12ae840e · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.773647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.773647Z digest=sha256:b475488ebee5b69585b107d2db4571f8b917aeec47302de09c5359f6729e6dd1

Observation 67b001ae-295d-45cf-906a-cafeb2cb178e · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.778621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.778621Z digest=sha256:de200cd03cccad835db2f040548f1df125ec36965b7988c6bfa44e807186621c

Observation a59232dd-a6e0-4986-affb-a787d45ad1d7 · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.784958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.784958Z digest=sha256:06daac64a8c740d9c2c1ebb852f4fb0f3136cd21b0af8994b2150557c2ab1799

Observation c7bf4ee0-a720-4031-a320-aec3d66ec9fa · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.788990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.788990Z digest=sha256:767bfb9fecbe8af11f1f0a3f614d0a9aff0f017985324b45b1734aa780e46184

Observation af78bd7e-562d-4997-bc73-5b3ee4c2cb8d · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.793581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.793581Z digest=sha256:d7920ce9134da63d28c6912ed6d6db522b6d328f85a513708f6e12604580ebdc

Observation a4a41ed7-6599-47f3-bd10-b788e8688ad2 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.Meta AI Blog.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.Meta AI Blog

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.798585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.798585Z digest=sha256:76fd463f74e068d63178ba3fd1e7dba8d0504ca54fa036ac8830b6ca5fc6ca40

Observation e6d8f519-7a59-4d74-9593-14d47ac90e97 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Learning transferable visual models from natural language supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.801951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.801951Z digest=sha256:7ca8f22c921d20ff4367b8c36e7c48fe7887689c7b9d01cb35979cbbee05c6db

Observation 0fdf2af3-937f-4bc1-ac7c-302b208475d2 · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 2019.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Language models are unsupervised multitask learners.OpenAI blog, 2019

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.525704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:13.809976Z digest=sha256:aa85ce64bb6183332015b3a97d7e46d9267f16b6d9b1a0cab01ae1189b1b8245

Observation 82f9ab2e-4bd1-4a08-917c-5cecb7d27d47 · outbound

This paper cites Searching for Activation Functions.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Searching for Activation Functions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.819653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.819653Z digest=sha256:408937cb2c78867014da550f32e2edf619f5f9cb6f36471b5f9ddf276594fda2

Observation b7754035-5226-457d-a5d6-8fd86eda5e15 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.831606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.831606Z digest=sha256:235af35272423f52b40a3c8e5ec893cb74a14a9d3966fc8c2f42f2c548a92c36

Observation 18327225-f708-490a-a6ed-ddbf008e6d03 · outbound

This paper cites Towards vqa models that can read.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Towards vqa models that can read

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.836198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.836198Z digest=sha256:55d7d8ff1d5fd1e9d017acfa4ca6a0dfc786615b9ec295d32742dd9031bd4ba3

Observation 63618e8b-195e-4ee8-969e-b82cdfb805f5 · outbound

This paper cites Flops profiler, 2025.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Flops profiler, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.505878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:13.843580Z digest=sha256:49b14a51142921251a805a23cc4b6dd337dde8ba7d077c5d1beb13fac26ded93

Observation 578a6673-3b6e-455d-8f3a-493a78cf7456 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Gemini: A Family of Highly Capable Multimodal Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.848825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.848825Z digest=sha256:701656f7a3f7b4fd069c294baa7b69daf6b13d28b940b63f0b6bad3e4ab3e9ca

Observation bb38303f-ae63-4e72-844c-9059286c1dbc · outbound

This paper cites Mlp- mixer: An all-mlp architecture for vision.Advances in neural information processing systems, 34:24261–24272, 2021.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mlp- mixer: An all-mlp architecture for vision.Advances in neural information processing systems, 34:24261–24272, 2021

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.852755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.852755Z digest=sha256:43cd85eb7ee3a7697700f0ed9c6e16978f1cef9c023bf19eb378bbd2fb32dad2

Observation ab8e50e7-81d8-4492-a41f-7c3711c74420 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.863089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.863089Z digest=sha256:b032a2d4081f02c65cd49427337606279d9457cf5241592b033fa0ed3e0a2b42

Observation 2debb5ab-3107-4b9d-bc1a-d9f60141ed15 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.878802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.878802Z digest=sha256:3bd170bd10df7491b2513fad8474b30501b0a142a80b1e51d2b9c7707ff722e2

Observation e808306e-e95a-40df-a2bf-51795d5012aa · outbound

This paper cites Patches Are All You Need?.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Patches Are All You Need?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.886946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.886946Z digest=sha256:841a458a34376f9f9770cec32c2a1dc39173971908452998f8e8b85f4b3965a9

Observation a476ad82-1e46-458e-a4f5-f3e88a671a9e · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.890551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.890551Z digest=sha256:cc23819899758bf9fb992b953b83c7ac0516e7efd1df15330c66138e90ea2a57

Observation fdea2b54-0b7f-49e1-965a-c3998b9c5e04 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.894002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.894002Z digest=sha256:08d19587501f3d0f6e1cae3ef85f0b25f7e5d290f9acf1f631041e948c75d291

Observation 556e51c2-2ca4-4449-9009-2ac112aa07e0 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Cogvlm: Visual expert for pretrained language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.475667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:13.903098Z digest=sha256:df8e7ed9c36a1399f82827e202b7e80835c8dec00fd9f8143ae2e4b7ad42587e

Observation 4716d017-f258-4481-801d-4b9bdde49b55 · outbound

This paper cites Qwen2 Technical Report.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Qwen2 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.912263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.912263Z digest=sha256:4e1d87d3d0c33b1a9c8fb281e72bb36df77d36849b9b71a380dba1bc48c44305

Observation 90372d15-99a1-4b77-8d32-c05862fa7573 · outbound

This paper cites mplug-owl3: Towards long image-sequence understanding in multi-modal large language models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation mplug-owl3: Towards long image-sequence understanding in multi-modal large language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.464007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:13.916176Z digest=sha256:6bc4605385371982db501891698e33ecb65b79bbb1190402910b0aa759a44b1f

Observation 1843186c-721c-4ef5-9ade-c338527de335 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.451218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:13.920019Z digest=sha256:69889d22c1f8561f501a8d52428103fde9b6b4c82135c10e176bf49b9a9101ec

Observation 9467c2b3-921c-48c6-a38b-87433cedbbc4 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.923332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.923332Z digest=sha256:0b3ae6a185f3eec2d909fb352f9c091cad49e09a8b042772c4c7b074b00fe598

Observation d2685d41-a199-451c-b56c-5dd04ca1235d · outbound

This paper cites Sigmoid loss for language image pre-training.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Sigmoid loss for language image pre-training

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.930971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.930971Z digest=sha256:75cb0218d065c023bdfdac7a987de4dfca95ba698133e21bc8c9c36842ca703b

Observation 601d00e5-a49a-4575-8f25-1aa2ee7a768e · outbound

This paper cites Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.947456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.947456Z digest=sha256:4b3db092baf01651268cba127c1df5061cd9a8b85326c4ec9c8726458d986d42

Observation de1d6570-519c-4776-876d-ecf848bd4c48 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.955406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.955406Z digest=sha256:e96720088535da9bb834f8e420190a9e5c67f783c58c49f02ff4a17598c1e14d

Observation ed01f621-792c-4699-afb1-f54e0c1f7f6d · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.967404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.967404Z digest=sha256:09401132a878ddcc6ab86b78e01248a1d406cab128931ca78b5d96c38e800ed2

Observation 35bc73c8-054a-4132-919a-d3a9d7e7a3e6 · outbound

This paper cites Long Context Transfer from Language to Vision.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Long Context Transfer from Language to Vision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.971332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.971332Z digest=sha256:a21941c10843fbca6036a29b2707c1dc0d334eacfdb45daa0b0770d52cdf74f5

Observation 149c2511-5fae-4c5e-9958-47f2b52cd5fb · outbound

This paper cites Wings: Learning Multimodal LLMs without Text-only Forgetting.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Wings: Learning Multimodal LLMs without Text-only Forgetting

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:42:14.042993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:13.975077Z digest=sha256:b8cd6186ab175088322a8e8a993a8f08bc4f9c83012a373923faa12fc48a866f

Observation 7abdeaac-5cb1-4a79-9983-d83daf5b6c7d · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation MLVU: Benchmarking Multi-task Long Video Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.979447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.979447Z digest=sha256:562a06d385897e1851307c997dcca4df54d07657fa39a6db52378a80e3fef6d2

Pith citing papers

Observation 29312d64-c8fe-498a-bfda-894e49983ae1 · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.181799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:6e2fffe56ee4ea0c71bcbd51306f7f76b70122357a41c77c6a9533753753192f

Observation a6862ffb-5dcc-4899-848e-b4b0aa1d8cdb · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:19:31.833880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:117fb1012d582508af1e79a87de7a35db98fb5347c58abc924a2ad2210256253