Pith. sign in

Paper Citation Record · LEDGER

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

As of 20 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 2 inbound Pith citation observations for arXiv:2506.16691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16691 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:42:13.979447Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T18:08:56.044278Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:31.831734Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9aee4c1c-04f1-4f91-87fa-21060a6a7576 · outbound

This paper cites GPT-4 Technical Report.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:11.467453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:11.467453Z digest=sha256:becc8f5a3a9a097b95d5180a91164cfe09eb6239931a8c7a7efb04a0af62c421

Observation 5ddc843a-52dd-4c47-bc13-381f8f0a0edb · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:11.568556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:11.568556Z digest=sha256:c2df859b635674908791decfb67c354a44c19c18ab9f27d73d31943a959edace

Observation 02645f41-f1cb-41ec-b2a7-d7dea623420d · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Bottom-up and top-down attention for image captioning and visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:11.685462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:11.685462Z digest=sha256:12fab33a45567f10a9ed77408960613714850b3cdcc1fef4eaa88f09b60ac6f9

Observation 9dc02497-c659-4e8c-8485-0fe1dc005226 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.Claude-3 Model Card, 2024.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation The claude 3 model family: Opus, sonnet, haiku.Claude-3 Model Card, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:11.820401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:11.820401Z digest=sha256:fd1991d2821efd98e87c5d2c900b8174ac5a4c81851d7c31d21f7007e2b68165

Observation a6e64210-2b0a-4bfd-980a-64efcfea572c · outbound

This paper cites Vqa: Visual question answering.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Vqa: Visual question answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:11.964755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:11.964755Z digest=sha256:49919c19dc356d98cb7099b7e181bd06f9043820a43ead7787e70576e0a7c3d9

Observation 2c3e6796-1489-41cb-9281-7828da2465a2 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.116980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.116980Z digest=sha256:d382e3103954fbc34629ad704ed1313bd04ddee578eb0c97f921512d58b056b1

Observation 88557c66-ee15-4964-a694-9fe9e21332f9 · outbound

This paper cites Layer Normalization.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Layer Normalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.195256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.195256Z digest=sha256:a35557a18069b16683a1f4454d5d13c387fa45ff16cc24d46fde90357eb92cd7

Observation b14f4bd6-4b83-419e-a955-1f6fc8c0609e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.353821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.353821Z digest=sha256:1c9a779cc078920bd62c0fa6ebba4796679994c0c800c17c37a8adbf466e25a1

Observation 20014592-ddc7-4c9d-ac48-4e5e6e5bce6b · outbound

This paper cites Nltk: the natural language toolkit.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Nltk: the natural language toolkit

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.761772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:42:12.478544Z digest=sha256:cb85d51641017a7c87c0748406e47171b0f6bcfdefe73e0ac3e92274af91d7d1

Observation 82b5761f-a244-4e85-a8c3-af36bc21f00c · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.608375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.608375Z digest=sha256:471f748520949fe28174eab5c80e7d2d49f022fe66337b6e65bc2e947ef37e3d

Observation 6c0cd4ff-faab-4b10-8712-113e8b02756c · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.745946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.745946Z digest=sha256:1419eb23da1a5f50af4c6fbb6bbcdc12bb633ac7bfede1b7b202c112ba55780b

Observation 7381ecb4-c8b7-408c-aeab-a652d13bdaa3 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.861061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.861061Z digest=sha256:aae645dd2f44d6834d2db8ab57df023ea88ce586e3d09c551c2b1ac029a8f75e

Observation 86ed57a2-71bd-40c1-853d-80c33517ed53 · outbound

This paper cites Uniter: Universal image-text representation learning.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Uniter: Universal image-text representation learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.735153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:42:12.987283Z digest=sha256:c49c800b6d4d4a262adeb6b7adcf8a01b708d3ee9cdea13a26c32a77fe0ab5bc

Observation 5bee69ee-7818-46a8-9df8-f73dd6b562f5 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101, 2024.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.104331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.104331Z digest=sha256:5c205adcf494be7312f15ae3f165f73416f96a11bc544ec42e82987fcd819e55

Observation 462315c5-df46-4cd5-902b-3e7e4a06829e · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.115317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.115317Z digest=sha256:c0dce885cc7658b35afadd1c6cfc7ea0800f637dc8cc3ecf86981c971336f565

Observation df06f042-e510-4e5b-9b2e-e0738b1bc496 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.221205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.221205Z digest=sha256:fa7cdc074d906a4cbc9c1004652f99d443c2c5bdce5e050cb7d01dd12d8ecb50

Observation 79fbf471-1e39-4852-ba00-21bc3a6d6044 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Gonzalez, Ion Stoica, and Eric P

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.312034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.312034Z digest=sha256:d4b21158903e332ff92404368eacc6087e115f6cc529dba131af5f07476bd622

Observation 0d6832b0-410d-4b1f-92d6-4cc10e3b1c2e · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.454102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.454102Z digest=sha256:83ab0abf32c05c0aa23c0e41fde201802eff34f1718e4a85a95def4c1f451f6e

Observation 449027db-9c5c-4f6d-8032-641bcdf7cb38 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.573381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.573381Z digest=sha256:d15dad0658e0c46dcb2fd98a1f30d3bac089cae11bb48feae0127cd96aed0294

Observation 13947f52-783b-454a-91fa-8130a6778368 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.643721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.643721Z digest=sha256:5f80688cd3df41d2580c7642f7f806c3e6bc6d61f3b64e506ac2c2cafa335907

Observation 1b2efe2f-aa52-4eb1-9131-c6e27d88536d · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.647604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.647604Z digest=sha256:354c2be9b78d02575352c717a23048449a590d5e0e6761e3e2fbd7845dcff703

Observation 95f8fe7e-d105-4ad8-81f1-b0a2a69cf37a · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.651057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.651057Z digest=sha256:1e1715e1c1a1baf59904082ce37d275024c506fa226ac5c24bdb183eb979aede

Observation 7e0b6d1c-7c87-40ec-9329-9ffcb2d1820f · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Vizwiz grand challenge: Answering visual questions from blind people

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.659269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.659269Z digest=sha256:803564dd6842f6a27b3fba07e3c0df4ed8da2c4b56ab5f46fc195362be9a3bee

Observation 78709f2e-cfab-4f01-bf8a-e7c4225d944f · outbound

This paper cites Measuring Massive Multitask Language Understanding.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Measuring Massive Multitask Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.664483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.664483Z digest=sha256:1449274faade32dda8bfd29d83f88d88603e1cd2de3f463fc080845540009dca

Observation 48cc5443-fa87-494f-b4f2-7827e6dc9681 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.668729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.668729Z digest=sha256:cb48dd06a4ce2ab115580cd70b7057c7f1a6200a6c610ce1a4d704ca444ebd49

Observation ae2c4f91-7e96-4078-b5a6-3d1d269e129a · outbound

This paper cites Revisiting visual question answering baselines.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Revisiting visual question answering baselines

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.677213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:42:13.673567Z digest=sha256:f8cd56234c8b9a2e6cd91d9b586bb23bfd67603bcf62764092ccc69955e8803b

Observation 12d634bf-4016-487f-8859-b9e3e94c7490 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.682523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.682523Z digest=sha256:4e58027aa4494b2ecdaef4b7fe232c88a899f6e5cb164ba3d5ee0e93b95c6849

Observation ce93ee01-fe91-472f-bcf2-f0a7db191003 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Seed-bench: Benchmarking multimodal large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.686298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.686298Z digest=sha256:18e46726a2cad8a179d7bf7ab1da8a2447a04017fd59b7506de7d1f73c05818e

Observation 723cb2b3-701b-41ba-b3d1-7188ab3ff7da · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.691664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.691664Z digest=sha256:4f2dcb03c2dda56e94fa29552a5c106d13514349f769e5ede5bc4d0c6d246a1b

Observation 30b5d714-c6a1-48e9-9f0c-1694aef87db8 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.704172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.704172Z digest=sha256:b77e082f46726e75c206d514af0d4acd9810b6beafe006ab6f0304859518467a

Observation d4b2980e-3966-49a0-8db7-1f6893e8df60 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation.Advances in neural information processing systems, 34:9694–9705, 2021.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Align before fuse: Vision and language representation learning with momentum distillation.Advances in neural information processing systems, 34:9694–9705, 2021

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.707741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.707741Z digest=sha256:c5b83e551582f2a06f198d836040a39baa16eb2df38a17b7198666b1cc30ba91

Observation 97785b96-6907-414c-b42a-18ef2769ef47 · outbound

This paper cites Videochat: Chat-centric video understanding, 2023.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Videochat: Chat-centric video understanding, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.611889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:42:13.711251Z digest=sha256:d732048ea9cf190785830fd1c36376aeb8b0ed104625a7a1dca619be2aaca675

Observation e2db0624-1235-4266-9b1c-cdaaacf9840a · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.717277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.717277Z digest=sha256:ba3f7529ad60c025892f47674f88c7cd69805550e073b69d41a7d78eabe4d7c3

Observation b47862b3-ec9f-4319-b291-e06a96870b66 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llama-vid: An image is worth 2 tokens in large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.725584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.725584Z digest=sha256:e3e45e3f188a798109f75af3501413fe07d670457940ecd3381d84c6d3aed7d2

Observation 8622997f-673b-4e3f-a8c5-87bf0c700d2f · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Evaluating Object Hallucination in Large Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.730319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.730319Z digest=sha256:0caf77369b15ab5d35d923accbf3ea7f5d2a39646b6da111cc714ca12f152c5c

Observation a231d075-a609-4905-b05a-9bab02935191 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.735467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.735467Z digest=sha256:c329add1170543712c180d58394339202781e4eb6fb021db4dbbb0913f3d46dd

Observation 5b8cf93e-0509-4b35-8c94-f8154c1d8458 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.740088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.740088Z digest=sha256:d805d04fb72ad8ed6bc4543797fd34e72b04b5ab424ade033a3a24802ea2c398

Observation ce0f5d7b-3421-45af-b0a9-5c6c2810d65c · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.746024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.746024Z digest=sha256:34326fa14607013ddcd68adf26e9e26e2338ec49d5755859499645b9e5e30445

Observation 019916fc-4086-4e9e-ba18-06309506835d · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Improved baselines with visual instruction tuning, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.754518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.754518Z digest=sha256:fd37f6288747a003335e48b7be58c9835f8036be8a6bd778339c1bea5ddf5a3d

Observation 403461a0-177e-4689-a44c-3525b9891f31 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.759294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.759294Z digest=sha256:f41eb66c158857eefb48b5d679adb0cabc9ec552f4bc6193d7cf28ab89e83091

Observation 297559bd-906f-46be-b760-ff7b5aa00b18 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.762862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.762862Z digest=sha256:278d82f9290c9243f1b5b13f9968929cfbd6515543eee44827cf34e3f172b068

Observation cdd77114-445b-4867-9373-2a5056bc7b38 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.770024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.770024Z digest=sha256:ffe368c3314d62d4d4748682167b2fd856de2724e47c1d92cb043e07097e9f95

Observation 8519c36c-eef2-429e-aa49-9d5a12ae840e · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.773647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.773647Z digest=sha256:8d66a8f39d824f18351fec2baff929757d7d4a871bdc302d25e440502b0c667f

Observation 67b001ae-295d-45cf-906a-cafeb2cb178e · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.778621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.778621Z digest=sha256:c550d9fbb25ee1cebc992cc450525d0d2406d38b23443c9a03daec3c6d5d7059

Observation a59232dd-a6e0-4986-affb-a787d45ad1d7 · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.784958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.784958Z digest=sha256:c59260ca50c7d8a26b7ac4756faa0b449dd20ee22e8a3b25a399c594f7a79c01

Observation c7bf4ee0-a720-4031-a320-aec3d66ec9fa · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.788990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.788990Z digest=sha256:383ee381d15a5600f167d6c73c2b4131d191181f0d43a4863ae83359f8e55417

Observation af78bd7e-562d-4997-bc73-5b3ee4c2cb8d · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.793581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.793581Z digest=sha256:c9a262b85fdffff78b786b28c7856d692ad1b7dc9718c139eedcf567df6f5650

Observation a4a41ed7-6599-47f3-bd10-b788e8688ad2 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.Meta AI Blog.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.Meta AI Blog

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.798585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.798585Z digest=sha256:83eb6c7dd21142ee7986decf7f21cdbcaaaa7392c8ab8ff8415fa265f7fd2992

Observation e6d8f519-7a59-4d74-9593-14d47ac90e97 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Learning transferable visual models from natural language supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.801951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.801951Z digest=sha256:13b6d6979d292a928ea98d9f20c9da2de15d78b02ad3c82f84a7555fa0c668a1

Observation 0fdf2af3-937f-4bc1-ac7c-302b208475d2 · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 2019.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Language models are unsupervised multitask learners.OpenAI blog, 2019

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.525704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:42:13.809976Z digest=sha256:9cf7149b066b5e26aba957d932ca59c59a1b46174085708bf4cd6031193e73e6

Observation 82f9ab2e-4bd1-4a08-917c-5cecb7d27d47 · outbound

This paper cites Searching for Activation Functions.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Searching for Activation Functions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.819653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.819653Z digest=sha256:30622f85dabe064841f22f7013a3520b1787da9ce7056e6f0569da3f4aad12c5

Observation b7754035-5226-457d-a5d6-8fd86eda5e15 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.831606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.831606Z digest=sha256:7862a07bdf98e90bcd146451bb5b356e920b2fc33de3f6cce4fa5b497f7d816d

Observation 18327225-f708-490a-a6ed-ddbf008e6d03 · outbound

This paper cites Towards vqa models that can read.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Towards vqa models that can read

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.836198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.836198Z digest=sha256:7e4d4f2a8c1c6dcc5479fb5aa5c8218f13ebeafe00564333b35427445f00d83f

Observation 63618e8b-195e-4ee8-969e-b82cdfb805f5 · outbound

This paper cites Flops profiler, 2025.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Flops profiler, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.505878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:42:13.843580Z digest=sha256:5afe3531861da34b7798a3bae1123b0a9fca834e3a69565619a6b544188720c3

Observation 578a6673-3b6e-455d-8f3a-493a78cf7456 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Gemini: A Family of Highly Capable Multimodal Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.848825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.848825Z digest=sha256:33715ffb0f72b751023ecd3ae26d23bca10b079b9874dd1bfaf90d10bae2fde1

Observation bb38303f-ae63-4e72-844c-9059286c1dbc · outbound

This paper cites Mlp- mixer: An all-mlp architecture for vision.Advances in neural information processing systems, 34:24261–24272, 2021.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mlp- mixer: An all-mlp architecture for vision.Advances in neural information processing systems, 34:24261–24272, 2021

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.852755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.852755Z digest=sha256:0467a71c0db1150812b900878c67c0f18be29ae31d7ab891a921c7b8fbff884d

Observation ab8e50e7-81d8-4492-a41f-7c3711c74420 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.863089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.863089Z digest=sha256:8f432aa53184b9d4d7432bda5492f1dc58721a71d1baeb41457603fc50573151

Observation 2debb5ab-3107-4b9d-bc1a-d9f60141ed15 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.878802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.878802Z digest=sha256:c8bcbf16b3c8d0f86d42879fc2288441961e758163b46e64cc0acb42de953d08

Observation e808306e-e95a-40df-a2bf-51795d5012aa · outbound

This paper cites Patches Are All You Need?.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Patches Are All You Need?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.886946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.886946Z digest=sha256:98fb2b4ecc97289bc53c089bd489e0ece8666fe3fe7a3f5de4bf8c2d034f668f

Observation a476ad82-1e46-458e-a4f5-f3e88a671a9e · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.890551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.890551Z digest=sha256:a713554e52795a76460fb05ff332058fbd312d96f312cad9d91c1670e7bc3804

Observation fdea2b54-0b7f-49e1-965a-c3998b9c5e04 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.894002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.894002Z digest=sha256:8b2b6c412f218615da2d9e79b7c19d6bfab37a33ca56b7d297fb296fb7718ab4

Observation 556e51c2-2ca4-4449-9009-2ac112aa07e0 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Cogvlm: Visual expert for pretrained language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.475667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:42:13.903098Z digest=sha256:247efa73c458df183656b945f6262710da16c95325db9b4f89ddfa74d4bb10c4

Observation 4716d017-f258-4481-801d-4b9bdde49b55 · outbound

This paper cites Qwen2 Technical Report.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Qwen2 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.912263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.912263Z digest=sha256:c316d995f626289082a90895e3b1aa084d4a451a85dfdbede81097419939bf38

Observation 90372d15-99a1-4b77-8d32-c05862fa7573 · outbound

This paper cites mplug-owl3: Towards long image-sequence understanding in multi-modal large language models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation mplug-owl3: Towards long image-sequence understanding in multi-modal large language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.464007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:42:13.916176Z digest=sha256:2b715810a7be13a224f8cff4439492d532d182148a693914d0b7a57e328f9191

Observation 1843186c-721c-4ef5-9ade-c338527de335 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:14.451218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:42:13.920019Z digest=sha256:157eac0e011355173b323981f8e5818e0f9670c989848f27c58e63801b8ea4e2

Observation 9467c2b3-921c-48c6-a38b-87433cedbbc4 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.923332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.923332Z digest=sha256:d5b993c213587d937357d48be61060e2e1d2429df3c714c4e4374097f6ab352f

Observation d2685d41-a199-451c-b56c-5dd04ca1235d · outbound

This paper cites Sigmoid loss for language image pre-training.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Sigmoid loss for language image pre-training

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.930971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.930971Z digest=sha256:a6a5d3f472f466cff985bb8fb0af68b4a51fd5448beda593eff9cc76a254893d

Observation 601d00e5-a49a-4575-8f25-1aa2ee7a768e · outbound

This paper cites Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.947456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.947456Z digest=sha256:2ffd5776ed0a78dc94a468ed286412f78a6a1be67b7122620036e71198600bc4

Observation de1d6570-519c-4776-876d-ecf848bd4c48 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.955406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.955406Z digest=sha256:59faca77b84f36313ec8c112f5912f1a378a77b10d628a968780829f47a43ca8

Observation ed01f621-792c-4699-afb1-f54e0c1f7f6d · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.967404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.967404Z digest=sha256:883fc576b42e6f0180209f072f1492882d4d2a8b95dac0138122b745c92f8f4f

Observation 35bc73c8-054a-4132-919a-d3a9d7e7a3e6 · outbound

This paper cites Long Context Transfer from Language to Vision.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Long Context Transfer from Language to Vision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.971332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.971332Z digest=sha256:815725e4d148bd8404ec3e0796e5b316cb47f0218a6842fbe961da27865a0548

Observation 149c2511-5fae-4c5e-9958-47f2b52cd5fb · outbound

This paper cites Wings: Learning Multimodal LLMs without Text-only Forgetting.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Wings: Learning Multimodal LLMs without Text-only Forgetting

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:42:14.042993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:42:13.975077Z digest=sha256:e768dda7ba3ab2e44d98bf226fb5f621e7a030ff86b3bec5a243da8c3d057218

Observation 7abdeaac-5cb1-4a79-9983-d83daf5b6c7d · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation MLVU: Benchmarking Multi-task Long Video Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.979447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.979447Z digest=sha256:53e3757ebddbf91c923b9d57b92900b8f34c0b9e45ee3c57cc0fe6007b381a6b

Pith citing papers

Observation 29312d64-c8fe-498a-bfda-894e49983ae1 · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.181799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:c793c742a01b3daf758e58ecd50d329054f4d1741ce4cab09ddbf84b29af9c69

Observation a6862ffb-5dcc-4899-848e-b4b0aa1d8cdb · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:19:31.833880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:ada8d1a3ebd043eae783d225cd16620205cc786f7367b90422fd81a9fa9de256