Pith. sign in

REVIEW 4 cited by

M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.04633 v3 pith:KT77VXKD submitted 2025-04-06 cs.CV cs.AI

classification cs.CVcs.AI
keywords lvlmsmultimodalrepresentationdemonstrationsengineeringin-contextlearningcross-modal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Multimodal in-context learning (ICL) equips Large Vision-language Models (LVLMs) with the ability to adapt to new tasks via multiple user-provided demonstrations, without requiring any model parameter updates. However, its effectiveness is constrained by the token-intensive nature of multimodal inputs and the complexity of cross-modal few-shot reasoning, which together hinder LVLMs from extracting useful patterns from demonstrations. To address these challenges, we propose \textbf{M$^2$IV}, a novel representation engineering approach that replaces explicit token-level demonstrations with a set of learnable Multimodal In-context Vectors directly injected into the residual streams of LVLMs. By analyzing the distinct roles of multi-head attention (MHA) and multi-layer perceptrons (MLP) in the ICL process, we design a training strategy that enables M$^2$IV to perform fine-grained semantic distillation and robust cross-modal representation learning. M$^2$IV not only improves performance across diverse tasks and LVLMs but also significantly reduces token overhead, enabling graceful scaling to many-shot scenarios. To further enhance usability, we introduce \textbf{VLibrary}, a repository that stores trained M$^2$IVs for flexible retrieval and injection. With VLibrary, users can steer pre-trained LVLMs in a customized manner that meets diverse requirements. Extensive experiments demonstrate that M$^2$IV consistently outperforms vanilla ICL and prior representation engineering baselines, achieving an average accuracy gain of 3.74\% with substantial improvements in overall efficiency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Zero-shot Cross-lingual NER via Mitigating Language Difference: An Entity-aligned Translation Perspective

    cs.CL 2025-09 conditional novelty 6.0 of 10

    An LLM-based dual-translation method with multi-round chain-of-thought and Wikipedia-based entity alignment fine-tuning improves zero-shot named entity recognition on non-Latin script languages.

  2. STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A structure-aware exemplar retriever with a hidden-state syntactic injection module improves in-context semantic parsing across four benchmarks.

  3. FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter

    cs.MM 2025-08 reject novelty 5.0 of 10

    FakeSV-VLM reaches 90.22% and 89.30% accuracy on FakeSV and FakeTT by adding a two-stage MoE adapter and contrastive alignment to InternVL2.5-8B.

  4. SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SurgBench assembles a 53-million-frame surgical video pretraining corpus from 16 sources plus a 72-task evaluation benchmark, and shows continual pretraining with VideoMAE improves accuracy by 7.9% top-3 over Kinetics...

Pith tools