REVIEW 6 cited by
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The rise of foundation models has transformed machine learning research, prompting efforts to uncover their inner workings and develop more efficient and reliable applications for better control. While significant progress has been made in interpreting Large Language Models (LLMs), multimodal foundation models (MMFMs) - such as contrastive vision-language models, generative vision-language models, and text-to-image models - pose unique interpretability challenges beyond unimodal frameworks. Despite initial studies, a substantial gap remains between the interpretability of LLMs and MMFMs. This survey explores two key aspects: (1) the adaptation of LLM interpretability methods to multimodal models and (2) understanding the mechanistic differences between unimodal language models and crossmodal systems. By systematically reviewing current MMFM analysis techniques, we propose a structured taxonomy of interpretability methods, compare insights across unimodal and multimodal architectures, and highlight critical research gaps.
Forward citations
Cited by 6 Pith papers
-
FairFlow: Demystifying and Mitigating Stereotype Bias in Text-to-Image Diffusion Transformers
Bias in MM-DiTs is mediated by sparse stage-wise semantic binding hubs, and sparse inference-time steering at those hubs mitigates gender, race, and intersectional stereotypes with low overhead.
-
Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations
Local low-rank Gaussian neighborhoods in VLM residual streams reveal model-specific fusion trajectories and serve as causal steering and retrieval units.
-
Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs
CMA uses diffusion-generated counterfactuals and two-player Shapley values to measure whether an image or the text drives a multimodal LLM's prediction, hitting 98% on synthetic biased benchmarks.
-
How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA
The paper proposes four operation-level VLM failure modes and a pathway dissociation, but the dissociation is not supported by the paper's own intervention statistics.
-
Unsupervised Features Mining via Activation Geometry
Prefix-induced activation shifts (MAG) yield model-relative reasoning directions that predict verdicts, support matched-format steering, and select transfer datasets at 94.7% Top-1 accuracy.
-
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
GrAInS uses Integrated Gradients to identify the most influential tokens, then builds layer-wise steering vectors that improve truthfulness, reduce hallucination, and preserve general capabilities in LLMs and VLMs.
Discussion (0). Continue with ORCID to comment.