Pith. sign in

REVIEW 4 cited by

MultiViz: Towards Visualizing and Understanding Multimodal Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.00056 v3 pith:SE4KIPDE submitted 2022-06-30 cs.LG cs.AIcs.CLcs.CVcs.MM

classification cs.LGcs.AIcs.CLcs.CVcs.MM
keywords modelsmultimodalmultivizmodelfeaturesinteractionsinternalprediction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The promise of multimodal models for real-world applications has inspired research in visualizing and understanding their internal mechanics with the end goal of empowering stakeholders to visualize model behavior, perform model debugging, and promote trust in machine learning models. However, modern multimodal models are typically black-box neural networks, which makes it challenging to understand their internal mechanics. How can we visualize the internal modeling of multimodal interactions in these models? Our paper aims to fill this gap by proposing MultiViz, a method for analyzing the behavior of multimodal models by scaffolding the problem of interpretability into 4 stages: (1) unimodal importance: how each modality contributes towards downstream modeling and prediction, (2) cross-modal interactions: how different modalities relate with each other, (3) multimodal representations: how unimodal and cross-modal interactions are represented in decision-level features, and (4) multimodal prediction: how decision-level features are composed to make a prediction. MultiViz is designed to operate on diverse modalities, models, tasks, and research areas. Through experiments on 8 trained models across 6 real-world tasks, we show that the complementary stages in MultiViz together enable users to (1) simulate model predictions, (2) assign interpretable concepts to features, (3) perform error analysis on model misclassifications, and (4) use insights from error analysis to debug models. MultiViz is publicly available, will be regularly updated with new interpretation tools and metrics, and welcomes inputs from the community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CMA uses diffusion-generated counterfactuals and two-player Shapley values to measure whether an image or the text drives a multimodal LLM's prediction, hitting 98% on synthetic biased benchmarks.

  2. Ego-centric Predictive Model Conditioned on Hand Trajectories

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Ego-PM predicts future hand trajectories and then uses them to condition latent diffusion video generation, jointly outputting actions and future frames in egocentric and robotic scenes.

  3. Partitioner Guided Modal Learning Framework

    cs.CL 2025-07 conditional novelty 6.0 of 10

    PgM segments multimodal representations into uni-modal and paired-modal features with cumulative-softmax gates and trains them with separate learners, reconstruction, and classification losses, yielding accuracy gains...

  4. I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts

    cs.LG 2025-05 conditional novelty 4.0 of 10

    I2MoE improves multimodal fusion by training interaction-specialized experts with perturbed-modality supervision and reweighting their outputs per sample.

Pith tools