Pith. sign in

REVIEW 2 cited by

On the Benefits of Early Fusion in Multimodal Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.07191 v1 pith:F3VXXHZH submitted 2020-11-14 cs.LG cs.AI

classification cs.LGcs.AI
keywords multimodalaudiofusionvisualearlyinputslearninginformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Intelligently reasoning about the world often requires integrating data from multiple modalities, as any individual modality may contain unreliable or incomplete information. Prior work in multimodal learning fuses input modalities only after significant independent processing. On the other hand, the brain performs multimodal processing almost immediately. This divide between conventional multimodal learning and neuroscience suggests that a detailed study of early multimodal fusion could improve artificial multimodal representations. To facilitate the study of early multimodal fusion, we create a convolutional LSTM network architecture that simultaneously processes both audio and visual inputs, and allows us to select the layer at which audio and visual information combines. Our results demonstrate that immediate fusion of audio and visual inputs in the initial C-LSTM layer results in higher performing networks that are more robust to the addition of white noise in both audio and visual inputs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation

    cs.LG 2025-09 conditional novelty 5.0 of 10

    MultiFair couples modality-balancing gradient modulation with group-AUC-based fairness scaling and reports improved balanced accuracy on two glaucoma datasets.

  2. I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts

    cs.LG 2025-05 conditional novelty 4.0 of 10

    I2MoE improves multimodal fusion by training interaction-specialized experts with perturbed-modality supervision and reweighting their outputs per sample.

Pith tools