Pith. sign in

REVIEW 8 cited by

Supervised Multimodal Bitransformers for Classifying Images and Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.02950 v2 pith:JFK4VRHC submitted 2019-09-06 cs.CL cs.CVcs.LGstat.ML

classification cs.CLcs.CVcs.LGstat.ML
keywords multimodalclassificationimagesinformationperformancesupervisedtaskstext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised bidirectional transformer models such as BERT have led to dramatic improvements in a wide variety of textual classification tasks. The modern digital world is increasingly multimodal, however, and textual information is often accompanied by other modalities such as images. We introduce a supervised multimodal bitransformer model that fuses information from text and image encoders, and obtain state-of-the-art performance on various multimodal classification benchmark tasks, outperforming strong baselines, including on hard test sets specifically designed to measure multimodal performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EPIC: Efficient Prompt Interaction for Text-Image Classification

    cs.CV 2025-07 conditional novelty 6.0 of 10

    EPIC uses similarity-gated cross-modal prompt interactions in frozen CLIP to improve text-image classification accuracy while training about 1% of parameters.

  2. MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection

    cs.CL 2025-07 conditional novelty 6.0 of 10

    MIND uses unlabeled similar memes, bidirectional AI insight derivation, and multi-agent debate to improve zero-shot harmful meme detection on HarM, FHM, and MAMI.

  3. AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness

    cs.CL 2025-07 conditional novelty 6.0 of 10

    AdamMeme is an adaptive, agent-based evaluation framework that iteratively refines meme text to expose model-specific weaknesses in multimodal models' understanding of meme harmfulness.

  4. Class Similarity-Based Multimodal Classification under Heterogeneous Category Sets

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The paper introduces MMHCL, a setting where modalities train on heterogeneous category sets, and CSCF, a model using semantic alignment, uncertainty-based dominance selection, and class-similarity fusion to recognize ...

  5. MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A new 1,565-item benchmark shows large vision-language models classify meme-context relationships far better than they infer the poster's intent.

  6. EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection

    cs.CV 2026-07 conditional novelty 5.0 of 10

    EVL-MCoT combines multiple hateful and benign chain-of-thought explanations with prototype-guided vision-language fusion, reporting state-of-the-art harmful meme detection accuracy on HatefulMemes and MultiOFF.

  7. Applying multimodal learning to Classify transient Detections Early (AppleCiDEr) I: Data set, methods, and infrastructure

    astro-ph.IM 2025-07 conditional novelty 4.0 of 10

    AppleCiDEr combines photometry, images, metadata, and spectra in one deep learning pipeline to classify ZTF transients and variable stars, with high accuracy on common classes but poor performance on tidal disruption events.

  8. Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A new multimodal fusion architecture reports small state-of-the-art gains on two offensive content benchmarks using co-attention, dimension-wise gating, and expert fusion.

Pith tools