Pith. sign in

REVIEW 2 cited by

MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.11215 v1 pith:2B7LCS6B submitted 2024-05-18 cs.CL cs.CY

classification cs.CLcs.CY
keywords mememqamemesmultimodalarsenalcommunicationexplanationsframeworkharm
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Memes have evolved as a prevalent medium for diverse communication, ranging from humour to propaganda. With the rising popularity of image-focused content, there is a growing need to explore its potential harm from different aspects. Previous studies have analyzed memes in closed settings - detecting harm, applying semantic labels, and offering natural language explanations. To extend this research, we introduce MemeMQA, a multimodal question-answering framework aiming to solicit accurate responses to structured questions while providing coherent explanations. We curate MemeMQACorpus, a new dataset featuring 1,880 questions related to 1,122 memes with corresponding answer-explanation pairs. We further propose ARSENAL, a novel two-stage multimodal framework that leverages the reasoning capabilities of LLMs to address MemeMQA. We benchmark MemeMQA using competitive baselines and demonstrate its superiority - ~18% enhanced answer prediction accuracy and distinct text generation lead across various metrics measuring lexical and semantic alignment over the best baseline. We analyze ARSENAL's robustness through diversification of question-set, confounder-based evaluation regarding MemeMQA's generalizability, and modality-specific assessment, enhancing our understanding of meme interpretation in the multimodal communication landscape.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering

    cs.CV 2025-04 conditional novelty 5.0 of 10

    An OCR, captioning, sub-label retrieval, and multi-turn VQA pipeline reaches 73.5 percent accuracy and 78.35 AUROC on Facebook Hateful Memes.

  2. Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

    cs.CL 2026-07 conditional novelty 4.0 of 10

    A systematic survey and cross-benchmark evaluation showing that multimodal LLMs can recognize humor artifacts but still struggle to interpret the intended meaning and mechanisms of visual humor.

Pith tools