Pith. sign in

REVIEW 1 cited by

Improving Hateful Meme Detection through Retrieval-Guided Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.08110 v3 pith:4AMFX7JW submitted 2023-11-14 cs.CL cs.CV

classification cs.CLcs.CV
keywords hatefulmemesdetectionsystemcontrastiveembeddinghatefulnessinternet
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hateful memes have emerged as a significant concern on the Internet. Detecting hateful memes requires the system to jointly understand the visual and textual modalities. Our investigation reveals that the embedding space of existing CLIP-based systems lacks sensitivity to subtle differences in memes that are vital for correct hatefulness classification. We propose constructing a hatefulness-aware embedding space through retrieval-guided contrastive training. Our approach achieves state-of-the-art performance on the HatefulMemes dataset with an AUROC of 87.0, outperforming much larger fine-tuned large multimodal models. We demonstrate a retrieval-based hateful memes detection system, which is capable of identifying hatefulness based on data unseen in training. This allows developers to update the hateful memes detection system by simply adding new examples without retraining, a desirable feature for real services in the constantly evolving landscape of hateful memes on the Internet.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Simple embedding fusion improves F1 by 9.9 points on the HateMM video dataset but reaches only 0.628 AUROC on the Hateful Memes dataset, showing that fusion methods do not transfer across modality types.

Pith tools