Pith. sign in

REVIEW 1 cited by

Memes in the Wild: Assessing the Generalizability of the Hateful Memes Challenge Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.04313 v1 pith:YESZ5ANB submitted 2021-07-09 cs.CV

classification cs.CV
keywords memeshatefulchallengedatasetwildcaptionscurrentfacebook
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hateful memes pose a unique challenge for current machine learning systems because their message is derived from both text- and visual-modalities. To this effect, Facebook released the Hateful Memes Challenge, a dataset of memes with pre-extracted text captions, but it is unclear whether these synthetic examples generalize to `memes in the wild'. In this paper, we collect hateful and non-hateful memes from Pinterest to evaluate out-of-sample performance on models pre-trained on the Facebook dataset. We find that memes in the wild differ in two key aspects: 1) Captions must be extracted via OCR, injecting noise and diminishing performance of multimodal models, and 2) Memes are more diverse than `traditional memes', including screenshots of conversations or text on a plain background. This paper thus serves as a reality check for the current benchmark of hateful meme detection and its applicability for detecting real world hate.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Representation Decomposition for Learning Similarity and Contrastness Across Modalities for Affective Computing

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A method that decomposes CLIP image and text features into a shared low-rank component and modality-specific sparse components, then uses an attention-weighted soft prompt to guide an LLM for sentiment, emotion, and h...

Pith tools