Pith. sign in

Paper Citation Record · LEDGER

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models

As of 19 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2505.00150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00150 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:55:28.327964Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 27a3026c-8019-4ea1-8265-bac079e7f749 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:29.141780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.088787Z digest=sha256:f2653c0c12a73acb4c9400285da13d2fa5d2205303ff0e2ec8f8d6d6571fcd63

Observation 715f2511-14f1-45ad-91ea-94d41e0d9a33 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.095027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.095027Z digest=sha256:28ce7836e04e859fd2098375276ea6a89680b12e0da4ba4e3516dc6bb3254bf5

Observation ab681ece-2041-4fd9-ac45-237358c941fb · outbound

This paper cites Prompting for Multimodal Hateful Meme Classification.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Prompting for Multimodal Hateful Meme Classification

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.101046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.101046Z digest=sha256:b1daa4cd0a692c24eb8a57d2017768603d855954cb464bfcbd1451dd06e2b739

Observation 4dae2445-6fe6-4f36-ac7e-5c76545edf3d · outbound

This paper cites Modularized networks for few-shot hateful meme detection.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Modularized networks for few-shot hateful meme detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:29.124204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.107326Z digest=sha256:af08bbdf15f2a04b984bdf3f6ca3b9810794da371747151e2efaf553784f3bab

Observation 8cc59cd8-b6d4-4ccb-979a-ad3a279ab324 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.112729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.112729Z digest=sha256:c536d2a8d3876b6f3932fa32bd38eaf5f5e4bc36ffa9536cc1908a3e73498776

Observation 17e15a67-b841-44b9-978a-f1a048b1bb29 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.118746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.118746Z digest=sha256:49e75f4020f8fc6251e4bbf15e61bdf40d43c42c3ebc311d8c4ec43b298f1df2

Observation 58749952-d8c6-4df9-8cc9-7cb8672a5cda · outbound

This paper cites Deep residual learning for image recognition.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Deep residual learning for image recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.124735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.124735Z digest=sha256:c7e5003befb1c75f6215fe3eed91d3a3b1974b361da450c285d7e687efab24ec

Observation d29ebf04-445a-4509-8709-b2dfbb4d99bb · outbound

This paper cites Recent advances in online hate speech moderation: Multimodality and the role of large models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Recent advances in online hate speech moderation: Multimodality and the role of large models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:29.093645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.129926Z digest=sha256:90bf66fe82d86387e456d0289ad5e4da61592edaf5bde84f6f2f20852acceef8

Observation b0a87666-f99c-40cb-ab81-2bd4a9bac7fc · outbound

This paper cites Openclip, 2021.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Openclip, 2021

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.134604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.134604Z digest=sha256:da656bd0114cb1590b091da1695c0bfe1cea0a9dce7488b4956ad03ce4bbf7e9

Observation c5d1c160-87dd-4d46-b1b1-721aacc15379 · outbound

This paper cites Capalign: Improving cross modal alignment via informative captioning for harmful meme detection.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Capalign: Improving cross modal alignment via informative captioning for harmful meme detection

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:29.064962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.139044Z digest=sha256:3526e3962764cf5e6d5acc8244fecde6e78415b2523dc55b25f07e39adac66d5

Observation 4784a6f9-9cad-4d84-9cc5-9d12b36057ed · outbound

This paper cites Supervised Multimodal Bitransformers for Classifying Images and Text.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Supervised Multimodal Bitransformers for Classifying Images and Text

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.143837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.143837Z digest=sha256:fc851d8582fc2ca95871ec8ea2f359435a62d88c65e788975fe8514329ac3e5b

Observation 90167aa6-9778-4423-819e-70e3b590095b · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.149189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.149189Z digest=sha256:251c6ca3f399e76d1a9724fd0eb83715f337d410a3b29a0c88a80370609fde03

Observation a6da361d-0501-424c-a64e-90b3f5094bed · outbound

This paper cites The hateful memes challenge: Competition report.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models The hateful memes challenge: Competition report

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:29.036156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.153924Z digest=sha256:848ec396348bb94f87e52a5bff90586ddb5fd38b83900c8422e617bb0af7b01e

Observation 83116132-e221-472d-99c3-835d39beed2a · outbound

This paper cites Why is it hate speech? masked ratio- nale prediction for explainable hate speech detection.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Why is it hate speech? masked ratio- nale prediction for explainable hate speech detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:29.017987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.158511Z digest=sha256:39f641b3f8f468454b83cab7b0dced5eb3648f296e160f117b5c238c52c04e2b

Observation cfe61be9-246c-491d-ba36-3ce256c73838 · outbound

This paper cites Segment anything.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Segment anything

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.163110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.163110Z digest=sha256:01706189cabf8f407d787de4977ab6952b64cf640f0d7ee506b63d3292676288

Observation c2ca12f7-87b6-4278-a7c3-a1028ff340de · outbound

This paper cites LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.168132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.168132Z digest=sha256:e0d8e5cd7716097d8a46cc92186d5df0a3971911a3cf181ed46f856c3ff9b465

Observation 9a673173-6d41-46a3-92b8-bc454f84d70f · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.173187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.173187Z digest=sha256:ba0f1fe6432d467509230ae358a0fe8de1b74df444b830cf6e05fb0e75f3feb5

Observation 27cd51f3-2628-4dbd-b8cf-95fce6661f8f · outbound

This paper cites Towards explainable harmful meme detection through multimodal debate between large language models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Towards explainable harmful meme detection through multimodal debate between large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.984925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.178135Z digest=sha256:831e8087a07937a5dc8dbe2fd9b5ad1f616f739345fb13a644c3c4a8ed2c5c67

Observation cb0c1590-2025-4a6d-91dd-c5cd6381f695 · outbound

This paper cites A Multimodal Framework for the Detection of Hateful Memes.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models A Multimodal Framework for the Detection of Hateful Memes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.182942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.182942Z digest=sha256:da356d63043dc6ab44b08ce0c319d6d2880ef1c52135e8fe10000525c67c776c

Observation 12607ca3-8f04-45db-b528-c90674b99931 · outbound

This paper cites Visual Instruction Tuning.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Visual Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.188523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.188523Z digest=sha256:664cd225a918b8d1a0793785719fa98b41ba4dfc59229f311e1b4817355bd33c

Observation 92183d9d-f78a-4287-ae26-597e4739640c · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.193622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.193622Z digest=sha256:6334f39337fbc858e577625f3e202a95ba84e558366ab1a78ad2abf24ed79e04

Observation 16d433d1-594f-41b7-90f4-22b6da9fb43c · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.198790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.198790Z digest=sha256:24a23c41ea8ded6ef4802fc53e9cac2f06b47121197ca6aea2648662f9a288d4

Observation 81d60a65-1c98-4d21-8c04-49174206c2b1 · outbound

This paper cites Hatexplain: A benchmark dataset for explainable hate speech detection.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Hatexplain: A benchmark dataset for explainable hate speech detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.966826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.203784Z digest=sha256:6cfeff278207647712584e9fb7318b8b84d916d3a1b26822cf7e63103a964aed

Observation 32d06113-ec91-4234-ae51-93dbe547cd37 · outbound

This paper cites ETHOS: an Online Hate Speech Detection Dataset.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models ETHOS: an Online Hate Speech Detection Dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.208724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.208724Z digest=sha256:d42d50f47263390d71e9bfd8c31e6be3b2f0c660396d5c57103ce4d679477b77

Observation 3e834cf9-0d87-4b47-b0bf-850ddca0a227 · outbound

This paper cites Vilio: State-of-the-art Visio-Linguistic Models applied to Hateful Memes.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Vilio: State-of-the-art Visio-Linguistic Models applied to Hateful Memes

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.213637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.213637Z digest=sha256:221bbc843f3ac61fee0c0e71b075a9f9f95068e30a1017acedb160ebe3cdb6fe

Observation ec4bab69-10c7-4529-92a0-14085d108fb3 · outbound

This paper cites GPT-4 Technical Report.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models GPT-4 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.218559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.218559Z digest=sha256:d6752d5472eb5a4b2b5730673d84aa0c805fe70ffc2d00a4066201ec23ef2976

Observation 0f8bd0e9-2177-4d03-b8dd-4e67ba36e8e1 · outbound

This paper cites Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.223211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.223211Z digest=sha256:3da742eb98177173627321ee56116a12ef2d1c9eb8ed6fdf61001068b424c51b

Observation 1da24017-0a3d-41c0-b2ae-fa22a565caab · outbound

This paper cites Learning transferable visual models from natural language supervision.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Learning transferable visual models from natural language supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.228722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.228722Z digest=sha256:169b501a15de17794bc71a99f932e7d87dabb5ace9c3d0616670fbd3487ee160

Observation 6d9b3b92-8ba6-4e34-8337-a62d3532c2dc · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.233478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.233478Z digest=sha256:a5edbe035393b94dc91ab222af53b8b5673288338b59010353bc65671a7f3307

Observation c781a117-7ef3-44fe-929b-d80fef9691a6 · outbound

This paper cites Detecting Hateful Memes Using a Multimodal Deep Ensemble.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Detecting Hateful Memes Using a Multimodal Deep Ensemble

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.238525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.238525Z digest=sha256:270366e93ad6771adf05adc8786bd9e7f525044a929473e1ad8a1f354bebc7da

Observation 9eda2e25-6c3e-42d7-8b47-906f7f2afa7b · outbound

This paper cites Detecting formal thought disorder by deep contextualized word representations.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Detecting formal thought disorder by deep contextualized word representations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.927220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.243316Z digest=sha256:3cdc4cdf10588e1dffd0957d53c9252b906a4e8c13dd0b99839565277383f0ce

Observation ecba3b8f-09f6-407b-979a-efc115661f23 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.248211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.248211Z digest=sha256:c875c4fc75c5a337edcc531282f64611c1174318733033bdaf48452038e342e6

Observation 709d1768-4ad3-4869-8d48-0abb639ea700 · outbound

This paper cites Large Language Models Encode Clinical Knowledge.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Large Language Models Encode Clinical Knowledge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.252964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.252964Z digest=sha256:cb32d293570e4780476b5daf60fabe62d1785972a4633f3813e026af6d057526

Observation 7e31eb3b-a8a7-4db6-b683-1e64fc33f7f8 · outbound

This paper cites Resolution-robust large mask inpainting with fourier convolutions.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Resolution-robust large mask inpainting with fourier convolutions

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.896919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.257972Z digest=sha256:0f957f7177cc2ad15113916d5dc98f25f26296cf33b9f56362acf25645ddd095

Observation e860ffdb-c39d-4dfa-878e-80d67f035e85 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.262753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.262753Z digest=sha256:b74b6a0cd52583cea63efc3bb8115a9dcd4c027d094ab4ee12c97a769fa6634e

Observation fe78932e-7f8f-4f0b-849d-807cb4ac6294 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.268336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.268336Z digest=sha256:98d593301e019ee4090928605b60a946acd24e1750bf7a5c05dcd155132d872c

Observation 2b00b850-8e44-4985-8f20-bf7c41cd78e9 · outbound

This paper cites On large visual language models for medical imaging analysis: An empirical study.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models On large visual language models for medical imaging analysis: An empirical study

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.877975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.273511Z digest=sha256:40d1e59025fbe62b4d215d204f2531f4cfda176445708c15be809b5c3fb256fc

Observation 68a69b57-5211-4cec-aab3-d4b763283f99 · outbound

This paper cites Detecting Hate Speech in Memes Using Multimodal Deep Learning Approaches: Prize-winning solution to Hateful Memes Challenge.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Detecting Hate Speech in Memes Using Multimodal Deep Learning Approaches: Prize-winning solution to Hateful Memes Challenge

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.278432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.278432Z digest=sha256:42716df48e1c21e98adcd6717b68a1d3937e8b9985f61f6ebc4c34af4497ef0d

Observation 176e6f5b-20fb-42bf-9ac9-65154d5a5c81 · outbound

This paper cites Memecraft: Contextual and stance-driven multimodal meme generation.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Memecraft: Contextual and stance-driven multimodal meme generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.859820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.283296Z digest=sha256:1d69408227e8418df3427dab836c0f08460dd0b8faa38676b087e61d9fd539ea

Observation d48cd019-2e12-462d-9843-8ecb8b589f24 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.288109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.288109Z digest=sha256:fb275e8d47e4e8a107d6ed1c0da34e02dfa18c1a2165ab60cee6baea8d722378

Observation 2c9ce84d-1427-4cd6-bc88-1e6c9aee62de · outbound

This paper cites Aggregated residual transformations for deep neural networks.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Aggregated residual transformations for deep neural networks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.293095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.293095Z digest=sha256:6d9f78c8e7d9f2d6ead7d5dc474721bb69926471b284c107f9ff7027771bbd9e

Observation 7c3ae4b9-e547-49c0-8cfa-c47ea4533354 · outbound

This paper cites Coded hate speech detection via contextual information.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Coded hate speech detection via contextual information

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.830858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.297774Z digest=sha256:424dce39aea1b0743e256e368fa313bcb24d581714fab1d5cddfd57251ce4c56

Observation 5bb76d33-83f3-434a-a6c0-ce0a5f822d11 · outbound

This paper cites An empirical study of gpt-3 for few-shot knowledge-based vqa.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models An empirical study of gpt-3 for few-shot knowledge-based vqa

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.813384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.302275Z digest=sha256:f73a7f8196d8e33ee1bfd4d70ef3c3b6374a5fdbf3935a30ab6d8b3ef2a10fd4

Observation 041eb16d-2f38-4db5-a17d-8deb554b4ecc · outbound

This paper cites Inpaint Anything: Segment Anything Meets Image Inpainting.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Inpaint Anything: Segment Anything Meets Image Inpainting

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.307144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.307144Z digest=sha256:5d80c772d301be43f96435584ae0d6ca6796c3ebd3a5f024a40fb28d60ea5492

Observation fa5ea7f7-4550-4bbf-a441-ac8e85e1a750 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.312338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.312338Z digest=sha256:dae735f228930b48adae93a9760b3f38e0ef412806b2a344456637a23fa2aa86

Observation 42b50c13-34e7-4099-9e99-690d94fb0514 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models OPT: Open Pre-trained Transformer Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.317391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.317391Z digest=sha256:3fcb33b54fa4ecdfcb9d65425348d4bbace112a8e6223b37b3f40475c996a40a

Observation 2e316939-a3d1-44fd-ab84-4eeaa6169eb7 · outbound

This paper cites Enhance Multimodal Transformer With External Label And In-Domain Pretrain: Hateful Meme Challenge Winning Solution.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Enhance Multimodal Transformer With External Label And In-Domain Pretrain: Hateful Meme Challenge Winning Solution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.322480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.322480Z digest=sha256:70c089fb533e40384d9052480a5960217c17ccdc4c8b4692cea33574cbd2e112

Observation d1570484-a50e-4715-a304-3ca31926a19e · outbound

This paper cites C.2 Multimodal: unimodal pretraining Multimodal models from unimodal pretraining typically combine the output or the features of vision and linguistic models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models C.2 Multimodal: unimodal pretraining Multimodal models from unimodal pretraining typically combine the output or the features of vision and linguistic models

Reference 768

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.795461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:55:28.327964Z digest=sha256:afc50fc4db44186b44acf0a2c18868894c21b0f6c5c06088649c9fd3de7553b6

Pith citing papers

No inbound Pith citation observations are available.