Pith. sign in

Paper Citation Record · LEDGER

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models

As of 17 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2505.00150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00150 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:55:28.327964Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 27a3026c-8019-4ea1-8265-bac079e7f749 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:29.141780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.088787Z digest=sha256:b4ea79c89579b20bafc708f1a680eb634a553ed77da9c21edb2a40cd8ba60b37

Observation 715f2511-14f1-45ad-91ea-94d41e0d9a33 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.095027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.095027Z digest=sha256:bc515083ba6a03496d86c83a1f2567d732b8a369d8b7fba6f0037c55c3b5edc1

Observation ab681ece-2041-4fd9-ac45-237358c941fb · outbound

This paper cites Prompting for Multimodal Hateful Meme Classification.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Prompting for Multimodal Hateful Meme Classification

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.101046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.101046Z digest=sha256:3d60e202a53cb33928713f53222e8367798410d16b2df91f8cae9ea13d99b5f6

Observation 4dae2445-6fe6-4f36-ac7e-5c76545edf3d · outbound

This paper cites Modularized networks for few-shot hateful meme detection.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Modularized networks for few-shot hateful meme detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:29.124204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.107326Z digest=sha256:62c9b6839825ebca0ffb6c8a02f5d8bff380a58d3b461564cd0af8c9da9be2ad

Observation 8cc59cd8-b6d4-4ccb-979a-ad3a279ab324 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.112729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.112729Z digest=sha256:8c6aa721f6b7b9589fada6e0e1fdb7be4a4d54bd5ec5c69c6c20eec29df71004

Observation 17e15a67-b841-44b9-978a-f1a048b1bb29 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.118746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.118746Z digest=sha256:ceb7691c88a53a118aaf4613c9f2592a57a1cdc3f9e640690f7331f138931746

Observation 58749952-d8c6-4df9-8cc9-7cb8672a5cda · outbound

This paper cites Deep residual learning for image recognition.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Deep residual learning for image recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.124735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.124735Z digest=sha256:7a910f49b0533e8b095e166633ae0fa689fd4be72fc425050dd4241afb291779

Observation d29ebf04-445a-4509-8709-b2dfbb4d99bb · outbound

This paper cites Recent advances in online hate speech moderation: Multimodality and the role of large models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Recent advances in online hate speech moderation: Multimodality and the role of large models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:29.093645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.129926Z digest=sha256:6f9cc0420a3bad645751c6a302f91b52b2b1a203088908251f69e232664fb6e0

Observation b0a87666-f99c-40cb-ab81-2bd4a9bac7fc · outbound

This paper cites Openclip, 2021.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Openclip, 2021

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.134604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.134604Z digest=sha256:6b66f5f2cb73d155f9079aaa6dde4095a34dc89eb1a5fe59794c89ff8fa2871b

Observation c5d1c160-87dd-4d46-b1b1-721aacc15379 · outbound

This paper cites Capalign: Improving cross modal alignment via informative captioning for harmful meme detection.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Capalign: Improving cross modal alignment via informative captioning for harmful meme detection

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:29.064962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.139044Z digest=sha256:4927ad1aed5de80366a815f9092ccdf74f0d14cc701d8b3c06236e52393333e7

Observation 4784a6f9-9cad-4d84-9cc5-9d12b36057ed · outbound

This paper cites Supervised Multimodal Bitransformers for Classifying Images and Text.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Supervised Multimodal Bitransformers for Classifying Images and Text

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.143837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.143837Z digest=sha256:c1666b2bfc9b55ba238db7e7e71c919762a921f837311017dc1ccb126d2f471b

Observation 90167aa6-9778-4423-819e-70e3b590095b · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.149189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.149189Z digest=sha256:62f871fe0ac96e81bf4f3fe8d183a44ee4c1cd7cde3793387a4078f67c6a0e1f

Observation a6da361d-0501-424c-a64e-90b3f5094bed · outbound

This paper cites The hateful memes challenge: Competition report.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models The hateful memes challenge: Competition report

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:29.036156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.153924Z digest=sha256:0d93166b3f1ef78305bee730a31bc32344fc790286e5f3ebf39d44eb154ac74d

Observation 83116132-e221-472d-99c3-835d39beed2a · outbound

This paper cites Why is it hate speech? masked ratio- nale prediction for explainable hate speech detection.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Why is it hate speech? masked ratio- nale prediction for explainable hate speech detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:29.017987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.158511Z digest=sha256:e3c703c299f58c90335123d6eba6ee01f25943919cdc5d790d2fb3e9ebcf962c

Observation cfe61be9-246c-491d-ba36-3ce256c73838 · outbound

This paper cites Segment anything.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Segment anything

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.163110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.163110Z digest=sha256:5e1a89a6b13d9d0c221463bc732c42749b28a72009b9dc177aba2f2bec76038d

Observation c2ca12f7-87b6-4278-a7c3-a1028ff340de · outbound

This paper cites LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.168132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.168132Z digest=sha256:be4b985e1b79365d9459e1e6d613cb4eaea7926b05a1b23391e6547fafd876ed

Observation 9a673173-6d41-46a3-92b8-bc454f84d70f · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.173187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.173187Z digest=sha256:ca182082ed555e7b78797349ebca033a390c2c6339d7c34496f5c48168b2a56f

Observation 27cd51f3-2628-4dbd-b8cf-95fce6661f8f · outbound

This paper cites Towards explainable harmful meme detection through multimodal debate between large language models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Towards explainable harmful meme detection through multimodal debate between large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.984925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.178135Z digest=sha256:6a86d02d13f623a3a2c400fdc4279d100cbfe37f1fdba81524c016eacebe3298

Observation cb0c1590-2025-4a6d-91dd-c5cd6381f695 · outbound

This paper cites A Multimodal Framework for the Detection of Hateful Memes.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models A Multimodal Framework for the Detection of Hateful Memes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.182942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.182942Z digest=sha256:1bc315142f078a857594f8a10a557134724278c554e3d0e46a98688e6ca1b4ef

Observation 12607ca3-8f04-45db-b528-c90674b99931 · outbound

This paper cites Visual Instruction Tuning.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Visual Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.188523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.188523Z digest=sha256:7159fa7711229d95bc3f7b2e6269ef57d5f908c95f04a3420961cb57705430bf

Observation 92183d9d-f78a-4287-ae26-597e4739640c · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.193622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.193622Z digest=sha256:590f9c4bcee0ffe9c29debd307a2612ab26ed9b12910a0d567c07f956f387642

Observation 16d433d1-594f-41b7-90f4-22b6da9fb43c · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.198790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.198790Z digest=sha256:66a5fd413df619d6778ce8af7a4cc01a81ed8bd9f05f75bd507cb062c6b5785c

Observation 81d60a65-1c98-4d21-8c04-49174206c2b1 · outbound

This paper cites Hatexplain: A benchmark dataset for explainable hate speech detection.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Hatexplain: A benchmark dataset for explainable hate speech detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.966826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.203784Z digest=sha256:d324dea8138c3d2a573393288f00fe8989f765ecf4d435c294a648bf0dc8995f

Observation 32d06113-ec91-4234-ae51-93dbe547cd37 · outbound

This paper cites ETHOS: an Online Hate Speech Detection Dataset.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models ETHOS: an Online Hate Speech Detection Dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.208724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.208724Z digest=sha256:92528c2017417299664a8d653c9ea0f58cd99e3a4cc8ff81a4a64f82395e350a

Observation 3e834cf9-0d87-4b47-b0bf-850ddca0a227 · outbound

This paper cites Vilio: State-of-the-art Visio-Linguistic Models applied to Hateful Memes.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Vilio: State-of-the-art Visio-Linguistic Models applied to Hateful Memes

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.213637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.213637Z digest=sha256:c6f376b8477146197607af72e0f7edc28fdb01d7e3663409cc09e5e7f616e01a

Observation ec4bab69-10c7-4529-92a0-14085d108fb3 · outbound

This paper cites GPT-4 Technical Report.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models GPT-4 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.218559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.218559Z digest=sha256:470e5111662b4364be8e29580abe5939c0e32b18f98eec3a54946d1ee68535e7

Observation 0f8bd0e9-2177-4d03-b8dd-4e67ba36e8e1 · outbound

This paper cites Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.223211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.223211Z digest=sha256:c576fde67d1c76d823c6bc49ab5bb7926ec3f3db5982be6e76026242432943ce

Observation 1da24017-0a3d-41c0-b2ae-fa22a565caab · outbound

This paper cites Learning transferable visual models from natural language supervision.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Learning transferable visual models from natural language supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.228722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.228722Z digest=sha256:ca13dfc912dd608933d8eb268f31232b572a139d82c7d33684d023f5a391f8a3

Observation 6d9b3b92-8ba6-4e34-8337-a62d3532c2dc · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.233478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.233478Z digest=sha256:ca1d16fdc51d03f75ec8e054834911002a4c1dcb18602fe8e48712df258e92f9

Observation c781a117-7ef3-44fe-929b-d80fef9691a6 · outbound

This paper cites Detecting Hateful Memes Using a Multimodal Deep Ensemble.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Detecting Hateful Memes Using a Multimodal Deep Ensemble

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.238525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.238525Z digest=sha256:bbe029074344d7d402f15543e5fb70e37d6eb588fcb56b25d1181a5a2a6b1b12

Observation 9eda2e25-6c3e-42d7-8b47-906f7f2afa7b · outbound

This paper cites Detecting formal thought disorder by deep contextualized word representations.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Detecting formal thought disorder by deep contextualized word representations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.927220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.243316Z digest=sha256:9250ca788910ec0e4d2db698fde2e61194eb94ca70c69673ea6b67b2dcb8e291

Observation ecba3b8f-09f6-407b-979a-efc115661f23 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.248211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.248211Z digest=sha256:f658becfd4a8b8e31704cd18fc70529ccd0aecab1134b306422f6c9b1de283dd

Observation 709d1768-4ad3-4869-8d48-0abb639ea700 · outbound

This paper cites Large Language Models Encode Clinical Knowledge.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Large Language Models Encode Clinical Knowledge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.252964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.252964Z digest=sha256:b5ca17465a5dcd56778e825c721b074f637cbe760002d613dd0b5740ac99c366

Observation 7e31eb3b-a8a7-4db6-b683-1e64fc33f7f8 · outbound

This paper cites Resolution-robust large mask inpainting with fourier convolutions.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Resolution-robust large mask inpainting with fourier convolutions

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.896919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.257972Z digest=sha256:4751ad8ed2ffac486777cdb457a8c73860ec1473ad1210f62d8fc91f8454b67b

Observation e860ffdb-c39d-4dfa-878e-80d67f035e85 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.262753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.262753Z digest=sha256:f0428e765d56ddaf0fbdc0eb2d3b04af96193f0aaae281fee793bc8ca9ed4861

Observation fe78932e-7f8f-4f0b-849d-807cb4ac6294 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.268336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.268336Z digest=sha256:e9db60abf5456ecbf246c293a882201353ffe1a19285685572c6054460beddde

Observation 2b00b850-8e44-4985-8f20-bf7c41cd78e9 · outbound

This paper cites On large visual language models for medical imaging analysis: An empirical study.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models On large visual language models for medical imaging analysis: An empirical study

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.877975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.273511Z digest=sha256:5257aa1f880a3dae4d4ccd9f09687a9ee224971060e11aaa1b45754c0f2ebfbb

Observation 68a69b57-5211-4cec-aab3-d4b763283f99 · outbound

This paper cites Detecting Hate Speech in Memes Using Multimodal Deep Learning Approaches: Prize-winning solution to Hateful Memes Challenge.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Detecting Hate Speech in Memes Using Multimodal Deep Learning Approaches: Prize-winning solution to Hateful Memes Challenge

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.278432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.278432Z digest=sha256:ebfb17b8ea4afe4b1b3b4e2d228ebdd72dcf7a64de364a363f4115db18025428

Observation 176e6f5b-20fb-42bf-9ac9-65154d5a5c81 · outbound

This paper cites Memecraft: Contextual and stance-driven multimodal meme generation.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Memecraft: Contextual and stance-driven multimodal meme generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.859820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.283296Z digest=sha256:dadedc76855433bf8dbdb913f0bdc0eac0c5f0b9a1446f8b3289efd1e4982714

Observation d48cd019-2e12-462d-9843-8ecb8b589f24 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.288109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.288109Z digest=sha256:2d9593df777efe571a5ad28b711e37c2c71bad8c2029c4badc999bf7ca466751

Observation 2c9ce84d-1427-4cd6-bc88-1e6c9aee62de · outbound

This paper cites Aggregated residual transformations for deep neural networks.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Aggregated residual transformations for deep neural networks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.293095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.293095Z digest=sha256:3bb3f48924e480a636c76a961bfc152231a793bf8f1a173a28223ce23024a451

Observation 7c3ae4b9-e547-49c0-8cfa-c47ea4533354 · outbound

This paper cites Coded hate speech detection via contextual information.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Coded hate speech detection via contextual information

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.830858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.297774Z digest=sha256:eea4cbebeb94a5a1a65ab85155530e29a307ba0d9702199cb08ce2a0894e4f3f

Observation 5bb76d33-83f3-434a-a6c0-ce0a5f822d11 · outbound

This paper cites An empirical study of gpt-3 for few-shot knowledge-based vqa.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models An empirical study of gpt-3 for few-shot knowledge-based vqa

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.813384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.302275Z digest=sha256:e3446b946680b107d3891a71e78001f49ab78b0c1df368d5c8c73cefaebf3a5c

Observation 041eb16d-2f38-4db5-a17d-8deb554b4ecc · outbound

This paper cites Inpaint Anything: Segment Anything Meets Image Inpainting.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Inpaint Anything: Segment Anything Meets Image Inpainting

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.307144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.307144Z digest=sha256:b45fdbcb072c8f89e4492c7c96182125f5d4bd72f87acc2c3188c3dbae303dc0

Observation fa5ea7f7-4550-4bbf-a441-ac8e85e1a750 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.312338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.312338Z digest=sha256:e1b06107b577ba27be5de6ae16a2a3e8adc9529165af7800420e2d90e1bed463

Observation 42b50c13-34e7-4099-9e99-690d94fb0514 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models OPT: Open Pre-trained Transformer Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.317391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.317391Z digest=sha256:bf377605f5ce4fbf8f61611de80f8a1222d4d529327035adbb668462ad379ed9

Observation 2e316939-a3d1-44fd-ab84-4eeaa6169eb7 · outbound

This paper cites Enhance Multimodal Transformer With External Label And In-Domain Pretrain: Hateful Meme Challenge Winning Solution.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models Enhance Multimodal Transformer With External Label And In-Domain Pretrain: Hateful Meme Challenge Winning Solution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.322480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.322480Z digest=sha256:fb3dfc4525ccc6c14be791daa57a8a8ecc80e13e307f30ee6c83ce193088e311

Observation d1570484-a50e-4715-a304-3ca31926a19e · outbound

This paper cites C.2 Multimodal: unimodal pretraining Multimodal models from unimodal pretraining typically combine the output or the features of vision and linguistic models.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models C.2 Multimodal: unimodal pretraining Multimodal models from unimodal pretraining typically combine the output or the features of vision and linguistic models

Reference 768

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:55:28.795461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:55:28.327964Z digest=sha256:0832ce6983a2910a50b2bf8c0eef668e52654e660f88271103ce3986650a25a9

Pith citing papers

No inbound Pith citation observations are available.