Pith. sign in

Paper Citation Record · LEDGER

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection

As of 8 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2509.13608.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.13608 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:30:29.434614Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1000e86f-98a5-4714-9d12-ac504861e9dc · outbound

This paper cites GPT-4o mini: advancing cost-efficient intelligence.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection GPT-4o mini: advancing cost-efficient intelligence

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:28.327387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:28.327387Z digest=sha256:29c191efdd8b0f8d081c49a3c722da859065bfc12a2f49683961b154a9a9c108

Observation 74634001-9b30-416c-b525-439e772a006a · outbound

This paper cites ChatGPT Statistics 2025.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection ChatGPT Statistics 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:28.378517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:28.378517Z digest=sha256:4a2a9c27337f216f847d9b4dd0b22a171a6ecf3583306288fee1d8659fc4031c

Observation 39719fdb-deff-4cba-b0b9-78b83161c184 · outbound

This paper cites The hateful memes challenge: Detecting hate speech in mul- timodal memes.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection The hateful memes challenge: Detecting hate speech in mul- timodal memes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:28.427662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:28.427662Z digest=sha256:1bc7fb08776e4015cfb5a6c169550e873a48db3f0cdb499ddf7527e75b1f003a

Observation 3e9bfc75-46d5-40c5-b377-9e8566541b5f · outbound

This paper cites Cooperative in- verse reinforcement learning.Advances in neural information processing systems, 29, 2016.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Cooperative in- verse reinforcement learning.Advances in neural information processing systems, 29, 2016

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:28.511564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:28.511564Z digest=sha256:f8f29c4d5a24f0028b1991ff5ad2ca0eb5e093112f16ca9bef32218a77497a81

Observation 89eeb79d-e1ac-45fe-a168-562b51cb1e58 · outbound

This paper cites Viking, 2019.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Viking, 2019

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:28.570854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:28.570854Z digest=sha256:66e32290622ab0a3f10a951063b1206f7f034c09268a4e8680978b2047684180

Observation 6250ec90-87ae-4f03-91a9-c28a2c735827 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:28.629504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:28.629504Z digest=sha256:a0a2b3f8ed5d20fae824f57f5b9fce7096a28c7b0de37fccfcf2c41eebd946d0

Observation 59801848-7a23-4750-840d-8d662980be26 · outbound

This paper cites Hate-CLIPper: Multimodal Hateful Meme Classification based on Cross-modal Interaction of CLIP Features.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Hate-CLIPper: Multimodal Hateful Meme Classification based on Cross-modal Interaction of CLIP Features

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:28.691674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:28.691674Z digest=sha256:d46c35c8497665c15e30dd88ba30380741e5745ae798e1788ce60ed309b38d12

Observation e09ab6d1-fbea-4974-bb98-1d679ce53eff · outbound

This paper cites Singh, and V.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Singh, and V

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:28.740712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:28.740712Z digest=sha256:0cd401375c682146d28b84d4834ec8a56a65722878f0b79ccbef09917195e6bd

Observation d027d822-aaae-4b4f-90a2-0caecebd5e1c · outbound

This paper cites Multimodal detection of hateful memes by applying a vision-language pre-training model.PLoS one, 16(9):e0257302, 2021.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Multimodal detection of hateful memes by applying a vision-language pre-training model.PLoS one, 16(9):e0257302, 2021

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:28.803755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:28.803755Z digest=sha256:aee2796bc202a9882c797071416d5c7e617b9efe18e5cc7a3f411e9e6fdf8ad1

Observation 6b1ac153-4b87-4865-9dfc-307f2ba5b179 · outbound

This paper cites Multimodal hate speech detection based on multi-task learning.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Multimodal hate speech detection based on multi-task learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:28.835192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:28.835192Z digest=sha256:26444ceb6beae6da0c0ffd5c202a8ec4c2336cbcb2d83ebbddad7f8ec1217fa3

Observation 4767ef8c-c91b-4531-a401-6adbfcbdb9b6 · outbound

This paper cites Demystifying Hateful Content: Leveraging Large Multimodal Models for Hateful Meme Detection with Explainable Decisions.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Demystifying Hateful Content: Leveraging Large Multimodal Models for Hateful Meme Detection with Explainable Decisions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:28.893257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:28.893257Z digest=sha256:8f537d644401f5c192d8d9273cc42b3dbeeacfae11df5e2c0c2540390930c8ef

Observation 72a80e90-7474-487a-9275-91f58e0d4e25 · outbound

This paper cites Toward Fairness via Maximum Mean Discrepancy Regularization on Logits Space.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Toward Fairness via Maximum Mean Discrepancy Regularization on Logits Space

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:28.962552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:28.962552Z digest=sha256:26f8967540f88cc68b02a03aedb10adf8b8d433adc6eea644f04d8339405acd5

Observation 2d31c79a-d3c7-462f-964f-b89a06b4b6b0 · outbound

This paper cites Classification of Carotid Plaque with Jellyfish Sign Through Convolutional and Recurrent Neural Networks Utilizing Plaque Surface Edges.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Classification of Carotid Plaque with Jellyfish Sign Through Convolutional and Recurrent Neural Networks Utilizing Plaque Surface Edges

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:29.002464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:29.002464Z digest=sha256:a294bb22298f455d3a388748d8d52cfa9f4f6ad4b4d0d94bf63cf2d7a5b11c44

Observation 569bbb60-5c61-4576-ae5d-01e637352e70 · outbound

This paper cites AI ”safety” vs ”control” vs ”alignment”.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection AI ”safety” vs ”control” vs ”alignment”

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:29.084744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:29.084744Z digest=sha256:840a0d0cab72155e39095b34b0b621dda07e318c84d74bd040938784b621f0ad

Observation 7193b448-e006-4a89-8b57-be90ae8a648e · outbound

This paper cites How AI is shaping the next generation of brand safety and suitabil- ity.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection How AI is shaping the next generation of brand safety and suitabil- ity

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:29.134117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:29.134117Z digest=sha256:98fc3b431dde7ba70399650e2f1a7339298d423f3bd43c3e8ebb28bc8e62e27b

Observation 4e3f0002-9a6e-4dd9-bfef-8b569985bc22 · outbound

This paper cites YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:29.203246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:29.203246Z digest=sha256:9d549c9251e49e3a9d008e6ec7a52d936617991f135e6d222cdd6a3acb788a2e

Observation 7b2d6dea-a34c-4679-af48-f1a3f2b8f599 · outbound

This paper cites SocialCounterfactuals: Probing and Mitigating Intersectional Social Biases in Vision-Language Models with Counterfactual Example.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection SocialCounterfactuals: Probing and Mitigating Intersectional Social Biases in Vision-Language Models with Counterfactual Example

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:29.267971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:29.267971Z digest=sha256:4ab42bdcab8ea2e1b2f5c1a9554d359e9db4de15b2de19389146dcadd800dd76

Observation b8ba8047-e7dc-4314-9691-3abd4de223af · outbound

This paper cites Topologically Faithful Multi-class Segmentation in Medical Images.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Topologically Faithful Multi-class Segmentation in Medical Images

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:29.331818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:29.331818Z digest=sha256:32025e3623da8fba956e60297f6a882e335614598ab807a684873f1764c6f5d1

Observation 3894a040-3054-4e18-a724-87ac4b6121f9 · outbound

This paper cites Safety of Multimodal Large Language Models on Images and Texts.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Safety of Multimodal Large Language Models on Images and Texts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:29.384240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:29.384240Z digest=sha256:aee9d01cb5ec5ad45adfe3b5cf1f6fd2be3140512a84fcc0b2574f927a90b671

Observation 305ba678-17fb-4ee7-891c-e7cc8f01e5e3 · outbound

This paper cites Safety of Multimodal Large Language Models on Images and Text.

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection Safety of Multimodal Large Language Models on Images and Text

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T16:30:29.434614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:30:29.434614Z digest=sha256:ebe125bf7e844618d19ce5bb2d35165a3f910bfc94dfbc5d3420bb45cd5298c5

Pith citing papers

No inbound Pith citation observations are available.