Pith. sign in

Paper Citation Record · LEDGER

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns

As of 22 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 1 inbound Pith citation observation for arXiv:2501.16750.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16750 v1

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T11:03:29.370944Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T01:02:51.934157Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T01:02:53.997511Z

Reference resolution

100 of 100 outbound references displayed

  • verified exact3
  • verified fuzzy53
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ad7f745-cbff-421d-9ec5-fc59bfb5fbbb · outbound

This paper cites https://en.wikipedia.org/wiki/ Coleman-Liau_index.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://en.wikipedia.org/wiki/ Coleman-Liau_index

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.902823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.902823Z digest=sha256:6b11086470ebb73d360e3c11db060683cb05da087ae380fa04a3df65cb2fd7ef

Observation 03b1fd37-1304-4c6b-ad09-a10ebfb26c4c · outbound

This paper cites https://github.com/unitaryai/detoxify.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://github.com/unitaryai/detoxify

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.907255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.907255Z digest=sha256:01e6a58f98071a31c9890430e4057de22f4ea2a927d1ac8503e79a2553be180a

Observation d4f50355-d933-4a08-88d7-26d878d1b1b5 · outbound

This paper cites https://gdpr-info.eu/.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://gdpr-info.eu/

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.911411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.911411Z digest=sha256:37a74b21f53bfc284fd669266b6e445dba25be13cdb4902cf7b5ac661b1cc061

Observation 6dcf60a8-89d3-48d1-939a-e30de2bde689 · outbound

This paper cites https://chatgpt.com/g/g-w0y3CvDM 9-freddy-griffin.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://chatgpt.com/g/g-w0y3CvDM 9-freddy-griffin

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.915052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.915052Z digest=sha256:d1f50cc7ac6f35601e00e66c899cb05d5d7c55e8dd81d65ff2a56cf28facb9eb

Observation a63bdbfd-cad7-412f-9a4b-8f93bc1429cd · outbound

This paper cites https://chatgpt.com/g/g-JlQ9WBdHB-hate /.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://chatgpt.com/g/g-JlQ9WBdHB-hate /

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.919260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.919260Z digest=sha256:17d8f80a023cb789fe2165c9aa31d4633cc55b7d3c9eb6aabe272fa9eed26a98

Observation fd1fe658-5211-4357-b5f7-f4cc5a9b7645 · outbound

This paper cites https://chatgpt.com/g/g-87uTmBE65-ru de-gpt.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://chatgpt.com/g/g-87uTmBE65-ru de-gpt

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.937731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.937731Z digest=sha256:fb55b636236603116441f0d0c74ecc612fd41fcadc5f2715f4cb8ca962ba2048

Observation 0bbde68a-1663-43f4-a4ce-e354f24e7abe · outbound

This paper cites https://www.perspectiveapi.com.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://www.perspectiveapi.com

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.967561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.967561Z digest=sha256:2f031039e89fde2ea2effe00c191010fec09915b769ea0be1ef0de7687c242a7

Observation 3ac119e8-d12a-4ca5-a0b5-a24c13c232be · outbound

This paper cites https://osf.io/edua3/.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://osf.io/edua3/

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.987546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.987546Z digest=sha256:ab6b258ad411bdf80b98c61ed6f37f42feea3fb7ca8c7e027007fe7685f765e3

Observation 8edefd31-f421-45fc-b7bc-f66f21388a73 · outbound

This paper cites https://lmsys.org/blog/2023-03-30-vicuna/.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://lmsys.org/blog/2023-03-30-vicuna/

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.018050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.018050Z digest=sha256:31ffea6f4aba51f7c45977527328450cd7118fa373f23b46e5094504c045a00e

Observation ebaeb6ad-bb19-4144-8aed-535777dd7068 · outbound

This paper cites ADL Task Force Issues Report Detailing Widespread Anti-Semitic Harassment of Journalists on Twitter During 2016 Campaign.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns ADL Task Force Issues Report Detailing Widespread Anti-Semitic Harassment of Journalists on Twitter During 2016 Campaign

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.048010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.048010Z digest=sha256:a7e5b06fb79c04d27a82d829ac896309cc3a718f32bab280319600f12a811340

Observation e659f29b-ae11-4dec-bf41-e49ed0e947be · outbound

This paper cites Online Hate and Harassment: The American Experi- ence 2023.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Online Hate and Harassment: The American Experi- ence 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.056730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.056730Z digest=sha256:84e759f97782b247985cba9097d0eaf9699e713d57e6aeec7e335c5bebe4d2dc

Observation 350e5bd4-1ec6-429b-b67f-34892d7950fc · outbound

This paper cites Aunties, Strangers, and the FBI: Online Privacy Concerns and Experiences of Muslim- American Women.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Aunties, Strangers, and the FBI: Online Privacy Concerns and Experiences of Muslim- American Women

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.094242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.094242Z digest=sha256:5922adfa2839766cfd20706c5df406926018273256b706f5809454ab0e2c09a1

Observation ef305288-9568-414e-bd57-75a81094b2ab · outbound

This paper cites Google’s Jigsaw was trying to fight toxic speech with AI.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Google’s Jigsaw was trying to fight toxic speech with AI

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.127008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.127008Z digest=sha256:95abdb91816b4347800af6614c81b339e4d75bcb8a8ca1560be7a4a435a86c28

Observation 8a0f5ff9-abec-4355-a5b8-1cdd73640df6 · outbound

This paper cites Robust Hate Speech Detection in Social Media: A Cross-Dataset Empirical Evaluation.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Robust Hate Speech Detection in Social Media: A Cross-Dataset Empirical Evaluation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-10T11:03:29.690850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.161185Z digest=sha256:3b7de5073796e5ea73da8859fe307163592727f644db9f5904d18df699d1b060

Observation 085a2934-d5cf-43a2-977e-d98ce24af45b · outbound

This paper cites METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.189110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.189110Z digest=sha256:6d49ef2f9c14bb3c3c2745d57b65fe02b9905db08671cacb42dea2cb245a9440

Observation 20622d68-2f80-4d1b-bf9d-8f72e7490f70 · outbound

This paper cites The Pushshift Reddit Dataset.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns The Pushshift Reddit Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.199948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.199948Z digest=sha256:b990abc126a52491f90b799a98110e44761c66d5dfd22c121a5f21a18613f6fb

Observation b2799bf0-0fce-4029-a73a-5d097b6c6e68 · outbound

This paper cites Nuanced Metrics for Measuring Unin- tended Bias with Real Data for Text Classification.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Nuanced Metrics for Measuring Unin- tended Bias with Real Data for Text Classification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.218992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.218992Z digest=sha256:ec535ed0297961b628fcfa85a80049592a969c4a3b950e74bd2189a63d54fd39

Observation 7793d8ca-21ff-48cf-997d-30f68d8a965d · outbound

This paper cites HateGAN: Adversarial Generative-Based Data Augmentation for Hate Speech Detec- tion.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns HateGAN: Adversarial Generative-Based Data Augmentation for Hate Speech Detec- tion

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.232006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.232006Z digest=sha256:438625888be09a982a08e84f833d46c44f4cc70a50a2caa9c2309ae181fa9b42

Observation e8e5201c-059b-43ca-bd76-21bab6bfc547 · outbound

This paper cites John, Noah Constant, Mario Guajardo- Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns John, Noah Constant, Mario Guajardo- Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.241154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.241154Z digest=sha256:4cc801b1757e8f00b60c534d3e776ae53fabc7476f6ef9a10d6bdecd244899b4

Observation 9d5aa832-e0aa-453f-b517-7e8ed0cccfec · outbound

This paper cites Hate is not Binary: Studying Abusive Behavior of #GamerGate on Twitter.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Hate is not Binary: Studying Abusive Behavior of #GamerGate on Twitter

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.254179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.254179Z digest=sha256:444cee2a93ee3719a96a45a0882022896f591f199f4ab1a13d734f244c921174

Observation a93cd6a6-24b0-4a66-8eb4-3a6261c457f8 · outbound

This paper cites Christiano, Jan Leike, Tom B.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Christiano, Jan Leike, Tom B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.269724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.269724Z digest=sha256:0fd5104ac21aaa7f3aa7616c6cb3897025453927a17a46071f7b7a04903e0eb0

Observation 2df97a1e-c6ba-4887-bd58-49efbe1fe76c · outbound

This paper cites Hate Campaign.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Hate Campaign

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.380797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.284063Z digest=sha256:b169567cf79bc96bd023f2d2b176e0209d0235c675fe0192d1ebd482502050b4

Observation 99b942b8-788a-41f9-a1d3-8136f5181f24 · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.368765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.296688Z digest=sha256:b50bcd906eb870276209f6b36fa6a9465e2e02a44b3673deab418c5aec952cbb

Observation 6bc594ad-b278-462e-b48c-96687bb8aec7 · outbound

This paper cites Toxicity in ChatGPT: Analyzing Persona-assigned Language Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Toxicity in ChatGPT: Analyzing Persona-assigned Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.304717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.304717Z digest=sha256:593769092bceb1ce8feebcb21f44f5d41a3fd7324a79d3134cda5417fc686fb8

Observation 28de4e90-701c-4e55-b3dc-17e14de7a540 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Trans- formers for Language Understanding.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns BERT: Pre-training of Deep Bidirectional Trans- formers for Language Understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.357619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.309122Z digest=sha256:547355836b485ae49a9da501fd10d6b1ae706621b42af3b786324075b016b30e

Observation eea97b84-2711-4c36-8827-81c04e587bf6 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:31.347120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.315782Z digest=sha256:31b654d8769a9d27e3b52467802fa2bbf2a799b82feaca4f0023ae98b54df9b2

Observation 613a99d5-4c1b-4084-b042-e22588cb4e20 · outbound

This paper cites Paraphrase a text.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Paraphrase a text

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.337317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.327174Z digest=sha256:34ff207a55652bce310787a6c7f798b88e5c1957b815708348ba256f9a472e7f

Observation 76ba4c38-f8d0-461e-b25b-b4f8cff0cc8f · outbound

This paper cites Guide: Large Language Models-Generated Fraud, Malware, and Vulnerabilities.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Guide: Large Language Models-Generated Fraud, Malware, and Vulnerabilities

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.326970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.333912Z digest=sha256:5484b57908663099bb6e09c943d66403184342adfca326a45d03c203106535e1

Observation 04699b60-c2af-4c3b-b9f2-8c5f1663c1fe · outbound

This paper cites Black-Box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Black-Box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.315813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.350922Z digest=sha256:54baf3172f863664dce57018c275f9cc1aeeffdf60b01ae77983116eec2afebd

Observation c7948ace-3c0f-4252-a6a9-4621348310dc · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.360191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.360191Z digest=sha256:6721eada002a4702936c42f8b96002be786edbd6e87de216e03d30d847836273

Observation 7352fa66-8c9f-435c-a322-6dd973b7e909 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:31.304561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.364370Z digest=sha256:30cd0693b001d7448688b9968453e0541f786ea63d22f48bab124b3bbf29d257

Observation 8a497de1-fa71-4e2f-a6e4-cbf5a18d9def · outbound

This paper cites Hancock, and Zakir Durumeric.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Hancock, and Zakir Durumeric

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.289569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.368006Z digest=sha256:404b2a8ebff56dcafad8eaa15998c5addb37e52b73b13358bc7b4047de299143

Observation 3fe4fc75-27d8-45af-ab7c-d413cf709c09 · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.267866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.372041Z digest=sha256:6cd2f6fdeb55df694014178a0850eeceac4e80bcc1a924f5c7bcf9c45dc540a5

Observation c02e3e85-f651-41d7-aad1-4cec723a7953 · outbound

This paper cites Hess, Kelley P.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Hess, Kelley P

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.213406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.378471Z digest=sha256:17d92b12276fceebb477a3bbaff26ab3b87ecc70c62cedebc00b3b2d317662d0

Observation cdf58a07-63fb-49cd-8294-29708f51959c · outbound

This paper cites Deceiving Google's Perspective API Built for Detecting Toxic Comments.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Deceiving Google's Perspective API Built for Detecting Toxic Comments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.386994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.386994Z digest=sha256:22bcc133f0e79fddbecd9d08ccad1362937192af78e54e5e85b6ba6406d7f182

Observation c7bd1e3f-72df-40d0-b70b-8a873d84cdfb · outbound

This paper cites Adversarial Example Generation with Syntactically Controlled Paraphrase Networks.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Adversarial Example Generation with Syntactically Controlled Paraphrase Networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.149871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.390871Z digest=sha256:bb6e946a875fef04d78767a3c01a78ed9bdba2e4ee262b69e0c28e5660425db5

Observation 867775fb-fa87-4bb4-ab05-25d298584f0d · outbound

This paper cites High Accuracy and High Fi- delity Extraction of Neural Networks.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns High Accuracy and High Fi- delity Extraction of Neural Networks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.082157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.416115Z digest=sha256:a68291ca8c4bc91233912d6b367161fe7846f212dec61ee90235827a5c5640a5

Observation 18b20293-f9a8-4252-9dcc-d9f66218911d · outbound

This paper cites Is BERT Really Robust? A Strong Baseline for Natural Lan- guage Attack on Text Classification and Entailment.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Is BERT Really Robust? A Strong Baseline for Natural Lan- guage Attack on Text Classification and Entailment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.447493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.447493Z digest=sha256:86597dd580a60694a54b8b0d4e1215542f9669c2f85bd4d8c783024c43ba9f2e

Observation 68f67a60-d589-40c7-854f-bb23bb8592e4 · outbound

This paper cites Toxic Comment Classification Challenge, 2017.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Toxic Comment Classification Challenge, 2017

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.055167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.484402Z digest=sha256:8706208b600e149f3fd6d00e37f41b138f068f9a1e9a31593a1d501e7b44f9e0

Observation 850c24bd-9013-43db-a4cb-0eb5d49bbb12 · outbound

This paper cites Jigsaw Unintended Bias in Toxicity Classification,.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Jigsaw Unintended Bias in Toxicity Classification,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.042736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.491766Z digest=sha256:5ff90dd368ec3dfb6a05f6e2cc0cda7c38e7628310325922b89ab1500a8a6ff1

Observation 3560f2f0-ae31-4139-b05c-a2b323d259dc · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:31.030756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.537604Z digest=sha256:1a2ed2701b7ba39da5865be4a33c3d26468df459c5235fcad26358c72c81d7af

Observation 1119b65d-bc8f-44bd-80f4-5b0d50529dbe · outbound

This paper cites Content Analysis: An Introduction to Its Methodology.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Content Analysis: An Introduction to Its Methodology

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.020120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.567684Z digest=sha256:e3bc43e40edb87581d1d625b3f425e86a712cb5236414a280baf173e6f146595

Observation fd7a6e8a-f462-44dc-b6ce-68e12e98ba1b · outbound

This paper cites Parikh, Nico- las Papernot, and Mohit Iyyer.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Parikh, Nico- las Papernot, and Mohit Iyyer

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.006299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.590059Z digest=sha256:59522266b99136e9b4cad3edbcf44bd82f608c51e91f04baa36eb26928788e3b

Observation af4b1e44-f981-4f32-92e8-7b956d279b36 · outbound

This paper cites TweetBLM: A Hate Speech Dataset and Analysis of Black Lives Matter-related Microblogs on Twitter.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns TweetBLM: A Hate Speech Dataset and Analysis of Black Lives Matter-related Microblogs on Twitter

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-10T11:03:29.570718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.620351Z digest=sha256:98375cac891111917646be7718cdca0b850f47ad45aa220b7b188838ed23d2d5

Observation 5ba05853-35ad-47aa-80bb-f9575da20453 · outbound

This paper cites TextBugger: Generating Adversarial Text Against Real-world Applications.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns TextBugger: Generating Adversarial Text Against Real-world Applications

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.995160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.652938Z digest=sha256:31776167029a4a11c5749cf0676aec3f2c457de0e81853153a563703a005af1b

Observation 4d76440f-0804-46ca-aff2-72fd0895ebf8 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.689441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.689441Z digest=sha256:59267758317e494df0beeb73981ddb17401ac6afb52268fec68aaf9707a0356b

Observation cd03df65-d1a3-4b55-95ec-c71b069e6fb4 · outbound

This paper cites A Holistic Approach to Undesired Content Detection in the Real World.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns A Holistic Approach to Undesired Content Detection in the Real World

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.983943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.710325Z digest=sha256:acb4c6d19f2df8f21671cdc7314bf7e3a335508f0bbd73e6aed9f87a0d499326

Observation 0a6cefe5-9337-45eb-991d-8652c1d33b41 · outbound

This paper cites HateX- plain: A Benchmark Dataset for Explainable Hate Speech De- 14 tection.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns HateX- plain: A Benchmark Dataset for Explainable Hate Speech De- 14 tection

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.958534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.727341Z digest=sha256:1a7f95f80ff11e73616feb719e1f9d5bbd292e636157505a462105c860755c52

Observation 515c39ea-1658-4802-a2eb-f09dec649daa · outbound

This paper cites AI Trained on 4Chan Becomes ‘Hate Speech Machine’.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns AI Trained on 4Chan Becomes ‘Hate Speech Machine’

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.930790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.739293Z digest=sha256:d8ad86c8dc0cc7ecc6e9fbbfeea334e565f55840de509ce7f98079392472a856

Observation fbfdfdfb-1768-4df7-a7e6-fe61556966a1 · outbound

This paper cites Mazurek, Florian Schaub, and Elissa M.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Mazurek, Florian Schaub, and Elissa M

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.892022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.749638Z digest=sha256:4fdf8fea4d906411f8e939c65015be837f593877f250ac1e86a9802da6fa8b38

Observation 76e81ca1-999c-42c7-93f7-be9c5dde4df2 · outbound

This paper cites The Challenge of Detecting Hate Speech.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns The Challenge of Detecting Hate Speech

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.796736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.766621Z digest=sha256:17e3fb152a4cb2eeefc730424e59bd08b6f3a644eb8cab461486a146d18fd0c9

Observation df5f1a68-3d31-4125-893d-ab986c018863 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:30.746708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.783465Z digest=sha256:2432e2acb30dd7129a2b134f57578be1de3b46ec067380800a30b0039fa5a430

Observation 54135c04-8df8-43c4-97a4-01d6cffc24ee · outbound

This paper cites What is hate speech.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns What is hate speech

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.735138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.800210Z digest=sha256:3ecff06fe6897d86fd8b9e33154fca38ca32cc70d92f66e2b032929a0d34cfc4

Observation 4c2f3670-5515-45eb-8bde-5f11cd7172ee · outbound

This paper cites Handling Disagreement in Hate Speech Modelling.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Handling Disagreement in Hate Speech Modelling

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.723835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.805787Z digest=sha256:1ac92a666a8cfb18e4497fe4d9f14b19e26c9ce2bcdc0e91766c9d9775620b7f

Observation e5466900-68dc-497c-b282-cf1c31e194ff · outbound

This paper cites I Know What You Trained Last Summer: A Survey on Stealing Ma- chine Learning Models and Defences.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns I Know What You Trained Last Summer: A Survey on Stealing Ma- chine Learning Models and Defences

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.712642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.810279Z digest=sha256:45f38e5299f5ca692f9626d7a0d1c14edb27fc105dd41f4b91e873b31617adfc

Observation dd69ecf4-ebe1-418a-9767-1f520db815f3 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:30.700183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.816793Z digest=sha256:760e4a200af99f09702ab06e368eef8c9077bb99f834acd13f1ca124183b26e3

Observation 566f0571-34f3-4fe8-8d60-856a0dd5e825 · outbound

This paper cites Introducing GPTs.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Introducing GPTs

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.688919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.830570Z digest=sha256:6a4ab9170f6da3ec12434ad603d565b75e107b83eafd4653be7dbb0972ef0412

Observation be5b8a04-8b95-4a9e-b688-049aef60b248 · outbound

This paper cites GPT-4 Technical Report.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns GPT-4 Technical Report

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.836125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.836125Z digest=sha256:41eefb8f5ac859d20feb627e275a591701f195c0389cb9b11ccc743f964dbb93

Observation 055fc79f-3b4f-4522-bd68-cddc1c2c34b6 · outbound

This paper cites Offensive and Hateful Text Multiclassification.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Offensive and Hateful Text Multiclassification

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.678911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.847019Z digest=sha256:8bd9cd6eea9a261518c272fed042dc46230a688621fda718115cde550021059f

Observation 4512db10-1eff-4c35-b2cd-94e30131da9d · outbound

This paper cites Facebook’s race-blind practices around hate speech came at the expense of Black users, new docu- ments show.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Facebook’s race-blind practices around hate speech came at the expense of Black users, new docu- ments show

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.658933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.857094Z digest=sha256:e0d8d759e7daecb2ceb19183ab0fb23cfdff80dc0635d65d5b15df67fc9fe83d

Observation cc9834e9-ce0b-4570-978c-f1483673bc5a · outbound

This paper cites UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.872472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.872472Z digest=sha256:2c7c946dafc9c90afd7a33346a4b54fd6e0e60c654d42a4135dc46f3bdf85604

Observation bc3954aa-717b-4681-abc0-a62a2463d96d · outbound

This paper cites Gener- ating Natural Language Adversarial Examples through Proba- bility Weighted Word Saliency.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Gener- ating Natural Language Adversarial Examples through Proba- bility Weighted Word Saliency

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.626013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.884560Z digest=sha256:4c1d4496b61537c0aa826baf7550de1831150437545f39e5b84c7ced01e072f0

Observation 5d432944-2249-4239-95ba-b70f0ef1aa70 · outbound

This paper cites Schuller.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Schuller

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.571183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.891454Z digest=sha256:a1bc2b3b9f48cd9919ddc62d2a953b890a758c431753188b79d2b5021e916396

Observation 10176289-a493-4eb0-9720-25beae53e1c8 · outbound

This paper cites Margetts, and Janet B.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Margetts, and Janet B

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.505330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.902216Z digest=sha256:f9765e144cfc7c95d537abe1330d7d3fad1c6df7b503db349134c4f676da9da3

Observation 5c772455-2cf9-45d1-aa32-3115a565f95c · outbound

This paper cites Sachdeva, Renata Barreto, Geoff Bacon, Alexander Sahn, Claudia von Vacano, and Chris J.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Sachdeva, Renata Barreto, Geoff Bacon, Alexander Sahn, Claudia von Vacano, and Chris J

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.443534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.906112Z digest=sha256:e1d5e379a55692aa4adc9cbb33780305d5f5ab3133616dd701558bbda6d44c53

Observation 574f0623-0853-4351-997b-924c24e0f4f9 · outbound

This paper cites TUBERAIDER: Attributing Coordinated Hate Attacks on YouTube Videos to their Source Communities.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns TUBERAIDER: Attributing Coordinated Hate Attacks on YouTube Videos to their Source Communities

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.373777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.914055Z digest=sha256:5629683cdc8ce5f5c6ce3710347722b21021ca6331c9ed582b04b32cbe4aed8e

Observation 3dd02bca-aeeb-470d-9dc9-b95fc2e2d14b · outbound

This paper cites Generative AI as a Vector for Harassment and Harm.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Generative AI as a Vector for Harassment and Harm

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.361183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.917853Z digest=sha256:d012f095d4125e16991d98873625c8e3458f72a61a2f895b2c33cf8e18f65ec5

Observation b43cb975-e63a-4db4-af32-ce6610a0568f · outbound

This paper cites Man who harassed black student online must deliver ‘sincere’ apology, renounce white supremacy.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Man who harassed black student online must deliver ‘sincere’ apology, renounce white supremacy

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.350491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.922110Z digest=sha256:672922d3da7f2abacaf345adf8a9e8ec819567c9720f60c93ce38441c2a81e58

Observation c945fcff-def5-4e11-b494-4b0f1d8c8bb8 · outbound

This paper cites Do Anything Now: Characterizing and Evaluat- ing In-The-Wild Jailbreak Prompts on Large Language Mod- els.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Do Anything Now: Characterizing and Evaluat- ing In-The-Wild Jailbreak Prompts on Large Language Mod- els

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.338709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.926135Z digest=sha256:ad5e6e34f2f7e0fe3aee0840f4386806a8e920cc072854ca4f2a6fcc12211ae7

Observation 8cd9120a-fe98-49f3-845c-8e79e8e9c624 · outbound

This paper cites On Xing Tian and the Perseverance of Anti-China Sentiment Online.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns On Xing Tian and the Perseverance of Anti-China Sentiment Online

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.327376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.930210Z digest=sha256:8d454442206954633ccb618dc5dcd98e5a19a3a0250d2342d551f83438f55b4c

Observation 50ec7bac-64e6-4f6e-8608-02e913eced0b · outbound

This paper cites Model Stealing Attacks Against Inductive Graph Neural Networks.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Model Stealing Attacks Against Inductive Graph Neural Networks

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.316013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.933692Z digest=sha256:1ee184df4e21e98e906b5ecf02ff0f913988507b951bc8fe992fd6530693d0d5

Observation efd0ea7a-3324-4732-bbf3-3ddaf119f503 · outbound

This paper cites Analyzing the Targets of Hate in Online Social Media.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Analyzing the Targets of Hate in Online Social Media

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.304658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.937591Z digest=sha256:26367c672ea52ec1074a19064b4a70560a1f2db0da654b13c4de08cb5d44c31e

Observation 4f3f5cb9-3a3b-48a8-877b-806ceca6bd09 · outbound

This paper cites Learning to summarize from human feedback.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Learning to summarize from human feedback

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.955630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.955630Z digest=sha256:8145838bc911f3a36e96136df9d97a5e3c344b74d3482da62084bb02a9f1eee3

Observation 2ef07ce3-c05c-4fd6-ac2a-f360ec628db1 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:30.290822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.980546Z digest=sha256:d55a898598e995b7845c1d6ad63c05d3ae4c7f244df0f4c979f586b5fed9b24a

Observation 0c967b31-044a-49d2-a191-b7678e5d6db1 · outbound

This paper cites Large-Scale Hate Speech Detection with Cross-Domain Transfer.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Large-Scale Hate Speech Detection with Cross-Domain Transfer

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.210478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.015186Z digest=sha256:3021ec4272f5af96180db1d7475cc3592f7d3c2eed9ee12ab99a2e3bb9c17996

Observation e1d50bf6-b164-47f5-9bf6-7978ac6476dc · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns LLaMA: Open and Efficient Foundation Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:29.053661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:29.053661Z digest=sha256:5dc2cc76102bc59bbb6d0a545b1ec466e283d189fcfc797b598410e2a0ba7af4

Observation 4913f0b0-d5dc-4d5d-8ad7-f434389b6bda · outbound

This paper cites Reiter, and Thomas Ristenpart.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Reiter, and Thomas Ristenpart

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.194371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.071955Z digest=sha256:a643aaf2943e37f1b61efecbd5bc12f2b629583d438e5d58bf5c0c4913105233

Observation a76556e8-7611-4441-91cb-99eca8b791e6 · outbound

This paper cites Visualizing Data using t-SNE.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Visualizing Data using t-SNE

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:29.094145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:29.094145Z digest=sha256:97d90f8e51c5b599a499ceace7766eb7423dfddbc40d885fcf7f5732e1be20a9

Observation 62f85124-fd18-4fc4-b752-6cbff76e2090 · outbound

This paper cites Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.175553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.123862Z digest=sha256:2fc3a2b5edad2dceb1f74379bd9d9fac253d83937d2beb8eb1a87d67b0eb4b45

Observation d3ad8727-8c8a-46d9-8470-645e194b289e · outbound

This paper cites Moderating New Waves of Online Hate with Chain-of- Thought Reasoning in Large Language Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Moderating New Waves of Online Hate with Chain-of- Thought Reasoning in Large Language Models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.164132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.159151Z digest=sha256:e3be1cd1a68a2d9cb2f732cd12ae018791ef6b36324d1143d4df828c49006a70

Observation a7abee6c-1f43-47e3-9d9c-7ef49454d838 · outbound

This paper cites Vu, Alice Hutchings, and Ross J.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Vu, Alice Hutchings, and Ross J

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.152770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.193217Z digest=sha256:9597aa5adb047baca7b82436a5e3663b83ade6cfcd56e464c96b202862ecb285

Observation 3ffcd2df-e5bf-4d0c-ba60-94956a9dc078 · outbound

This paper cites There’s so much responsibility on users right now: Expert Advice for Staying Safer From Hate and Harassment.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns There’s so much responsibility on users right now: Expert Advice for Staying Safer From Hate and Harassment

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.125214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.206641Z digest=sha256:c027c68d8e6c4f04db88c25a29d23e1fbfced1922178516eabd41f8cd56b1057

Observation ee32c18a-1f36-4b04-b4aa-bc5d59673f41 · outbound

This paper cites Challenges in Detoxifying Language Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Challenges in Detoxifying Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:29.226885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:29.226885Z digest=sha256:acaca7ae37f9562926b429f4b9528097c359307208a7512cf5637268c7c3da38

Observation 10dcf881-889b-49a2-bdee-0601e963f804 · outbound

This paper cites Not All Asians are the Same: A Disaggregated Approach to Identify- ing Anti-Asian Racism in Social Media.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Not All Asians are the Same: A Disaggregated Approach to Identify- ing Anti-Asian Racism in Social Media

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.097329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.232894Z digest=sha256:b4e2a7248cb8163fa91aa1c166f0085235825f2579bc98092620a654499202b2

Observation cc1edeb2-f072-4d71-9e79-b69876745678 · outbound

This paper cites Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.058739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.239336Z digest=sha256:8edd4dd78dcbb7a37b0594bb403816e8b6783c1aaadfdf5b2278cadabe97197e

Observation afb446ee-7248-4722-8ea3-e0267bce9bd5 · outbound

This paper cites Fight Fire with Fire: Fine-tuning Hate Detectors using Large Samples of Generated Hate Speech.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Fight Fire with Fire: Fine-tuning Hate Detectors using Large Samples of Generated Hate Speech

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:29.928514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.244624Z digest=sha256:10af5db90ffc642e7afca3dc99812689173c7a9dc6eff5443d7d72b75b0fa19a

Observation 7d101714-7083-4b0b-9768-8b2465c0bf0d · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Baichuan 2: Open Large-scale Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:29.254537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:29.254537Z digest=sha256:6683008b919b438b420558f9d8a37e9c244d6c3a46db03e60fa6cc6419e0f04f

Observation 5e69bce8-4c74-4ab1-b41a-2832fbc0011e · outbound

This paper cites GPT-4chan: This is the worst AI ever.https: //tinyurl.com/2s4jh5p4, 2022.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns GPT-4chan: This is the worst AI ever.https: //tinyurl.com/2s4jh5p4, 2022

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:29.916299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.263941Z digest=sha256:f916c1b4e03f7b2b3250dba8e943c95ee3fb268594ffdd5326b94e59906358b3

Observation 854195b2-ce6d-410d-84fb-d210bef4997a · outbound

This paper cites OpenAttack: An Open-source Textual Adversarial At- tack Toolkit.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns OpenAttack: An Open-source Textual Adversarial At- tack Toolkit

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:29.900311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.276493Z digest=sha256:3b925b630e65c0dfb98b2ccb526ad9c5bf1e980b26cede17372ca26a225d67f1

Observation 19b25f2e-afc1-4a23-8ff5-27052e604aa4 · outbound

This paper cites SecurityNet: Assess- ing Machine Learning Vulnerabilities on Public Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns SecurityNet: Assess- ing Machine Learning Vulnerabilities on Public Models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:29.883125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.293814Z digest=sha256:6477828b63968ada44611ac94240f4c4166befa43b04d04232479b8a7693c663

Observation c9b65d46-ebc3-4c3d-990a-fc43901af425 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns OPT: Open Pre-trained Transformer Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:29.306584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:29.306584Z digest=sha256:33c33476384fa1b9a20ea068195aa30527d38382b220671d499463e6e497c964

Observation cb337bfa-4dfa-43ab-a3e2-8bd60141f990 · outbound

This paper cites Generating Natural Adversarial Examples.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Generating Natural Adversarial Examples

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:29.870122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.319248Z digest=sha256:50a5e7461faa9b1a4ca5d04d7571a76179bf0519c7f784531cf4b3509eed852e

Observation 2d3426a3-2e04-4116-adc9-8c80586222b1 · outbound

This paper cites PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:29.331421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:29.331421Z digest=sha256:96652294136410e9ce47d8a57e6693df4695f4887a7c01f331f5e1566e8c7946

Observation 04791f2d-d5d0-434a-8917-1add3bcf9edc · outbound

This paper cites 2, 3, 5, 17, 18.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns 2, 3, 5, 17, 18

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.390138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:28.910119Z digest=sha256:58108c185922cdefcf709042dea7924e14036ff04abb230fd524a2c66694d76a

Observation 9cc12a26-8c36-4a77-8b09-b58c130cb60a · outbound

This paper cites Racism is a Virus: Anti-Asian Hate and Counterspeech in Social Media during the COVID-19 Crisis.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Racism is a Virus: Anti-Asian Hate and Counterspeech in Social Media during the COVID-19 Crisis

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-08-10T11:03:29.417012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.345163Z digest=sha256:fe3e710f26a31ea2c6003a1c67718012c170de3fd65aec6fbcda1812cc021890

Observation f44ee80e-a821-4bb2-b532-13a3b18b682a · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:29.845892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.359653Z digest=sha256:a48e417100d9667a8d55c8efdadf7929734171db59530edefb7a92364f58b428

Observation e8b5c314-e0f9-4f8d-847e-cd06e98960e6 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:29.832288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.363030Z digest=sha256:e2772ed48170c5cdcc12591e3c05c87b1a4725eec73d58e21cd376ee851871b7

Observation 988831f8-7a0a-4902-a5af-3abb1bb49485 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:29.788575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.367165Z digest=sha256:1078a2880e5fc6ff39491da37b228d77c454244547490a22aea2b61bb39c4734

Observation 81c29158-d397-4a44-9abc-f140e901b128 · outbound

This paper cites identity attack.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns identity attack

Reference 100

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T11:03:29.738939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.370944Z digest=sha256:b52cc3eafd7aace1e9cace5185bc765658678fc18e5f4e694ae1f6b1efcfd7b8

Observation c3846f5f-3496-42a5-8729-02c84077f5c3 · outbound

This paper cites toxicity,.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns toxicity,

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:29.858217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T11:03:29.353557Z digest=sha256:146069bc4c00c0e3a5cbf08a85e4fb6bde05aa0ec5d0cf991aded163428b6a09

Pith citing papers

Observation 16bb6919-7b9a-4b14-bf1b-9cb3e6c48e0e · inbound

Are Today's LLMs Ready to Explain Well-Being Concepts? cites this paper.

Are Today's LLMs Ready to Explain Well-Being Concepts? HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T01:02:54.096747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T01:02:51.934157Z digest=sha256:373918c3fa4dfa91380e63f1376a6bcb2ab35400ebcec46572ffe1f984546285