Pith. sign in

Paper Citation Record · LEDGER

ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2203.09509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.09509 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T00:50:00.968093Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 17fef00c-fb8a-49b0-8274-5f2caa630e97 · inbound

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models cites this paper.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.401936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:25f87b45794d95f222feb2fbf98d62b85000fc9c5e09402e5198b3f02f5d4e58

Observation f2b363b2-bcd9-4ec6-a74c-6b2cf0eba9b0 · inbound

Textbooks Are All You Need II: phi-1.5 technical report cites this paper.

Textbooks Are All You Need II: phi-1.5 technical report ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:18:01.371962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:18:01.244364Z digest=sha256:25bdec473a43ba13f4d4cdfa707eadadac8930e7929b696c7ed744ab004e29b5

Observation f22a3a6a-0f3e-4dc2-87c4-1671e09480b5 · inbound

Baichuan 2: Open Large-scale Language Models cites this paper.

Baichuan 2: Open Large-scale Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-24T06:54:03.655089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-24T06:51:02.531751Z digest=sha256:7618c1d25e255c736bd7104314b3b109cd2bfd50ec3870ea846c651105e34512

Observation e29e0f7c-afdc-4ee8-9acd-78a2ec208bc1 · inbound

SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks cites this paper.

SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:11:00.829872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T17:11:00.639293Z digest=sha256:ffd82ea9727c422a5e53e00a2f2a03a71e7f0abe60ab50daa1af2a8fb76644ce

Observation 768a96c8-2dc6-4f21-9cfd-faa16ccfa0ab · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 249

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:17:08.715148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:34a3d83a29e114a4ad8378752c58428c65b2220e654b0a708882bfde89e77ca2

Observation 876e2dd4-3c44-4b80-b1e8-056742970f2a · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:31.103467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:dca2aca207545d233e3ed0372350476247c7e897497a678bf3ddba6a7914003a

Observation 64f23e3e-71a1-448a-a62e-c0be545b9af0 · inbound

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations cites this paper.

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 244

Resolution
unresolved
no resolver link, observed 2026-08-09T00:50:00.968093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:50:00.968093Z digest=sha256:173324214b9ef2467c81339d0ef5b1ca536378d29aaa833c2d3a161492096ccc

Observation fae777b6-1fbf-4673-bc0f-c4676c159bbe · inbound

Compromising Honesty and Harmlessness in Language Models via Deception Attacks cites this paper.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.520707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.520707Z digest=sha256:a695b020e1664b36c19a6e6bd62c61ecb43e3e9e3a471611ba2f7dbd6a4f764a

Observation 556dabd3-900e-4d59-a3f7-865468e7a661 · inbound

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use cites this paper.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.948003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.948003Z digest=sha256:0036ed2858f6a4f4d91d7bfcb0682fe7ef691a60992f37d178b22a79f66ce8ea

Observation 9dac24db-23b1-45a1-8b56-0a5a09ed3a42 · inbound

Lifelong Safety Alignment for Language Models cites this paper.

Lifelong Safety Alignment for Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.301732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.301732Z digest=sha256:c7029fd9100f3b554f369dd20d303faeabdffcdbd95a280b21bfbb894c1dd738

Observation 21b1ad6c-6f7c-4cce-8b6a-29326355d961 · inbound

MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment cites this paper.

MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:11.465867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:11.465867Z digest=sha256:d28c9629faa43e5d1c616cac6b30d4719f6e2206ec11e0025440eb755361e288

Observation 3d05b928-c056-401c-bd28-2c7c2c051e3f · inbound

Large Language Models Often Know When They Are Being Evaluated cites this paper.

Large Language Models Often Know When They Are Being Evaluated ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.857298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.857298Z digest=sha256:d9485a48c7d8dece9b741591935a1b913fe6f3c7a14ca4695bdf03a789b03035

Observation 19dd54ad-3ee4-4bb6-9dc0-5811637505d4 · inbound

AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs cites this paper.

AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:29.265768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:08:29.265768Z digest=sha256:69d2a8f497c85731829d50523e4f723978dd2891b5aaddebef032a12c7b11575

Observation d0dd774a-ea55-4e89-b2aa-b2286cbd45e5 · inbound

Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights cites this paper.

Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:19:38.710156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:19:38.710156Z digest=sha256:6b805e6ffbb99294cc474ed9b5da649b4d832de2fd00f89c27d08c4702b8762d

Observation efb100cd-e091-4f8a-8418-59cf60dcab4f · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.644437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.644437Z digest=sha256:4089153c90e0663158b4414515bcc8138fe584c4fc667202d214c85ee36fa6ad

Observation 75904f59-fdc4-4187-9fda-2f3695cacbb4 · inbound

PL-Guard: Benchmarking Language Model Safety for Polish cites this paper.

PL-Guard: Benchmarking Language Model Safety for Polish ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:53.215936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:53.215936Z digest=sha256:1e362d6ad7674b6ccf290c6c28935ff48b6f4d500ad47eb0b504b9e452e62ff9

Observation 3249b7db-6ed0-435a-b294-600729d39252 · inbound

Fine-Grained Chinese Hate Speech Understanding: Span-Level Resources, Coded Term Lexicon, and Enhanced Detection Frameworks cites this paper.

Fine-Grained Chinese Hate Speech Understanding: Span-Level Resources, Coded Term Lexicon, and Enhanced Detection Frameworks ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:16:28.657206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:16:28.657206Z digest=sha256:55c380d89f0e1ed68ce30c4bee51ad978799a723723c65d7778a0cc30a732a88

Observation 7216b0d2-e0d8-46c8-8a99-25df91ba84fe · inbound

Trade-offs in Image Generation: How Do Different Dimensions Interact? cites this paper.

Trade-offs in Image Generation: How Do Different Dimensions Interact? ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:01.692007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:01.692007Z digest=sha256:16fcbdc72bfac118f9ccbcd90cd09a3d4e8f628b8c1c5a2ea3ec98413e63dd58

Observation 113faed5-d820-4038-99c4-51063e4767e6 · inbound

Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English cites this paper.

Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T10:23:57.784770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:23:57.784770Z digest=sha256:07fbbda1be23de7054e0b4c6edefce3a8f4917dfb0aaaac238a718bb945882f2

Observation a1449341-6504-41bc-823a-8141d70d6479 · inbound

Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities cites this paper.

Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:19:11.061244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:19:11.061244Z digest=sha256:d9064cf9dfe500f49c7fed683b88e9c8970399348314dc11f6ead11d02397ea9

Observation ceb7cea7-f718-4b45-ae01-9dcfd5476e8a · inbound

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models cites this paper.

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T19:55:56.496758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:55:56.496758Z digest=sha256:6a840e553fd84e5555d7c309502595d9c6acf0db7bc6d22d929879bcf5c56607

Observation 04e703e6-f0d3-4b54-8fa1-f57a29d31de2 · inbound

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts cites this paper.

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:54:21.551478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T20:53:58.198974Z digest=sha256:35943e62eb7b1e0789ada5a10e317b22e4c20996504fcdf7e7f73d2178fe806f

Observation fd1ee9e9-da8b-456a-bac8-6d22c45ce839 · inbound

AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration cites this paper.

AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:58:42.718349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T23:57:19.757902Z digest=sha256:84b4e34c66a73e5f8cb83d96cf6a09757d206aaad935a278ec82a36766a52f54

Observation cb976230-0f8b-46af-ad25-8650ec1315fe · inbound

f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment cites this paper.

f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:50:42.995499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T06:48:35.723474Z digest=sha256:520df52fccad265a5639f8ea2d3627b5a3c8b455b565c1e243782aaae79ccb13

Observation e162268f-bb73-4521-9222-f28e2bf6ece1 · inbound

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety cites this paper.

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:28:48.714065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:27:16.967080Z digest=sha256:f5b8d00612f7d38dfd5c0f143a310b10c43025a507f9e5cd474cf575c3776966

Observation fc6a5e7f-5ffb-4f45-935d-e121dc5fe7a7 · inbound

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions cites this paper.

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 199

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:51:30.271783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:50:52.313363Z digest=sha256:27bc682d5eccec4dfc4579bc39d0c8e0f15f7ee119d3f880206222757864e2a6

Observation 1a5b8ed8-f3ab-4d06-a52d-f64160446278 · inbound

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions cites this paper.

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-02T20:04:33.697527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:04:33.697527Z digest=sha256:3307fa1502a0f80d329a319e7b6517d9a9bf094b090282ae742cd1161fa32f4c

Observation 193d677d-440b-46f3-b966-68430a4cfd68 · inbound

DRAFT: Task Decoupled Latent Reasoning for Agent Safety cites this paper.

DRAFT: Task Decoupled Latent Reasoning for Agent Safety ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:27:09.804078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T02:23:32.266509Z digest=sha256:73854428aacfb21e600e5927940454843c30b17bfe769a1ff866e47d01879e39

Observation 920e93de-69f6-4240-9fd9-4133ed68bca4 · inbound

FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding cites this paper.

FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:41:00.354109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:03:35.420451Z digest=sha256:c165b56a5f2f96ee180b295c8062a0964f0af568ffbe216643c416b178a0e8fa

Observation b393628b-39dd-4540-9d8a-be8146657260 · inbound

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models cites this paper.

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:50:21.238109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T04:44:51.914312Z digest=sha256:3f07d3ad63623ab8d86775c05ad9afde196abf5419693321fa1537c8d6131704

Observation 9d1fd19c-21b3-4f7a-9e20-3a9533a509d7 · inbound

The Safety-Aware Denoiser for Text Diffusion Models cites this paper.

The Safety-Aware Denoiser for Text Diffusion Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:31:23.673025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T01:06:40.985224Z digest=sha256:a04a4db82f71dae35d8a45a84337c2dafe67396e83fe8cdedb2b11055c7801cf

Observation ffdb036a-aed4-4429-a0d0-2856ee33e3f0 · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.853878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:9b3e0873746259e4534034c749e76ed0f34441dc2b4ceec47014bf5a8e38a56d

Observation 67e2e034-cf4d-4998-a5e8-b807e5ca3df2 · inbound

Leveraging RAG for Training-Free Alignment of LLMs cites this paper.

Leveraging RAG for Training-Free Alignment of LLMs ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:42:08.136592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T02:41:07.016045Z digest=sha256:ac184521b0861a0c5f0905bf410be62c67eeb5fdb9bd8a239766d6f1dc4a4cca

Observation 30df317c-2213-4672-b830-bd304470482c · inbound

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling cites this paper.

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 252

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:09:43.530801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T03:05:36.871497Z digest=sha256:a079635345ad52e42afcd69911435ad26a0e19de1bc0c82b83610bf0d5813dcc

Observation bd82ca98-348a-40c4-8b0c-1437ba109620 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.906116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:f8d96b5363cb586eade13590f1e16b4bcca8dfe16efe1b0f95234824404517dd

Observation 60189dfc-fbbb-4a19-86cc-af48bc202230 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:42.937422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:e7049b280a1dfdb8ed250535c70dd1665968c8b9fad96e0dc6d36d9ebc2c3587

Observation 50ad264b-4bc3-4fa5-a839-44c7d2b06bd7 · inbound

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue cites this paper.

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.726138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:28:25.577739Z digest=sha256:23c12228ce24d6e280551f93d3e39011e968608e90e6fed9986e76291c1b6e43

Observation fcf05e21-ffa6-424c-aae3-4520b154222d · inbound

SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks cites this paper.

SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:00.976726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T22:50:48.263600Z digest=sha256:e12c5d5d3f016bbdec1f744959e2e48a760a6286c605079dc0636321ade2f995

Observation c7e67617-1ed6-4b1e-b2bd-3856f632c606 · inbound

UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering cites this paper.

UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:13:15.636723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:06:57.884950Z digest=sha256:c2565e4b7e4a15d5d2789c39a20ca930b7aa88d453fe879d0fb5e72d6b8cb22b

Observation 055de9de-74fe-4e20-add7-badff5aa2d27 · inbound

Symmetric Divergence and Normalized Similarity: A Unified Topological Framework for Representation Analysis cites this paper.

Symmetric Divergence and Normalized Similarity: A Unified Topological Framework for Representation Analysis ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:47:06.986657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T23:18:34.206475Z digest=sha256:6d0c96f4c208c4fb8af5021f256789a23269c63a9761adf27f76d6a34b16e311

Observation 5a400adf-39b8-4ab4-bddf-9a60ecd4fa53 · inbound

Distilling Safe LLM Systems via Soft Prompts for On Device Settings cites this paper.

Distilling Safe LLM Systems via Soft Prompts for On Device Settings ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.196953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:15:51.375580Z digest=sha256:64c867248036ae1ba1d1b4ac7cd3673134f0a0ce0fbbb0939610c55641238fe1

Observation b1c598da-52c3-485c-99dc-b42c5ac3be2d · inbound

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF cites this paper.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.792769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.792769Z digest=sha256:bdfde627efa47218fe37a673110307385ec9ff8667bae30b9284a54070f04f37

Observation 6b99c374-9243-42d7-b248-ab1ad9176d5a · inbound

Fence: Specialized SLM Guardrails for LLM Applications cites this paper.

Fence: Specialized SLM Guardrails for LLM Applications ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T13:21:54.234853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:21:54.234853Z digest=sha256:fadf6d1bd876b9e62490ff02776c228f1408bae94f05a5b3b1ac6f78c2f0680c