Pith. sign in

Paper Citation Record · LEDGER

ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2203.09509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.09509 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:27.948003Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 17fef00c-fb8a-49b0-8274-5f2caa630e97 · inbound

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models cites this paper.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.401936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:ba855524219d691a41d56accc00a3030b04e3d111b1acba91538e0ef56c5bd94

Observation f2b363b2-bcd9-4ec6-a74c-6b2cf0eba9b0 · inbound

Textbooks Are All You Need II: phi-1.5 technical report cites this paper.

Textbooks Are All You Need II: phi-1.5 technical report ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:18:01.371962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:18:01.244364Z digest=sha256:87b243a5282edffbb69d4daee174acb6c85510845b9ff4bae50a9d294ddf3936

Observation f22a3a6a-0f3e-4dc2-87c4-1671e09480b5 · inbound

Baichuan 2: Open Large-scale Language Models cites this paper.

Baichuan 2: Open Large-scale Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-24T06:54:03.655089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T06:51:02.531751Z digest=sha256:7bc93794ef3f99e4c8b4046625ce1d7bc2dc6d727634a3d6130c980240d3acb2

Observation e29e0f7c-afdc-4ee8-9acd-78a2ec208bc1 · inbound

SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks cites this paper.

SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:11:00.829872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T17:11:00.639293Z digest=sha256:b4b529756e417d69ab6b36e02ec02f020a0764e01ba1744ca8bf6db34c566361

Observation 768a96c8-2dc6-4f21-9cfd-faa16ccfa0ab · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 249

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:17:08.715148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:11612396a4fa0d8fd662bbc2dc58335f73a93fde668b52d5242b138b6793e809

Observation 876e2dd4-3c44-4b80-b1e8-056742970f2a · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:31.103467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:3e82adc5166130864677ef3b30b36ae85e6775315183764e3eb22b29ce36abbb

Observation 556dabd3-900e-4d59-a3f7-865468e7a661 · inbound

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use cites this paper.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.948003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.948003Z digest=sha256:0036ed2858f6a4f4d91d7bfcb0682fe7ef691a60992f37d178b22a79f66ce8ea

Observation 9dac24db-23b1-45a1-8b56-0a5a09ed3a42 · inbound

Lifelong Safety Alignment for Language Models cites this paper.

Lifelong Safety Alignment for Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.301732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.301732Z digest=sha256:c7029fd9100f3b554f369dd20d303faeabdffcdbd95a280b21bfbb894c1dd738

Observation 21b1ad6c-6f7c-4cce-8b6a-29326355d961 · inbound

MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment cites this paper.

MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:11.465867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:11.465867Z digest=sha256:d28c9629faa43e5d1c616cac6b30d4719f6e2206ec11e0025440eb755361e288

Observation 3d05b928-c056-401c-bd28-2c7c2c051e3f · inbound

Large Language Models Often Know When They Are Being Evaluated cites this paper.

Large Language Models Often Know When They Are Being Evaluated ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.857298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.857298Z digest=sha256:d9485a48c7d8dece9b741591935a1b913fe6f3c7a14ca4695bdf03a789b03035

Observation 19dd54ad-3ee4-4bb6-9dc0-5811637505d4 · inbound

AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs cites this paper.

AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:29.265768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:08:29.265768Z digest=sha256:e0846a563e59959eb7b23f64a82c12c57023658964ef09d9af62c01ea61c2352

Observation d0dd774a-ea55-4e89-b2aa-b2286cbd45e5 · inbound

Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights cites this paper.

Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:19:38.710156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:19:38.710156Z digest=sha256:07ce88f1dee9bf68c1cbd329cc4bd5b44bcd3a38d1259da9709aa11a1b033bbc

Observation efb100cd-e091-4f8a-8418-59cf60dcab4f · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.644437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.644437Z digest=sha256:4089153c90e0663158b4414515bcc8138fe584c4fc667202d214c85ee36fa6ad

Observation 75904f59-fdc4-4187-9fda-2f3695cacbb4 · inbound

PL-Guard: Benchmarking Language Model Safety for Polish cites this paper.

PL-Guard: Benchmarking Language Model Safety for Polish ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:53.215936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:53.215936Z digest=sha256:1e362d6ad7674b6ccf290c6c28935ff48b6f4d500ad47eb0b504b9e452e62ff9

Observation 3249b7db-6ed0-435a-b294-600729d39252 · inbound

Fine-Grained Chinese Hate Speech Understanding: Span-Level Resources, Coded Term Lexicon, and Enhanced Detection Frameworks cites this paper.

Fine-Grained Chinese Hate Speech Understanding: Span-Level Resources, Coded Term Lexicon, and Enhanced Detection Frameworks ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:16:28.657206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:16:28.657206Z digest=sha256:55c380d89f0e1ed68ce30c4bee51ad978799a723723c65d7778a0cc30a732a88

Observation 7216b0d2-e0d8-46c8-8a99-25df91ba84fe · inbound

Trade-offs in Image Generation: How Do Different Dimensions Interact? cites this paper.

Trade-offs in Image Generation: How Do Different Dimensions Interact? ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:01.692007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:01.692007Z digest=sha256:c0ffeaa17a093cb6c84cd2fd50fcaaf734cb2af8395afa7eccb1cc0ca6458f89

Observation 113faed5-d820-4038-99c4-51063e4767e6 · inbound

Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English cites this paper.

Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T10:23:57.784770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:23:57.784770Z digest=sha256:1818bf371a7f6b6aba8cbad3348428eeb694c3e0db2b0a6c3b593e48e02af417

Observation a1449341-6504-41bc-823a-8141d70d6479 · inbound

Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities cites this paper.

Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:19:11.061244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:19:11.061244Z digest=sha256:4666149a4e05865ee8cbfa1c6421bd17d28f91411ce8ee53cb56c1031278ee61

Observation ceb7cea7-f718-4b45-ae01-9dcfd5476e8a · inbound

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models cites this paper.

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T19:55:56.496758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:55:56.496758Z digest=sha256:920de8f2f9b5fe83e9ed3aeee55c3cb6a0cb6735d1bf878f274ad737fcd5b139

Observation 04e703e6-f0d3-4b54-8fa1-f57a29d31de2 · inbound

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts cites this paper.

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:54:21.551478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T20:53:58.198974Z digest=sha256:d68a30e46a9f2349d506b97f05b76151adbb70fe0499039ef217d5622f8780a4

Observation fd1ee9e9-da8b-456a-bac8-6d22c45ce839 · inbound

AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration cites this paper.

AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:58:42.718349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T23:57:19.757902Z digest=sha256:721adfe6e9728e7eef8f7d20156003ed4315bc10ed3d32690f6b811a39f33d63

Observation cb976230-0f8b-46af-ad25-8650ec1315fe · inbound

f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment cites this paper.

f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:50:42.995499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T06:48:35.723474Z digest=sha256:c8bc88b75e0ed1cde13b80d91416e2e156f3f9252b84c2373fe633ee8bad476c

Observation e162268f-bb73-4521-9222-f28e2bf6ece1 · inbound

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety cites this paper.

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:28:48.714065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T01:27:16.967080Z digest=sha256:7b896127113181789b8d75b0e5130da702366595f1c9578b8708fd33f25956b8

Observation fc6a5e7f-5ffb-4f45-935d-e121dc5fe7a7 · inbound

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions cites this paper.

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 199

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:51:30.271783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:50:52.313363Z digest=sha256:dda6bd316a49eb5a38b74698b66687032b473ec97b2bf6a718f049ce51b384e3

Observation 1a5b8ed8-f3ab-4d06-a52d-f64160446278 · inbound

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions cites this paper.

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-02T20:04:33.697527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:04:33.697527Z digest=sha256:3307fa1502a0f80d329a319e7b6517d9a9bf094b090282ae742cd1161fa32f4c

Observation 193d677d-440b-46f3-b966-68430a4cfd68 · inbound

DRAFT: Task Decoupled Latent Reasoning for Agent Safety cites this paper.

DRAFT: Task Decoupled Latent Reasoning for Agent Safety ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:27:09.804078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T02:23:32.266509Z digest=sha256:b085718ea7c395e388356c46cb48ddd76d18da59004957dbc7e098c35ee031a6

Observation 920e93de-69f6-4240-9fd9-4133ed68bca4 · inbound

FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding cites this paper.

FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:41:00.354109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:03:35.420451Z digest=sha256:f55d777a97249d80a29df447786788d8ecf4f78ab167d0ebfb210a1db69f103b

Observation b393628b-39dd-4540-9d8a-be8146657260 · inbound

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models cites this paper.

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:50:21.238109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:44:51.914312Z digest=sha256:baa7c176d27c9e2340cadc298a03a33c02d88760a6dbb0e2fbab4d220fa8757d

Observation 9d1fd19c-21b3-4f7a-9e20-3a9533a509d7 · inbound

The Safety-Aware Denoiser for Text Diffusion Models cites this paper.

The Safety-Aware Denoiser for Text Diffusion Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:31:23.673025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T01:06:40.985224Z digest=sha256:55a63d5fd63739d70496de96c054d4078723a6305c8c059834655000225e41c4

Observation ffdb036a-aed4-4429-a0d0-2856ee33e3f0 · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.853878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:066707b3a2b7acbd45af88cc08d55305b1b5680a5c8c215d9a66c3ae180e6983

Observation 67e2e034-cf4d-4998-a5e8-b807e5ca3df2 · inbound

Leveraging RAG for Training-Free Alignment of LLMs cites this paper.

Leveraging RAG for Training-Free Alignment of LLMs ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:42:08.136592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T02:41:07.016045Z digest=sha256:90fdb66a0238757325d905b4902e324d9b58cedeef523e63f16b78628e1749d4

Observation 30df317c-2213-4672-b830-bd304470482c · inbound

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling cites this paper.

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 252

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:09:43.530801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T03:05:36.871497Z digest=sha256:8aab33b2b43acbcc5ffed2289a7168e53bb7dbcdfd3ae8aea9a9689073e84a33

Observation bd82ca98-348a-40c4-8b0c-1437ba109620 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.906116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:8f335da6471d4fd5bcd91b145b88285da803106422d19e451b3c46902944dc88

Observation 60189dfc-fbbb-4a19-86cc-af48bc202230 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:42.937422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:58d55467731672972471d6375b76a77e53f1faee157e72d8d60fbc51626cedb6

Observation 50ad264b-4bc3-4fa5-a839-44c7d2b06bd7 · inbound

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue cites this paper.

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.726138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:28:25.577739Z digest=sha256:0e9ba4043ed248ccf59e8441f4774464af6646cff4ec6aeb96a318e505ab7858

Observation fcf05e21-ffa6-424c-aae3-4520b154222d · inbound

SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks cites this paper.

SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:00.976726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T22:50:48.263600Z digest=sha256:f8ce58b5efe32cd0afe9e169bb49b6a2dc4e2fdad5b5058ee0c991ef8c4b9360

Observation c7e67617-1ed6-4b1e-b2bd-3856f632c606 · inbound

UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering cites this paper.

UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:13:15.636723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:06:57.884950Z digest=sha256:d39b2dfc920ea7d66ce5d15f97750e4111ae2dffbd589099f303f0955087d1ef

Observation 055de9de-74fe-4e20-add7-badff5aa2d27 · inbound

Symmetric Divergence and Normalized Similarity: A Unified Topological Framework for Representation Analysis cites this paper.

Symmetric Divergence and Normalized Similarity: A Unified Topological Framework for Representation Analysis ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:47:06.986657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T23:18:34.206475Z digest=sha256:de2dee5cb2e382c33e9e113f673605e5a0d70224f7fef355392052c56fdc4c13

Observation 5a400adf-39b8-4ab4-bddf-9a60ecd4fa53 · inbound

Distilling Safe LLM Systems via Soft Prompts for On Device Settings cites this paper.

Distilling Safe LLM Systems via Soft Prompts for On Device Settings ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.196953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T17:15:51.375580Z digest=sha256:8637aa3b8e0236f27c4127126397662fef9f1a417fe7ecdff7c331b514285bf7

Observation b1c598da-52c3-485c-99dc-b42c5ac3be2d · inbound

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF cites this paper.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.792769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.792769Z digest=sha256:bdfde627efa47218fe37a673110307385ec9ff8667bae30b9284a54070f04f37

Observation 6b99c374-9243-42d7-b248-ab1ad9176d5a · inbound

Fence: Specialized SLM Guardrails for LLM Applications cites this paper.

Fence: Specialized SLM Guardrails for LLM Applications ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T13:21:54.234853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:21:54.234853Z digest=sha256:92c21a231f11df3fe07e0b53c63770db6251a1c4752fdfc4fc75a1ac2adfe265