Pith. sign in

Paper Citation Record · LEDGER

ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2310.17389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.17389 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:38:43.475947Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.654187Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4ae480b6-9c61-470c-a90d-6399246e790d · inbound

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs cites this paper.

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:25:14.866587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T16:25:14.744887Z digest=sha256:a8adcfe6a059cf9381446dbf35dd1c6740156015652f66759c886fb04b9c4607

Observation a6049c2b-aeaa-44a8-8685-d434ac56dbb4 · inbound

ShieldGemma: Generative AI Content Moderation Based on Gemma cites this paper.

ShieldGemma: Generative AI Content Moderation Based on Gemma ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.523130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T13:17:39.444002Z digest=sha256:2233e16ec438d68dc044c24dda33fbd1f63754614aa1c7bb83185bd7c3bcedae

Observation 116905c3-65b8-486b-8be9-7045ea8274e6 · inbound

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models cites this paper.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.475947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.475947Z digest=sha256:5fad98db83d67dab36b42b3bb481abb4015e52feef481bd752ef36505fb8051d

Observation 1272fbd4-66a2-4e77-b78a-bbad6b853592 · inbound

Improved Large Language Model Jailbreak Detection via Pretrained Embeddings cites this paper.

Improved Large Language Model Jailbreak Detection via Pretrained Embeddings ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:21:32.488307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:21:32.488307Z digest=sha256:95327f5a08fc8ea1e40791914b5f4a1a5be28940c78b3329d7df184b0114ec2c

Observation c8b0fca0-7505-4e7a-b79e-7f55e6f44299 · inbound

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models cites this paper.

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:52.197451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:52.197451Z digest=sha256:a8253baf77f65df66ce3e94eab957cefa2327f08d7985d44a3ac81db7a9706e6

Observation 198ee191-dbd8-4346-be03-fbbedc16531f · inbound

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs cites this paper.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.152264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.152264Z digest=sha256:668244e1bf83814f5a15bf63c26074293b44bfceb38ae5b446c2b61ee428c30f

Observation bfd16ff5-243f-48be-8b4e-a9f4127b063a · inbound

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch cites this paper.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.236804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.236804Z digest=sha256:20745f229fe695b8ef33af93b971e5fa5cb1caa051e917aa0e6c8b92dd35201c

Observation 39a67614-0899-461c-905f-e26885284d9b · inbound

Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails cites this paper.

Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:54.750339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:15:54.750339Z digest=sha256:4e11e7a044850110868e99772040a2746b59f0314ba1c7ae0ac57e839c544fa4

Observation 57b6ae12-f0d5-406f-b47f-1ed7be08b990 · inbound

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences cites this paper.

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T10:23:06.168953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:23:06.168953Z digest=sha256:c3f76bd50f8bb92919fe973f44aa602d2f69ac9bbdeb168e195c2194d072f2f2

Observation e773cb32-b65a-4db0-8695-1cd4132b6e08 · inbound

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use cites this paper.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.909375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.909375Z digest=sha256:a1542f36d9ad4775972aeb2a12dcf61bad4cc2ff8bd3580b8fefb419b030e182

Observation 96d8ff20-356b-4be2-9aba-bb9b38955dea · inbound

Large Language Models Often Know When They Are Being Evaluated cites this paper.

Large Language Models Often Know When They Are Being Evaluated ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.795060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.795060Z digest=sha256:043d6321782f096d73a58a8aabbeff3282a45f39269409bab917c3d06aae57e6

Observation c0f879a3-92f5-496f-abec-9b8f05638b4d · inbound

From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law cites this paper.

From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:54.802416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:54.802416Z digest=sha256:aa0ea30b3455a76ef24afe7d6dbd75cd50b6c7fd288a8a8b191380ca7b8eff05

Observation 899debdd-8dbc-4652-b50f-8fe1d3245bd9 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.647489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.647489Z digest=sha256:ec892dd33fe8089d3a3bcdf5ce8f3e46750512641406f5935ead84fd53203f30

Observation a59b22d2-ab12-433d-8fca-fcf93f1968f3 · inbound

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments cites this paper.

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:02.699513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:02.699513Z digest=sha256:18cb39c95989d0dd99e3388086c269d0f97e194dbd1a968bb6df76a4053cb215

Observation 8945f4db-2a41-4c61-9a11-ef0ac1d66d3e · inbound

LLMs Encode Harmfulness and Refusal Separately cites this paper.

LLMs Encode Harmfulness and Refusal Separately ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:03:15.999906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:03:15.999906Z digest=sha256:4a45083d9e2465c91923250184c1f027aa31894e5e9ae3902afdec1bc752bb18

Observation 7d08457f-5a13-4da3-af0c-6a428b8406fa · inbound

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails cites this paper.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:44:22.044971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T20:42:40.823721Z digest=sha256:24f88c6f8e28211a35783d1f3fa1a30833e068d322c4ead352d240375db67bf7

Observation fe7c1cb3-98e9-4202-af1b-118f250474bf · inbound

GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio cites this paper.

GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T05:07:09.235559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:07:09.235559Z digest=sha256:880a8ba36ad074faa371289aa1f0fd5271d34e0c412918c8973c5cde8d8acf5f

Observation 86848d71-d41c-4229-aecd-7ea9e6c58cbc · inbound

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders cites this paper.

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:12.009201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:12.009201Z digest=sha256:63e3b063912fe60c85efe49f438affd7481ba4feab6d489ba7b0d5b2c41dc26d

Observation f7e3cf6a-63e5-4fe4-b5bc-3396ff7d194c · inbound

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization cites this paper.

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:56:01.782343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:22:40.613937Z digest=sha256:975ec03a8cc82c9f1234dea11d3341069d90a512a77ba3552e72659251a9e642

Observation 14c65156-fe1a-4726-8f6c-d6fff637a2ff · inbound

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs cites this paper.

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:46:37.658425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:27:13.339411Z digest=sha256:19d13d014491f41aa86805eabc32c62a091739e30e8bf048168858f1047e8370

Observation f3e2d616-a3f3-4103-a614-aa060b0a2e05 · inbound

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts cites this paper.

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:08:25.512197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T09:07:57.713675Z digest=sha256:84fb9174829c86bd5e77519ce30c2c1513880ea8897787ad677bff72987d295f

Observation d5b2a56b-da7c-4d7e-99ac-83fb41c9db39 · inbound

LLM Safety From Within: Detecting Harmful Content with Internal Representations cites this paper.

LLM Safety From Within: Detecting Harmful Content with Internal Representations ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:4fef9024661c243482dd3cfd0bf6dcc9f2bfe168b6a6699a72b402937edc4180

Observation e8af5d2c-560f-4e6f-aa37-e9d5cb70633c · inbound

Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs cites this paper.

Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:46:05.501341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T00:32:00.789721Z digest=sha256:b4950a4f8bd07b4eb503a8a6c8dbdbf6464ccf2e7533fb4875ac3fbee8307a93

Observation 2d758f06-8adc-4c14-8c4d-2f363f1f20ad · inbound

Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems cites this paper.

Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T19:06:09.741391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T12:41:53.950049Z digest=sha256:3af27feb3bde90616330d4649555515435f02ee314c913ca9a0f13fb6bf5f776

Observation 01f579b9-361c-4712-a23f-9db28b5a9661 · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.946872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:47a551978cdcfa2c50d17ccadafd017218927692425a0acb1af1cab96672295f

Observation 700c0b50-026d-4736-b3d9-31e4fa83de6f · inbound

When Youth Enter the Algorithmic Wild: Discovering and Understanding Potentially Harmful Teen Videos on Douyin and Kwai cites this paper.

When Youth Enter the Algorithmic Wild: Discovering and Understanding Potentially Harmful Teen Videos on Douyin and Kwai ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:25:19.788374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T04:20:27.456558Z digest=sha256:e65d0a737918a2d11e12c33ebe6a3a876000a8089666ab1d0092f6aaabe2cf55

Observation c61f411c-1406-4b77-8914-6f10225690b0 · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.897253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:baec3bc82e6ca1509eeda8a2564197c807c2b1ae2eb0d7efbfb742f4a2d64459

Observation 59558a22-481c-44f1-a762-506f6272a304 · inbound

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety cites this paper.

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.655714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-03T14:38:55.045628Z digest=sha256:c76d51b849d142fa2e3b7a4927c3138f29a8bce660e5939c51e54fcb415994e7

Observation 048e4edc-5e2d-46cd-aa71-38c0bc50b6ab · inbound

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B cites this paper.

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T14:52:12.591220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:52:12.591220Z digest=sha256:93184cbee3c6e1835ce8872a08446ba5c95cbf13bd81017d678cd17195403e0a

Observation 2fd0eeb0-ba99-451f-a1ed-ebdbb1a87b24 · inbound

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) cites this paper.

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:07:16.653669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:07:16.653669Z digest=sha256:8e1e05e62753d205b0632daee1dbc12da448cccc8ef053d0f06d47a93084ed16