Pith. sign in

Paper Citation Record · LEDGER

ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2310.17389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.17389 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:23:06.168953Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.654187Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4ae480b6-9c61-470c-a90d-6399246e790d · inbound

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs cites this paper.

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:25:14.866587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T16:25:14.744887Z digest=sha256:c2aecb6155649de3568b71f63c5ccd6c6bd0472ee2a52a05c5ed959495ef5647

Observation a6049c2b-aeaa-44a8-8685-d434ac56dbb4 · inbound

ShieldGemma: Generative AI Content Moderation Based on Gemma cites this paper.

ShieldGemma: Generative AI Content Moderation Based on Gemma ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.523130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:17:39.444002Z digest=sha256:581c86dd498c4f9ca47ef6dec91e87fff585bd678b5a5c2aaf7dfaa241888314

Observation 57b6ae12-f0d5-406f-b47f-1ed7be08b990 · inbound

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences cites this paper.

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T10:23:06.168953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:23:06.168953Z digest=sha256:9624a9c56dfc075f31a0b3e1be2cad380a179f60e693bf412307842762145765

Observation e773cb32-b65a-4db0-8695-1cd4132b6e08 · inbound

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use cites this paper.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.909375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.909375Z digest=sha256:82036c0af1fa3c865373cd86edc577e533465fdaa308c3a67c82014d9fdf63ab

Observation 96d8ff20-356b-4be2-9aba-bb9b38955dea · inbound

Large Language Models Often Know When They Are Being Evaluated cites this paper.

Large Language Models Often Know When They Are Being Evaluated ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.795060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.795060Z digest=sha256:4bdebc64aaf034cc6774ae5fbb7946573b84e6bfbd1d898cb773f17c16f5f2d5

Observation c0f879a3-92f5-496f-abec-9b8f05638b4d · inbound

From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law cites this paper.

From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:54.802416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:54.802416Z digest=sha256:48e3342fbb74225da5d9b68a51723713e66a9e7bdfaebdda17d90d3b6528ea7c

Observation 899debdd-8dbc-4652-b50f-8fe1d3245bd9 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.647489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.647489Z digest=sha256:3b00593451c8669d2cbe2413b31f96c2f8515d2b9813791e43795e6cd3897c8f

Observation a59b22d2-ab12-433d-8fca-fcf93f1968f3 · inbound

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments cites this paper.

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:02.699513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:02.699513Z digest=sha256:56fbc8bb30f0ba83c19c95c507c52b68b3b1a768981bddbea43ff3c04c14f642

Observation 8945f4db-2a41-4c61-9a11-ef0ac1d66d3e · inbound

LLMs Encode Harmfulness and Refusal Separately cites this paper.

LLMs Encode Harmfulness and Refusal Separately ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:03:15.999906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:03:15.999906Z digest=sha256:247796c6e56ab771d68acb786e5cf86718a523b3eeb57276cb607f3fcc5ba81d

Observation 7d08457f-5a13-4da3-af0c-6a428b8406fa · inbound

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails cites this paper.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:44:22.044971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T20:42:40.823721Z digest=sha256:19145a69e48e1379f38e674dea6ad7bff5f78cd94a0a8715c22e5acc6ca12d39

Observation fe7c1cb3-98e9-4202-af1b-118f250474bf · inbound

GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio cites this paper.

GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T05:07:09.235559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:07:09.235559Z digest=sha256:ac90c3d8529d83de6d4b2391fc5c8dea4fccedd85acd4525aa0ddc2925befde2

Observation 86848d71-d41c-4229-aecd-7ea9e6c58cbc · inbound

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders cites this paper.

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:12.009201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:12.009201Z digest=sha256:e191c895c155f6d21af736182044b30f7d4a2b037dcb10ee197f1b5bcdf9678d

Observation f7e3cf6a-63e5-4fe4-b5bc-3396ff7d194c · inbound

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization cites this paper.

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:56:01.782343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:22:40.613937Z digest=sha256:987f457d9217c712e7a3c2a2f563fc85078f76196e0ec33b391180c39f9aacc9

Observation 14c65156-fe1a-4726-8f6c-d6fff637a2ff · inbound

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs cites this paper.

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:46:37.658425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:27:13.339411Z digest=sha256:baf172ff3b23892f90227efd0c38008c2b9b3b33dbd48177134b112b3e993596

Observation f3e2d616-a3f3-4103-a614-aa060b0a2e05 · inbound

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts cites this paper.

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:08:25.512197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T09:07:57.713675Z digest=sha256:3666cae62eae5a4fc871af947159a3d8645e82012bf7b2ec5c4662c69acffeaa

Observation d5b2a56b-da7c-4d7e-99ac-83fb41c9db39 · inbound

LLM Safety From Within: Detecting Harmful Content with Internal Representations cites this paper.

LLM Safety From Within: Detecting Harmful Content with Internal Representations ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:9d18c48dd15113ae9a88ac348fce7d98eb41cea9bb1c471d5322b80ad965ffa4

Observation e8af5d2c-560f-4e6f-aa37-e9d5cb70633c · inbound

Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs cites this paper.

Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:46:05.501341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:32:00.789721Z digest=sha256:007ef0e978587e5f97b031fe54c69de77fb82486db680b4bb86c85d6654f5455

Observation 2d758f06-8adc-4c14-8c4d-2f363f1f20ad · inbound

Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems cites this paper.

Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T19:06:09.741391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T12:41:53.950049Z digest=sha256:f04ed5373492c9189dff46d478a02ef14a50d1e1646f6eb22ab67d85129d6d40

Observation 01f579b9-361c-4712-a23f-9db28b5a9661 · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.946872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:3be24adb622e5bcd76b5d22d41f9625394d2655a1aab87b1db28e48426144667

Observation 700c0b50-026d-4736-b3d9-31e4fa83de6f · inbound

When Youth Enter the Algorithmic Wild: Discovering and Understanding Potentially Harmful Teen Videos on Douyin and Kwai cites this paper.

When Youth Enter the Algorithmic Wild: Discovering and Understanding Potentially Harmful Teen Videos on Douyin and Kwai ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:25:19.788374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:20:27.456558Z digest=sha256:b99b6fb97ab6d52568d6a7db43da9ab522b38fd3e155880e9e82538284083526

Observation c61f411c-1406-4b77-8914-6f10225690b0 · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.897253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:20d6f62855d65311adebeb809c64d99ec8e8d76af9d715f3e1a6785b653c3f49

Observation 59558a22-481c-44f1-a762-506f6272a304 · inbound

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety cites this paper.

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.655714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T14:38:55.045628Z digest=sha256:fec45a28a33550159d488b732c5b6483129a1bbaec84a456c06ce7e8ed0c0a93

Observation 048e4edc-5e2d-46cd-aa71-38c0bc50b6ab · inbound

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B cites this paper.

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T14:52:12.591220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:52:12.591220Z digest=sha256:b081539fea95eb6b8235b0b76c23ac6bb30e781b021d2b94a5e7f0feb3615c36

Observation 2fd0eeb0-ba99-451f-a1ed-ebdbb1a87b24 · inbound

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) cites this paper.

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:07:16.653669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:07:16.653669Z digest=sha256:351a3ca9a1eadb6060d38da5bd48e828bf44fcc7ef1f71cecba5c54e4de799f1