Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T16:19:48.082702Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2501.13302.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T16:19:48.082702Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:41:54.081628Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T00:05:50.801220Z
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 957accdd-a6af-408f-9936-4feae9a131a4 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994b7894-04d6-4678-b01f-86d99dcb43e4 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers PaLM 2 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdf210e6-af58-4c6f-9439-bbbd01208ec2 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b387d232-52ee-40ff-98bc-53a7a506ca64 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 94bda893-9bd1-4600-a460-f0db83ae7d79 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Measuring Political Bias in Large Language Models: What Is Said and How It Is Said
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f5ed728-a3d1-4782-85af-0560f4be0bde · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6aaee4bf-019f-4f79-b7d8-6cd9bc434e03 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Machine Learning Robustness: A Primer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ba0447-1e27-4fa7-986e-7cfc3f37fe33 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9a4ae2d-2986-44c0-86ae-2c1bfcfde9ab · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5f466b0d-35e0-453c-9b36-1ab2527b282c · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6d38a130-b64e-4e27-bce1-906b147fc153 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6d49dfe9-edeb-466e-88ca-990fdb6f3091 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers what data benefits my classifier?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 01d6af94-8c05-4917-8193-7f0a93522893 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0b158943-e49e-4194-a188-d673ef496a69 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Towards F air V ideo S ummarization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 469152dd-23c9-4741-bc33-b57b8f86194a · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers https://clarifai.com/clarifai/main/models/moderation-english-text-classification
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 833e1be4-6313-4157-b8b2-71756fda40d8 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eb2e53b-6f52-4a39-9029-3be0d4989878 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 37171bea-4513-4777-a01f-b934a5d8eb45 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e01e04b-652c-40d9-b9fe-89ec5d468370 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd54c1ef-1759-49df-99fb-9369e50a2a02 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Building Guardrails for Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c56b1aa-31a9-4a15-ab87-aa275dbc12f6 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Towards Measuring the Representation of Subjective Global Opinions in Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffe43d74-258e-45a0-ba09-5ebafb5619d2 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40265bf1-2d53-4cea-b645-f9dc54041843 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f7fc283-4ac4-4a98-8f4a-69a3027413e7 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3120bf2e-c8d9-4288-96e8-ce3b560ca2d2 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b5a3957c-06e0-43b3-a3ec-8e97510de234 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers BERTopic: Neural topic modeling with a class-based TF-IDF procedure
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e0b635-8d6d-444f-89b2-965b6abd3673 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d24dd804-9b07-43ef-b486-3422c813e1bc · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f120699-f8b7-414f-992a-24852e974628 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3b0804cd-4ab9-4b74-986a-422a78b03bca · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 735f0aa3-ae13-4d38-a391-cdb011a368c0 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e62d96b7-2b9c-4e4c-a03b-dbf15871af17 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3ca1292-2bac-4aa4-91bc-0f9598203d2e · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c83c018-6e34-4c78-8582-53e0e3149d44 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cc62cd94-0d3a-422c-b00f-eebbb86aed88 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Toxic Bias: Perspective API Misreads German as More Toxic
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7957712b-2d95-451e-83a7-b446cda18662 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers https://platform.openai.com/docs/guides/moderation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8a19fc6-0376-434a-b9a3-6de65f6886d0 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 81c6a5ad-2387-4b42-af6e-661b23d00fd4 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87807929-10c4-4f5a-acdf-3ba5921ceb0e · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be32b94c-a11a-42f4-91ad-6f8b2e6f96a6 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers The Woman Worked as a Babysitter: On Biases in Language Generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5501aa05-f5be-4440-b8c3-fc48a7402fea · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8308824a-6f8c-404f-a3d3-d3f10e695b8d · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c7a6349-470a-42d5-a062-b6dd14613b53 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 232e39aa-c4af-45bb-a4fa-bde25f02c480 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a1c3282-f29a-4686-a57d-3e52fbfca083 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers URL: " 'urlintro :=
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d0b8269-2142-4f6b-92f0-ff69ef57c879 · outbound
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers write newline
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cae08519-4457-440f-886f-b526fbb18938 · inbound
To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.