Pith. sign in

Paper Citation Record · LEDGER

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers

As of 11 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2501.13302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13302 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:19:48.082702Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:41:54.081628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:05:50.801220Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 957accdd-a6af-408f-9936-4feae9a131a4 · outbound

This paper cites GPT-4 Technical Report.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.910889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.910889Z digest=sha256:5dade66be5c356f902171a26b53e4cb65abb71e7253febb07fe2dddeb5017eee

Observation 994b7894-04d6-4678-b01f-86d99dcb43e4 · outbound

This paper cites PaLM 2 Technical Report.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers PaLM 2 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.915189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.915189Z digest=sha256:96ef8704cb577d26a9d9b6a6a3224dc4e53122cae720cb84d95e3b43113b83f6

Observation fdf210e6-af58-4c6f-9439-bbbd01208ec2 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.555144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.919481Z digest=sha256:601524ab1bbfde3631be08d1f6013f0bfd98e2fa3a835dcda33c421b170d2071

Observation b387d232-52ee-40ff-98bc-53a7a506ca64 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.543327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.923447Z digest=sha256:4dbb01c35747ef9fc7573af7aca4d782b52c793ad358d383268487692737bf10

Observation 94bda893-9bd1-4600-a460-f0db83ae7d79 · outbound

This paper cites Measuring Political Bias in Large Language Models: What Is Said and How It Is Said.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Measuring Political Bias in Large Language Models: What Is Said and How It Is Said

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.927230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.927230Z digest=sha256:09c09ac6761883c6153c645843b1605f8e9a8fa3ed237b0f66f73ae525e69cb7

Observation 6f5ed728-a3d1-4782-85af-0560f4be0bde · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.531178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.931343Z digest=sha256:bade775d2ce61a8d3ec9774b317b0590b0b549790bb953c8891b0361477e90cb

Observation 6aaee4bf-019f-4f79-b7d8-6cd9bc434e03 · outbound

This paper cites Machine Learning Robustness: A Primer.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Machine Learning Robustness: A Primer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.935427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.935427Z digest=sha256:b8c6b02095f0cd2339399df4ae00d46f87b60f204ca40d5e0514c97a26aa9436

Observation 40ba0447-1e27-4fa7-986e-7cfc3f37fe33 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.939518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.939518Z digest=sha256:b509f7739a688b2983f9abb7ec5342338193a69886f676d901634ef5acd24045

Observation b9a4ae2d-2986-44c0-86ae-2c1bfcfde9ab · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.511482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.942816Z digest=sha256:d5b567159495ceec317aeafbccdd0b33e3b63242efd492cdcce759967dc6d0d1

Observation 5f466b0d-35e0-453c-9b36-1ab2527b282c · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.499247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.946449Z digest=sha256:eaecd657e5fea2ee9140194454558cd21119930183305bac3cdb3e6f1b070fbe

Observation 6d38a130-b64e-4e27-bce1-906b147fc153 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.486852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.949889Z digest=sha256:ddca759abcceb3f4846d718b8b389a7caafb1c19f3d1e26c7e405ae72ad311d9

Observation 6d49dfe9-edeb-466e-88ca-990fdb6f3091 · outbound

This paper cites what data benefits my classifier?.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers what data benefits my classifier?

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:19:48.475115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.953417Z digest=sha256:ab972096124ef8e725985bb9b72bf73c377408b597ccc100a2f445a584c49b00

Observation 01d6af94-8c05-4917-8193-7f0a93522893 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.463442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.956787Z digest=sha256:d0b219aeb7dcf948b37670b983466fd8ef6f90b87b246549ff248d1b22c0a9cf

Observation 0b158943-e49e-4194-a188-d673ef496a69 · outbound

This paper cites Towards F air V ideo S ummarization.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Towards F air V ideo S ummarization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:19:48.451843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.960199Z digest=sha256:d36fbe381bcde40a7de27dd09f8caa4101bd37b1b58b6589b73960b358df1a65

Observation 469152dd-23c9-4741-bc33-b57b8f86194a · outbound

This paper cites https://clarifai.com/clarifai/main/models/moderation-english-text-classification.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers https://clarifai.com/clarifai/main/models/moderation-english-text-classification

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:19:48.439695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.963889Z digest=sha256:4c4559590547fd6f283233d648c2d95f61ebd72c2849453d874f843ebd202204

Observation 833e1be4-6313-4157-b8b2-71756fda40d8 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.967288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.967288Z digest=sha256:c3481015c23b2f9bf4b0f86b002acbc3dcc9aa05efc047dfcccd8eb942217ead

Observation 7eb2e53b-6f52-4a39-9029-3be0d4989878 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.420477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.970743Z digest=sha256:9c542996da7f64e72dcaa2f0cde9ca090ccc3e6502c406867ad8194a040a786b

Observation 37171bea-4513-4777-a01f-b934a5d8eb45 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.974101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.974101Z digest=sha256:762496ca9d714fdc371a68f9a91e2226640651e84844dde7a7bdc15e0ab1966a

Observation 4e01e04b-652c-40d9-b9fe-89ec5d468370 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.977419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.977419Z digest=sha256:a72f7a6b059aefbe62a64a25fe201d82d706021b2a53a5b2cb2fd358a0482d25

Observation fd54c1ef-1759-49df-99fb-9369e50a2a02 · outbound

This paper cites Building Guardrails for Large Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Building Guardrails for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.980860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.980860Z digest=sha256:d8aab48502a7057f1647bb99843b117bdb5f1810efc05044a808b85249d8e9f8

Observation 1c56b1aa-31a9-4a15-ab87-aa275dbc12f6 · outbound

This paper cites Towards Measuring the Representation of Subjective Global Opinions in Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Towards Measuring the Representation of Subjective Global Opinions in Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.984874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.984874Z digest=sha256:fb84fcd3d0a41f7d6003f1ce6c2ab71c85a3fcde9531008a85212ec50513c64a

Observation ffe43d74-258e-45a0-ba09-5ebafb5619d2 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.988841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.988841Z digest=sha256:d1ecf7b9dee106b719e3b93c7034c8948bb4c3f25b8482f2b3cddfb0aea253e5

Observation 40265bf1-2d53-4cea-b645-f9dc54041843 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.992480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.992480Z digest=sha256:8e0dc3d7c3c31841a06d182a7b35170c9a7c47dc7160944dfa668cf8ebd355e1

Observation 4f7fc283-4ac4-4a98-8f4a-69a3027413e7 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.389553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.996348Z digest=sha256:41cca4f42a41e1c5eaa9d3951dc6395d5a80c2decda78c7f393ba340cadf19db

Observation 3120bf2e-c8d9-4288-96e8-ce3b560ca2d2 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.378264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.999917Z digest=sha256:2c6b5902575798e71b34e20b3c962ce626e07f2b48b229a7a1352e4f6d00d9a9

Observation b5a3957c-06e0-43b3-a3ec-8e97510de234 · outbound

This paper cites BERTopic: Neural topic modeling with a class-based TF-IDF procedure.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers BERTopic: Neural topic modeling with a class-based TF-IDF procedure

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.003786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.003786Z digest=sha256:465937fcb8d51027b13335b56f93cedb87a8376a516ecd82ca4a46fe0da82ebf

Observation c3e0b635-8d6d-444f-89b2-965b6abd3673 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.367010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.007666Z digest=sha256:58ad90f58927455dc3734e418274530e19b50a2e5d58ebf98dc3c47e5e7a37d7

Observation d24dd804-9b07-43ef-b486-3422c813e1bc · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.011118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.011118Z digest=sha256:b5170137e5b8742a1e5f9a054f73f0862a2760c6b08327d9c542fb6a6877fb42

Observation 9f120699-f8b7-414f-992a-24852e974628 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.349006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.014641Z digest=sha256:c8b870537011c12ae2daca60ab25c1304759da65514d326a7f4c978694deffc1

Observation 3b0804cd-4ab9-4b74-986a-422a78b03bca · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.337606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.018074Z digest=sha256:bf03ef9ed4a267194b49aa829e73503b765b51275ab44e735d7d964651e3bc44

Observation 735f0aa3-ae13-4d38-a391-cdb011a368c0 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.024738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.024738Z digest=sha256:74aedd6ddb39d302739f008f8b48f9859d8d9553e9bb4b5b2a757aca865e682b

Observation e62d96b7-2b9c-4e4c-a03b-dbf15871af17 · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.028776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.028776Z digest=sha256:9d837b3fe1571260196ee5b6fe06401688ff123f3063073f321d868b17af7a44

Observation f3ca1292-2bac-4aa4-91bc-0f9598203d2e · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.032831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.032831Z digest=sha256:a8614b4f886f871b408880b60b0b5ec23441cafead1ea332a80814cd8b89a906

Observation 4c83c018-6e34-4c78-8582-53e0e3149d44 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.320838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.036579Z digest=sha256:5d2da9d207c0c9baa2a7cf1fdb2aa7a68c947cc510017f2e1232468b902ceffd

Observation cc62cd94-0d3a-422c-b00f-eebbb86aed88 · outbound

This paper cites Toxic Bias: Perspective API Misreads German as More Toxic.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Toxic Bias: Perspective API Misreads German as More Toxic

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.040346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.040346Z digest=sha256:0c93e7161b696ed0a13ce01523b151b21620b0f4fd9d923aa15f42f52d32009e

Observation 7957712b-2d95-451e-83a7-b446cda18662 · outbound

This paper cites https://platform.openai.com/docs/guides/moderation.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers https://platform.openai.com/docs/guides/moderation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:19:48.309591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.044231Z digest=sha256:bf7c2daf6a8036267c54803ccd56d75cdaf60255a3932831b75fe686be2c9cd1

Observation c8a19fc6-0376-434a-b9a3-6de65f6886d0 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.298317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.048144Z digest=sha256:494a077f4d1cd234bc6959a0618184c8706eddccae7204b2b9c848d625357566

Observation 81c6a5ad-2387-4b42-af6e-661b23d00fd4 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.051712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.051712Z digest=sha256:d0da2b31870469e2242c493971e78e4f0875096eaa806d8a80f5173145be5f0c

Observation 87807929-10c4-4f5a-acdf-3ba5921ceb0e · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.055948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.055948Z digest=sha256:9ade7284189abfe6d599ea0c9f6eb4bb6a8453cd1337b5a0f5322debc2c2e385

Observation be32b94c-a11a-42f4-91ad-6f8b2e6f96a6 · outbound

This paper cites The Woman Worked as a Babysitter: On Biases in Language Generation.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers The Woman Worked as a Babysitter: On Biases in Language Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.059776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.059776Z digest=sha256:e0aea6c4026d8a55be407df5fe7e676a46569d335a7bae08a3c2fd03def63305

Observation 5501aa05-f5be-4440-b8c3-fc48a7402fea · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.287544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.063813Z digest=sha256:14d16159643f81955ddb3cde99a44b31c8765b474e74cb7433ec6d51f0eadcd3

Observation 8308824a-6f8c-404f-a3d3-d3f10e695b8d · outbound

This paper cites A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.067303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.067303Z digest=sha256:3916ae12d6d239092f2c69817202dad7bd88805af99b07f1e512bf720ede11bd

Observation 1c7a6349-470a-42d5-a062-b6dd14613b53 · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.071026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.071026Z digest=sha256:51bc484896372586c973c829229053cd1a40b97aaf18d73bd997c80108df45df

Observation 232e39aa-c4af-45bb-a4fa-bde25f02c480 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.074807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.074807Z digest=sha256:bc8ac6d6e5244e2b45e08da3818774d142ad01eefbeea9c8cdc032b611cdaa32

Observation 4a1c3282-f29a-4686-a57d-3e52fbfca083 · outbound

This paper cites URL: " 'urlintro :=.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers URL: " 'urlintro :=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.078510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.078510Z digest=sha256:7d44a09d680b9087cfab2bd9271aa624b97760b4dc66e6774fb5a6464178c75b

Observation 8d0b8269-2142-4f6b-92f0-ff69ef57c879 · outbound

This paper cites write newline.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers write newline

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.082702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.082702Z digest=sha256:1a0ff27713611187a9b18acef8f5dd21e37ca02b6218b3c2092ef93b3a2d8401

Pith citing papers

Observation cae08519-4457-440f-886f-b526fbb18938 · inbound

To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs cites this paper.

To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:50.803437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T18:41:54.081628Z digest=sha256:e5c58fdcb20a9513236084b437d701b373a0340ed08cc79ec55430ed02bb0272