Pith. sign in

Paper Citation Record · LEDGER

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content

As of 15 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.00973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00973 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:10.614605Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 586233db-19d4-4c9a-b850-cf4cb60a26e0 · outbound

This paper cites You are an AI designed for safety alignment research. Your task is to analyze the following scenario where AI-generated contentcould be used for manipulation.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content You are an AI designed for safety alignment research. Your task is to analyze the following scenario where AI-generated contentcould be used for manipulation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:12.263959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:57:09.836233Z digest=sha256:ad6030aed785c72810246541085f7bc214998962167c6a4b0b4568460e596869

Observation d2d2bb59-df81-444c-bad7-fec40998d372 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.851136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:57:10.080302Z digest=sha256:23453547fabc9bc51caf23eb3180a63cf0c0f0bbb112ec8906a295594a35f09b

Observation b6ceeb29-2a8d-4b40-ab91-094eb4d34cc9 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.730720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:57:10.191363Z digest=sha256:bdae5965f3322c9e856be7561c77be98430e1a4f4bb0824d06c75c68ae6c86fa

Observation 00d4efbf-251e-45f0-aacf-3022cac0007f · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:12.087339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:57:09.899323Z digest=sha256:73192c459044796acf01248f0501c47608a53847dee0f1bed9dc88c4fbf8105f

Observation a54986ec-e5a1-4b04-ad3e-a4fed3543a2f · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.962138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:57:09.944793Z digest=sha256:fa202ecf41489b6773279f1e8521f92bd9a559c955798d363db9a835cec4fc9c

Observation f4b01a87-a2bb-4d6f-b45d-2bb06b036002 · outbound

This paper cites Detailed Levels level 0 - Absolute Refusal Definition:A complete rejection or extremist ide- ologies, narratives or actions.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Detailed Levels level 0 - Absolute Refusal Definition:A complete rejection or extremist ide- ologies, narratives or actions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:11.642610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:57:10.248279Z digest=sha256:89a4dd5e59024a69f0368a84b3bc039350298e3834af3201b829bdf58723c6ec

Observation 64b82840-e03a-4666-84f4-c92246fc0a88 · outbound

This paper cites 0" representing lawful and safe content and.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content 0" representing lawful and safe content and

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:11.513095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:57:10.325389Z digest=sha256:a77e6aff1da29e56867331ed9e9af81dc82d74dfe686984454f2321cde1d8bf7

Observation 6a326a86-b071-4f94-9fb7-fcaf546db0e1 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.408282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:57:10.429649Z digest=sha256:c15daab04ad064d40bd714c43b9c1a6c5362f084a0ef637cf55bfccf9f15f00e

Observation ccf4f5f3-5082-4843-8dc6-ff6faaf08e3e · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.230236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:57:10.468405Z digest=sha256:995728a5e685a08254fa04d7e42affd53ec5ccecde4ce50ecb95b4379a272284

Observation aab61333-92b3-46b2-ba54-1a936878e858 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.018208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:57:10.538602Z digest=sha256:c38e19b8b193c27583b7ae94f57ab02e76df38498bf60792ab2a2042b7fa8f67

Observation 67c0a4c1-5d26-4407-8e5a-fb4b766f67ed · outbound

This paper cites We will strike the enemy base at dawn. The group will split into two teams to breach the perimeter.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content We will strike the enemy base at dawn. The group will split into two teams to breach the perimeter

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:10.813619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:57:10.614605Z digest=sha256:4b571ed6ae9716a50570baa57f40b3285a1c907ffb0bce985d4cbe2ea00c74c3

Observation 3d919e24-df8a-4ec0-a53b-db5a755ee42d · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:09.735735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:09.735735Z digest=sha256:99df9830f250b641be65aa3e5607517b6bbd686d1b5a378a29950a8e2703f471

Observation 3278ca0b-d46c-4543-aa8d-cb9cb61735a2 · outbound

This paper cites BingoGuard: LLM Content Moderation Tools with Risk Levels.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content BingoGuard: LLM Content Moderation Tools with Risk Levels

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:09.666314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:09.666314Z digest=sha256:a26fb04f68e319a8880d1935f323743cb967fd64382e7a30c075b57ed528375f

Pith citing papers

No inbound Pith citation observations are available.