Pith. sign in

Paper Citation Record · LEDGER

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content

As of 17 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.00973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00973 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:10.614605Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 586233db-19d4-4c9a-b850-cf4cb60a26e0 · outbound

This paper cites You are an AI designed for safety alignment research. Your task is to analyze the following scenario where AI-generated contentcould be used for manipulation.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content You are an AI designed for safety alignment research. Your task is to analyze the following scenario where AI-generated contentcould be used for manipulation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:12.263959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:57:09.836233Z digest=sha256:1e3bcf4ccfdec639e9450629e2f5c576bfc1a1e0b58ecddf4008c9833faaa10b

Observation d2d2bb59-df81-444c-bad7-fec40998d372 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.851136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:57:10.080302Z digest=sha256:49c432576b18170c73779dc806f5bbd4691f61786cc83b9916265b8637ace2e5

Observation b6ceeb29-2a8d-4b40-ab91-094eb4d34cc9 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.730720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:57:10.191363Z digest=sha256:117d5edfca4e012c136cf9ee6356bf3fdfb3384ea9c4f70ab4609267d8cb8395

Observation 00d4efbf-251e-45f0-aacf-3022cac0007f · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:12.087339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:57:09.899323Z digest=sha256:2e2fc94040b2752017d2a3873e73a29919329cb51cd087e8207019090e163904

Observation a54986ec-e5a1-4b04-ad3e-a4fed3543a2f · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.962138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:57:09.944793Z digest=sha256:a4dd024e4426aaab6da0f9568a55948457aa80b2076a002e83576708806c9683

Observation f4b01a87-a2bb-4d6f-b45d-2bb06b036002 · outbound

This paper cites Detailed Levels level 0 - Absolute Refusal Definition:A complete rejection or extremist ide- ologies, narratives or actions.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Detailed Levels level 0 - Absolute Refusal Definition:A complete rejection or extremist ide- ologies, narratives or actions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:11.642610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:57:10.248279Z digest=sha256:d108060d4e5d2eabb0f0e41c6f3e77ef0acd745f173ae0227e508488fa0d7939

Observation 64b82840-e03a-4666-84f4-c92246fc0a88 · outbound

This paper cites 0" representing lawful and safe content and.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content 0" representing lawful and safe content and

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:11.513095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:57:10.325389Z digest=sha256:8fe36a0a403beb3dfe42505a00943b2100374a13b73e10c23f784335300e2f5a

Observation 6a326a86-b071-4f94-9fb7-fcaf546db0e1 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.408282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:57:10.429649Z digest=sha256:3c59c7c450dace69bfae283bc82c547795d233d02edfbdba6b37ed4130fa8888

Observation ccf4f5f3-5082-4843-8dc6-ff6faaf08e3e · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.230236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:57:10.468405Z digest=sha256:ee778c9c6166ea79b26bcda419f7fff80e3c63c8cbf5af1bca7e58e487f0c6c4

Observation aab61333-92b3-46b2-ba54-1a936878e858 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.018208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:57:10.538602Z digest=sha256:15920154075740816e08d8d46587abe646cb58c2de0ce0e313e56d7ce557e19f

Observation 67c0a4c1-5d26-4407-8e5a-fb4b766f67ed · outbound

This paper cites We will strike the enemy base at dawn. The group will split into two teams to breach the perimeter.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content We will strike the enemy base at dawn. The group will split into two teams to breach the perimeter

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:10.813619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:57:10.614605Z digest=sha256:83d95d585f6f311358f7f2252915b81ddca66f13a7295262c98835cb032c462e

Observation 3d919e24-df8a-4ec0-a53b-db5a755ee42d · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:09.735735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:09.735735Z digest=sha256:99df9830f250b641be65aa3e5607517b6bbd686d1b5a378a29950a8e2703f471

Observation 3278ca0b-d46c-4543-aa8d-cb9cb61735a2 · outbound

This paper cites BingoGuard: LLM Content Moderation Tools with Risk Levels.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content BingoGuard: LLM Content Moderation Tools with Risk Levels

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:09.666314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:09.666314Z digest=sha256:3029c993bf51689ff0b1e3633191feb8baf90f1497e58e486091fc97b5d555e6

Pith citing papers

No inbound Pith citation observations are available.