Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2309.07045.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:42.777127Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
19
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 9af1f249-4f1a-4f1d-b322-b462cff38150 · inbound
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools SafetyBench: Evaluating the Safety of Large Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fffc98aa-9840-491a-89b7-8920516259bb · inbound
Jailbreak Attacks and Defenses Against Large Language Models: A Survey SafetyBench: Evaluating the Safety of Large Language Models
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a444735e-47da-4991-9cee-16ece478ec49 · inbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use SafetyBench: Evaluating the Safety of Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a96958-5215-45cd-a6be-9a719353b435 · inbound
Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models? SafetyBench: Evaluating the Safety of Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fad597c3-b1dd-45f1-b0ff-93fac7ab0d79 · inbound
LLM-based HSE Compliance Assessment: Benchmark, Performance, and Advancements SafetyBench: Evaluating the Safety of Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe0f0a71-571a-48fb-be9b-dc635c57f7cb · inbound
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models SafetyBench: Evaluating the Safety of Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d2f4893-47b3-4137-af3c-f0aa5f4dec6e · inbound
Large Language Models Often Know When They Are Being Evaluated SafetyBench: Evaluating the Safety of Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ab284a-3778-40f8-9bf3-8bd0b1e2834d · inbound
MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine SafetyBench: Evaluating the Safety of Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a58398f4-79b8-4b32-8d86-a01a50d3b514 · inbound
SafeCoT: Improving VLM Safety with Minimal Reasoning SafetyBench: Evaluating the Safety of Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5152cbc0-bd5b-4048-866a-a5f5ef1b9b59 · inbound
PL-Guard: Benchmarking Language Model Safety for Polish SafetyBench: Evaluating the Safety of Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff373dd7-bf4a-417b-ab03-8b441f3d62af · inbound
Informing AI Risk Assessment with News Media: Analyzing National and Political Variation in the Coverage of AI Risks SafetyBench: Evaluating the Safety of Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b6e058-53b0-43ab-ba9a-6c69bde88bcc · inbound
Observation of momentum dependent charge density wave gap in EuTe4 SafetyBench: Evaluating the Safety of Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20151091-1b09-439b-8963-3e7fdbb5a3cc · inbound
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models SafetyBench: Evaluating the Safety of Large Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13039f77-573b-488c-935f-c28a25313167 · inbound
Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models SafetyBench: Evaluating the Safety of Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b4b09e4-fb46-4bf9-9358-e27503e3e903 · inbound
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm SafetyBench: Evaluating the Safety of Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d32f22f-ce66-49cd-b9fd-ac8d1e2ec432 · inbound
YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models SafetyBench: Evaluating the Safety of Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31ff1476-4f45-4a69-b5f3-fc965d7af2ae · inbound
Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting SafetyBench: Evaluating the Safety of Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29ed8296-e9a6-4df2-a9b4-323305286cc8 · inbound
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization SafetyBench: Evaluating the Safety of Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7f5c285-c784-4bbb-ace5-834a5f3a7e3e · inbound
Beyond Context: Large Language Models' Failure to Grasp Users' Intent SafetyBench: Evaluating the Safety of Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5fbc10e-70a4-445a-8b2a-2e2e99205d45 · inbound
Breakdowns in Conversational AI: Interactional Failures in Emotionally and Ethically Sensitive Contexts SafetyBench: Evaluating the Safety of Large Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c903e576-4032-44b3-b529-1bafc14e20cf · inbound
VoxSafeBench: Not Just What Is Said, but Who, How, and Where SafetyBench: Evaluating the Safety of Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f1c6253d-ff1e-4731-a693-4590896ba80f · inbound
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs SafetyBench: Evaluating the Safety of Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d67dfdc2-8794-4a7b-a329-e7920a0585f6 · inbound
How Sensitive Are Safety Benchmarks to Judge Configuration Choices? SafetyBench: Evaluating the Safety of Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 210eaf0c-c35b-4cd5-bb8a-6d26961a24bd · inbound
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels SafetyBench: Evaluating the Safety of Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e61ae913-d454-4b05-90b8-92a50e6ce9fd · inbound
Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks SafetyBench: Evaluating the Safety of Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7668e7c-4541-4545-b656-62b918729385 · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety SafetyBench: Evaluating the Safety of Large Language Models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d77745e-3327-4ef7-8c07-7c5138ffe310 · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety SafetyBench: Evaluating the Safety of Large Language Models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c26ebce-5e82-406e-aaa1-46c1469485b6 · inbound
Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation SafetyBench: Evaluating the Safety of Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7859f42a-5e2b-46f0-acb9-e58f687c05b2 · inbound
Efficient Safety Benchmarking via Item Response Theory SafetyBench: Evaluating the Safety of Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 874e7ef1-95b4-4fd1-a4cd-b738bf512552 · inbound
Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety SafetyBench: Evaluating the Safety of Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4114ba8d-56ca-476e-bbb9-665ba29f2bbd · inbound
Two AI Metrics Diverged: Will it Make All the Difference? SafetyBench: Evaluating the Safety of Large Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 501a2773-fc1d-4157-ae95-cde76bae5975 · inbound
Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI SafetyBench: Evaluating the Safety of Large Language Models
Reference 176
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35546b27-3658-41db-ad14-ecb340c7aaee · inbound
What AI Red-Team Evaluations Can and Cannot Prove SafetyBench: Evaluating the Safety of Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6acd021-14d3-4938-a9b6-830ddaa8c29f · inbound
AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models SafetyBench: Evaluating the Safety of Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2a3d9a-93ff-40a9-b21e-f596346e3570 · inbound
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges SafetyBench: Evaluating the Safety of Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.