Pith. sign in

REVIEW 16 cited by

SafetyBench: Evaluating the Safety of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.07045 v2 pith:FNA6XVC5 submitted 2023-09-13 cs.CL

classification cs.CL
keywords safetyllmssafetybenchhttpsevaluatingevaluationabilitiesavailable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the rapid development of Large Language Models (LLMs), increasing attention has been paid to their safety concerns. Consequently, evaluating the safety of LLMs has become an essential task for facilitating the broad applications of LLMs. Nevertheless, the absence of comprehensive safety evaluation benchmarks poses a significant impediment to effectively assess and enhance the safety of LLMs. In this work, we present SafetyBench, a comprehensive benchmark for evaluating the safety of LLMs, which comprises 11,435 diverse multiple choice questions spanning across 7 distinct categories of safety concerns. Notably, SafetyBench also incorporates both Chinese and English data, facilitating the evaluation in both languages. Our extensive tests over 25 popular Chinese and English LLMs in both zero-shot and few-shot settings reveal a substantial performance advantage for GPT-4 over its counterparts, and there is still significant room for improving the safety of current LLMs. We also demonstrate that the measured safety understanding abilities in SafetyBench are correlated with safety generation abilities. Data and evaluation guidelines are available at \url{https://github.com/thu-coai/SafetyBench}{https://github.com/thu-coai/SafetyBench}. Submission entrance and leaderboard are available at \url{https://llmbench.ai/safety}{https://llmbench.ai/safety}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 19 citations worldwide. Full citation record

  1. Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm

    cs.CR 2025-09 conditional novelty 7.0 of 10

    A trigger-tag watermark embedded by fine-tuning lets modified LLMs mark their own phishing outputs for cheap detection.

  2. What AI Red-Team Evaluations Can and Cannot Prove

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A single closed-form expression determines how much evidence a safety benchmark's clean result carries; above a calculable harm rate it certifies safety, below it no feasible passive benchmark can.

  3. Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

    cs.CL 2026-05 conditional novelty 6.0 of 10

    Per-bias selection of a cross-family LLM auditor lifts biased-judgment accuracy from 0.805/0.824 baselines to 0.884.

  4. Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting

    cs.AI 2025-10 conditional novelty 6.0 of 10

    An early-exit rule with a zero-shot fallback, calibrated by Learn-then-Test risk control, keeps the average loss from corrupted in-context demonstrations under a preset bound.

  5. YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models

    cs.HC 2025-09 conditional novelty 6.0 of 10

    Introduces YAIR, a youth-GenAI risk benchmark, and YouthSafe, a fine-tuned classifier with AUPRC 0.94 on it, though it compares a trained model to untrained baselines.

  6. Informing AI Risk Assessment with News Media: Analyzing National and Political Variation in the Coverage of AI Risks

    cs.CY 2025-07 conditional novelty 6.0 of 10

    AI risk coverage in the news varies by country and by U.S. outlet political bias, with right-leaning outlets emphasizing malicious actors and political-culture risks.

  7. PL-Guard: Benchmarking Language Model Safety for Polish

    cs.CL 2025-06 reject novelty 6.0 of 10

    A small Polish BERT classifier proved more robust than larger fine-tuned LLMs at classifying safe versus unsafe Polish content, including under character-level adversarial perturbations.

  8. MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MTCMB is a 12-dataset benchmark for evaluating LLMs on Traditional Chinese Medicine knowledge, reasoning, and safety, with results showing models still fail at clinical reasoning and safe prescriptions.

  9. Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

    cs.CL 2026-04 conditional novelty 5.5 of 10

    Qwen Guard (4B) reaches 83.97% recall on a 79k NIST-aligned safety benchmark while larger models such as Llama Guard 12B and GPT-OSS 20B miss up to 75% of unsafe content; model size does not predict detection performance.

  10. AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models

    cs.AI 2026-07 conditional novelty 5.0 of 10

    AIR-BENCH Live autonomously extends a safety benchmark using new regulations and regenerates multilingual prompts, showing the new prompts are harder and non-English prompts expose weaker safety.

  11. Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models

    cs.CR 2025-09 conditional novelty 5.0 of 10

    A benchmark of 500 camouflaged jailbreak prompts finds open-weight LLMs comply with 94% of harmful requests, but the result is confounded by task complexity and an overly permissive compliance metric.

  12. SafeCoT: Improving VLM Safety with Minimal Reasoning

    cs.AI 2025-06 conditional novelty 5.0 of 10

    Training vision-language models to emit a short rule-based reasoning chain before refusing improves the safety-usefulness balance, with reported gains even at 100 training samples.

  13. LLM-based HSE Compliance Assessment: Benchmark, Performance, and Advancements

    cs.CL 2025-05 conditional novelty 5.0 of 10

    HSE-Bench is a new 1,020-question LLM benchmark for HSE compliance reasoning, and the paper claims LLMs rely on semantic matching rather than structured legal reasoning.

  14. Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

    cs.CY 2026-07 accept novelty 4.0 of 10

    AI safety is a systems-governance problem: six recurring organizational failure patterns from past disasters remain unlearned in AI development, so component-level fixes like benchmarks and alignment cannot deliver safety.

  15. Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization

    cs.LG 2025-10 reject novelty 4.0 of 10

    A linear-programming 'safety game' selects among LLM candidate answers to maximize helpfulness under a self-reported risk cap, improving safety-benchmark accuracy over reranking baselines in multiple-choice settings.

  16. Observation of momentum dependent charge density wave gap in EuTe4

    cond-mat.mes-hall 2025-08 unverdicted novelty 4.0 of 10

    EuTe4 shows a momentum-dependent charge density wave gap at the Fermi level, largest along Gamma-Y and smallest along Gamma-X, plus a low-temperature magnetic phase diagram near TN = 6.9 K.

Pith tools