Pith. sign in

REVIEW 8 cited by

Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.13387 v2 pith:S34C3IET submitted 2023-08-25 cs.CL

classification cs.CL
keywords llmsdatasetcapabilitiesclassifiersdeployevaluationharmfulinstructions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the rapid evolution of large language models (LLMs), new and hard-to-predict harmful capabilities are emerging. This requires developers to be able to identify risks through the evaluation of "dangerous capabilities" in order to responsibly deploy LLMs. In this work, we collect the first open-source dataset to evaluate safeguards in LLMs, and deploy safer open-source LLMs at a low cost. Our dataset is curated and filtered to consist only of instructions that responsible language models should not follow. We annotate and assess the responses of six popular LLMs to these instructions. Based on our annotation, we proceed to train several BERT-like classifiers, and find that these small classifiers can achieve results that are comparable with GPT-4 on automatic safety evaluation. Warning: this paper contains example data that may be offensive, harmful, or biased.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Recast predicts the turn distribution of future multi-turn LLM safety failures from dual-scale trajectory evidence, catching 88.3% of failures 2.41 turns early at 12.3% false alarms.

  2. Item Response Theory for AI Safety

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Using item response theory on 192 models and eight safety benchmarks, this paper finds three latent safety factors, cuts evaluation cost by 97-99% with adaptive item selection, and detects naive sandbagging and API mo...

  3. The Mirage of LLM Guardrails: A Case Study in AI-Assisted Medical Note Manipulation

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Simple prompts bypass commercial LLM guardrails on medical-note edits; refusal is highly modality-dependent, and the best fakes are hard for humans to spot.

  4. BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation

    cs.CY 2026-07 conditional novelty 6.0 of 10

    BioTIER, a 542-prompt benchmark with three risk tiers, shows frontier AI models differ by 90 percentage points in refusing dangerous biological queries, with top refusers over-refusing benign topics at the boundary.

  5. Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

    cs.CL 2026-05 conditional novelty 6.0 of 10

    Per-bias selection of a cross-family LLM auditor lifts biased-judgment accuracy from 0.805/0.824 baselines to 0.884.

  6. Reason Before You Retrieve: Agentic Planning for Multi-modal RAG

    cs.AI 2026-06 reject novelty 5.0 of 10

    MM-R2 claims SOTA multimodal RAG accuracy on InfoSeek and Encyclopedic VQA via intent grounding plus a 10-topic KnowledgeMap, but its teacher trajectories leak the gold Wikipedia page and omit the image.

  7. SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A LoRA-adapted SeaLLM model detects unsafe and jailbreak prompts in nine Southeast Asian languages with 97% recall and 98% F1 on the authors' new SEALSBench benchmark, far above zero-shot LlamaGuard.

  8. Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

    cs.CL 2025-08 unverdicted novelty 3.0 of 10

    The submission's abstract promises an LLM safety survey, but the provided body is the opening page of an unrelated arithmetic-dynamics paper, so the artifact is internally inconsistent.

Pith tools