REVIEW 8 cited by
Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the rapid evolution of large language models (LLMs), new and hard-to-predict harmful capabilities are emerging. This requires developers to be able to identify risks through the evaluation of "dangerous capabilities" in order to responsibly deploy LLMs. In this work, we collect the first open-source dataset to evaluate safeguards in LLMs, and deploy safer open-source LLMs at a low cost. Our dataset is curated and filtered to consist only of instructions that responsible language models should not follow. We annotate and assess the responses of six popular LLMs to these instructions. Based on our annotation, we proceed to train several BERT-like classifiers, and find that these small classifiers can achieve results that are comparable with GPT-4 on automatic safety evaluation. Warning: this paper contains example data that may be offensive, harmful, or biased.
Forward citations
Cited by 8 Pith papers
-
Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions
Recast predicts the turn distribution of future multi-turn LLM safety failures from dual-scale trajectory evidence, catching 88.3% of failures 2.41 turns early at 12.3% false alarms.
-
Item Response Theory for AI Safety
Using item response theory on 192 models and eight safety benchmarks, this paper finds three latent safety factors, cuts evaluation cost by 97-99% with adaptive item selection, and detects naive sandbagging and API mo...
-
The Mirage of LLM Guardrails: A Case Study in AI-Assisted Medical Note Manipulation
Simple prompts bypass commercial LLM guardrails on medical-note edits; refusal is highly modality-dependent, and the best fakes are hard for humans to spot.
-
BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation
BioTIER, a 542-prompt benchmark with three risk tiers, shows frontier AI models differ by 90 percentage points in refusing dangerous biological queries, with top refusers over-refusing benign topics at the boundary.
-
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges
Per-bias selection of a cross-family LLM auditor lifts biased-judgment accuracy from 0.805/0.824 baselines to 0.884.
-
Reason Before You Retrieve: Agentic Planning for Multi-modal RAG
MM-R2 claims SOTA multimodal RAG accuracy on InfoSeek and Encyclopedic VQA via intent grounding plus a 10-topic KnowledgeMap, but its teacher trajectories leak the gold Wikipedia page and omit the image.
-
SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems
A LoRA-adapted SeaLLM model detects unsafe and jailbreak prompts in nine Southeast Asian languages with 97% recall and 98% F1 on the authors' new SEALSBench benchmark, far above zero-shot LlamaGuard.
-
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
The submission's abstract promises an LLM safety survey, but the provided body is the opening page of an unrelated arithmetic-dynamics paper, so the artifact is internally inconsistent.
Discussion (0). Sign in to comment.