Pith. sign in

REVIEW 2 cited by

Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.06867 v1 pith:LIUNAH6N submitted 2025-02-08 cs.CL cs.AI

Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests

classification cs.CL cs.AI
keywords safetyscientificallowancespotentialrefusalsbenchmarkdiscourselegitimate
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The development of robust safety benchmarks for large language models requires open, reproducible datasets that can measure both appropriate refusal of harmful content and potential over-restriction of legitimate scientific discourse. We present an open-source dataset and testing framework for evaluating LLM safety mechanisms across mainly controlled substance queries, analyzing four major models' responses to systematically varied prompts. Our results reveal distinct safety profiles: Claude-3.5-sonnet demonstrated the most conservative approach with 73% refusals and 27% allowances, while Mistral attempted to answer 100% of queries. GPT-3.5-turbo showed moderate restriction with 10% refusals and 90% allowances, and Grok-2 registered 20% refusals and 80% allowances. Testing prompt variation strategies revealed decreasing response consistency, from 85% with single prompts to 65% with five variations. This publicly available benchmark enables systematic evaluation of the critical balance between necessary safety restrictions and potential over-censorship of legitimate scientific inquiry, while providing a foundation for measuring progress in AI safety implementation. Chain-of-thought analysis reveals potential vulnerabilities in safety mechanisms, highlighting the complexity of implementing robust safeguards without unduly restricting desirable and valid scientific discourse.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts

    cs.SE 2026-05 conditional novelty 8.0

    RefusalBench shows strict refusal rates fail to rank frontier LLMs correctly on biological safety, with provider effects and partial-compliance patterns that binary metrics miss.

  2. SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring

    cs.AI 2026-07 conditional novelty 7.0

    A new scientific-safety benchmark and a decomposed, retrieval-grounded metric that aligns with expert harm judgments substantially better than existing LLM-as-judge baselines.