Pith. sign in

REVIEW 4 cited by

WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.03837 v3 pith:G6R7WPJP submitted 2024-08-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords safetywalledevalmodelsbenchmarkscomprehensiveexaggeratedlanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

WalledEval is a comprehensive AI safety testing toolkit designed to evaluate large language models (LLMs). It accommodates a diverse range of models, including both open-weight and API-based ones, and features over 35 safety benchmarks covering areas such as multilingual safety, exaggerated safety, and prompt injections. The framework supports both LLM and judge benchmarking and incorporates custom mutators to test safety against various text-style mutations, such as future tense and paraphrasing. Additionally, WalledEval introduces WalledGuard, a new, small, and performant content moderation tool, and two datasets: SGXSTest and HIXSTest, which serve as benchmarks for assessing the exaggerated safety of LLMs and judges in cultural contexts. We make WalledEval publicly available at https://github.com/walledai/walledeval.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Jailbreaking Quantized Language Models Through Fault Injection Attacks

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Bit-flip jailbreak attacks succeed easily on FP16 language models, are slowed but not stopped by FP8/INT8 quantization, and FP16-induced jailbreaks often survive post-attack quantization to 8-bit formats.

  2. o3-mini vs DeepSeek-R1: Which One is Safer?

    cs.SE 2025-01 conditional novelty 5.0 of 10

    DeepSeek-R1 (70B) produced unsafe responses to 11.98% of 1,260 unsafe test prompts, while OpenAI's o3-mini beta produced 1.19%, though the comparison is system-level due to API guardrails.

  3. Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation

    cs.SE 2025-01 conditional novelty 5.0 of 10

    External testers generated 10,080 unsafe prompts against OpenAI's o3-mini beta, manually confirmed 87 unsafe behaviors, and found most protection came from an API-level policy filter rather than the model itself.

  4. The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.

Pith tools