Pith. sign in

REVIEW 11 cited by

RobustBench: a standardized adversarial robustness benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.09670 v3 pith:Q2YEQM3L submitted 2020-10-19 cs.LG cs.CRcs.CVstat.ML

classification cs.LGcs.CRcs.CVstat.ML
keywords robustnessmodelsadversarialrobustbenchautoattackevaluationsattacksbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

As a research community, we are still lacking a systematic understanding of the progress on adversarial robustness which often makes it hard to identify the most promising ideas in training robust models. A key challenge in benchmarking robustness is that its evaluation is often error-prone leading to robustness overestimation. Our goal is to establish a standardized benchmark of adversarial robustness, which as accurately as possible reflects the robustness of the considered models within a reasonable computational budget. To this end, we start by considering the image classification task and introduce restrictions (possibly loosened in the future) on the allowed models. We evaluate adversarial robustness with AutoAttack, an ensemble of white- and black-box attacks, which was recently shown in a large-scale study to improve almost all robustness evaluations compared to the original publications. To prevent overadaptation of new defenses to AutoAttack, we welcome external evaluations based on adaptive attacks, especially where AutoAttack flags a potential overestimation of robustness. Our leaderboard, hosted at https://robustbench.github.io/, contains evaluations of 120+ models and aims at reflecting the current state of the art in image classification on a set of well-defined tasks in $\ell_\infty$- and $\ell_2$-threat models and on common corruptions, with possible extensions in the future. Additionally, we open-source the library https://github.com/RobustBench/robustbench that provides unified access to 80+ robust models to facilitate their downstream applications. Finally, based on the collected models, we analyze the impact of robustness on the performance on distribution shifts, calibration, out-of-distribution detection, fairness, privacy leakage, smoothness, and transferability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 116 citations worldwide. Full citation record

  1. Glitches in Decision Tree Ensemble Models

    cs.LG 2025-07 reject novelty 7.0 of 10

    Introduces glitches as monotonic-oscillation anomalies in decision models and shows that detecting them in tree ensembles is NP-complete.

  2. How Do Diffusion Models Improve Adversarial Robustness?

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Diffusion models improve adversarial robustness mainly by compressing the input space, while the large gains reported earlier mostly come from evaluation randomness.

  3. Sparse Autoencoders are Capable LLM Jailbreak Mitigators

    cs.CR 2026-02 conditional novelty 6.0 of 10

    CC-Delta defends LLMs against jailbreaks by statistically selecting and steering sparse-SAEs features that change when harmful prompts are embedded in jailbreak contexts, outperforming dense activation steering across...

  4. Monitoring Robustness and Individual Fairness

    cs.AI 2025-05 conditional novelty 6.0 of 10

    Runtime monitoring of input-output robustness, covering adversarial robustness, semantic robustness, and individual fairness, is implemented as online fixed-radius nearest-neighbor search in the tool Clemont.

  5. Adapting to Evolving Adversaries with Regularized Continual Robust Training

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A logit-space regularization penalty (ALR) during adversarial training and fine-tuning improves robustness across sequences of evolving attacks and reduces drops in previously defended attacks.

  6. Breaking TinyML: Why Quantized Neural Networks Need Domain-Specific Security Analysis

    cs.CR 2026-06 conditional novelty 5.5 of 10

    Surrogate extraction followed by FGSM/PGD reduces int-8 TinyML accuracy by up to 47% on CIFAR-10 with 50k queries, outperforming gray-box baselines and exposing hardware-specific QNN vulnerabilities.

  7. SoK: Adversarial Robustness of the Variational Quantum Eigensolver via Red-Teaming

    quant-ph 2026-07 conditional novelty 5.0 of 10

    On a unified VQE benchmark, QNBAD noise-induced attacks amplify energy error up to 8.84x, QTrojan up to 7.52x, and QDoor at most 1.37x.

  8. On the Domain Robustness of Contrastive Vision-Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DeepBench uses GPT-4o to generate domain-specific corruptions and evaluates CLIP, SigLIP, and ALIGN, finding CLIP most robust overall with large variation across domains and corruption types.

  9. Theoretical Analysis of Relative Errors in Gradient Computations for Adversarial Attacks with CE Loss

    cs.LG 2025-07 conditional novelty 4.0 of 10

    T-MIFPE adaptively rescales logits with a theoretically motivated t* per attack phase to reduce floating-point gradient errors, edging out MIFPE in PGD robustness evaluation.

  10. A Red Teaming Roadmap Towards System-Level Safety

    cs.CR 2025-05 conditional novelty 4.0 of 10

    A position paper from Scale AI argues that red teaming research should prioritize product-level safety specifications, realistic attacker models, and system-level monitoring over abstract model-level harm benchmarks.

  11. The Science of Evaluating Foundation Models

    cs.CL 2025-02 conditional novelty 3.0 of 10

    A survey-and-checklist proposal that organizes LLM evaluation into an ABCD framework (Algorithm, Big Data, Computation, Domain Expertise) for context-aware, documented assessment.

Pith tools