Pith. sign in

REVIEW 6 cited by

Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.02840 v2 pith:U6EAZM6S submitted 2021-11-04 cs.CL cs.CRcs.LG

classification cs.CLcs.CRcs.LG
keywords adversariallanguagemodelsadvgluebenchmarkattacksgluerobustness
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale pre-trained language models have achieved tremendous success across a wide range of natural language understanding (NLU) tasks, even surpassing human performance. However, recent studies reveal that the robustness of these models can be challenged by carefully crafted textual adversarial examples. While several individual datasets have been proposed to evaluate model robustness, a principled and comprehensive benchmark is still missing. In this paper, we present Adversarial GLUE (AdvGLUE), a new multi-task benchmark to quantitatively and thoroughly explore and evaluate the vulnerabilities of modern large-scale language models under various types of adversarial attacks. In particular, we systematically apply 14 textual adversarial attack methods to GLUE tasks to construct AdvGLUE, which is further validated by humans for reliable annotations. Our findings are summarized as follows. (i) Most existing adversarial attack algorithms are prone to generating invalid or ambiguous adversarial examples, with around 90% of them either changing the original semantic meanings or misleading human annotators as well. Therefore, we perform a careful filtering process to curate a high-quality benchmark. (ii) All the language models and robust training methods we tested perform poorly on AdvGLUE, with scores lagging far behind the benign accuracy. We hope our work will motivate the development of new adversarial attacks that are more stealthy and semantic-preserving, as well as new robust language models against sophisticated adversarial attacks. AdvGLUE is available at https://adversarialglue.github.io.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A procedural-generation benchmark (PDE) shows that depth models are surprisingly vulnerable to camera changes and occlusion, while resisting lighting changes.

  2. UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation

    cs.LG 2026-02 unverdicted novelty 6.0 of 10

    UniComp finds that LLM compression preserves factual recall but degrades multi-step reasoning, multilingual ability, and reliability, while task-specific calibration recovers up to 50% of lost reasoning performance in...

  3. SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds

    cs.LG 2025-08 conditional novelty 5.0 of 10

    SALMAN ranks each text sample's fragility via the distortion between input and output embedding distances and uses the ranking to improve attack success rates and fine-tuning robustness.

  4. TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    TReB evaluates 26 large language models on 26 table reasoning subtasks using textual, programmatic, and interleaved reasoning modes, finding that the best model reaches only about 70 on a 0-100 judging scale.

  5. PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training

    cs.CR 2025-07 reject novelty 3.0 of 10

    A PRM-free alignment pipeline combining genetic algorithm red teaming and multi-objective adversarial training is claimed to beat PRM-based methods at 61% lower cost, but the experiments are unverifiable.

  6. The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.

Pith tools