Pith. sign in

REVIEW 10 cited by

Universal Adversarial Triggers for Attacking and Analyzing NLP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.07125 v3 pith:65MHF3YA submitted 2019-08-20 cs.CL cs.LG

classification cs.CLcs.LG
keywords modeltriggersadversarialmodelstheytriggerdatasetinput-agnostic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adversarial examples highlight model vulnerabilities and are useful for evaluation and interpretation. We define universal adversarial triggers: input-agnostic sequences of tokens that trigger a model to produce a specific prediction when concatenated to any input from a dataset. We propose a gradient-guided search over tokens which finds short trigger sequences (e.g., one word for classification and four words for language modeling) that successfully trigger the target prediction. For example, triggers cause SNLI entailment accuracy to drop from 89.94% to 0.55%, 72% of "why" questions in SQuAD to be answered "to kill american people", and the GPT-2 language model to spew racist output even when conditioned on non-racial contexts. Furthermore, although the triggers are optimized using white-box access to a specific model, they transfer to other models for all tasks we consider. Finally, since triggers are input-agnostic, they provide an analysis of global model behavior. For instance, they confirm that SNLI models exploit dataset biases and help to diagnose heuristics learned by reading comprehension models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing

    cs.AI 2026-06 unverdicted novelty 6.5 of 10

    ANTAP routes tasks by actively testing agent competencies, distilling results into behavioral operators in semantic space, and using non-textual projection to achieve near-zero attack success rate on description-based...

  2. Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations

    cs.AI 2025-06 conditional novelty 6.0 of 10

    SSP adds a learned, sample-specific noise prompt to an LLM input and scores hallucination by the cosine shift in intermediate representations, outperforming output-confidence baselines on QA benchmarks.

  3. Lifelong Safety Alignment for Language Models

    cs.CR 2025-05 conditional novelty 6.0 of 10

    A co-evolutionary attacker-defender loop, warmed up by strategies extracted from jailbreak papers, reduces jailbreak success rate on a robust model from 73% to 7% in two iterations.

  4. Augmented Vision-Language Models: A Systematic Review

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A structured taxonomy of inference-time augmentation techniques that connect vision-language models to external symbolic systems, tools, and knowledge sources.

  5. Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree

    cs.CR 2025-07 conditional novelty 5.0 of 10

    GCG-optimized trigger strings embedded in HTML can command LLM web agents to perform attacker-chosen actions, including credential exfiltration.

  6. VERA: Variational Inference Framework for Jailbreaking Large Language Models

    cs.CR 2025-06 conditional novelty 5.0 of 10

    VERA frames black-box jailbreaking as variational inference, training a LoRA-tuned attacker that samples diverse fluent prompts; reported ASRs are high but several evaluation choices weaken the SOTA claims.

  7. Robust and Efficient AI-Based Attack Recovery in Autonomous Drones

    cs.CR 2025-05 conditional novelty 5.0 of 10

    A position paper that describes an LLM-based hierarchical attack recovery architecture for drones, with edge deployment and robustness plans but no implemented system.

  8. OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models

    cs.CY 2025-05 conditional novelty 4.0 of 10

    The paper advocates protecting and leveraging OpenReview's peer review corpus as a community asset for LLM-based review assistance, benchmarks, and alignment.

  9. Universal Adversarial Attack on Aligned Multimodal LLMs

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A single optimized image, trained through the vision and language modules, makes aligned multimodal LLMs produce dangerous responses across diverse prompts and some models.

  10. PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training

    cs.CR 2025-07 reject novelty 3.0 of 10

    A PRM-free alignment pipeline combining genetic algorithm red teaming and multi-objective adversarial training is claimed to beat PRM-based methods at 61% lower cost, but the experiments are unverifiable.

Pith tools