REVIEW 10 cited by
Universal Adversarial Triggers for Attacking and Analyzing NLP
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Adversarial examples highlight model vulnerabilities and are useful for evaluation and interpretation. We define universal adversarial triggers: input-agnostic sequences of tokens that trigger a model to produce a specific prediction when concatenated to any input from a dataset. We propose a gradient-guided search over tokens which finds short trigger sequences (e.g., one word for classification and four words for language modeling) that successfully trigger the target prediction. For example, triggers cause SNLI entailment accuracy to drop from 89.94% to 0.55%, 72% of "why" questions in SQuAD to be answered "to kill american people", and the GPT-2 language model to spew racist output even when conditioned on non-racial contexts. Furthermore, although the triggers are optimized using white-box access to a specific model, they transfer to other models for all tasks we consider. Finally, since triggers are input-agnostic, they provide an analysis of global model behavior. For instance, they confirm that SNLI models exploit dataset biases and help to diagnose heuristics learned by reading comprehension models.
Forward citations
Cited by 10 Pith papers
-
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
ANTAP routes tasks by actively testing agent competencies, distilling results into behavioral operators in semantic space, and using non-textual projection to achieve near-zero attack success rate on description-based...
-
Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
SSP adds a learned, sample-specific noise prompt to an LLM input and scores hallucination by the cosine shift in intermediate representations, outperforming output-confidence baselines on QA benchmarks.
-
Lifelong Safety Alignment for Language Models
A co-evolutionary attacker-defender loop, warmed up by strategies extracted from jailbreak papers, reduces jailbreak success rate on a robust model from 73% to 7% in two iterations.
-
Augmented Vision-Language Models: A Systematic Review
A structured taxonomy of inference-time augmentation techniques that connect vision-language models to external symbolic systems, tools, and knowledge sources.
-
Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
GCG-optimized trigger strings embedded in HTML can command LLM web agents to perform attacker-chosen actions, including credential exfiltration.
-
VERA: Variational Inference Framework for Jailbreaking Large Language Models
VERA frames black-box jailbreaking as variational inference, training a LoRA-tuned attacker that samples diverse fluent prompts; reported ASRs are high but several evaluation choices weaken the SOTA claims.
-
Robust and Efficient AI-Based Attack Recovery in Autonomous Drones
A position paper that describes an LLM-based hierarchical attack recovery architecture for drones, with edge deployment and robustness plans but no implemented system.
-
OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
The paper advocates protecting and leveraging OpenReview's peer review corpus as a community asset for LLM-based review assistance, benchmarks, and alignment.
-
Universal Adversarial Attack on Aligned Multimodal LLMs
A single optimized image, trained through the vision and language modules, makes aligned multimodal LLMs produce dangerous responses across diverse prompts and some models.
-
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
A PRM-free alignment pipeline combining genetic algorithm red teaming and multi-objective adversarial training is claimed to beat PRM-based methods at 61% lower cost, but the experiments are unverifiable.
Discussion (0). Continue with ORCID to comment.