REVIEW 14 cited by
Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) are swiftly advancing in architecture and capability, and as they integrate more deeply into complex systems, the urgency to scrutinize their security properties grows. This paper surveys research in the emerging interdisciplinary field of adversarial attacks on LLMs, a subfield of trustworthy ML, combining the perspectives of Natural Language Processing and Security. Prior work has shown that even safety-aligned LLMs (via instruction tuning and reinforcement learning through human feedback) can be susceptible to adversarial attacks, which exploit weaknesses and mislead AI systems, as evidenced by the prevalence of `jailbreak' attacks on models like ChatGPT and Bard. In this survey, we first provide an overview of large language models, describe their safety alignment, and categorize existing research based on various learning structures: textual-only attacks, multi-modal attacks, and additional attack methods specifically targeting complex systems, such as federated learning or multi-agent systems. We also offer comprehensive remarks on works that focus on the fundamental sources of vulnerabilities and potential defenses. To make this field more accessible to newcomers, we present a systematic review of existing works, a structured typology of adversarial attack concepts, and additional resources, including slides for presentations on related topics at the 62nd Annual Meeting of the Association for Computational Linguistics (ACL'24).
Forward citations
Cited by 14 Pith papers
-
SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks
Multi-turn LLM jailbreaks succeed based on how harmful intent is organized across turns, not on interaction length, and detection should shift to session and cross-session scope.
-
Understanding the Supply Chain and Risks of Large Language Model Applications
A new benchmark dataset traces dependencies across 3,859 LLM applications, 109,211 models, 2,474 datasets, and 8,862 libraries, and finds widespread known vulnerabilities in application dependencies.
-
Just Testing, Move Along: Evasion of LLM-based System Log Interpretation by Prompt Injection
Prompt injections framing malicious system logs as authorized testing can flip multiple SOTA LLMs from attack to benign classifications, while their explanations often expose the manipulation.
-
SkillGuard: A Permission-Centric Framework for Agent Skill Security
SkillGuard presents a dual-plane permission framework for agent skills that achieves 99.76% taxonomy coverage and reduces attack success rates in evaluations on 315 skills.
-
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges
Per-bias selection of a cross-family LLM auditor lifts biased-judgment accuracy from 0.805/0.824 baselines to 0.884.
-
SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
A systemization of LLM jailbreak security that adds linked taxonomies, an evaluation platform, and JailbreakDB, while its main attack–defense comparison results remain deferred.
-
Explicit Vulnerability Generation with LLMs: An Investigation Beyond Adversarial Attacks
Open-source 7B LLMs frequently produce requested C vulnerabilities when explicitly prompted, but the reported rates exclude most model outputs and the claimed persona effects are inconsistent.
-
CluCERT: Certifying LLM Robustness via Clustering-Guided Denoising Smoothing
A clustering-based denoising smoothing method claims tighter certified robustness bounds and lower cost for LLMs, but its central certificate depends on a fitted stability factor and an assumed cluster shift.
-
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
MoGUv2 embeds small routers in the deeper layers of LLMs to dynamically blend a helpful variant and a refusal variant, improving safety against jailbreak and fine-tuning attacks while preserving usability.
-
SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
SALMAN ranks each text sample's fragility via the distortion between input and output embedding distances and uses the ranking to improve attack success rates and fine-tuning robustness.
-
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.
-
An Empirical Study of Vulnerable Package Dependencies in LLM Repositories
In 52 open-source LLM projects, 75.8% of those with dependency configs use at least one vulnerable package, and half of supply chain vulnerabilities stay undisclosed for over 56 months.
-
Bridging Robustness and Generalization Against Word Substitution Attacks in NLP via the Growth Bound Matrix Approach
A Jacobian-magnitude regularization called GBM improves empirical robustness of CNN/LSTM/S4 text classifiers to synonym-substitution attacks, but the claimed certified robustness is not delivered.
-
Securing AI Systems: A Guide to Known Attacks and Impacts
A practitioner-oriented review that organizes known adversarial attacks on predictive and generative AI systems into eleven types mapped to confidentiality, integrity, and availability impacts.
Discussion (0). Sign in to comment.