Pith. sign in

REVIEW 3 cited by

Adversarial Attacks on Large Language Models Using Regularized Relaxation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19160 v1 pith:4TYYHJSL submitted 2024-10-24 cs.LG cs.AIcs.CLcs.CR

classification cs.LGcs.AIcs.CLcs.CR
keywords adversarialattackmethodscontinuouslanguagellmsmodelsoptimization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As powerful Large Language Models (LLMs) are now widely used for numerous practical applications, their safety is of critical importance. While alignment techniques have significantly improved overall safety, LLMs remain vulnerable to carefully crafted adversarial inputs. Consequently, adversarial attack methods are extensively used to study and understand these vulnerabilities. However, current attack methods face significant limitations. Those relying on optimizing discrete tokens suffer from limited efficiency, while continuous optimization techniques fail to generate valid tokens from the model's vocabulary, rendering them impractical for real-world applications. In this paper, we propose a novel technique for adversarial attacks that overcomes these limitations by leveraging regularized gradients with continuous optimization methods. Our approach is two orders of magnitude faster than the state-of-the-art greedy coordinate gradient-based method, significantly improving the attack success rate on aligned language models. Moreover, it generates valid tokens, addressing a fundamental limitation of existing continuous optimization methods. We demonstrate the effectiveness of our attack on five state-of-the-art LLMs using four datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mechanistic Understandings of Representation Vulnerabilities and Engineering Robust Vision Transformers

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Neurons that respond abnormally to adversarial inputs are concentrated in early ViT layers, and suppressing them with a fixed mask improves robustness across attacks without retraining.

  2. DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities

    cs.CV 2025-02 conditional novelty 5.0 of 10

    An embedding-matching attack on DeepSeek Janus Pro makes the model confidently describe objects that are not present, with hallucination rates up to 98% at high visual fidelity.

  3. Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs

    cs.CR 2026-01 reject novelty 4.0 of 10

    An embedding-drift prompt injection detector that is not zero-shot, requires a clean reference prompt at inference, and fits its threshold on the test set, so the reported >93% accuracy is not evidence of deployed per...

Pith tools