Pith. sign in

REVIEW 2 cited by

Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.12321 v2 pith:JQ6CZIYL submitted 2023-12-19 cs.CR cs.AIcs.CLcs.LG

classification cs.CRcs.AIcs.CLcs.LG
keywords attacksllmssafetytrainingattackopen-sourceprimingalignment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

With the recent surge in popularity of LLMs has come an ever-increasing need for LLM safety training. In this paper, we investigate the fragility of SOTA open-source LLMs under simple, optimization-free attacks we refer to as $\textit{priming attacks}$, which are easy to execute and effectively bypass alignment from safety training. Our proposed attack improves the Attack Success Rate on Harmful Behaviors, as measured by Llama Guard, by up to $3.3\times$ compared to baselines. Source code and data are available at https://github.com/uiuc-focal-lab/llm-priming-attacks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Saffron-1: Safety Inference Scaling

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A multifurcation reward model that scores all next-token candidates in one call makes inference-time safety scaling with tree search far more compute-efficient than best-of-N sampling.

  2. Fast Proxies for LLM Robustness Evaluation

    cs.CR 2025-02 conditional novelty 5.0 of 10

    Simple prompt-based and embedding-space attacks predict, with rank correlations up to 0.94, how open-source LLMs fare against a six-attack red-teaming ensemble, at roughly one thousandth of the compute.

Pith tools