Pith. sign in

REVIEW 2 cited by

WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.14023 v1 pith:VWRHULXZ submitted 2024-05-22 cs.LG

classification cs.LG
keywords attackllmsalignmentcontentobfuscationqueryresponsesafety
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent breakthrough in large language models (LLMs) such as ChatGPT has revolutionized production processes at an unprecedented pace. Alongside this progress also comes mounting concerns about LLMs' susceptibility to jailbreaking attacks, which leads to the generation of harmful or unsafe content. While safety alignment measures have been implemented in LLMs to mitigate existing jailbreak attempts and force them to become increasingly complicated, it is still far from perfect. In this paper, we analyze the common pattern of the current safety alignment and show that it is possible to exploit such patterns for jailbreaking attacks by simultaneous obfuscation in queries and responses. Specifically, we propose WordGame attack, which replaces malicious words with word games to break down the adversarial intent of a query and encourage benign content regarding the games to precede the anticipated harmful content in the response, creating a context that is hardly covered by any corpus used for safety alignment. Extensive experiments demonstrate that WordGame attack can break the guardrails of the current leading proprietary and open-source LLMs, including the latest Claude-3, GPT-4, and Llama-3 models. Further ablation studies on such simultaneous obfuscation in query and response provide evidence of the merits of the attack strategy beyond an individual attack.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

    cs.CL 2025-05 conditional novelty 5.0 of 10

    An automated two-stage rewriting framework, IntentPrompt, bypasses LLM content guardrails with 80-98% success by turning harmful asks into declarative outlines.

  2. Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses

    cs.CL 2025-06 reject novelty 4.0 of 10

    GCG+PAIR and GCG+WordGame hybrids show mixed attack-success improvements, but methodological flaws, including a questionable GCG loss term and a pre-generated baseline, undermine the paper's main claims.

Pith tools