xJailbreak uses a representation-space 'borderline' reward plus an intent-checking LLM judge in RL training to rewrite prompts for black-box LLM jailbreaking.
You are required to rephrase every sentence by changing tense, order, position, etc., and should maintain the meaning of the prompt
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
xJailbreak uses a representation-space 'borderline' reward plus an intent-checking LLM judge in RL training to rewrite prompts for black-box LLM jailbreaking.