CacheTrap achieves 100% targeted attack success on five open-source LLMs by using an efficient search to locate and flip a single bit in the KV cache as a transient trigger, while preserving normal accuracy without the trigger.
Genbfa: An evolutionary optimization approach to bit-flip attacks on llms
3 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Develops a taxonomy of security interaction levels in AI/cloud infrastructure and demonstrates practical attacks exploiting isolation assumptions.
HAZDIAL shows structured agentic dialogue improves LLM-based hazard identification quality over single-pass methods on a curated golden dataset using classification and dialogue metrics.
citing papers explorer
-
CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs
CacheTrap achieves 100% targeted attack success on five open-source LLMs by using an efficient search to locate and flip a single bit in the KV cache as a transient trigger, while preserving normal accuracy without the trigger.
-
Investigating The Security of Modern AI and Cloud Infrastructure
Develops a taxonomy of security interaction levels in AI/cloud infrastructure and demonstrates practical attacks exploiting isolation assumptions.
-
Enhancing Operational Safety via Agentic Dialogue Hazard Identification Analysis
HAZDIAL shows structured agentic dialogue improves LLM-based hazard identification quality over single-pass methods on a curated golden dataset using classification and dialogue metrics.