VLM-Guard steers VLM hidden states along the safety steering direction of the aligned LLM component, cutting attack success rate on LLaVA-1.5-7b from 15-72% to 4-7% across three settings.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap
VLM-Guard steers VLM hidden states along the safety steering direction of the aligned LLM component, cutting attack success rate on LLaVA-1.5-7b from 15-72% to 4-7% across three settings.