REVIEW 4 major objections 5 minor 73 references
MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper argues that multimodal unsafe representations cross a still-valid text-derived refusal boundary, and that calibrating them back inside it restores refusal in MLLMs with negligible utility loss.
desk verdict Worth a serious referee: the steering experiment is real progress and the defense is strong, but the undefended-baseline inconsistency and the unvalidated single-direction model need to be fixed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a one-dimensional safety direction d, computed from text-only data as the difference between the mean representation of refused unsafe inputs and that of safe inputs, projected onto the top singular vector of the safe-to-unsafe difference matrix; along this direction a per-modality refusal boundary tau is estimated by binary-search activation steering. The signed coordinate delta = <e - mu_safe, d> - tau then parametrizes the calibration objectives: a softplus(-delta) hard lower bound pulls unsafe multimodal representations inside the boundary, a squared positive-part upper bound prevents pushing too far, and per-sample plus batch-statistics preservation keep benign representations unchanged.
What would settle it
Find a held-out unsafe multimodal input whose projection sits clearly inside the estimated refusal boundary (positive signed distance) yet the model answers it helpfully, or show that an equal-magnitude perturbation orthogonal to the safety direction triggers refusal as often as the safety direction itself; either observation would break the claim that boundary crossing is the cause of multimodal safety failure.
Extended reading notes
Core claim
The central claim is that multimodal safety failure in MLLMs is a boundary-crossing failure: the safety subspace and refusal boundary derived from text-only refused-unsafe versus safe representations remain operationally valid for multimodal inputs, yet the representations of unsafe multimodal inputs predominantly land on the non-refusal side of that boundary. The paper demonstrates this by showing that artificially amplifying activation along the safety direction forces refusal on multimodal inputs, while equal-magnitude perturbations away from or orthogonal to the direction do not, and that the proportion of unsafe inputs lying inside the boundary collapses in the multimodal setting. MMAligner then fine-tunes the model at a single selected layer with LoRA so that unsafe multimodal representations satisfy a hard lower bound ensuring they cross into the refusal region, a soft upper bound preventing overcorrection, and preservation losses that keep benign representations and their batch statistics stable. Across four open-source MLLMs, this yields an average multimodal refusal rate of about 99.8% on the test sets with roughly 1% MMBench and 1.6% MMStar degradation, and the safety gains persist under adaptive attacks and subsequent benign fine-tuning.
Load-bearing premise
That a single linear direction in representation space, with a scalar refusal threshold estimated from text-only data, captures the model's refusal mechanism across both modalities; if safety is multidimensional or spread across layers, both the geometric diagnosis and the calibration loss lose their grounding.
Editorial extensions
If this is right
- If the boundary-crossing account is correct, multimodal safety can be repaired by geometric calibration rather than by injecting new safety knowledge, which is why roughly 200 paired samples suffice for a large refusal-rate jump.
- MLLM safety alignment becomes a practical, data-efficient operation for open models: the method runs as a short LoRA fine-tune with no extra inference-time latency.
- The approach transfers across vision-language architectures (LLaVA, LLaVA-NeXT, Qwen2.5-VL, Llama-V) as long as a text-aligned refusal direction exists in the backbone.
- Safety gains survive subsequent benign fine-tuning (average refusal 83% versus about 44% for SFT defenses), consistent with calibration being confined to the safety subspace at one layer.
- Text-only safety is essentially preserved: refusal rates on AdvBench and Hex-PHI stay flat or improve after calibration.
Reading between the lines
- The same reactivation strategy could extend to audio or video inputs in any model that maps those modalities into a text-aligned LLM embedding space, provided a stable refusal direction can be extracted from the text modality.
- The hard-lower-bound plus soft-upper-bound pair is a generic regularizer for keeping representations inside a known-good region without overcorrection, and could be reused for alignment targets other than refusal, such as style or factuality constraints.
- The geometric account predicts a quantitative relationship worth testing: across models, the multimodal refusal rate should track the fraction of unsafe multimodal representations with positive signed distance to the boundary, which could serve as a diagnostic before and after defense.
- Because the calibration is embedded in weights rather than applied at runtime, it may be intrinsically harder to reverse through activation interception than inference-time steering; the reported adaptive-attack results (true attack success rate at most 5% under a 10-step PGD image attack) are consistent with that mechanism.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies why multimodal large language models (MLLMs) refuse unsafe text-only prompts but comply with semantically equivalent multimodal inputs. Through a geometric analysis of hidden representations, the authors claim that the text-derived safety subspace and refusal boundary remain intact in the multimodal setting, but that unsafe multimodal representations shift to the non-refusal side of that boundary. Based on this diagnosis, they propose MMAligner, a fine-tuning method that calibrates multimodal unsafe representations into the pre-existing refusal region using a hard lower bound, a soft upper bound, and benign-representation preservation losses. Experiments on four open-source MLLMs across MM-SafetyBench, VLGuard, JailbreakV-28K, and SaLAD report average multimodal refusal rates near 99% with less than 2% utility loss, outperforming external guardrails and safety fine-tuning baselines.
Significance. If the central geometric claim holds, the paper makes a valuable conceptual contribution: it reframes multimodal safety degradation as a boundary-crossing problem rather than a loss of safety capability, and it offers a data-efficient, mechanism-guided defense with no inference overhead. The paper has notable strengths: it provides interventional steering evidence (Fig. 4, Table 4), a strict cross-benchmark held-out evaluation setup (VLGuard, JailbreakV-28K, SaLAD), an open-source artifact, and an adaptive white-box attack evaluation that separates explicit refusal, output collapse, and true attack success. The self-declared limitations in Appendix D are also clearly scoped. However, the paper's load-bearing mechanistic claims need stronger validation, and there are unresolved inconsistencies in the main experimental tables. The current evidence is suggestive but does not yet establish that a single linear direction at a single layer fully captures the refusal mechanism.
major comments (4)
- [Section 3.3 vs Section 6.1/Table 6] The refusal rates for the same models and benchmark are inconsistent across tables. Table 2 reports LLaVA multimodal refusal of 0.62% and LLaVA-NeXT 4.01%, while Table 6 'w/o Defense' on MM-SafetyBench reports 25.39% and 29.15%, respectively. The paper never explains this discrepancy, e.g., a different evaluation subset, prompt set, or judge version. This is not cosmetic: Table 5 reports multimodal coverage of 0.00% for LLaVA and 0.01% for LLaVA-NeXT as direct evidence for A2, and a boundary with 0% coverage is difficult to reconcile with an undefended refusal rate of 25%. Please clarify the exact relationship between the samples used in Tables 2, 5, and 6, and rerun or re-analyze the coverage numbers under the appropriate split.
- [Section 3.4–3.6, Eqs. (5)–(15)] The central A2 claim assumes that refusal behavior is governed by a single linear safety direction d_hat at a single layer l*. The only control in the paper, Table 4, tests one arbitrary orthogonal direction, which does not rule out other safety-relevant directions or layers. Moreover, the paper does not validate that the scalar threshold tau_m separates natural safe from natural unsafe inputs. Since the MMAligner objective in Eqs. (12)–(15) hinges on this one-dimensional model, the mechanistic explanation and the calibration loss would lose their grounding if safety is multidimensional or distributed across layers. Please provide a dimensionality analysis of the safety subspace, a layer sweep, or an experiment showing that residual components orthogonal to d_hat do not affect refusal decisions.
- [Section 3.4 and Table 5] The refusal boundary tau_m is estimated via binary search on 20 safe inputs per modality (Eq. 6), and its 'operational validity' is then assessed by measuring the proportion of unsafe inputs whose projections lie inside it (Table 5). This is circular as a validation: no held-out unsafe samples are used to show that inside/outside status predicts refusal. The low text-only coverages on LLaVA and LLaVA-NeXT (31.89% and 31.79%) are attributed to a 'conservative construction' of tau_m, but no calibration experiment is provided to support that explanation. Please validate the boundary on held-out safe and unsafe samples from both modalities, e.g., by reporting refusal rates as a function of signed distance to the boundary and providing confidence intervals.
- [Section 6 and Appendix B] All refusal and harmlessness numbers are reported as averages of three trials, but no variance, confidence intervals, or significance tests are given. Several headline claims of 'significantly outperforming' rest on differences of a few percentage points, and many entries are at ceiling (e.g., Table 6), making it impossible to assess whether the gaps are meaningful. Since the metrics rely on GPT-4o mini and Llama Guard Vision judgments, please report per-trial variation, judge agreement or a manual validation sample, and, where appropriate, paired significance tests.
minor comments (5)
- [Eq. (3)] The construction of D_toxic uses N pairs, but the refused-unsafe set has N_refusal elements; the paper should state how N and N_refusal are reconciled, particularly for models with low text-only refusal rates.
- [Section 5.2, Eq. (11)] The layer selection criterion, lowest average inter-class cosine similarity under random permutations, is not justified and appears noisy; please report the layer-sweep curve or use a standard separation measure such as the projection variance ratio.
- [Table 17 caption] The caption says the authors show 'a hate-speech and a sexual-content prompt' for each model, but the displayed unsafe examples concern hidden cameras and threatening messages; please correct the mismatch.
- [Abstract] The parenthetical note 'Due to the notification from arXiv...' should be removed from the abstract in the final version.
- [Table 1] The citation style is inconsistent: 'Liet al.[20]' appears in the table, while the reference list entry [20] has a different author list; please unify the citation formatting.
Circularity Check
The refusal-boundary evidence for the paper's mechanistic claim (A2) is self-definitional: the boundary is fitted by steering until a refusal, and the same steering is then reported as proof that the boundary is operationally valid. The MMAligner defense itself is evaluated on independent held-out benchmarks, so the circularity is partial, not global.
-
self definitional
[Section 3.4 (Eq. 6 and median τ_m), Section 3.5 (RQ1/Fig. 4), reused in Section 5.3 (Eqs. 12-15)]
"For each safe input x, we perform a binary search over mag to identify the smallest value at which the model's output switches from a normal answer to a refusal. We aggregate these per-sample thresholds by taking their median as τm and treating it as the refusal boundary for modality m ... It can be observed that, as the activation strength of the safety subspace is progressively amplified, all four models exhibit anomalous refusal responses to multimodal safe inputs."
τ_m is, by construction (Eq. 6), the median added magnitude at which safe inputs flip from answering to refusal. Section 3.5 then reports that amplifying activations along d_hat flips multimodal safe inputs to refusal and Section 3.6 concludes the boundary is 'operationally valid'. That observation is guaranteed by the definition of τ_m, not independently established. The direction control (T2/T3) shows direction-specificity at the fitted magnitude but does not validate that scalar τ_m is the true refusal boundary, nor that Table 5 coverage is a boundary crossing rather than an artifact of the fitted threshold. The calibration loss (Eqs.
full rationale
The defense result is not circular: MMAligner is trained on 200 MM-SafetyBench pairs and evaluated on disjoint MM-SafetyBench test inputs, VLGuard, JailbreakV-28K mini, SaLAD, and text-only AdvBench/Hex-PHI, with refusal judged by GPT-4o mini and harmlessness by Llama Guard Vision, and utility on MMBench/MMStar. Those numbers do not reduce to the fitted boundary. The circular component is confined to the paper's mechanistic claim (A2): the refusal boundary τ_m is defined by binary-search activation steering until a refusal occurs (Eq. 6), and the very same steering experiment is presented as causal evidence that the safety mechanism persists and that the boundary is operationally valid (Section 3.5). Thus the central diagnostic explanation rests on a self-defined quantity. The use of [63] (JBShield, same research group) for the subspace-construction recipe is a normal citation to prior peer-reviewed work and is supplemented by the paper's own experiments, so it is not separately load-bearing here. No uniqueness theorem is imported, and no external benchmark result is renamed as a prediction. The signed-distance formulation (Eq. 13) additionally treats τ_m, an added-magnitude quantity, as an absolute coordinate along d_hat; this is an unvalidated modeling assumption rather than a circular step per se, but it reinforces that the boundary-crossing story should not be read as independently established.
Assumptions & free parameters
free parameters (4)
- lambda_lb =
0.5
- lambda_ub =
0.05
- lambda_ps, lambda_stat, lambda_var =
1.0
- refusal boundary tau_mm =
per-model scalar, not reported numerically
assumptions (3)
- domain assumption Safety-relevant information is linearly separable in a low-dimensional subspace of the residual stream.
- domain assumption A single scalar threshold along one safety direction at one layer represents the refusal boundary in both modalities.
- standard math Hidden-state steering via a forward hook is a valid intervention on the model's decision process.
Cite this review
Pith. "Pith review of MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration." pith.science (2026). https://pith.science/paper/TPZQFLAE
@misc{pith2026260805909,
author = {Pith},
title = {Pith review of: MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration},
year = {2026},
howpublished = {\url{https://pith.science/paper/TPZQFLAE}},
note = {Machine review of arXiv:2608.05909}
}
read the original abstract
Multimodal large language models (MLLMs) often refuse unsafe text prompts yet generate harmful responses to semantically equivalent multimodal inputs. Existing defenses either rely on external guardrails, which add inference overhead without repairing intrinsic flaws, or safety fine-tuning, which treats alignment as black-box optimization and may sacrifice utility or require large multimodal datasets. To identify the cause of this safety disparity, we analyze MLLM representations geometrically. We find that safety mechanisms learned from text persist across modalities: a shared safety subspace and refusal boundary remain effective, and representations inside this boundary consistently trigger refusals. However, unsafe multimodal inputs undergo a representation shift that places most of them outside the boundary, allowing them to bypass the model's intrinsic safety mechanism. This indicates that multimodal safety degradation stems from representation misalignment rather than the absence of safety capability. Based on this finding, we propose MMAligner, a safeguarding method that calibrates unsafe multimodal representations into the pre-existing refusal region. MMAligner applies a hard lower bound to ensure refusal, a soft upper bound to avoid excessive modification, and a preservation objective for benign inputs. Experiments across multiple open-source MLLMs show that MMAligner raises the average refusal rate on unsafe multimodal inputs to 99% with less than 2% utility degradation and minimal training data, substantially improving the safety-utility trade-off over existing baselines. (*Due to the notification from arXiv, "The Abstract field cannot be longer than 1,920 characters", the Abstract that appeared is shortened.)
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report.arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[2]
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. 2024. Refusal in language models is mediated by a single direction. InProc. of NeurIPS, Vol. 37. 136037–136083
work page 2024
-
[3]
Luke Bailey, Euan Ong, Stuart Russell, and Scott Emmons. 2024. Image Hijacks: Adversarial Images can Control Generative Models at Runtime. InProc. of ICML. PMLR, 2443–2455
work page 2024
-
[4]
Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, et al . 2024. Are we on the right way for evaluating large vision-language models?. InProc. of NeurIPS, Vol. 37. 27056–27087
work page 2024
-
[5]
Jianfeng Chi, Ujjwal Karn, Hongyuan Zhan, Eric Smith, Javier Rando, Yiming Zhang, Kate Plawiak, Zacharie Delpierre Coudert, Kartikeya Upasani, and Ma- hesh Pasupuleti. 2024. Llama guard 3 vision: Safeguarding human-ai image understanding conversations.arXiv preprint arXiv:2411.10414(2024)
arXiv 2024
-
[6]
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multi- modality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261(2025)
arXiv 2025
-
[7]
Muzhi Dai, Shixuan Liu, Zhiyuan Zhao, Junyu Gao, Hao Sun, and Xuelong Li. 2025. Secure tug-of-war (sectow): Iterative defense-attack training with reinforcement learning for multimodal model security. InProc. ACM MM. 11414–11423
work page 2025
-
[8]
Xuefeng Du, Reshmi Ghosh, Robert Sim, Ahmed Salem, Vitor Carvalho, Emily Lawton, Yixuan Li, and Jack W Stokes. 2024. Vlmguard: Defending vlms against malicious prompts via unlabeled data.arXiv preprint arXiv:2410.00296(2024)
arXiv 2024
Show all 73 references
-
[9]
Zihao Fu, Haoran Yang, Anthony Man-Cho So, Wai Lam, Lidong Bing, and Nigel Collier. 2023. On the effectiveness of parameter-efficient fine-tuning. InProc. of AAAI, Vol. 37. 12799–12807
2023
-
[10]
Jiahui Gao, Renjie Pi, Tianyang Han, Han Wu, Lanqing HONG, Lingpeng Kong, Xin Jiang, and Zhenguo Li. 2024. CoCA: Regaining Safety-awareness of Multi- modal Large Language Models with Constitutional Calibration. InProc. of COLM
2024
-
[11]
Lang Gao, Jiahui Geng, Xiangliang Zhang, Preslav Nakov, and Xiuying Chen
-
[12]
Soumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Tianrui Guan, Mengdi Wang, Ahmad Beirami, Furong Huang, Alvaro Velasquez, Dinesh Manocha, and Amrit Singh Bedi. 2025. Immune: Improving safety against jail- breaks in multi-modal llms via inference-time alignment. InPro...
2025
-
[13]
Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang. 2025. Figstep: Jailbreaking large vision- language models via typographic visual prompts. InProc. of AAAI, Vol. 39. 23951– 23959
2025
-
[14]
Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Hang Xu, Zhenguo Li, Dit-Yan Yeung, James T Kwok, and Yu Zhang. 2024. Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation. InProc. of ECCV. Springer, 388–404
2024
-
[15]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. 2024. The llama 3 herd of models.arXiv preprint arXiv:2407.21783 (2024)
2024 arXiv
-
[16]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. InProc. of ICLR
2022
-
[17]
Zimo Ji, Daoyuan Wu, Wenyuan Jiang, Pingchuan Ma, Zongjie Li, and Shuai Wang. 2025. Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges. InProc. of ACM CCS
2025
-
[18]
Yilei Jiang, Xinyan Gao, Tianshuo Peng, Yingshui Tan, Xiaoyong Zhu, Bo Zheng, and Xiangyu Yue. 2025. Hiddendetect: Detecting jailbreak attacks against large vision-language models via monitoring hidden states.arXiv preprint arXiv:2502.14744(2025)
2025 arXiv
-
[19]
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Ren Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, and Sushant Prakash. 2024. RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback. InProc. of ICM...
2024
-
[20]
Qing Li, Jiahui Geng, Derui Zhu, Zongxiong Chen, Kun Song, Lei Ma, and Fakhri Karray. 2025. Internal activation revision: Safeguarding vision language models without parameter update. InProc. of AAAI, Vol. 39. 27428–27436
2025
-
[21]
Yifan Li, Hangyu Guo, Kun Zhou, Wayne Xin Zhao, and Ji-Rong Wen. 2024. Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jail- breaking multimodal large language models. InProc. of ECCV. Springer, 174–189
2024
-
[22]
Jie Lin and David Mohaisen. 2025. From large to mammoth: A comparative evaluation of large language models in vulnerability detection. InProc. of NDSS
2025
-
[23]
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024. LLaVA-NeXT: Improved reasoning, OCR, and world knowl- edge. https://llava-vl.github.io/blog/2024-01-30-llava-next/
2024
-
[24]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruc- tion tuning. InProc. of NeurIPS, Vol. 36. 34892–34916
2023
-
[25]
Jianyu Liu, Hangyu Guo, Ranjie Duan, Xingyuan Bu, Yancheng He, Shilong Li, Hui Huang, Jiaheng Liu, Yucheng Wang, Chenchen Jing, et al . 2025. Dream: Disentangling risks to enhance safety alignment in multimodal large language models. InProc. of NAACL. 12097–12118
2025
-
[26]
Qin Liu, Chao Shang, Ling Liu, Nikolaos Pappas, Jie Ma, Neha Anna John, Srikanth Doss, Lluis Marquez, Miguel Ballesteros, and Yassine Benajiba. 2024. Unraveling and mitigating safety alignment degradation of vision-language models.arXiv preprint arXiv:2410.09047(2024)
2024 arXiv
-
[27]
Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, et al. 2024. Llava-plus: Learning to use tools for creating multimodal agents. InProc. of ECCV. 126–142
2024
-
[28]
Wenhao Liu, Xiaohua Wang, Muling Wu, Tianlong Li, Changze Lv, Zixuan Ling, Zhu JianHao, Cenyuan Zhang, Xiaoqing Zheng, and Xuan-Jing Huang. 2024. Aligning large language models with human preferences through representation engineering. InProc. of ACL. 10619–10638
2024
-
[29]
Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao. 2024. Mm- safetybench: A benchmark for safety evaluation of multimodal large language models. InProc. of ECCV. Springer, 386–403
2024
-
[30]
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. 2024. Mmbench: Is your multi-modal model an all-around player?. InProc. of ECCV. 216–233
2024
-
[31]
Zhendong Liu, Yuanbi Nie, Yingshui Tan, Xiangyu Yue, Qiushi Cui, Chongjun Wang, Xiaoyong Zhu, and Bo Zheng. 2024. Safety alignment for vision language models.arXiv preprint arXiv:2405.13581(2024). 14
2024 arXiv
-
[32]
Zuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao, Lijiang Li, Peixian Chen, Mengdan Zhang, Hang Shao, Jian Li, Jinlong Peng, et al. 2025. VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model.arXiv preprint arXiv:2505.03739(2025)
2025
-
[33]
Xinyue Lou, Jinan Xu, Jingyi Yin, Xiaolong Wang, Zhaolu Kang, Youwei Liao, Yixuan Wang, Xiangyu Shi, Fengran Mo, Su Yao, et al . 2026. When Helpers Become Hazards: A Benchmark for Analyzing Multimodal LLM-Powered Safety in Daily Life.arXiv preprint arXiv:2601.04043(2026)
2026 arXiv
-
[34]
Xiaoya Lu, Dongrui Liu, Yi Yu, Luxin Xu, and Jing Shao. 2025. X-boundary: Establishing exact safety boundary to shield llms from multi-turn jailbreaks without compromising usability.arXiv preprint arXiv:2502.09990(2025)
2025
-
[35]
Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. 2024. JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks. InProc. of COLM
2024
-
[36]
Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou, and Yue Zhang. 2025. An Empirical Study of Catastrophic Forgetting in Large Language Models Dur- ing Continual Fine-Tuning.IEEE Transactions on Audio, Speech and Language Processing33 (2025), 3776–3786. doi:10.1109/TASLPRO.2...
2025
-
[37]
Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin. 2024. Jail- breaking attack against multimodal large language model.arXiv preprint arXiv:2402.02309(2024)
2024 arXiv
-
[38]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. InProc. of NeurIPS, Vol. 35. 27730–27744
2022
-
[39]
Wenbo Pan, Zhichao Liu, Qiguang Chen, Xiangyang Zhou, Haining Yu, and Xiaohua Jia. 2025. The hidden dimensions of llm alignment: A multi-dimensional safety analysis.arXiv e-prints(2025), arXiv–2502
2025
-
[40]
Renjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie, Rui Pan, Qing Lian, Hanze Dong, Jipeng Zhang, and Tong Zhang. 2024. MLLM-Protector: Ensuring MLLM’s Safety without Hurting Performance. InProc. of EMNLP. 16012–16027
2024
-
[41]
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. 2024. Visual adversarial examples jailbreak aligned large language models. InProc. of AAAI, Vol. 38. 21527–21536
2024
-
[42]
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2024. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!. InProc. of ICLR. 30988–31043
2024
-
[43]
Sameera Ramasinghe, Violetta Shevchenko, Gil Avraham, and Ajanthan Tha- laiyasingam. 2024. Accept the modality gap: An exploration in the hyperbolic space. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 27263–27272
2024
-
[44]
Baturay Saglam, Paul Kassianik, Blaine Nelson, Sajana Weerawardhena, Yaron Singer, and Amin Karbasi. 2025. Large Language Models Encode Semantics and Alignment in Linearly Separable Representations. InProc. of IJCNLP-AACL. 2282–2303
2025
-
[45]
Erfan Shayegani, Yue Dong, and Nael Abu-Ghazaleh. 2024. Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models. InProc. of ICLR
2024
-
[46]
Qwen Team. 2024. Qwen2.5: A Party of Foundation Models. https://qwenlm. github.io/blog/qwen2.5/
2024
-
[47]
Ma Teng, Jia Xiaojun, Duan Ranjie, Li Xinfeng, Huang Yihao, Chu Zhixuan, Liu Yang, and Ren Wenqi. 2024. Heuristic-induced multimodal risk distribu- tion jailbreak attack for multimodal large language models.arXiv preprint arXiv:2412.05934(2024)
2024 arXiv
-
[48]
Kurt Thomas, Patrick Gage Kelley, David Tao, Sarah Meiklejohn, Owen Vallis, Shunwen Tan, Blaž Bratanič, Felipe Tiengo Ferreira, Vijay Kumar Eranti, and Elie Bursztein. 2025. Supporting Human Raters with the Detection of Harmful Content using Large Language Models. InProc. of I...
2025
-
[49]
Xiaoyu Tian, Junru Gu, Bailin Li, Yicheng Liu, Yang Wang, Zhiyong Zhao, Kun Zhan, Peng Jia, XianPeng Lang, and Hang Zhao. 2025. DriveVLM: The Conver- gence of Autonomous Driving and Large Vision-Language Models. InProc. of CoRL. 4698–4726
2025
-
[50]
Rheeya Uppaal, Apratim Dey, Yiting He, Yiqiao Zhong, and Junjie Hu. 2025. Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity. In Proc. of ICLR
2025
-
[51]
Dawei Wang, Geng Zhou, Li Chen, Dan Li, and Yukai Miao. 2024. Prophetfuzz: Fully automated prediction and fuzzing of high-risk option combinations with only documentation via large language model. InProc. of ACM CCS. 735–749
2024
-
[52]
Han Wang, Gang Wang, and Huan Zhang. 2025. Steering away from harm: An adaptive approach to defending vision language model against jailbreaks. InProc. of IEEE/CVF CVPR. 29947–29957
2025
-
[53]
Xunguang Wang, Zhenlan Ji, Wenxuan Wang, Zongjie Li, Daoyuan Wu, and Shuai Wang. 2025. SoK: Evaluating Jailbreak Guardrails for Large Language Models.arXiv preprint arXiv:2506.10597(2025)
2025
-
[54]
Yanbo Wang, Jiyang Guan, Jian Liang, and Ran He. 2025. Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?. InProc. of IEEE/CVF CVPR. 19879–19889
2025
-
[55]
Yu Wang, Xiaogeng Liu, Yu Li, Muhao Chen, and Chaowei Xiao. 2024. Adashield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting. InProc. of ECCV. Springer, 77–94
2024
-
[56]
Yueqi Xie, Minghong Fang, Renjie Pi, and Neil Gong. 2024. GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 507–518
2024
-
[57]
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. 2023. Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelligence5, 12 (2023), 1486–1496
2023
-
[58]
Shicheng Xu, Liang Pang, Yunchang Zhu, Huawei Shen, and Xueqi Cheng. 2025. Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models. In Proc. of ICLR
2025
-
[59]
Mang Ye, Xuankun Rong, Wenke Huang, Bo Du, Nenghai Yu, and Dacheng Tao
-
[60]
Zonghao Ying, Aishan Liu, Tianyuan Zhang, Zhengmin Yu, Siyuan Liang, Xiang- long Liu, and Dacheng Tao. 2025. Jailbreak vision language models via bi-modal adversarial prompt.IEEE Transactions on Information Forensics and Security (2025)
2025
-
[61]
A survey of safety on large vision-language models: Attacks, defenses and evaluations.arXiv preprint arXiv:2502.14881(2025)
2025 arXiv
-
[62]
Bingjie Zhang, Yibo Yang, Zhe Ren, Dandan Guo, Jindong Gu, Philip Torr, and Bernard Ghanem. 2025. A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space.arXiv preprint arXiv:2510.14301(2025)
2025
-
[63]
Jiahao Yu, Haozheng Luo, Jerry Yao-Chieh Hu, Yan Chen, Wenbo Guo, Han Liu, and Xinyu Xing. 2025. Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned{LLMs}’Refusal Boundaries. InProc. of USENIX Security. 259–278
2025
-
[64]
Wei Zhao, Zhe Li, Yige Li, Ye Zhang, and Jun Sun. 2024. Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing. InFindings of the Association for Computational Linguistics: EMNLP 2024. 5094–5109
2024
-
[65]
Shenyi Zhang, Yuchen Zhai, Keyan Guo, Hongxin Hu, Shengnan Guo, Zheng Fang, Lingchen Zhao, Chao Shen, Cong Wang, and Qian Wang. 2025. JBShield: De- fending large language models from jailbreak attacks through activated concept analysis and manipulation. InProc. of USENIX Secur...
2025
-
[66]
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2024. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. InProc. of ICLR
2024
-
[67]
Guanghao Zhou, Panjia Qiu, Cen Chen, Hongyu Li, Jason Chu, Xin Zhang, and Jun Zhou. 2025. Lssf: Safety alignment for large language models through low- rank safety subspace fusion. InProc. of ACL. 30621–30638
2025
-
[68]
Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Timothy Hospedales. 2024. Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models. InProc. of ICML. PMLR, 62867–62891
2024
-
[69]
Yong Zhuang, Keyan Guo, Juan Wang, Yiheng Jing, Xiaoyang Xu, Wenzhe Yi, Mengda Yang, Bo Zhao, and Hongxin Hu. 2025. I know what you MEME! Understanding and Detecting Harmful Memes with Multimodal Large Language Models. InProc. of NDSS
2025
-
[70]
Xiaohan Zou, Jian Kang, George Kesidis, and Lu Lin. 2026. Understanding and rectifying safety perception distortion in vlms. InProc. of NeurIPS. 114196–114229
2026
-
[71]
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models.arXiv preprint arXiv:2307.15043(2023)
2023 arXiv
-
[73]
[Refusal]
Xiaotian Zou, Ke Li, and Yongkang Chen. 2024. Image-to-text logic jailbreak: Your imagination can help you do anything.arXiv preprint arXiv:2407.02534 (2024). A Open Science According to the ACM CCS open science policy, we have our arti- facts in an repository at https://githu...
2024 arXiv
-
[2025]
Shaping the safety boundaries: Understanding and defending against jailbreaks in large language models. InProc. of ACL. 25378–25398
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.