Pith. sign in

REVIEW 1 cited by

ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs

T0 review · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read ShieldVLM detects multimodal implicit toxicity through deliberate cross-modal reasoning, outperforming existing moderation APIs and models on the new MMIT benchmark.

arxiv 2505.14035 v1 pith:KVEL644M submitted 2025-05-20 cs.MM cs.CL

classification cs.MMcs.CL
keywords multimodaltoxicityimplicitdetectionpromptsshieldvlmstatementsappears
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Many posts pair an image with a short text. Sometimes each one is harmless by itself, but together they deliver a hostile or dangerous message. The authors call this multimodal implicit toxicity: the risk lives in the way the two pieces combine. To map the problem, they define five ways a combination can become toxic: a word shifting meaning, a situation changing context, a metaphor, an unstated implication, or shared cultural knowledge. They built a dataset of 2,100 such text-image pairs, spanning seven risk categories, using a mix of existing benchmarks, GPT-4o-assisted generation, image synthesis, and human review. They then fine-tuned an open 7B vision-language model, Qwen2.5-VL, into ShieldVLM. Before giving a safety label, ShieldVLM writes a step-by-step analysis of what the text and image each suggest, what they imply together, and which rule they break. On the authors' test set, ShieldVLM beats OpenAI's moderation API, GPT-4o, Claude, and Llama Guard in detecting both implicit and explicit toxicity. It also does well on out-of-distribution benchmarks for explicit content. The main limits are that no code or data are released in this version, and the in-domain test labels were made with the same GPT-4o-assisted pipeline that created the training material, so part of the advantage may be alignment with GPT-4o's judgments.
Extended reading notes

Core claim

The experiments show that ShieldVLM outperforms existing strong baselines in detecting both implicit and explicit toxicity (Abstract). Section 6.4 states: "ShieldVLM demonstrates the highest accuracy in detecting both implicit and explicit toxicity across three forms of multimodal content." If true, a 7B-parameter open model fine-tuned with deliberative reasoning beats closed large models and dedicated moderation tools on the MMIT benchmark and on OOD explicit-toxicity benchmarks.

Load-bearing premise

The MMIT ground-truth labels and reasoning analyses, produced by a GPT-4o-driven generation, verification, and reasoning-annotation pipeline and reviewed by the authors and colleagues, are a reliable standard for what counts as multimodal implicit toxicity. If GPT-4o's safety judgments are systematically different from a broader notion of toxicity, the in-domain evaluation in Tables 3 and 5 measures agreement with GPT-4o rather than true detection ability. This assumption enters at Sec 4.3 (automatic safety check and refinement) and Sec 5.2 (reasoning generation).

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central empirical claim does not rest on a mathematical derivation with fitted constants. It rests on the adequacy of the dataset and labels: GPT-4o's safety judgments, the five-mode taxonomy, and human annotation by the authors. These are domain assumptions, not standard mathematical axioms, and they are the load-bearing premises that future work should test by releasing artifacts and obtaining independent labels.

assumptions (4)
  • domain assumption GPT-4o's safety judgments are a valid proxy for ground-truth implicit toxicity in data generation, verification, and reasoning annotation.
    Invoked in Sec 4.3 Step 3 for automatic safety checks and iterative refinement, and in Sec 5.2 for generating reasoning analyses. The model is trained on these judgments and evaluated on test labels produced by the same pipeline.
  • domain assumption The five cross-modal correlation modes (Semantic Drift, Contextualization, Metaphor, Implication, Knowledge) are an adequate decomposition of how text-image combinations become implicitly toxic.
    The taxonomy is asserted in Sec 3.2 and used both to guide data generation (Sec 4.3 Step 2) and to analyze model performance (Sec 6.5.2); no empirical or theoretical argument establishes completeness.
  • domain assumption Human annotators, mainly the authors and their colleagues, provide reliable safety labels and reasoning reviews.
    Sec 4.4 describes three rounds of quality control by the authors and colleagues; the paper does not report inter-annotator agreement.
  • domain assumption The stated safety criteria for safe text and safe image (Sec 3.1) are an accepted standard for judging unimodal safety.
    The criteria are adapted from prior safety research cited in Sec 3.1, but the paper treats them as given rather than validating them.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs." pith.science (2026). https://pith.science/paper/KVEL644M

@misc{pith2026250514035,
  author       = {Pith},
  title        = {Pith review of: ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KVEL644M}},
  note         = {Machine review of arXiv:2505.14035}
}
read the original abstract

Toxicity detection in multimodal text-image content faces growing challenges, especially with multimodal implicit toxicity, where each modality appears benign on its own but conveys hazard when combined. Multimodal implicit toxicity appears not only as formal statements in social platforms but also prompts that can lead to toxic dialogs from Large Vision-Language Models (LVLMs). Despite the success in unimodal text or image moderation, toxicity detection for multimodal content, particularly the multimodal implicit toxicity, remains underexplored. To fill this gap, we comprehensively build a taxonomy for multimodal implicit toxicity (MMIT) and introduce an MMIT-dataset, comprising 2,100 multimodal statements and prompts across 7 risk categories (31 sub-categories) and 5 typical cross-modal correlation modes. To advance the detection of multimodal implicit toxicity, we build ShieldVLM, a model which identifies implicit toxicity in multimodal statements, prompts and dialogs via deliberative cross-modal reasoning. Experiments show that ShieldVLM outperforms existing strong baselines in detecting both implicit and explicit toxicity. The model and dataset will be publicly available to support future researches. Warning: This paper contains potentially sensitive contents.

Figures

Figures reproduced from arXiv: 2505.14035 by the authors.

Figure 1
Figure 1. Examples for multimodal implicit toxicity in forms [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Performance gap of representative moderation [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Illustration to the cross-modal correlation modes. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Illustration to the format, reasoning process and construction of ShieldVLM. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Model performances across correlation modes. [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 7
Figure 7. Figure 7: Case Study. performances in categories of Illegal Activities, Physical Harm and Provicy & Property Damage. We speculate that this is due to these risks are usually expressed more straightforwardly, with over 30% of instances falling under contextualization. Meanwhile, …

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

    cs.CL 2025-09 conditional novelty 4.0 of 10

    A structured literature survey concluding that reasoning capabilities do not automatically make LLMs more trustworthy and can introduce new vulnerabilities in safety, robustness, and privacy.

Reference graph

Works this paper leans on

55 extracted references · 25 canonical work pages · cited by 1 Pith paper

  1. [1]

    Stability AI. [n. d.].Stable-Diffusion-3.5-Medium. https://huggingface.co/ stabilityai/stable-diffusion-3.5-medium

  2. [2]

    Amazon. [n. d.].Amazon Rekognition Content Moderation. https://aws.amazon. com/rekognition/content-moderation/

  3. [3]

    2024.Claude 3.5 Sonnet

    Anthropic. 2024.Claude 3.5 Sonnet. https://www.anthropic.com/news/claude-3- 5-sonnet

  4. [4]

    Azure. [n. d.].Azure AI Content Safety Image Moderation. https://learn.microsoft. com/en-us/shows/responsible-ai/azure-ai-content-safety-image-moderation

  5. [5]

    2023.Azure AI Content Safety

    Azure. 2023.Azure AI Content Safety. https://azure.microsoft.com/en-us/ products/ai-services/ai-content-safety

  6. [6]

    2024.Analyze multimodal content (preview)

    Azure. 2024.Analyze multimodal content (preview). https://learn.microsoft.com/ en-us/azure/ai-services/content-safety/quickstart-multimodal

  7. [7]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Ming-Hsuan Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical ...

  8. [8]

    Meghana Moorthy Bhat, Saghar Hosseini, Ahmed Hassan Awadallah, Paul Ben- nett, and Weisheng Li. 2021. Say ‘YES’ to Positivity: Detecting Toxic Language in Workplace Communications. InFindings of the Association for Computational Linguistics: EMNLP 2021. 2017–2029. doi:10.18653/v1/2021.findings-emnlp.173

Show all 55 references
  1. [9]

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, Lixin Gu, Xuehui Wang, Qingyun Li, Yimin Ren, Zixuan Chen, Jiapeng Luo, Jiahao Wang, Tan Jiang, Bo Wang, Conghui He, Botian Shi, Xingcheng Zhang, Han Lv, Yi...

  2. [10]

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Zhong Muyan, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. 2023. Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.2024 I...

  3. [11]

    Upasani, and Mahesh Pa- supuleti

    Jianfeng Chi, Ujjwal Karn, Hongyuan Zhan, Eric Smith, Javier Rando, Yiming Zhang, Kate Plawiak, Zacharie Delpierre Coudert, K. Upasani, and Mahesh Pa- supuleti. 2024. Llama Guard 3 Vision: Safeguarding Human-AI Image Understand- ing Conversations.ArXivabs/2411.10414 (2024). ht...

  4. [12]

    Shiyao Cui, Zhenyu Zhang, Yilong Chen, Wenyuan Zhang, Tianyun Liu, Siqi Wang, and Tingwen Liu. 2023. FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity.CoRRabs/2311.18580 (2023). doi:10.48550/ARXIV.2311.18580 arXiv:2311.18580

  5. [13]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ah- mad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sra- vankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston...

  6. [14]

    2024.Fine-Tuned Vision Transformer (ViT) for NSFW Image Classifica- tion

    Falcons˙ai. 2024.Fine-Tuned Vision Transformer (ViT) for NSFW Image Classifica- tion. https://huggingface.co/Falconsai/nsfw_image_detection

  7. [15]

    Shaona Ghosh, Prasoon Varshney, Erick Galinkin, and Christopher Parisien

  8. [16]

    Shaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar, Aishwarya Padmakumar, Traian Rebedea, Jibin Rajan Varghese, and Christopher Parisien

  9. [17]

    Raul Gomez, Jaume Gibert, Lluís Gómez, and Dimosthenis Karatzas. 2020. Exploring Hate Speech Detection in Multimodal Publications. InIEEE Win- ter Conference on Applications of Computer Vision, W ACV 2020. 1459–1467. doi:10.1109/WACV45572.2020.9093414

  10. [18]

    Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. 2024. WildGuard: Open One-stop Mod- eration Tools for Safety Risks, Jailbreaks, and Refusals of LLMs. InAdvances in Neural Information Processing Systems 38: An...

  11. [19]

    Xuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang, and Jing Shao. 2024. VLSBench: Unveiling Visual Leakage in Multimodal Safety.CoRRabs/2411.19939 (2024). doi:10.48550/ARXIV.2411.19939 arXiv:2411.19939

  12. [20]

    Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. 2023. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.CoRRabs/2312.06674 (2023). doi:10...

  13. [21]

    Tran, Yi Tay, Jeffrey Sorensen, Jai Prakash Gupta, Donald Metzler, and Lucy Vasserman

    Alyssa Lees, Vinh Q. Tran, Yi Tay, Jeffrey Sorensen, Jai Prakash Gupta, Donald Metzler, and Lucy Vasserman. 2022. A New Generation of Perspective API: Efficient Multilingual Character-level Transformers. InKDD ’22: The 28th ACM SIGKDD Conference on Knowledge Discovery and Data...

  14. [22]

    Hongzhan Lin, Ziyang Luo, Bo Wang, Ruichao Yang, and Jing Ma. 2024. GOAT- Bench: Safety Insights to Large Multimodal Models through Meme-Based Social Abuse.CoRRabs/2401.01523 (2024). doi:10.48550/ARXIV.2401.01523 arXiv:2401.01523

  15. [23]

    Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

    Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. InEuropean Conference on Computer Vision. https://api. semanticscholar.org/CorpusID:14113767

  16. [24]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruc- tion Tuning. InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, Alic...

  17. [25]

    Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao. 2024. MM- SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models. InComputer Vision - ECCV 2024 - 18th European Conference (Lecture Notes in Computer Science, Vol. 15114), Ales Leo...

  18. [26]

    Yue Liu, Hongcheng Gao, Shengfang Zhai, Jun Xia, Tianyi Wu, Zhiwei Xue, Yulin Chen, Kenji Kawaguchi, Jiaheng Zhang, and Bryan Hooi. 2025. GuardReasoner: Towards Reasoning-based LLM Safeguards.CoRRabs/2501.18492 (2025). doi:10. 48550/ARXIV.2501.18492 arXiv:2501.18492

  19. [27]

    Yida Lu, Jiale Cheng, Zhexin Zhang, Shiyao Cui, Cunxiang Wang, Xiaotao Gu, Yuxiao Dong, Jie Tang, Hongning Wang, and Minlie Huang. 2025. LongSafety: Evaluating Long-Context Safety of Large Language Models.CoRRabs/2502.16971 (2025). doi:10.48550/ARXIV.2502.16971 arXiv:2502.16971

  20. [28]

    Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. 2024. JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks.CoRRabs/2404.03027 (2024). doi:10. 48550/ARXIV.2404.03027 arXiv:2404.03027

  21. [29]

    Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng. 2023. A Holistic Approach to Undesired Content Detection in the Real World. InThirty-Seventh AAAI Conference on Artificial Intelligence, AAAI2023...

  22. [30]

    2024.Llama 3 ˙2: Revolutionizing edge AI and vision with open, customizable models

    Meta. 2024.Llama 3 ˙2: Revolutionizing edge AI and vision with open, customizable models. https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile- devices/

  23. [31]

    Yingqian Min, Zhipeng Chen, Jinhao Jiang, Jie Chen, Jia Deng, Yiwen Hu, Yiru Tang, Jiapeng Wang, Xiaoxue Cheng, Huatong Song, Wayne Xin Zhao, Zheng Liu, Zhongyuan Wang, and Ji-Rong Wen. 2024. Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning ...

  24. [32]

    Sy, Ahmad Albunni, Myles Joshua Toledo Tan, and Nouar Aldahoul

    Mhd Adel Momo, Hezerul Bin Abdul Karim, Michael Aaron G. Sy, Ahmad Albunni, Myles Joshua Toledo Tan, and Nouar Aldahoul. 2023. Evaluation of Convolution and Attention Networks for Nudity and Pornography Detection in Sketch Images. 2023 IEEE Symposium on Computers & Informatics...

  25. [33]

    2024.GPT-4o system card

    OpenAI. 2024.GPT-4o system card. https://openai.com/index/gpt-4o-system- card/

  26. [34]

    2024.GPT-4V(ision) system card

    OpenAI. 2024.GPT-4V(ision) system card. https://openai.com/index/gpt-4v- system-card/

  27. [35]

    2024.Moderate images and text

    OpenAI. 2024.Moderate images and text. https://openai.com/index/upgrading- the-moderation-api-with-our-new-multimodal-moderation-model/

  28. [36]

    2025.Qwen2.5-VL-7B-Instruct

    Qwen Team. 2025.Qwen2.5-VL-7B-Instruct. https://huggingface.co/Qwen/Qwen2. 5-VL-7B-Instruct

  29. [37]

    Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2021. Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detec- tion. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International...

  30. [38]

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. Qwen2-VL: Enhancing Vision-Language Mode...

  31. [39]

    Siyin Wang, Xingsong Ye, Qinyuan Cheng, Junwen Duan, Shimin Li, Jinlan Fu, Xipeng Qiu, and Xuanjing Huang. 2024. Cross-Modality Safety Alignment.CoRR abs/2406.15279 (2024). doi:10.48550/ARXIV.2406.15279

  32. [40]

    Xiaofei Wen, Wenxuan Zhou, Wenjie Jacky Mo, and Muhao Chen. 2025. Think- Guard: Deliberative Slow Thinking Leads to Cautious Guardrails.CoRR abs/2502.13458 (2025). doi:10.48550/ARXIV.2502.13458 arXiv:2502.13458

  33. [41]

    Mengyang Wu, Yuzhi Zhao, Jialun Cao, Mingjie Xu, Zhongming Jiang, Xuehui Wang, Qinbin Li, Guangneng Hu, Shengchao Qin, and Chi-Wing Fu. 2024. ICM- Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation.CoRRabs/2412.18...

  34. [42]

    Chunpu Xu, Hanzhuo Tan, Jing Li, and Piji Li. 2022. Understanding Social Media Cross-Modality Discourse in Linguistic Space. InFindings of the Association for Computational Linguistics: EMNLP 2022. 2459–2471. doi:10.18653/v1/2022. findings-emnlp.182

  35. [43]

    Fan Yin, Philippe Laban, Xiangyu Peng, Yilun Zhou, Yixin Mao, Vaibhav Vats, Linnea Ross, Divyansh Agarwal, Caiming Xiong, and Chien-Sheng Wu. 2025. BingoGuard: LLM Content Moderation Tools with Risk Levels. InICLR 2025

  36. [44]

    Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, Joe Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush Kumar, Bhaktipriya Rad- harapu, Olivia Sturman, and Oscar Wahltinez. 2024. ShieldGemma: Genera- tive AI Content Moderation Based on Gemma.CoRRabs/2407.2177...

  37. [45]

    Yongting Zhang, Lu Chen, Guodong Zheng, Yifeng Gao, Rui Zheng, Jinlan Fu, Zhenfei Yin, Senjie Jin, Yu Qiao, Xuanjing Huang, Feng Zhao, Tao Gui, and Jing Shao. 2024. SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model.CoRRabs/2406.12030 (2024)....

  38. [46]

    Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2024. SafetyBench: Evaluating the Safety of Large Language Models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguisti...

  39. [47]

    Zhexin Zhang, Yida Lu, Jingyuan Ma, Di Zhang, Rui Li, Pei Ke, Hao Sun, Lei Sha, Zhifang Sui, Hongning Wang, and Minlie Huang. 2024. ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors. InFindings of the Association for Computational Linguistics:...

  40. [48]

    Qinyu Zhao, Ming Xu, Kartik Gupta, Akshay Asthana, Liang Zheng, and Stephen Gould. 2024. The First to Know: How Token Distributions Reveal Hidden Knowl- edge in Large Vision-Language Models?. InComputer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-...

  41. [49]

    Kaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Anderson Compalas, Dawn Song, and Xin Eric Wang. 2024. Multimodal Situational Safety.CoRRabs/2410.06172 (2024). doi:10.48550/ARXIV.2410.06172 arXiv:2410.06172

  42. [50]

    Ru Zhou, Wenya Guo, Xumeng Liu, Shenglong Yu, Ying Zhang, and Xiaojie Yuan

  43. [51]

    Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny

  44. [55]

    InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024

    MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. https://openreview.net/forum?id=1tZbq88f27

  45. [2023]

    InFindings of the Association for Computational Linguistics: ACL 2023

    AoM: Detecting Aspect-oriented Information for Multimodal Aspect-Based Sentiment Analysis. InFindings of the Association for Computational Linguistics: ACL 2023. 8184–8196. doi:10.18653/v1/2023.findings-acl.519

  46. [2024]

    doi:10.48550/ARXIV.2404.05993 arXiv:2404.05993

    AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts.CoRRabs/2404.05993 (2024). doi:10.48550/ARXIV.2404.05993 arXiv:2404.05993

  47. [2025]

    doi:10.48550/ARXIV.2501.09004 arXiv:2501.09004

    Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails.CoRRabs/2501.09004 (2025). doi:10.48550/ARXIV.2501.09004 arXiv:2501.09004

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.