Pith. sign in

REVIEW 2 cited by

SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.08317 v1 pith:OFOP3N2N submitted 2024-05-14 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords modelsadversarialattackattacksjailbreakingsafetyslmsspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Integrated Speech and Large Language Models (SLMs) that can follow speech instructions and generate relevant text responses have gained popularity lately. However, the safety and robustness of these models remains largely unclear. In this work, we investigate the potential vulnerabilities of such instruction-following speech-language models to adversarial attacks and jailbreaking. Specifically, we design algorithms that can generate adversarial examples to jailbreak SLMs in both white-box and black-box attack settings without human involvement. Additionally, we propose countermeasures to thwart such jailbreaking attacks. Our models, trained on dialog data with speech instructions, achieve state-of-the-art performance on spoken question-answering task, scoring over 80% on both safety and helpfulness metrics. Despite safety guardrails, experiments on jailbreaking demonstrate the vulnerability of SLMs to adversarial perturbations and transfer attacks, with average attack success rates of 90% and 10% respectively when evaluated on a dataset of carefully designed harmful questions spanning 12 different toxic categories. However, we demonstrate that our proposed countermeasures reduce the attack success significantly.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff

    cs.IT 2026-04 unverdicted novelty 6.0 of 10

    Synset-based reconstruction and synonymous variational inference are claimed to derive the distributional divergence in RDP and unify it with classical rate-distortion theory.

  2. Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A representation-space reshaping method improves LALM safety against harmful audio queries while keeping over-rejection low.

Pith tools