Pith. sign in

REVIEW 6 cited by

Did you hear that? Adversarial Examples Against Automatic Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1801.00554 v1 pith:P46VGQAT submitted 2018-01-02 cs.CL cs.CR

classification cs.CLcs.CR
keywords attacksspeechadversarialrecognitionaudioautomaticcliphuman
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Speech is a common and effective way of communication between humans, and modern consumer devices such as smartphones and home hubs are equipped with deep learning based accurate automatic speech recognition to enable natural interaction between humans and machines. Recently, researchers have demonstrated powerful attacks against machine learning models that can fool them to produceincorrect results. However, nearly all previous research in adversarial attacks has focused on image recognition and object detection models. In this short paper, we present a first of its kind demonstration of adversarial attacks against speech classification model. Our algorithm performs targeted attacks with 87% success by adding small background noise without having to know the underlying model parameter and architecture. Our attack only changes the least significant bits of a subset of audio clip samples, and the noise does not change 89% the human listener's perception of the audio clip as evaluated in our human study.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative Testing of Automated Speech Recognition Systems

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Phoneme-level latent interpolation in a TTS model yields ~98% black-box ASR failures with higher naturalness than waveform attacks and quality competitive with white-box PGD.

  2. ASRJam: Human-Friendly AI Speech Jamming to Prevent Automated Phone Scams

    cs.CL 2025-06 reject novelty 6.0 of 10

    EchoGuard adds echo-like acoustic distortions to outgoing speech that confuse scam bots' speech recognition while leaving human callers able to understand.

  3. WhisperFlow: speech foundation models in real time

    cs.SD 2024-12 conditional novelty 6.0 of 10

    WhisperFlow combines a learned 'hush word', beam pruning, and CPU/GPU pipelining to cut streaming Whisper latency by 1.6x-4.7x on client devices with near-unchanged accuracy.

  4. Universal Adversarial Audio Perturbations

    cs.LG 2019-08 conditional novelty 6.0 of 10

    Universal adversarial audio perturbations, found by a penalty-based optimizer, misclassify over 85% of test sounds across several 1D CNN audio classifiers, in both targeted and untargeted settings.

  5. Random Directional Attack for Fooling Deep Neural Networks

    cs.CR 2019-08 conditional novelty 6.0 of 10

    A hill-climbing search over randomly rotated directions generates adversarial examples with success rates competitive with gradient-based attacks, including in black-box settings.

  6. Imperio: Robust Over-the-Air Adversarial Examples for Automatic Speech Recognition Systems

    cs.CR 2019-08 conditional novelty 6.0 of 10

    Imperio generates targeted over-the-air adversarial audio for a hybrid ASR system by optimizing against many simulated room impulse responses, and achieves some 0% WER transcriptions in real rooms.

Pith tools