{"id":"5317d3e2-6819-4bc9-a96e-8d1d8df4cc19","arxiv_id":"2411.10034","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Injecting inaudible low-frequency adversarial perturbations into loudspeaker audio can block vibration-based side-channel speech eavesdropping with over 97% classifier protection while preserving perceived audio quality.","lead":"EveGuard is a software defense that adds specially shaped, low-frequency noise to audio before playback, so vibration-based eavesdroppers (radar, accelerometer, laser) capture garbage while human listeners hear almost no difference. It works by training a generative model to mimic what an eavesdropper's sensor would capture, then optimizing the audio perturbation against that surrogate.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The defense's core frequency-response premise is untested for wideband vibrometers; an LV-FS01-class sensor could bypass the <500 Hz LFAP and recover speech from untouched high-frequency bands.","rationale":"The reader's surrogate-fidelity concern is plausible but partially answered by the physical evaluations: the PGM's outputs are played through real radar, IMU, and laser channels, so end-to-end transfer is tested for those specific devices. What is not tested is the class of vibrometers whose bandwidth invalidates the underlying premise. Because the LFAP is deliberately limited to below 500 Hz and Table 11 shows it is the main driver of the increased MCD, any sensor that can exploit higher frequencies sidesteps the defense. This is a scope/correctness risk rather than an internal inconsistency, and it can be settled by one additional hardware evaluation. I would keep the CONDITIONAL verdict, with the condition made explicit: either evaluate against a wideband vibrometer or restrict the claimed defense to low-frequency-limited side channels.","tokens_in":26388,"tokens_out":11027,"duration_ms":128563,"concrete_test":"Obtain an LV-FS01 or equivalent laser Doppler vibrometer with a flat response to at least 8 kHz, place it in the Figure 9(b) setup, play EveGuard-perturbed utterances from the Table 5 test set, and measure MCD, WER, and DDR both on the raw recovered signal and after a 500 Hz high-pass filter followed by the attacker's cGAN enhancement and speech recognition. If WER falls below about 30% or MCD below 8 in either condition, the central claim that EveGuard defeats vibration-based SSEAs is overbroad; if WER stays above approximately 68%, the wideband-sensor objection does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"EveGuard's mechanism depends on the claim (Sec. 4.1) that vibrometry side channels have SNR near 0 dB above roughly 2 kHz, so inaudible LFAPs below 500 Hz dominate the eavesdropped signal. The paper itself concedes at Sec. 4.1 that a high-end laser vibrometer (LV-FS01) can enhance the frequency response, and Table 1 lists VibSpeech at up to 16 kHz sampling. For such a wideband sensor, an attacker can high-pass the recovered signal above 500 Hz and reconstruct speech from the mid/high bands that EveGuard deliberately leaves intact; the LFAP then reduces to a weak low-frequency noise component. The ablation in Table 11 shows LFAP is the component driving MCD up (13.1 vs 4.8 for FIR alone), so the defense would lose most of its effectiveness if high-frequency bands become usable. EveGuard is only evaluated on a 77/60 GHz mmWave radar, a 500 Hz accelerometer, and a low-end laser microphone, all of which match the poor-high-frequency-response premise. No experiment includes a wideband vibrometer, yet the abstract and Table 1 frame EveGuard as protecting against vibration-based SSEAs generally. This unvalidated frequency-response premise is load-bearing: if a capable sensor can hear above 500 Hz, the core attack surface is not removed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes EveGuard, a software-only defense that adds imperceptible low-frequency adversarial perturbations to loudspeaker audio before playback, aiming to make vibration-based side-channel eavesdropping (mmWave radar, accelerometer, optical/laser sensor) fail to recover intelligible speech. The defense is a two-stage Perturbation Generator Model (PGM) combining FIR filtering with low-frequency adversarial perturbations (LFAPs), trained end-to-end through a learned differentiable surrogate of the eavesdropping channel called Eve-GAN, which uses few-shot unpaired audio-to-SSEA translation to reduce data collection cost. Evaluation on re-implemented SSEAs across sensors, distances, angles, materials, sampling rates, and adaptive attacks reports MCD above 13.4, WER above 68%, DDR at or below 3%, PESQ 3.42, and a 24-participant user study.","tokens_in":26620,"tokens_out":6702,"duration_ms":71625,"significance":"If the frequency-response premise holds, EveGuard is a valuable contribution: it is a software-only, hardware-free defense against multiple vibration side channels, with a plausible transfer mechanism and unusually broad experimental coverage including unseen rooms, materials, and adaptive attackers. The few-shot Eve-GAN is a practical modeling contribution, and the explicit comparison against Gaussian noise and vanilla adversarial perturbations is informative. The paper also ships a public demo audio page, which strengthens the perceptual claims. However, the claimed generality over all vibration-based SSEAs is not yet established because the central premise that side channels are insensitive above roughly 2 kHz is untested for wideband vibrometers; with that caveat, the contribution is significant but needs additional validation.","major_comments":[{"comment":"The defense's load-bearing premise is that vibration-based side channels have near-zero SNR above roughly 2 kHz, so an inaudible LFAP below 500 Hz dominates the eavesdropped signal. This premise is not tested for wideband sensors. The paper itself concedes that a high-end laser vibrometer (LV-FS01) can enhance the frequency response (Section 4.1), and Table 1 lists VibSpeech with up to 16 kHz sampling. All EveGuard experiments use a 77/60 GHz mmWave radar with a 10 or 12 kHz chirp, a 500 Hz accelerometer, and a low-end optical sensor; no wideband vibrometer is evaluated. Since Table 11 shows that LFAP is the component responsible for the MCD increase (13.1 vs. 4.8 for FIR alone), an attacker who can recover the untouched bands above 500 Hz could bypass most of the defense. Please add an evaluation against a wideband sensor (e.g., a VibSpeech-class or LV-FS01-class device), or explicitly restrict the threat model to sensors with poor high-frequency response.","section":"Section 4.1, Table 1, Table 11"},{"comment":"The main quantitative claims are reported as single-point MCD, WER, and DDR values without error bars, standard deviations, or test-set sizes. Because the central claim of more than 97% protection and the comparisons to Gaussian and VAP baselines depend on these numbers, the paper should report means and confidence intervals over repeated sensor captures and over the test utterances. Without these, it is difficult to assess whether the differences between conditions and baselines are significant, and the reader cannot gauge run-to-run variability of the sensor channel.","section":"Section 6.4, Tables 5 and 9"},{"comment":"The fidelity of the surrogate Eve-GAN is quantified only by SSIM on full spectrograms (94.02% for mmWave and 96.51% for accelerometer), which does not directly constrain the sub-500 Hz band where LFAP energy is placed. The end-to-end transfer experiments in Section 6.4 provide indirect evidence that the surrogate works for the specific sensor hardware evaluated, but they do not establish gradient alignment for unseen wideband hardware. If the surrogate diverges in the LFAP band for a different sensor, the optimized perturbations may not transfer. Please report a spectral fidelity metric focused on the LFAP band, or evaluate with a second independently built surrogate channel, to strengthen this link.","section":"Section 5.3, Section 6.7"}],"minor_comments":[{"comment":"There is a typo: \"EvGuard\" should be \"EveGuard\".","section":"Section 6.1"},{"comment":"The text says \"WER of up to 85.5\" in the caption; this should be \"85.5%\".","section":"Figure 12(b)"},{"comment":"The symbol Y^k_s is used for a probability of a class label; please use a probability notation such as P^k_s for consistency with P r^k_s.","section":"Equation (10)"},{"comment":"Appendix H contains the program committee meta-review from a prior venue. This is not part of the scientific content of a journal submission and should be removed before publication.","section":"Appendix H"},{"comment":"The objective in Equation (1) is written as an arg max with hard constraints, while the training loss in Equation (7) is a weighted sum. Please clarify how the constraints are enforced or map the Lagrangian weights to the constraints.","section":"Section 5.1, Equation (1)"},{"comment":"Please report whether all 24 participants heard the same 30 audio pairs and how the Likert-scale responses were aggregated into the reported averages.","section":"Section 6.6"}],"recommendation":"major_revision","confidential_remarks":"The main substantive concern is the unvalidated wideband-vibrometer premise; this is fixable within the manuscript's scope by adding one targeted experiment or by narrowing the threat model. The inclusion of Appendix H (the prior meta-review) is unusual and should be removed. Code and data release would substantially increase confidence in the quantitative claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: EveGuard is the first software-only defense I've seen against vibration-based side-channel speech eavesdropping, and it's largely well done. The core trick—exploiting the fact that vibrometry sensors have poor high-frequency response while human hearing is insensitive below 500 Hz—is physically plausible, and the experiments back it up for the sensors they tested. The Eve-GAN few-shot translator is a nice piece of engineering that makes end-to-end training feasible without huge paired datasets.\n\nWhat the paper does well: the evaluation is broad (mmWave at two frequencies, accelerometer at multiple sampling rates, optical sensor, varying distances, angles, materials, volumes), and the adaptive attack analysis is more thorough than most. The ablation in Table 11 is honest and useful: it shows the LFAP component is what pushes MCD from 4.8 to 13.1, while FIR alone still raises WER to 52.5%. The user study adds real evidence for the inaudibility claim.\n\nNow the soft spots, in order of severity. The wideband vibrometer gap is the one that matters. The whole defense leans on the claim that side-channel SNR drops to near zero above ~2 kHz, so a low-frequency perturbation dominates. But they only test sensors that have poor high-frequency response: a 77/60 GHz mmWave radar, a 500 Hz accelerometer, and a cheap laser microphone. They cite VibSpeech and mention the LV-FS01 vibrometer as able to enhance frequency response, but they don't test it. An attacker with a wideband sensor could high-pass the recovered signal and keep the mid/high speech bands that EveGuard leaves intact. To be fair, the FIR component still gives partial protection (Table 11), so it wouldn't be a total bypass, but the headline 97% would likely drop. This needs an experiment or a clear scope limitation in the abstract.\n\nThe other issues are conventional: no code/data release, no error bars on the main metrics, and the LFAP SNR rho is tuned on the same tradeoff it's evaluated on. Those are addressable and don't undermine the core result.\n\nOverall, this is a serious paper with a real contribution. It deserves peer review, and with a wideband-vibrometer experiment (or an honest scope statement) it could be much stronger. I'd engage with it.","headline":"A well-engineered software-only defense against vibration eavesdropping, but the core frequency-response assumption is only validated on low-bandwidth sensors; a wideband vibrometer could bypass the main perturbation component.","tokens_in":27211,"tokens_out":2945,"would_cite":true,"duration_ms":31395,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"EveGuard scrambles speech for vibration spies, not for human ears.","keywords":["adversarial audio perturbations","side-channel eavesdropping","mmWave radar","voice privacy","generative adversarial network","few-shot domain translation","low-frequency perturbations","speech recognition resistance"],"falsifier":"Run EveGuard against a side channel whose sensor has a flat, high-sensitivity response below 500 Hz, such as a high-end laser vibrometer or an accelerometer mounted directly on the speaker enclosure, and measure whether reconstructed speech still exceeds the reported WER and MCD thresholds; if the eavesdropper recovers intelligible speech, the surrogate-transfer assumption fails.","tokens_in":26135,"feed_emoji":"🛡️","tokens_out":5058,"duration_ms":51327,"temperature":0.7,"pith_summary":"The paper tries to establish that vibration-based side-channel eavesdropping—recovering speech from mmWave radar, accelerometers, or optical sensors that pick up sound-induced vibrations—can be defeated in software before the audio is even played. EveGuard adds perturbations concentrated below 500 Hz, where these sensors listen but human hearing is least sensitive, so the eavesdropper's reconstruction becomes noise while the legitimate listener still hears clean speech. To train those perturbations, the paper builds Eve-GAN, a differentiable few-shot translator that simulates the eavesdropping channel from unpaired audio and sensor samples, enabling end-to-end optimization against an ensemble of surrogate attack models. If the central claim is right, this would give loudspeaker owners a practical, hardware-free privacy defense that works across attack distances, materials, and sensor types.","feed_headline":"EveGuard scrambles speech for vibration spies, not for human ears","feed_subtitle":"Software adds inaudible low-frequency noise that wrecks radar, laser, and accelerometer audio recovery.","key_machinery":"Two learned components carry the argument. Eve-GAN is a few-shot unpaired audio-to-side-channel translator: given clean audio and one reference eavesdropped sample, it outputs what that audio would sound like after passing through a particular sensor channel, making the attacker's reconstruction differentiable for gradient-based optimization. The Perturbation Generator Model (PGM) is a two-stage network: an FIR generator produces time-varying frequency-domain filters through a VAE-GAN bottleneck for spectral reshaping, and an LFAP generator emits random low-frequency (below 500 Hz) noise normalized by SNR; a discriminator keeps the output close to natural speech and diversifies the perturbations. Together they convert the defense into an end-to-end optimizable problem whose solution is imperceptible to humans but destructive to vibration-based eavesdroppers.","core_discovery":"On its own terms, the paper's central claim is that the frequency-response gap between vibrometry sensors and microphones is a reliable handle for privacy: side channels capture mostly low-frequency vibration while microphones and ears perceive the full band, with minimal human sensitivity below 500 Hz. EveGuard exploits this by generating two kinds of perturbations—a VAE-GAN-driven FIR filter that reshapes the audio spectrum and low-frequency adversarial perturbations—trained end-to-end through Eve-GAN and an ensemble of ten surrogate speech recognizers and classifiers. The reported result is that eavesdropped audio becomes unintelligible: Mel-Cepstral Distortion rises from roughly 3.4 to 13.6, Word Error Rate exceeds 68 percent, digit-classifier accuracy drops below 5 percent, while PESQ stays at 3.42, meaning the perturbed audio is perceptually close to the original, confirmed by a user study. EveGuard also reports robustness to adversarial training, perturbation removal, and speech transformations, and low enough latency for VoIP. The paper explicitly scopes EveGuard to loudspeaker playback and notes that it does not cover throat-vibration eavesdropping on live human speech.","pith_inferences":["The same low-frequency insertion principle could be ported to electromagnetic side channels and other vibration sensors if they exhibit a comparable low-frequency-dominant response, a direction the paper itself flags as future work.","The method's success hinges on low-frequency perturbations surviving the acoustic path to the vibrating object; testing in noisier, more reverberant, or more damped rooms where LFAP energy dissipates would be a natural stress test not covered by the reported scenarios.","If Eve-GAN's surrogate underestimates the true channel's sensitivity at higher frequencies for some sensor, the defense may need per-sensor calibration or an expanded ensemble to stay effective.","The few-shot, unpaired translation recipe could be reused to model other sensing channels, such as through-wall radar or lidar, for other privacy defenses without collecting paired clean and eavesdropped audio."],"forward_implications":["Loudspeaker privacy can be protected without hardware jammers, shields, or vibration motors, simply by preprocessing audio before playback.","The same defense transfers across mmWave radar frequencies, sampling rates, antenna configurations, distances, angles, volumes, insulators, and reverberating materials, because perturbations are placed in the band every tested sensor hears.","Perturbation diversity from random latent vectors makes it hard for an attacker who knows about EveGuard to train a denoiser or robust recognizer, since the perturbation distribution is not fixed.","EveGuard is deployable in real-time voice pipelines: 50 ms segments process in 2.7 to 11.7 ms on tested hardware, under the 150 ms VoIP budget.","Because the defense acts before playback, it also works for optical and accelerometer side channels, not just mmWave radar."],"supporting_citations":[{"why":"Supplies the mmWave radar acoustic eavesdropping method and the frequency-response behavior EveGuard defends against.","marker":"[23]"},{"why":"Provides the cGAN-based speech enhancement and unconstrained-vocabulary eavesdropping pipeline used in the evaluation.","marker":"[24]"},{"why":"Establishes the accelerometer-based side-channel attack and the low sampling-rate vibration channel.","marker":"[25]"},{"why":"Provides the mmWave phased-MIMO eavesdropping method and the digit classifier used as an attack model.","marker":"[59]"},{"why":"Demonstrates zero-permission motion-sensor eavesdropping that motivates the accelerometer defense.","marker":"[63]"},{"why":"Supplies the FIR-filter perturbation idea that PGM extends with VAE-GAN and side-channel constraints.","marker":"[47]"},{"why":"Provides the few-shot unsupervised image-to-image translation basis used by Eve-GAN.","marker":"[42]"},{"why":"Supplies the multi-period discriminator architecture used in Eve-GAN training.","marker":"[37]"},{"why":"Supplies the VAE-GAN mechanism used to generate diverse FIR filters in the PGM.","marker":"[22]"}],"fun_headline_variants":["EveGuard hides speech from laser spies without ruining audio","Sonic shield: EveGuard blocks vibration eavesdropping invisibly","EveGuard: 97% privacy win against vibration-based spying","EveGuard defeats laser eavesdroppers with inaudible noise","EveGuard: audio that sounds normal but blocks vibration spies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole defense rests on the assumption that the differentiable Eve-GAN surrogate faithfully matches the real eavesdropping channel in the low-frequency band where perturbations live, so that perturbations optimized against it transfer to real sensors.","fun_headline_variants_meta":{"raw":{"variants":["EveGuard hides speech from laser spies without ruining audio","Sonic shield: EveGuard blocks vibration eavesdropping invisibly","EveGuard: 97% privacy win against vibration-based spying","EveGuard defeats laser eavesdroppers with inaudible noise","EveGuard: audio that sounds normal but blocks vibration spies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000917,"raw_usage":{"total_tokens":3976,"prompt_tokens":1024,"completion_tokens":2952,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":2872}},"tokens_in":640,"tokens_out":2952,"duration_ms":19577,"temperature":1.0,"reasoning_tokens":2872,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:02:06.764758+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run EveGuard against a side channel whose sensor has a flat, high-sensitivity response below 500 Hz, such as a high-end laser vibrometer or an accelerometer mounted directly on the speaker enclosure, and measure whether reconstructed speech still exceeds the reported WER and MCD thresholds; if the eavesdropper recovers intelligible speech, the surrogate-transfer assumption fails.","supporting_citations":[{"cited_title":"mmecho: A mmwave-based acoustic eavesdropping method","cited_arxiv_id":null,"evidence_quote":"Supplies the mmWave radar acoustic eavesdropping method and the frequency-response behavior EveGuard defends against."},{"cited_title":"Milliear: Millimeter-wave acoustic eavesdropping with unconstrained vocabulary","cited_arxiv_id":null,"evidence_quote":"Provides the cGAN-based speech enhancement and unconstrained-vocabulary eavesdropping pipeline used in the evaluation."},{"cited_title":"Accear: Accelerometer acoustic eavesdropping with unconstrained vocabu- lary","cited_arxiv_id":null,"evidence_quote":"Establishes the accelerometer-based side-channel attack and the low sampling-rate vibration channel."},{"cited_title":"Privacy leakage via speech-induced vibrations on room objects through remote sensing based on phased-mimo","cited_arxiv_id":null,"evidence_quote":"Provides the mmWave phased-MIMO eavesdropping method and the digit classifier used as an attack model."},{"cited_title":"Stealthyimu: Stealing permission-protected private information from smartphone voice assistant using zero-permission sensors","cited_arxiv_id":null,"evidence_quote":"Demonstrates zero-permission motion-sensor eavesdropping that motivates the accelerometer defense."},{"cited_title":"V oiceblock: Privacy through real-time adversarial attacks with audio-to-audio models","cited_arxiv_id":null,"evidence_quote":"Supplies the FIR-filter perturbation idea that PGM extends with VAE-GAN and side-channel constraints."},{"cited_title":"Few-shot unsupervised image-to- image translation","cited_arxiv_id":null,"evidence_quote":"Provides the few-shot unsupervised image-to-image translation basis used by Eve-GAN."},{"cited_title":"Hierarchical patch vae-gan: Generating diverse videos from a single sample","cited_arxiv_id":null,"evidence_quote":"Supplies the VAE-GAN mechanism used to generate diverse FIR filters in the PGM."}],"review_version":1}