Pith. sign in

REVIEW 4 cited by

Making Acoustic Side-Channel Attacks on Noisy Keyboards Viable with LLM-Assisted Spectrograms' "Typo" Correction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.11622 v1 pith:IJGV6E25 submitted 2025-04-15 cs.CR cs.SDeess.AS

classification cs.CRcs.SDeess.AS
keywords ascasinformationllmsmodelmodelsperformancesolutionsacoustic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The large integration of microphones into devices increases the opportunities for Acoustic Side-Channel Attacks (ASCAs), as these can be used to capture keystrokes' audio signals that might reveal sensitive information. However, the current State-Of-The-Art (SOTA) models for ASCAs, including Convolutional Neural Networks (CNNs) and hybrid models, such as CoAtNet, still exhibit limited robustness under realistic noisy conditions. Solving this problem requires either: (i) an increased model's capacity to infer contextual information from longer sequences, allowing the model to learn that an initially noisily typed word is the same as a futurely collected non-noisy word, or (ii) an approach to fix misidentified information from the contexts, as one does not type random words, but the ones that best fit the conversation context. In this paper, we demonstrate that both strategies are viable and complementary solutions for making ASCAs practical. We observed that no existing solution leverages advanced transformer architectures' power for these tasks and propose that: (i) Visual Transformers (VTs) are the candidate solutions for capturing long-term contextual information and (ii) transformer-powered Large Language Models (LLMs) are the candidate solutions to fix the ``typos'' (mispredictions) the model might make. Thus, we here present the first-of-its-kind approach that integrates VTs and LLMs for ASCAs. We first show that VTs achieve SOTA performance in classifying keystrokes when compared to the previous CNN benchmark. Second, we demonstrate that LLMs can mitigate the impact of real-world noise. Evaluations on the natural sentences revealed that: (i) incorporating LLMs (e.g., GPT-4o) in our ASCA pipeline boosts the performance of error-correction tasks; and (ii) the comparable performance can be attained by a lightweight, fine-tuned smaller LLM (67 times smaller than GPT-4o), using...

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DECKER: Domain-invariant Embedding for Cross-Keyboard Extraction and Recognition

    cs.CR 2026-05 unverdicted novelty 7.0 of 10

    DECKER improves cross-keyboard and cross-user keystroke identification from audio using domain-invariant embeddings on the new diverse HEAR dataset, with additional gains from language model sequence correction.

  2. DECKER: Domain-invariant Embedding for Cross-Keyboard Extraction and Recognition

    cs.CR 2026-05 unverdicted novelty 6.0 of 10

    DECKER is a domain-invariant four-stage framework (keyboard normalization, adversarial disentanglement, cross-keyboard contrastive alignment, acoustic style randomization) plus LLM post-processing that improves keystr...

  3. Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack

    cs.CR 2026-06 unverdicted novelty 5.0 of 10

    KEYAC dataset created; KAN fine-tuning achieves SOTA on acoustic side-channel keystroke recognition from speech representations under zero-shot and partial fine-tuning.

  4. Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack

    cs.CR 2026-06 unverdicted novelty 5.0 of 10

    KEYAC dataset benchmarks speech models for keyboard acoustic side-channel attacks, with KAN fine-tuning setting new SOTA by addressing nonlinear feature interactions.

Pith tools