Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Advances in Small-Footprint Keyword Spotting: A Comprehensive Review of Efficient Models and Algorithms

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Review maps small-footprint keyword spotting into seven technique families and shows Int8 quantization shrinks models by up to 69%.

desk verdict Useful survey with a quantitative claim that its own Table 7 contradicts; worth revising, not rejecting. read the letter →

arxiv 2506.11169 v1 pith:72KZ5CB4 submitted 2025-06-12 eess.AS cs.SD

classification eess.AScs.SD
keywords keywordspottingsmall-footprintTinyMLTensorFlowLitemodelquantizationneuralarchitecturesearchspeechcommandsdataset
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper attempts to establish that the large body of research on small-footprint keyword spotting can be organized into seven technique families, and that these techniques can be combined with TinyML frameworks to deploy models on low-power edge devices. The authors support this claim with their own case study on the Google Speech Commands Dataset, reporting that converting Keras models to TensorFlow Lite Int8 format reduces model size by up to 69% while keeping accuracy nearly intact. They also show that multi-objective optimization methods such as Bayesian optimization, simulated annealing, and NSGA-II can find Pareto-optimal trade-offs between accuracy and model size. If these findings hold, practitioners get both a structured map of the field and a concrete recipe for building efficient keyword spotters for microcontrollers.

What carries the argument

The central organizing device is the seven-category taxonomy of SF-KWS techniques, which structures the entire survey and yields the observation that architecture innovations account for 37.3% of surveyed work, followed by neural architecture search at 15.7%. The experimental cornerstone is the Keras-to-TensorFlow-Lite Int8 quantization pipeline applied to eight model variants on the Google Speech Commands Dataset, alongside an Edge Impulse pipeline and the MicroNet architectures. This pipeline makes the survey's qualitative recommendations concrete and provides measurable footprint numbers for real deployment. The multi-objective optimization experiments then extend the same logic to automatic architecture selection under accuracy and size constraints.

What would settle it

A different review team could re-analyze the same set of roughly 250 papers using a pre-defined categorization rubric and produce substantially different category assignments or distribution percentages, which would cast doubt on the taxonomy's stability; likewise, reproducing the case study with a different model family or version of TensorFlow Lite could yield a size reduction far from the reported 69%.

Watch

Extended reading notes

Core claim

The paper claims that the SF-KWS literature is best understood through seven categories: model architecture, learning techniques, model compression, attention-aware architecture, feature optimization, neural architecture search, and hybrid approaches. On the experimental side, the paper reports that after converting DNN, CNN, DS-CNN, and MicroNet models to TensorFlow Lite Int8, model size drops by up to 69% with negligible accuracy loss, and that small models such as CNN-S and DS-CNN-S fit easily in typical microcontroller flash memory. Among the optimized models, MicroNet-S achieves the highest accuracy (95.3%). In a second experiment, tuning CNN, CRNN, and DS-CNN hyperparameters with Bayesian optimization, simulated annealing, and NSGA-II, the paper finds that Bayesian optimization yields the best balance between accuracy and model size when both are weighted equally.

Load-bearing premise

The taxonomy assumes that the seven chosen categories are a faithful and reproducible way to organize the SF-KWS literature, but the paper does not specify a systematic protocol for assigning works to categories.

Editorial extensions

If this is right

  • Int8 quantization is a workable default for small-footprint keyword spotting: the paper's numbers imply that roughly two-thirds of a model's footprint can be removed without significantly changing accuracy.
  • Small quantized convolutional models such as CNN-S and DS-CNN-S can be deployed directly in the flash memory of typical microcontroller boards, which is the practical precondition for always-on wake-word detection.
  • When accuracy and model size receive equal weight, Bayesian optimization is the paper's recommended route for tuning KWS hyperparameters, producing the best compromise among the three tested optimizers.
  • The taxonomy predicts where the field's effort is concentrated: most work has gone into architecture innovation, leaving attention-aware architectures and feature optimization as comparatively less explored territory.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the 69% size-reduction figure is tied to the specific Keras baselines and TensorFlow Lite conversion settings used here, so it should be treated as a representative magnitude rather than a universal guarantee across all keyword-spotting models.
  • Editorial inference: because the taxonomy's category assignment protocol is not specified, the reported category distribution could shift if another group reclassified the same papers; the survey's map is therefore most safely read as a useful heuristic rather than a precise census.
  • Editorial inference: the same two-case-study methodology could be extended to test newer architectures such as lightweight transformers or spiking networks, where the paper only surveys prior work and does not yet provide its own quantization or Pareto-front numbers.
  • Editorial inference: a reader wanting to act on the paper could take the best quantized small model and verify on their own board that end-to-end latency, RAM use, and false-alarm rate meet product requirements, since the reported metrics cover size and accuracy but not full system behavior in noisy streaming conditions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents a structured review of small-footprint keyword spotting (SF-KWS), organizing the literature into seven categories: model architecture, learning techniques, model compression, attention-aware architecture, feature optimization, neural architecture search, and hybrid approaches. It also reports two experimental case studies: (i) converting Keras models to TensorFlow Lite Int8 format and measuring model size, accuracy, and inference time, and (ii) multi-objective optimization (MOO) of CNN, CRNN, and DS-CNN hyperparameters using simulated annealing, Bayesian optimization, and NSGA-II. The manuscript claims that quantization yields a model-size reduction of up to 69% with minimal accuracy loss, and that Bayesian optimization produces the best Pareto-optimal models. The review portion draws on roughly 250 papers and includes tables summarizing architectures, datasets, TinyML frameworks, and benchmark results.

Significance. As a compilation, the survey could serve as a useful entry point to SF-KWS; the bibliography is broad and the seven-category taxonomy, although not fully reproducible from the text, covers the main methodological families. The experimental results would add practical value if they were reproducible, but the paper's own Table 7 contradicts the headline size-reduction claim for DNN-S, and the MOO experiments are presented without seeds, repetitions, or hardware/toolchain versions. The novelty is also partly inherited from the authors' earlier conference paper [20], which supplies Figures 18 and 19. The review's value should therefore be assessed primarily on curation and synthesis, not on the new quantitative evidence.

major comments (3)
  1. [Section 6.1, Table 7] The claim that converting Keras models to quantized Int8 with TensorFlow Lite 'led to a remarkable reduction in model size, up to 69%' is not supported by the table. DNN-S grows from 999 KB to 2147.9 KB, an increase of about 115%; DS-CNN-S shrinks from 480.6 KB to 98.7 KB, a reduction of about 79.5%, which exceeds the stated 69% upper bound; only CNN-S (922.4 KB to 280.4 KB) is close to 69%. The text must either correct the table, qualify the claim as holding for a subset of architectures, or explain the DNN-S anomaly. As printed, the central quantitative claim of the case study is internally inconsistent.
  2. [Section 6.2, Table 9 and Figure 21] The multi-objective optimization comparison is not reproducible from the manuscript. There are no random seeds, no repeated runs, no confidence intervals, and no statement of the training/evaluation budget or hardware/toolchain versions. In addition, the 'scalarization score' in Table 9 mixes units: 0.5 × Accuracy (%) minus 0.5 × Model Size (MB) means the accuracy term dominates the score (e.g., MOBO CNN: 43.775 vs NSGA-II CNN: 38.855), so the ranking is essentially by accuracy. The conclusion that 'Bayesian optimization performs better than NSGA-II' is not supported by a single run without variance or statistical testing.
  3. [Section 4, Table 3] The seven-category taxonomy is asserted without a reproducible curation protocol. The text states that 'almost 250 papers' were reviewed, but Table 3 lists only a small subset, and no inclusion/exclusion criteria, search strings, predefined category definitions, or independent labeling procedure are reported. Because the categorization is the paper's main organizational contribution, the absence of a defined protocol makes the map of the field difficult to verify or update systematically.
minor comments (5)
  1. [Abstract and throughout] There are grammatical errors and typos, including 'a efficient' in the abstract, 'T nyML' in Figure 17, and an incomplete sentence in Section 3.3 beginning 'Because waiting for additional context...'; these should be corrected.
  2. [Table 7] The header 'Inference Accuracy Time(ms)' is ambiguous; the columns should be clearly separated, and the units for inference time should be stated consistently. The text should also clarify whether GSCD v1 or v2 is used in Section 6.1 and which class split is used, since Section 6.2 specifies 10 classes.
  3. [Figures 18 and 19] Figures 18 and 19 are taken from the authors' prior conference paper [20]; the manuscript should state this explicitly and clarify which experimental results are new in this submission.
  4. [References] Reference formatting is inconsistent: [63] in Table 3 is attributed to 'Yusuf Goren' while the reference list points to Mishchenko et al.; [70] lists 'S. . Ark'; and [97] names 'Yundong Zhang' while the reference list entry is by Y. Zhang. These should be corrected.
  5. [Table 4] Several entries in Table 4 do not state the exact dataset version or configuration used for the reported accuracy; please add a column specifying the evaluation setting (e.g., GSCD v1 vs v2, 12-class vs 10-class, and any post-processing).

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the taxonomy and the TinyML/MOO experiments rest on external literature and direct measurements; self-citations such as [20] are not load-bearing.

full rationale

This paper is a survey with two small experimental case studies, and neither contains a derivation that reduces to its own inputs. The seven-category taxonomy in Section 4 is an organizational scheme applied to the external literature, and the category-distribution figures are descriptive summaries rather than predictions. The TinyML case study in Section 6.1 reports measured Keras-to-TFLite Int8 conversion results on the Google Speech Commands Dataset, including model size, inference time, and accuracy; the textual claim of 'up to 69%' size reduction is a summary of those measurements, not a quantity fitted to itself. The inconsistency between that summary and the DNN-S row of Table 7 (2147.9 KB vs 999 KB after conversion) is an internal-data-quality or reporting issue, not a circularity. The multi-objective optimization case study in Section 6.2 uses SA, BO, and NSGA-II to search hyperparameters and reports the resulting accuracy/size trade-offs; these are evaluation outcomes of an optimization procedure, not predictions forced by the optimizer's own objective. Self-citations exist, most noticeably the authors' earlier concise overview [20], which provides Figures 18 and 19 and some experimental framing. However, the present paper also reports the underlying data in Table 7 and points to an open-source implementation, so the central review conclusions do not depend on an unverified self-citation chain. No equation or fitted parameter is defined in terms of the very quantity the paper claims to establish, and no 'prediction' is a renamed input. The correct finding is therefore no significant circularity, with the table-vs-text size discrepancy flagged as a correctness issue rather than a circular one.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The paper's central claims rest on the taxonomy assumption and the experimental assumptions about benchmark representativeness and proxy metrics.

free parameters (1)
  • scalarization weight w = 0.5
    Used in Table 9 ranking objective 0.5*Accuracy - 0.5*Model Size; chosen by hand, not derived.
assumptions (3)
  • domain assumption Accuracy on Google Speech Commands is a representative proxy for SF-KWS performance.
    Used in experiments in Section 6; if the benchmark is not representative of real-world keyword spotting, the reported comparisons may not generalize.
  • domain assumption Model size in megabytes is a faithful proxy for deployability on microcontrollers.
    Used in MOO objectives in Section 6.2; ignores other footprint factors like RAM and latency.
  • domain assumption The reviewed papers are accurately interpreted and the field is well represented by the selected repositories.
    Section 2.1 states papers were curated but gives no reproducible selection protocol, so the review's coverage depends on the authors' choices.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Advances in Small-Footprint Keyword Spotting: A Comprehensive Review of Efficient Models and Algorithms." pith.science (2026). https://pith.science/paper/72KZ5CB4

@misc{pith2026250611169,
  author       = {Pith},
  title        = {Pith review of: Advances in Small-Footprint Keyword Spotting: A Comprehensive Review of Efficient Models and Algorithms},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/72KZ5CB4}},
  note         = {Machine review of arXiv:2506.11169}
}
read the original abstract

Small-Footprint Keyword Spotting (SF-KWS) has gained popularity in today's landscape of smart voice-activated devices, smartphones, and Internet of Things (IoT) applications. This surge is attributed to the advancements in Deep Learning, enabling the identification of predefined words or keywords from a continuous stream of words. To implement the SF-KWS model on edge devices with low power and limited memory in real-world scenarios, a efficient Tiny Machine Learning (TinyML) framework is essential. In this study, we explore seven distinct categories of techniques namely, Model Architecture, Learning Techniques, Model Compression, Attention Awareness Architecture, Feature Optimization, Neural Network Search, and Hybrid Approaches, which are suitable for developing an SF-KWS system. This comprehensive overview will serve as a valuable resource for those looking to understand, utilize, or contribute to the field of SF-KWS. The analysis conducted in this work enables the identification of numerous potential research directions, encompassing insights from automatic speech recognition research and those specifically pertinent to the realm of spoken SF-KWS.

Figures

Figures reproduced from arXiv: 2506.11169 by the authors.

Figure 1
Figure 1. System overview: on edge KWS vs. cloud service based ASR [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Year-wise distribution of keyword spotting publications (2017-2024) categorized by publication source. The chart [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Distribution of deep learning architectures used in recent KWS literature. [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (18 more)
Figure 4
Figure 4. Figure 4: The typical workflow of a deep learning-based spoken keyword spotting system involves three main steps: (i) [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: (a) The traditional procedure for obtaining Mel-frequency cepstral and log-Mel spectral speech features involves [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 6
Figure 6. Figure 6: Different types of the convolution operation. (a) Basic convolution. (b) Depthwise separable convolution. (c) Residual [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]
Figure 7
Figure 7. Figure 7: Framework of a typical Query-by-Example based KWS [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]
Figure 8
Figure 8. Figure 8: A broad classification of efficiency metrics. When deploying a model, its feasibility is typically assessed based on [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: Pareto optimal front for accuracy vs. model size. The red points (x) represent the optimal trade-offs, where accuracy [PITH_FULL_IMAGE:figures/full_fig_p021_9.png]
Figure 10
Figure 10. Figure 10: Distribution of SF-KWS research works based on algorithmic categories. [PITH_FULL_IMAGE:figures/full_fig_p022_10.png]
Figure 11
Figure 11. Figure 11: A simplified example highlighting the difference between 2D convolution and temporal convolution: (A) [PITH_FULL_IMAGE:figures/full_fig_p025_11.png]
Figure 12
Figure 12. Figure 12: ResNet architecture, with a magnified residual block. [PITH_FULL_IMAGE:figures/full_fig_p026_12.png]
Figure 13
Figure 13. Figure 13: Knowledge Distillation Process: A pre-trained Teacher model generates soft labels (probabilities) through a softmax [PITH_FULL_IMAGE:figures/full_fig_p029_13.png]
Figure 14
Figure 14. Figure 14: (a) Standard Self-Attention use three weight matrices, Wq, Wk, and Wv, each transforming the input X into queries (Q), key (K), and values (V). (b) Shared-Weight Self-Attention uses a single weight matrix of Wshared for all three transformations. This reduces the numb…
Figure 15
Figure 15. Figure 15: On-edge deployment of Fast Keyword Spotting System [PITH_FULL_IMAGE:figures/full_fig_p039_15.png]
Figure 16
Figure 16. Figure 16: Traditional NAS components: The algorithm explores the search space by generating candidate models and itera [PITH_FULL_IMAGE:figures/full_fig_p040_16.png]
Figure 17
Figure 17. Figure 17: Representative TnyML devices (MCUs) supported by TensorFlow Lite [9] [PITH_FULL_IMAGE:figures/full_fig_p044_17.png]
Figure 18
Figure 18. Figure 18: Comparison of Model Size vs. Accuracy for Software Models (Keras) and their TF-Lite Variants, highlighting [PITH_FULL_IMAGE:figures/full_fig_p048_18.png]
Figure 19
Figure 19. Figure 19: Inference Time Comparison of Software Models (Keras) vs. TF-Lite Models, showcasing the reduction in execution [PITH_FULL_IMAGE:figures/full_fig_p048_19.png]
Figure 20
Figure 20. Figure 20: Tradeoff between model performance and footprint. [PITH_FULL_IMAGE:figures/full_fig_p049_20.png]
Figure 21
Figure 21. Figure 21: Comparison of Accuracy and Model Size across different CNN, CRNN, and DS-CNN models obtained using three [PITH_FULL_IMAGE:figures/full_fig_p051_21.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OASI: Objective-Aware Surrogate Initialization for Multi-Objective Bayesian Optimization in TinyML Keyword Spotting

    cs.LG 2025-12 conditional novelty 5.0 of 10

    OASI, a simulated-annealing-based, objective-aware initialization for multi-objective Bayesian optimization, improves hypervolume and memory-feasible deployment for TinyML keyword spotting models, though the statistic...

Reference graph

Works this paper leans on

148 extracted references · 66 canonical work pages · cited by 1 Pith paper

  1. [20]

    Garai, S

    S. Garai, S. Samui, Exploring tinyml frameworks for small-footprint keyword spotting: A concise overview, in: 2024 International Conference on Signal Processing and Communications (SPCOM), IEEE, 2024, pp. 1–5

  2. [1]

    M. B. Hoy, Alexa, siri, cortana, and more: an introduction to voice assistants, Medical reference services quarterly 37 (1) (2018) 81–88

  3. [2]

    Tristan, S

    S. Tristan, S. Sharma, R. Gonzalez, Alexa/google home forensics, Digital Forensic Education: An Experiential Learning Approach (2020) 101–121. 53

  4. [3]

    A. H. Michaely, X. Zhang, G. Simko, C. Parada, P. Aleksic, Keyword spotting for google assistant using contextual speech recognition, in: 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), IEEE, 2017, pp. 272–278

  5. [4]

    L´ opez-Espejo, Z.-H

    I. L´ opez-Espejo, Z.-H. Tan, J. H. Hansen, J. Jensen, Deep spoken keyword spotting: An overview, IEEE Access 10 (2021) 4169–4199

  6. [5]

    R. C. Rose, D. B. Paul, A hidden markov model based keyword recognition system, in: International Conference on Acoustics, Speech, and Signal Processing, IEEE, 1990, pp. 129–132

  7. [6]

    J. G. Wilpon, L. G. Miller, P. Modi, Improvements and applications for key word recognition using hidden markov modeling techniques, in: [Proceedings] ICASSP 91: 1991 International Conference on Acoustics, Speech, and Signal Processing, IEEE, 1991, pp. 309–312

  8. [7]

    Motlicek, F

    P. Motlicek, F. Valente, I. Szoke, Improving acoustic based keyword spotting using lvcsr lattices, in: 2012 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2012, pp. 4413–4416

Show all 148 references
  1. [9]

    S. S. Saha, S. S. Sandha, M. Srivastava, Machine learning for microcontroller-class hardware-a review, IEEE Sensors Journal (2022)

  2. [10]

    Prabhavalkar, R

    R. Prabhavalkar, R. Alvarez, C. Parada, P. Nakkiran, T. N. Sainath, Automatic gain control and multi-style training for robust small-footprint keyword spotting with deep neural networks, in: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)...

  3. [11]

    B. D. Scott, M. E. Rafn, Suspending noise cancellation using keyword spotting, uS Patent 9,398,367 (Jul. 19 2016)

  4. [12]

    Rybakov, N

    O. Rybakov, N. Kononenko, N. Subrahmanya, M. Visontai, S. Laurenzo, Streaming Keyword Spotting on Mobile Devices, in: Proc. Interspeech 2020, 2020, pp. 2277–2281. doi:10.21437/Interspeech.2020-1003

  5. [13]

    Sandler, A

    M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, L.-C. Chen, Mobilenetv2: Inverted residuals and linear bottlenecks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4510–4520

  6. [14]

    F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, K. Keutzer, Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size, arXiv preprint arXiv:1602.07360 (2016)

  7. [15]

    Warden, D

    P. Warden, D. Situnayake, Tinyml: Machine learning with tensorflow lite on arduino and ultra-low-power microcon- trollers, O’Reilly Media, 2019

  8. [16]

    J. Lin, L. Zhu, W.-M. Chen, W.-C. Wang, S. Han, Tiny machine learning: Progress and futures [feature], IEEE Circuits and Systems Magazine 23 (3) (2023) 8–34

  9. [17]

    Shafique, T

    M. Shafique, T. Theocharides, V. J. Reddy, B. Murmann, Tinyml: Current progress, research challenges, and future roadmap, in: 2021 58th ACM/IEEE Design Automation Conference (DAC), IEEE, 2021, pp. 1303–1306

  10. [18]

    J. S. P. Giraldo, M. Verhelst, Hardware acceleration for embedded keyword spotting: Tutorial and survey, ACM Trans- actions on Embedded Computing Systems (TECS) 20 (6) (2021) 1–25

  11. [19]

    Tabibian, A survey on structured discriminative spoken keyword spotting, Artificial Intelligence Review 53 (4) (2020) 2483–2520

    S. Tabibian, A survey on structured discriminative spoken keyword spotting, Artificial Intelligence Review 53 (4) (2020) 2483–2520

  12. [21]

    K. T. Chitty-Venkata, A. K. Somani, Neural architecture search survey: A hardware perspective, ACM Computing Surveys 55 (4) (2022) 1–36

  13. [22]

    Menghani, Efficient deep learning: A survey on making deep learning models smaller, faster, and better, ACM Computing Surveys 55 (12) (2023) 1–37

    G. Menghani, Efficient deep learning: A survey on making deep learning models smaller, faster, and better, ACM Computing Surveys 55 (12) (2023) 1–37

  14. [23]

    Warden, Launching the speech commands dataset, Google Research Blog (2017)

    P. Warden, Launching the speech commands dataset, Google Research Blog (2017). 54

  15. [24]

    Liang, X

    J. Liang, X. Ban, K. Yu, B. Qu, K. Qiao, C. Yue, K. Chen, K. C. Tan, A survey on evolutionary constrained multiobjective optimization, IEEE Transactions on Evolutionary Computation 27 (2) (2022) 201–221

  16. [25]

    T. N. Sainath, C. Parada, Convolutional neural networks for small-footprint keyword spotting, in: Proc. Interspeech 2015, 2015, pp. 1478–1482. doi:10.21437/Interspeech.2015-352

  17. [26]

    Li, A lightweight architecture for query-by-example keyword spotting on low-power iot devices, IEEE Transactions on Consumer Electronics 69 (1) (2022) 65–75

    M. Li, A lightweight architecture for query-by-example keyword spotting on low-power iot devices, IEEE Transactions on Consumer Electronics 69 (1) (2022) 65–75

  18. [27]

    Gong, Y.-A

    Y. Gong, Y.-A. Chung, J. Glass, Ast: Audio spectrogram transformer, Proc. Interspeech 2021 (2021)

  19. [28]

    A. Berg, M. OConnor, M. T. Cruz, Keyword Transformer: A Self-Attention Model for Keyword Spotting, in: Proc. Interspeech 2021, 2021, pp. 4249–4253. doi:10.21437/Interspeech.2021-1286

  20. [29]

    Samui, S

    S. Samui, S. Garai, Time-frequency domain speech enhancement framework using audio spectrogram transformer with masked multi-head attention, in: 2023 8th International Conference on Computers and Devices for Communication (CODEC), 2023, pp. 1–2. doi:10.1109/CODEC60112.2023.10465846

  21. [30]

    David, J

    R. David, J. Duke, A. Jain, V. Janapa Reddi, N. Jeffries, J. Li, N. Kreeger, I. Nappier, M. Natraj, T. Wang, et al., Tensorflow lite micro: Embedded machine learning for tinyml systems, Proceedings of Machine Learning and Systems 3 (2021) 800–811

  22. [31]

    Hinton, O

    G. Hinton, O. Vinyals, J. Dean, Distilling the knowledge in a neural network, arXiv preprint arXiv:1503.02531 (2015)

  23. [32]

    L. Lei, G. Yuan, H. Yu, D. Kong, Y. He, Multilingual customized keyword spotting using similar-pair contrastive learning, IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023) 2437–2447

  24. [33]

    Chakravarthi, S.-C

    B. Chakravarthi, S.-C. Ng, M. Ezilarasan, M.-F. Leung, Eeg-based emotion recognition using hybrid cnn and lstm classification, Frontiers in computational neuroscience 16 (2022) 1019776

  25. [34]

    Leroy, A

    D. Leroy, A. Coucke, T. Lavril, T. Gisselbrecht, J. Dureau, Federated learning for keyword spotting, in: ICASSP 2019- 2019 IEEE international conference on acoustics, speech and signal processing (ICASSP), IEEE, 2019, pp. 6341–6345

  26. [35]

    van Esch, E

    D. van Esch, E. Sarbar, T. Lucassen, J. O’Brien, T. Breiner, M. Prasad, E. Crew, C. Nguyen, F. Beaufays, Writing across the world’s languages: Deep internationalization for gboard, the google keyboard, arXiv preprint arXiv:1912.01218 (2019)

  27. [36]

    Warden, Speech commands: A dataset for limited-vocabulary speech recognition, arXiv preprint arXiv:1804.03209 (2018)

    P. Warden, Speech commands: A dataset for limited-vocabulary speech recognition, arXiv preprint arXiv:1804.03209 (2018)

  28. [37]

    G. Chen, C. Parada, G. Heigold, Small-footprint keyword spotting using deep neural networks, in: 2014 IEEE interna- tional conference on acoustics, speech and signal processing (ICASSP), IEEE, 2014, pp. 4087–4091

  29. [38]

    M. Sun, A. Raju, G. Tucker, S. Panchapagesan, G. Fu, A. Mandal, S. Matsoukas, N. Strom, S. Vitaladevuni, Max-pooling loss training of long short-term memory networks for small-footprint keyword spotting, in: 2016 IEEE spoken language technology workshop (SLT), IEEE, 2016, pp. 474–480

  30. [39]

    Kumar, V

    R. Kumar, V. Yeruva, S. Ganapathy, On convolutional lstm modeling for joint wake-word detection and text dependent speaker verification., in: Interspeech, 2018, pp. 1121–1125

  31. [40]

    P. M. Sørensen, B. Epp, T. May, A depthwise separable convolutional neural network for keyword spotting on an embedded system, EURASIP Journal on Audio, Speech, and Music Processing 2020 (1) (2020) 1–14

  32. [41]

    Davis, P

    S. Davis, P. Mermelstein, Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences, IEEE transactions on acoustics, speech, and signal processing 28 (4) (1980) 357–366

  33. [42]

    Vitolo, R

    P. Vitolo, R. Liguori, L. Di Benedetto, A. Rubino, G. D. Licciardo, Automatic audio feature extraction for keyword spotting, IEEE Signal Processing Letters (2023)

  34. [43]

    LeCun, Y

    Y. LeCun, Y. Bengio, G. Hinton, Deep learning, nature 521 (7553) (2015) 436–444

  35. [44]

    F. Chen, S. Li, J. Han, F. Ren, Z. Yang, Review of lightweight deep convolutional neural networks., Archives of Compu- tational Methods in Engineering 31 (4) (2024)

  36. [45]

    A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, H. Adam, Mobilenets: Efficient 55 convolutional neural networks for mobile vision applications, arXiv preprint arXiv:1704.04861 (2017)

  37. [46]

    Goodfellow, Y

    I. Goodfellow, Y. Bengio, A. Courville, Deep learning, Vol. 1, 2016

  38. [47]

    Lebedev, Y

    V. Lebedev, Y. Ganin, M. Rakhuba, I. Oseledets, V. Lempitsky, Speeding-up convolutional neural networks using fine- tuned cp-decomposition, in: 3rd International Conference on Learning Representations, ICLR 2015-Conference Track Proceedings, 2015

  39. [48]

    M. M. H. Shuvo, S. K. Islam, J. Cheng, B. I. Morshed, Efficient acceleration of deep learning inference on resource- constrained edge devices: A review, Proceedings of the IEEE 111 (1) (2022) 42–91

  40. [49]

    Sigtia, J

    S. Sigtia, J. Bridle, H. Richards, P. Clark, E. Marchi, V. Garg, Progressive voice trigger detection: Accuracy vs latency, in: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2021, pp. 6843–6847

  41. [50]

    G.-S. Fu, T. Senechal, A. Challenner, T. Zhang, Unified speculation, detection, and verification keyword spotting, in: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2022, pp. 7557–7561

  42. [51]

    J. Wang, M. Xu, J. Hou, B. Zhang, X.-L. Zhang, L. Xie, F. Pan, Wekws: A production first small-footprint end-to- end keyword spotting toolkit, in: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2023, pp. 1–5

  43. [52]

    E. Park, D. Ahn, H. Kim, Reptor: Re-parameterizable temporal convolution for keyword spotting via differentiable kernel search, in: Proc. Interspeech 2024, 2024, pp. 4518–4522

  44. [53]

    X. Ding, X. Zhang, N. Ma, J. Han, G. Ding, J. Sun, Repvgg: Making vgg-style convnets great again, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 13733–13742

  45. [54]

    H. Yang, Z. Yang, L. Wan, B. Zhang, Y. Shi, Y. Huang, I. Enchev, L. Tang, R. Alvarez, M. Sun, et al., Lico-net: Linearized convolution network for hardware-efficient keyword spotting, arXiv preprint arXiv:2211.04635 (2022)

  46. [55]

    Huang, N

    Y. Huang, N. Hou, N. F. Chen, Progressive continual learning for spoken keyword spotting, in: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2022, pp. 7552–7556

  47. [56]

    Snell, K

    J. Snell, K. Swersky, R. Zemel, Prototypical networks for few-shot learning, Advances in neural information processing systems 30 (2017)

  48. [57]

    Rusci, T

    M. Rusci, T. Tuytelaars, On-device customization of tiny deep learning models for keyword spotting with few examples, Ieee Micro (2023)

  49. [58]

    Y. Wang, P. Getreuer, T. Hughes, R. F. Lyon, R. A. Saurous, Trainable frontend for robust and far-field keyword spotting, in: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2017, pp. 5670–5674

  50. [59]

    Samui, I

    S. Samui, I. Chakrabarti, S. K. Ghosh, Time–frequency masking based supervised speech enhancement framework using fuzzy deep belief network, Applied Soft Computing 74 (2019) 583–602

  51. [60]

    I. J. Goodfellow, J. Shlens, C. Szegedy, Explaining and harnessing adversarial examples, arXiv preprint arXiv:1412.6572 (2014)

  52. [61]

    Zhang, J

    Z. Zhang, J. Geiger, J. Pohjalainen, A. E.-D. Mousa, W. Jin, B. Schuller, Deep learning for environmentally robust speech recognition: An overview of recent developments, ACM Transactions on Intelligent Systems and Technology (TIST) 9 (5) (2018) 1–28

  53. [62]

    J. Du, X. Na, X. Liu, H. Bu, Aishell-2: Transforming mandarin asr research into industrial scale, arXiv preprint arXiv:1808.10583 (2018)

  54. [63]

    Mishchenko, Y

    Y. Mishchenko, Y. Goren, M. Sun, C. Beauchene, S. Matsoukas, O. Rybakov, S. N. P. Vitaladevuni, Low-bit quantization and quantization-aware training for small-footprint keyword spotting, in: 2019 18th IEEE International Conference On Machine Learning And Applications (ICMLA), ...

  55. [64]

    Higuchi, M

    T. Higuchi, M. Ghasemzadeh, K. You, C. Dhir, Stacked 1d convolutional networks for end-to-end small footprint voice trigger detection, Proc. Interspeech 2020 (2020)

  56. [65]

    B. Kim, M. Lee, J. Lee, Y. Kim, K. Hwang, Query-by-example on-device keyword spotting, in: 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), IEEE, 2019, pp. 532–538

  57. [66]

    J. Hou, Y. Shi, M. Ostendorf, M. Hwang, L. Xie, Region proposal network based small-footprint keyword spotting, IEEE Signal Process. Lett. 26 (10) (2019) 1471–1475. URL https://doi.org/10.1109/LSP.2019.2936282

  58. [67]

    Mazumder, S

    M. Mazumder, S. Chitlangia, C. Banbury, Y. Kang, J. M. Ciro, K. Achorn, D. Galvez, M. Sabini, P. Mattson, D. Kanter, et al., Multilingual spoken words corpus, in: Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021

  59. [68]

    X. Qin, H. Bu, M. Li, Hi-mia : A far-field text-dependent speaker verification database and the baselines (2019). arXiv:1912.01231

  60. [69]

    Ghandoura, F

    A. Ghandoura, F. Hjabo, O. Al Dakkak, Building and benchmarking an arabic speech commands dataset for small- footprint keyword spotting, Engineering Applications of Artificial Intelligence 102 (2021) 104267. doi:https://doi.org/ 10.1016/j.engappai.2021.104267. URL https://www....

  61. [70]

    S. . Ark, M. Kliegl, R. Child, J. Hestness, A. Gibiansky, C. Fougner, R. Prenger, A. Coates, Convolutional Recurrent Neural Networks for Small-Footprint Keyword Spotting, in: Proc. Interspeech 2017, 2017, pp. 1606–1610. doi:10.21437/ Interspeech.2017-1737

  62. [71]

    R. Tang, J. Lin, Deep residual learning for small-footprint keyword spotting, in: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2018, pp. 5484–5488

  63. [72]

    S. Choi, S. Seo, B. Shin, H. Byun, M. Kersner, B. Kim, D. Kim, S. Ha, Temporal convolution for real-time keyword spotting on mobile devices, Proc. INTERSPEECH 2019 (2019)

  64. [73]

    X. Chen, S. Yin, D. Song, P. Ouyang, L. Liu, S. Wei, Small-footprint keyword spotting with graph convolutional network, in: 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), IEEE, 2019, pp. 539–546

  65. [74]

    X. Li, X. Wei, X. Qin, Small-Footprint Keyword Spotting with Multi-Scale Temporal Convolution, in: Proc. Interspeech 2020, 2020, pp. 1987–1991. doi:10.21437/Interspeech.2020-3177

  66. [75]

    Majumdar, B

    S. Majumdar, B. Ginsburg, MatchboxNet: 1D Time-Channel Separable Convolutional Neural Network Architecture for Speech Commands Recognition, in: Proc. Interspeech 2020, 2020, pp. 3356–3360. doi:10.21437/Interspeech.2020-1058

  67. [76]

    B. Kim, S. Chang, J. Lee, D. Sung, Broadcasted residual learning for efficient keyword spotting, Proceedings of INTER- SPEECH 2021 (2021)

  68. [77]

    Chaudhary, V

    A. Chaudhary, V. Abrol, Towards on-device keyword spotting using low-footprint quaternion neural models, in: 2023 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (W ASPAA), IEEE, 2023, pp. 1–5

  69. [78]

    Akhtar, M

    Z. Akhtar, M. O. Khursheed, D. Du, Y. Liu, Small-footprint slimmable networks for keyword spotting, in: IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5

  70. [79]

    Tucker, M

    G. Tucker, M. Wu, M. Sun, S. Panchapagesan, G. Fu, S. Vitaladevuni, Model Compression Applied to Small-Footprint Keyword Spotting, in: Proc. Interspeech 2016, 2016, pp. 1878–1882. doi:10.21437/Interspeech.2016-1393

  71. [80]

    C. Gao, Y. Gu, F. Caliva, Y. Liu, Self-supervised speech representation learning for keyword-spotting with light- weight transformers, in: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2023, pp. 1–5

  72. [81]

    G.-P. Yang, Y. Gu, Q. Tang, D. Du, Y. Liu, On-device constrained self-supervised speech representation learning for keyword spotting via knowledge distillation, in: INTERSPEECH, 2023

  73. [82]

    Macha, O

    S. Macha, O. Oza, A. Escott, F. Caliva, R. Armitano, S. K. Cheekatmalla, S. H. K. Parthasarathi, Y. Liu, Fixed-point 57 quantization aware training for on-device keyword-spotting, in: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (IC...

  74. [83]

    M. Sun, D. Snyder, Y. Gao, V. Nagaraja, M. Rodehorst, S. Panchapagesan, N. Strom, S. Matsoukas, S. Vitaladevuni, Compressed Time Delay Neural Network for Small-Footprint Keyword Spotting, in: Proc. Interspeech 2017, 2017, pp. 3607–3611. doi:10.21437/Interspeech.2017-480

  75. [84]

    M. Luo, D. Wang, X. Wang, S. Qiao, Y. Zhou, Error-diffusion based speech feature quantization for small-footprint keyword spotting, IEEE Signal Processing Letters 29 (2022) 1357–1361

  76. [85]

    G.-P. Yang, Y. Gu, S. Macha, Q. Tang, Y. Liu, On-device constrained self-supervised learning for keyword spotting via quantization aware pre-training and fine-tuning, in: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, ...

  77. [86]

    C. Shan, J. Zhang, Y. Wang, L. Xie, Attention-based End-to-End Models for Small-Footprint Keyword Spotting, in: Proc. Interspeech 2018, 2018, pp. 2037–2041. doi:10.21437/Interspeech.2018-1777

  78. [87]

    Y. Bai, J. Yi, J. Tao, Z. Wen, Z. Tian, C. Zhao, C. Fan, A time delay neural network with shared weight self-attention for small-footprint keyword spotting., in: INTERSPEECH, 2019, pp. 2190–2194

  79. [88]

    E. A. Ibrahim, J. Huisken, H. Fatemi, J. P. de Gyvez, Keyword spotting using time-domain features in a temporal convolutional network, in: 2019 22nd Euromicro Conference on Digital System Design (DSD), IEEE, 2019, pp. 313–319

  80. [89]

    Mittermaier, L

    S. Mittermaier, L. K¨ urzinger, B. Waschneck, G. Rigoll, Small-footprint keyword spotting on raw audio data with sinc-convolutions, in: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2020, pp. 7454–7458

  81. [90]

    Riviello, J.-P

    A. Riviello, J.-P. David, Binary speech features for keyword spotting tasks., in: INTERSPEECH, 2019, pp. 3460–3464

  82. [91]

    Anderson, J

    A. Anderson, J. Su, R. Dahyot, D. Gregg, Performance-oriented neural architecture search, in: 2019 International Conference on High Performance Computing & Simulation (HPCS), IEEE, 2019, pp. 177–184

  83. [92]

    V´ eniat, O

    T. V´ eniat, O. Schwander, L. Denoyer, Stochastic adaptive neural architecture search for keyword spotting, in: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2019, pp. 2842– 2846

  84. [93]

    T. Mo, Y. Yu, M. Salameh, D. Niu, S. Jui, Neural Architecture Search for Keyword Spotting, in: Proc. Interspeech 2020, 2020, pp. 1982–1986. doi:10.21437/Interspeech.2020-3132

  85. [94]

    Zhang, W

    B. Zhang, W. Li, Q. Li, W. Zhuang, X. Chu, Y. Wang, Autokws: Keyword spotting with differentiable architecture search, in: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2021, pp. 2830–2834

  86. [95]

    Banbury, C

    C. Banbury, C. Zhou, I. Fedorov, R. Matas, U. Thakker, D. Gope, V. Janapa Reddi, M. Mattina, P. Whatmough, Micronets: Neural network architectures for deploying tinyml applications on commodity microcontrollers, Proceedings of Machine Learning and Systems 3 (2021) 517–532

  87. [96]

    Busia, G

    P. Busia, G. Deriu, L. Rinelli, C. Chesta, L. Raffo, P. Meloni, Target-aware neural architecture search and deployment for keyword spotting, IEEE Access 10 (2022) 40687–40700

  88. [97]

    Zhang, N

    Y. Zhang, N. Suda, L. Lai, V. Chandra, Hello edge: Keyword spotting on microcontrollers, arXiv preprint arXiv:1711.07128 (2017)

  89. [98]

    Peter, W

    D. Peter, W. Roth, F. Pernkopf, End-to-end keyword spotting using neural architecture search and quantization, in: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2022, pp. 3423–3427

  90. [99]

    L. T´ oth, Combining time-and frequency-domain convolution in convolutional neural network-based phone recognition, in: 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2014, pp. 190–194

  91. [100]

    Abdel-Hamid, A.-r

    O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, G. Penn, Applying convolutional neural networks concepts to hybrid nn- 58 hmm model for speech recognition, in: 2012 IEEE international conference on Acoustics, speech and signal processing (ICASSP), IEEE, 2012, pp. 4277–4280

  92. [101]

    Z. Song, Q. Liu, Q. Yang, Y. Peng, H. Li, Ed-skws: Early-decision spiking neural networks for rapid, and energy-efficient keyword spotting, in: Proc. Interspeech 2024, 2024, pp. 4528–4532

  93. [102]

    S. Wang, D. Zhang, K. Shi, Y. Wang, W. Wei, J. Wu, M. Zhang, Global-local convolution with spiking neural networks for energy-efficient keyword spotting, in: Proc. Interspeech 2024, 2024, pp. 4523–4527

  94. [103]

    Coucke, M

    A. Coucke, M. Chlieh, T. Gisselbrecht, D. Leroy, M. Poumeyrol, T. Lavril, Efficient keyword spotting using dilated con- volutions and gating, in: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2019, pp. 6351–6355

  95. [104]

    Baevski, Y

    A. Baevski, Y. Zhou, A. Mohamed, M. Auli, wav2vec 2.0: A framework for self-supervised learning of speech represen- tations, Advances in neural information processing systems 33 (2020) 12449–12460

  96. [105]

    Sze, Y.-H

    V. Sze, Y.-H. Chen, T.-J. Yang, J. S. Emer, Efficient processing of deep neural networks: A tutorial and survey, Proceedings of the IEEE 105 (12) (2017) 2295–2329

  97. [106]

    T.-J. Y. J. S. E. Vivienne Sze, Yu-Hsin Chen, Efficient Processing of Deep Neural Networks, Synthesis Lectures on Computer Architecture, Springer Cham, 2020. URL https://doi.org/10.1007/978-3-031-01766-7

  98. [107]

    LeCun, J

    Y. LeCun, J. Denker, S. Solla, Optimal brain damage, Advances in neural information processing systems 2 (1989)

  99. [108]

    Hassibi, D

    B. Hassibi, D. G. Stork, G. J. Wolff, Optimal brain surgeon and general network pruning, in: IEEE international conference on neural networks, IEEE, 1993, pp. 293–299

  100. [109]

    Jacob, S

    B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, D. Kalenichenko, Quantization and training of neural networks for efficient integer-arithmetic-only inference, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 2704–2713

  101. [110]

    Nahshan, B

    Y. Nahshan, B. Chmiel, C. Baskin, E. Zheltonozhskii, R. Banner, A. M. Bronstein, A. Mendelson, Loss aware post- training quantization, Machine Learning 110 (11) (2021) 3245–3262

  102. [111]

    Gholami, S

    A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, K. Keutzer, A survey of quantization methods for efficient neural network inference, in: Low-power computer vision, Chapman and Hall/CRC, 2022, pp. 291–326

  103. [112]

    Ostromoukhov, A simple and efficient error-diffusion algorithm, in: Proceedings of the 28th annual conference on Computer graphics and interactive techniques, 2001, pp

    V. Ostromoukhov, A simple and efficient error-diffusion algorithm, in: Proceedings of the 28th annual conference on Computer graphics and interactive techniques, 2001, pp. 567–572

  104. [113]

    K. Ding, M. Zong, J. Li, B. Li, Letr: A lightweight and efficient transformer for keyword spotting, in: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2022, pp. 7987–7991

  105. [114]

    Z. Niu, G. Zhong, H. Yu, A review on the attention mechanism of deep learning, Neurocomputing 452 (2021) 48–62

  106. [115]

    Lpez-Espejo, Z.-H

    I. Lpez-Espejo, Z.-H. Tan, J. Jensen, An Experimental Study on Light Speech Features for Small-Footprint Keyword Spotting , in: Proc. IberSPEECH 2022, 2022, pp. 131–135. doi:10.21437/IberSPEECH.2022-27

  107. [116]

    Benmeziane, K

    H. Benmeziane, K. El Maghraoui, H. Ouarnoughi, S. Niar, M. Wistuba, N. Wang, A comprehensive survey on hardware- aware neural architecture search, Ph.D. thesis, LAMIH, Universit´ e Polytechnique des Hauts-de-France (2021)

  108. [117]

    M. Tan, B. Chen, R. Pang, V. Vasudevan, M. Sandler, A. Howard, Q. V. Le, Mnasnet: Platform-aware neural architecture search for mobile, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 2820–2828

  109. [118]

    Ravanelli, Y

    M. Ravanelli, Y. Bengio, Speaker recognition from raw waveform with sincnet, in: 2018 IEEE spoken language technology workshop (SLT), IEEE, 2018, pp. 1021–1028

  110. [119]

    Shrivastava, A

    A. Shrivastava, A. Kundu, C. Dhir, D. Naik, O. Tuzel, Optimize what matters: Training dnn-hmm keyword spotting model using end metric, in: ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 4000–4004. doi:10.1109/ICA...

  111. [120]

    N. L. Gimnez, F. Freitag, J. Lee, H. Vandierendonck, Comparison of two microcontroller boards for on-device model training in a keyword spotting task, in: 2022 11th Mediterranean Conference on Embedded Computing (MECO), 2022, pp. 1–4. doi:10.1109/MECO55406.2022.9797171

  112. [121]

    H. Ren, D. Anicic, T. A. Runkler, Tinyol: Tinyml with online-learning on microcontrollers, in: 2021 International Joint Conference on Neural Networks (IJCNN), IEEE, 2021, pp. 1–8

  113. [122]

    J. Lin, L. Zhu, W.-M. Chen, W.-C. Wang, C. Gan, S. Han, On-device training under 256kb memory, Advances in Neural Information Processing Systems 35 (2022) 22941–22954

  114. [123]

    Embedded learning library (ell), https://microsoft.github.io/ELL/, accessed: 12 January 2024

  115. [124]

    ARM-NN, https://github.com/ARM-software/armnn, accessed: 12 January 2024

  116. [125]

    CMSIS-NN, https://arm-software.github.io/CMSIS_5/NN/html, accessed: 12 January 2024

  117. [126]

    STM32Cube.AI, https://www.st.com/en/embedded-software/x-cube-ai.html , accessed: 12 January 2024

  118. [127]

    AIfES, https://github.com/Fraunhofer-IMS/AIfES_for_Arduino, accessed: 12 January 2024

  119. [128]

    uTensor, https://github.com/uTensor/uTensor, accessed: 12 January 2024

  120. [129]

    TinyMLgen, https://github.com/eloquentarduino/tinymlgen, accessed: 12 January 2024

  121. [130]

    Cmix-nn, https://github.com/EEESlab/CMix-NN, accessed: 12 January 2024

  122. [131]

    Edge Impulse, https://edgeimpulse.com/, accessed: 12 January 2024

  123. [132]

    Miettinen, Nonlinear multiobjective optimization, Vol

    K. Miettinen, Nonlinear multiobjective optimization, Vol. 12, Springer Science & Business Media, 1999

  124. [133]

    K. Deb, Multi-objective optimisation using evolutionary algorithms: an introduction, in: Multi-objective evolutionary optimisation for product design and manufacturing, Springer, 2011, pp. 3–34

  125. [134]

    R. L. Rardin, R. Uzsoy, Experimental evaluation of heuristic optimization algorithms: A tutorial, Journal of Heuristics 7 (2001) 261–304

  126. [135]

    Bandyopadhyay, S

    S. Bandyopadhyay, S. Saha, U. Maulik, K. Deb, A simulated annealing-based multiobjective optimization algorithm: Amosa, IEEE transactions on evolutionary computation 12 (3) (2008) 269–283

  127. [136]

    G¨ ulc¨ u, Z

    A. G¨ ulc¨ u, Z. Ku¸ s, Multi-objective simulated annealing for hyper-parameter optimization in convolutional neural networks, PeerJ Computer Science 7 (2021) e338

  128. [137]

    K. Deb, A. Pratap, S. Agarwal, T. Meyarivan, A fast and elitist multiobjective genetic algorithm: Nsga-ii, IEEE trans- actions on evolutionary computation 6 (2) (2002) 182–197

  129. [138]

    A. A. Shaikh, A. K. Mukhopadhyay, S. Poddar, S. Samui, Toward robust and accurate myoelectric controller design based on multiobjective optimization using evolutionary computation, IEEE Sensors Journal 24 (5) (2024) 6418–6429. doi:10.1109/JSEN.2023.3347949

  130. [139]

    Parsa, J

    M. Parsa, J. P. Mitchell, C. D. Schuman, R. M. Patton, T. E. Potok, K. Roy, Bayesian multi-objective hyperparameter optimization for accurate, fast, and efficient neural network accelerator design, Frontiers in neuroscience 14 (2020) 667

  131. [140]

    Alibrahim, S

    H. Alibrahim, S. A. Ludwig, Hyperparameter optimization: Comparing genetic algorithm against grid search and bayesian optimization, in: 2021 IEEE congress on evolutionary computation (CEC), IEEE, 2021, pp. 1551–1559

  132. [141]

    Jin, Multi-objective machine learning, Vol

    Y. Jin, Multi-objective machine learning, Vol. 16, Springer Science & Business Media, 2007

  133. [142]

    Liberis, L

    E. Liberis, L. Dudziak, N. D. Lane, µnas: Constrained neural architecture search for microcontrollers, in: Proceedings of the 1st Workshop on Machine Learning and Systems, 2021, pp. 70–79

  134. [143]

    L. Ma, N. Li, G. Yu, X. Geng, S. Cheng, X. Wang, M. Huang, Y. Jin, Pareto-wise ranking classifier for multi-objective evolutionary neural architecture search, IEEE Transactions on Evolutionary Computation (2023)

  135. [144]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, Advances in neural information processing systems 30 (2017)

  136. [145]

    Gong, C.-I

    Y. Gong, C.-I. Lai, Y.-A. Chung, J. Glass, Ssast: Self-supervised audio spectrogram transformer, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, 2022, pp. 10699–10709. 60

  137. [146]

    Y. Sun, X. Wang, Z. Liu, J. Miller, A. Efros, M. Hardt, Test-time training with self-supervision for generalization under distribution shifts, in: International conference on machine learning, PMLR, 2020, pp. 9229–9248

  138. [147]

    Samui, I

    S. Samui, I. Chakrabarti, S. K. Ghosh, Tensor-train long short-term memory for monaural speech enhancement, arXiv preprint arXiv:1812.10095 (2018)

  139. [148]

    Fedorov, R

    I. Fedorov, R. P. Adams, M. Mattina, P. Whatmough, Sparse: Sparse architecture search for cnns on resource-constrained microcontrollers, Advances in Neural Information Processing Systems 32 (2019)

  140. [149]

    T. Wang, K. Wang, H. Cai, J. Lin, Z. Liu, H. Wang, Y. Lin, S. Han, Apq: Joint search for network architecture, pruning and quantization policy, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 2078–2087. 61

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.