Pith. sign in

REVIEW

Encoder-Decoder Neural Architecture Optimization for Keyword Spotting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.02738 v1 pith:VCTS2Q7C submitted 2021-06-04 cs.LG cs.MM

classification cs.LGcs.MM
keywords keywordneuralarchitecturespottingmodelsearchconvolutionalencoder-decoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Keyword spotting aims to identify specific keyword audio utterances. In recent years, deep convolutional neural networks have been widely utilized in keyword spotting systems. However, their model architectures are mainly based on off-the shelfbackbones such as VGG-Net or ResNet, instead of specially designed for the task. In this paper, we utilize neural architecture search to design convolutional neural network models that can boost the performance of keyword spotting while maintaining an acceptable memory footprint. Specifically, we search the model operators and their connections in a specific search space with Encoder-Decoder neural architecture optimization. Extensive evaluations on Google's Speech Commands Dataset show that the model architecture searched by our approach achieves a state-of-the-art accuracy of over 97%.

Discussion (0). Sign in to comment.

Pith tools