Pith. sign in

REVIEW 3 cited by

LEAF: A Learnable Frontend for Audio Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.08596 v1 pith:44F2GNAT submitted 2021-01-21 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords audiolearnablefrontendmel-filterbanksclassificationfeaturesoutperformssystem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mel-filterbanks are fixed, engineered audio features which emulate human perception and have been used through the history of audio understanding up to today. However, their undeniable qualities are counterbalanced by the fundamental limitations of handmade representations. In this work we show that we can train a single learnable frontend that outperforms mel-filterbanks on a wide range of audio signals, including speech, music, audio events and animal sounds, providing a general-purpose learned frontend for audio classification. To do so, we introduce a new principled, lightweight, fully learnable architecture that can be used as a drop-in replacement of mel-filterbanks. Our system learns all operations of audio features extraction, from filtering to pooling, compression and normalization, and can be integrated into any neural network at a negligible parameter cost. We perform multi-task training on eight diverse audio classification tasks, and show consistent improvements of our model over mel-filterbanks and previous learnable alternatives. Moreover, our system outperforms the current state-of-the-art learnable frontend on Audioset, with orders of magnitude fewer parameters.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Structural Bottlenecks on Frequency Representation in End-to-End Audio Models

    cs.SD 2026-07 conditional novelty 7.0 of 10

    State-of-the-art strided audio encoders impose predictable alias-collapse and resolution bottlenecks on frequency primitives; Gabor Latent Refactorization recovers much of the lost separability post-hoc.

  2. Semantic Sampling via Learnable Observation Front Ends

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Learnable acoustic filterbanks plus constrained mixing and temporal readout produce more informative low-rate observations for speech reconstruction than fixed waveform sampling at the same budget.

  3. Acoustic Classification of Maritime Vessels using Learnable Filterbanks

    cs.SD 2025-05 conditional novelty 6.0 of 10

    CATFISH, a learnable Gabor filterbank model with attention pooling and optional environmental data, achieves 96.63% test accuracy on the multi-scenario VTUAD underwater vessel classification benchmark.

Pith tools