Pith. sign in

REVIEW 2 cited by

E-PANNs: Sound Recognition Using Efficient Pre-trained Audio Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.18665 v1 pith:VDBCISOS submitted 2023-05-30 cs.SD cs.AIeess.ASeess.SP

classification cs.SDcs.AIeess.ASeess.SP
keywords pannsmodelsoundaudioneuralbeene-pannsnetworks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sounds carry an abundance of information about activities and events in our everyday environment, such as traffic noise, road works, music, or people talking. Recent machine learning methods, such as convolutional neural networks (CNNs), have been shown to be able to automatically recognize sound activities, a task known as audio tagging. One such method, pre-trained audio neural networks (PANNs), provides a neural network which has been pre-trained on over 500 sound classes from the publicly available AudioSet dataset, and can be used as a baseline or starting point for other tasks. However, the existing PANNs model has a high computational complexity and large storage requirement. This could limit the potential for deploying PANNs on resource-constrained devices, such as on-the-edge sound sensors, and could lead to high energy consumption if many such devices were deployed. In this paper, we reduce the computational complexity and memory requirement of the PANNs model by taking a pruning approach to eliminate redundant parameters from the PANNs model. The resulting Efficient PANNs (E-PANNs) model, which requires 36\% less computations and 70\% less memory, also slightly improves the sound recognition (audio tagging) performance. The code for the E-PANNs model has been released under an open source license.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Real-Time Emergency Vehicle Siren Detection with Efficient CNNs on Embedded Hardware

    cs.SD 2025-07 conditional novelty 4.0 of 10

    A fine-tuned CNN (E2PANNs) deployed on a Raspberry Pi 5 detects emergency vehicle sirens in real time with adaptive frame sizing and post-processing, reporting up to 78% framewise F1 on a corrected AudioSet-Strong subset.

  2. From Large-scale Audio Tagging to Real-Time Explainable Emergency Vehicle Sirens Detection

    cs.SD 2025-06 conditional novelty 4.0 of 10

    By fine-tuning a pruned PANNs CNN on a newly curated AudioSet subset, the authors build E2PANNs, a real-time emergency vehicle siren detector that runs on a Raspberry Pi 5 and is claimed to be state of the art.

Pith tools