Pith. sign in

REVIEW 2 cited by

Where's That Voice Coming? Continual Learning for Sound Source Localization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.03661 v3 pith:TWTGY62M submitted 2024-07-04 eess.AS cs.SD

classification eess.AScs.SD
keywords cl-ssllearningacousticacrossapplicationscontinualdataenvironments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sound source localization (SSL) is essential for many speech-processing applications. Deep learning models have achieved high performance, but often fail when the training and inference environments differ. Adapting SSL models to dynamic acoustic conditions faces a major challenge: catastrophic forgetting. In this work, we propose an exemplar-free continual learning strategy for SSL (CL-SSL) to address such a forgetting phenomenon. CL-SSL applies task-specific sub-networks to adapt across diverse acoustic environments while retaining previously learned knowledge. It also uses a scaling mechanism to limit parameter growth, ensuring consistent performance across incremental tasks. We evaluated CL-SSL on simulated data with varying microphone distances and real-world data with different noise levels. The results demonstrate CL-SSL's ability to maintain high accuracy with minimal parameter increase, offering an efficient solution for SSL applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Class-Incremental Learning for Sound Event Localization and Detection

    eess.AS 2024-11 conditional novelty 5.0 of 10

    An incremental learning method with MSE distillation lets a SELD model add four new sound classes after eight while roughly matching the performance of a model trained on all twelve at once.

  2. Enhancing Stereo Sound Event Detection with BiMamba and Pretrained PSELDnet

    eess.AS 2025-07 conditional novelty 4.0 of 10

    Replacing the Conformer decoder in pretrained PSELDnet with a bidirectional Mamba plus asymmetric convolution reports 39.6% versus 38.2% stereo SELD F20 on the DCASE2025 development set, using 76M versus 210M parameters.

Pith tools