Pith. sign in

REVIEW 1 cited by

IPDnet: A Universal Direct-Path IPD Estimation Network for Sound Source Localization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.07021 v1 pith:RVMEASOO submitted 2024-05-11 eess.AS cs.SD

classification eess.AScs.SD
keywords dp-ipdlocalizationmicrophoneproposedsoundnetworksourcearray
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Extracting direct-path spatial feature is crucial for sound source localization in adverse acoustic environments. This paper proposes the IPDnet, a neural network that estimates direct-path inter-channel phase difference (DP-IPD) of sound sources from microphone array signals. The estimated DP-IPD can be easily translated to source location based on the known microphone array geometry. First, a full-band and narrow-band fusion network is proposed for DP-IPD estimation, in which alternating narrow-band and full-band layers are responsible for estimating the rough DP-IPD information in one frequency band and capturing the frequency correlations of DP-IPD, respectively. Second, a new multi-track DP-IPD learning target is proposed for the localization of flexible number of sound sources. Third, the IPDnet is extend to handling variable microphone arrays, once trained which is able to process arbitrary microphone arrays with different number of channels and array topology. Experiments of multiple-moving-speaker localization are conducted on both simulated and real-world data, which show that the proposed full-band and narrow-band fusion network and the proposed multi-track DP-IPD learning target together achieves excellent sound source localization performance. Moreover, the proposed variable-array model generalizes well to unseen microphone arrays.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Two-Step Learning Framework for Enhancing Sound Event Localization and Detection

    cs.SD 2025-07 conditional novelty 4.0 of 10

    A two-step SELD framework with separate DoA and SED training, trackwise label reordering, and beamformed feature fusion achieves a 0.3891 SELD score on the 2023 DCASE Task 3 development test set.

Pith tools