REVIEW 3 cited by
Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the DCASE 2025 Challenge has recently introduced a task for spatial semantic segmentation of sound scenes (S5), which focuses on detecting and separating sound events in spatial sound scenes. In this paper, we explore methods for addressing the S5 task. Specifically, we present baseline S5 systems that combine audio tagging (AT) and label-queried source separation (LSS) models. We investigate two LSS approaches based on the ResUNet architecture: a) extracting a single source for each detected event and b) querying multiple sources concurrently. Since each separated source in S5 is identified by its sound event class label, we propose new class-aware metrics to evaluate both the sound sources and labels simultaneously. Experimental results on first-order ambisonics spatial audio demonstrate the effectiveness of the proposed systems and confirm the efficacy of the metrics.
Forward citations
Cited by 3 Pith papers
-
On Temporal Guidance and Iterative Refinement in Audio Source Separation
A DCASE 2025 Task 4 system that guides audio source separation with frame-level event detection and iterative refinement, reaching second place.
-
Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4
A combination of extra audio features, curated training data, and agent-based label correction raises CA-SDRi from 11.088 dB to 12.726 dB on DCASE 2025 Task 4.
-
Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
A new shared benchmark task and dataset for spatial semantic segmentation of sound scenes, with baseline ResUNet and ResUNetK results.
Discussion (0). Continue with ORCID to comment.