REVIEW 3 cited by
Enhanced Sound Event Localization and Detection in Real 360-degree audio-visual soundscapes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This technical report details our work towards building an enhanced audio-visual sound event localization and detection (SELD) network. We build on top of the audio-only SELDnet23 model and adapt it to be audio-visual by merging both audio and video information prior to the gated recurrent unit (GRU) of the audio-only network. Our model leverages YOLO and DETIC object detectors. We also build a framework that implements audio-visual data augmentation and audio-visual synthetic data generation. We deliver an audio-visual SELDnet system that outperforms the existing audio-visual SELD baseline.
Forward citations
Cited by 3 Pith papers
-
Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
A Cross-Modal Conformer that fuses CLAP audio and OWL-ViT visual embeddings with a CNN-Conformer SELD backbone, trained on large synthetic data, ranks second in DCASE 2025 Task 3 Track B.
-
Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos
Fusing frozen CLAP and OWL-ViT embeddings via a Cross-Modal Conformer, plus autocorrelation-based features, improves stereo SELD over DCASE 2025 baselines.
-
MVANet: Multi-Stage Video Attention Network for Sound Event Localization and Detection with Source Distance Estimation
MVANet applies multi-stage audio-guided video attention to improve audio-visual 3D sound event localization and source distance estimation on STARSS23.
Discussion (0). Continue with ORCID to comment.