Temporal attention pooling, combining time attention, velocity attention, and average pooling, improves frequency dynamic convolution for sound event detection by about 3% average PSDS1 and reaches a maximum PSDS1 of 0.459 when combined with multi-dilated frequency dynamic convolution.
Semi-supervised learning-based sound event detection using frequency dynamic convolution with large kernel attention for DCASE challenge 2023 task 4,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection
Temporal attention pooling, combining time attention, velocity attention, and average pooling, improves frequency dynamic convolution for sound event detection by about 3% average PSDS1 and reaches a maximum PSDS1 of 0.459 when combined with multi-dilated frequency dynamic convolution.