Pith. sign in

REVIEW 1 cited by

Permutation Invariant Recurrent Neural Networks for Sound Source Tracking Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.08510 v1 pith:ZJKZ5SPF submitted 2023-06-14 eess.AS cs.LGcs.SDeess.SP

Permutation Invariant Recurrent Neural Networks for Sound Source Tracking Applications

classification eess.AS cs.LGcs.SDeess.SP
keywords recurrentinputnetworksneuralstatetrackingvectorinformation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Many multi-source localization and tracking models based on neural networks use one or several recurrent layers at their final stages to track the movement of the sources. Conventional recurrent neural networks (RNNs), such as the long short-term memories (LSTMs) or the gated recurrent units (GRUs), take a vector as their input and use another vector to store their state. However, this approach results in the information from all the sources being contained in a single ordered vector, which is not optimal for permutation-invariant problems such as multi-source tracking. In this paper, we present a new recurrent architecture that uses unordered sets to represent both its input and its state and that is invariant to the permutations of the input set and equivariant to the permutations of the state set. Hence, the information of every sound source is represented in an individual embedding and the new estimates are assigned to the tracked trajectories regardless of their order.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection

    cs.SD 2026-06 unverdicted novelty 4.0

    AT2SELD extends pretrained audio tagging backbones to SELD via FOA descriptors, track-wise processing, permutation-aware supervision, and staged NAS on multiple datasets.