Pith. sign in

REVIEW 7 cited by

Can Masked Autoencoders Also Listen to Birds?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.12880 v4 pith:P23XUSCG submitted 2025-04-17 cs.LG cs.SDeess.AS

classification cs.LGcs.SDeess.AS
keywords birdsetaudiobird-maeclassificationfine-tuningprobingprototypicalautoencoders
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Masked Autoencoders (MAEs) learn rich semantic representations in audio classification through an efficient self-supervised reconstruction task. However, general-purpose models fail to generalize well when applied directly to fine-grained audio domains. Specifically, bird-sound classification requires distinguishing subtle inter-species differences and managing high intra-species acoustic variability, revealing the performance limitations of general-domain Audio-MAEs. This work demonstrates that bridging this domain gap domain gap requires full-pipeline adaptation, not just domain-specific pretraining data. We systematically revisit and adapt the pretraining recipe, fine-tuning methods, and frozen feature utilization to bird sounds using BirdSet, a large-scale bioacoustic dataset comparable to AudioSet. Our resulting Bird-MAE achieves new state-of-the-art results in BirdSet's multi-label classification benchmark. Additionally, we introduce the parameter-efficient prototypical probing, enhancing the utility of frozen MAE representations and closely approaching fine-tuning performance in low-resource settings. Bird-MAE's prototypical probes outperform linear probing by up to 37 percentage points in mean average precision and narrow the gap to fine-tuning across BirdSet downstream tasks. Bird-MAE also demonstrates robust few-shot capabilities with prototypical probing in our newly established few-shot benchmark on BirdSet, highlighting the potential of tailored self-supervised learning pipelines for fine-grained audio domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MetaPerch: Learning from metadata for bioacoustics foundation models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Adding location, season, and background-species prediction as auxiliary training tasks improves bioacoustic species identification transfer across acoustic, species, and geographic domain shifts, with modest average g...

  2. A Self-Supervised Approach for Minimal-Annotation Hydroacoustic Data Exploration

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Event-level MAE embeddings plus UMAP/HDBSCAN or K-Means clustering recover 15 hydroacoustic classes from multi-year Mayotte data with ~1 hour of annotation and detector-comparable F1.

  3. Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

    cs.SD 2025-09 conditional novelty 6.0 of 10

    A binarized prototypical probing method that pools per-class evidence from patch tokens substantially outperforms [cls]-token and attentive probes on multi-label audio classification benchmarks.

  4. Adversarial Training Improves Generalization Under Distribution Shifts in Bioacoustics

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Output-space adversarial training improved clean-data performance and adversarial robustness of two bird sound classifiers across seven soundscape test sets, and stabilized prototype-based explanations.

  5. Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026

    cs.SD 2026-07 conditional novelty 5.0 of 10

    For BirdCLEF+ 2026, a frozen Perch-v2 probe plus a trained HGNetV2-B0 SED net and non-bird prototype heads reach private LB 0.936, while WavTokenizer codec tokens collapse and four general audio transformers lag under...

  6. Deformation Driven Suction Cups: A Mechanics-Based Approach to Wearable Electronics

    physics.med-ph 2025-08 unverdicted novelty 5.0 of 10

    Suction adhesion on soft skin depends on cup geometry relative to substrate compliance: wide flat cups lose suction on skin, narrow tall domes retain volume and stick better.

  7. Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types

    cs.SD 2026-07 conditional novelty 4.5 of 10

    On WiWa, factorised multi-task bird species and call-type classification with adaptive loss balancing improves call-type recognition most consistently, while preferred weighting and adaptation depth depend on backbone...

Pith tools