Pith. sign in

REVIEW 5 cited by

NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.08887 v1 pith:MCPGLY4A submitted 2024-01-16 cs.SD cs.AIcs.CLeess.AS

classification cs.SDcs.AIcs.CLeess.AS
keywords challengedatasetsdistantnotsofar-1tasksacousticbaselinebenchmarking
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce the first Natural Office Talkers in Settings of Far-field Audio Recordings (``NOTSOFAR-1'') Challenge alongside datasets and baseline system. The challenge focuses on distant speaker diarization and automatic speech recognition (DASR) in far-field meeting scenarios, with single-channel and known-geometry multi-channel tracks, and serves as a launch platform for two new datasets: First, a benchmarking dataset of 315 meetings, averaging 6 minutes each, capturing a broad spectrum of real-world acoustic conditions and conversational dynamics. It is recorded across 30 conference rooms, featuring 4-8 attendees and a total of 35 unique speakers. Second, a 1000-hour simulated training dataset, synthesized with enhanced authenticity for real-world generalization, incorporating 15,000 real acoustic transfer functions. The tasks focus on single-device DASR, where multi-channel devices always share the same known geometry. This is aligned with common setups in actual conference rooms, and avoids technical complexities associated with multi-device tasks. It also allows for the development of geometry-specific solutions. The NOTSOFAR-1 Challenge aims to advance research in the field of distant conversational speech recognition, providing key resources to unlock the potential of data-driven methods, which we believe are currently constrained by the absence of comprehensive high-quality training and benchmarking datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SDBench: A Comprehensive Benchmark Suite for Speaker Diarization

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SDBench provides a reproducible 13-dataset benchmark for speaker diarization, and its companion SpeakerKit achieves a claimed 9.6x speedup over Pyannote v3.1 with comparable DER.

  2. SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A pseudo-supervised method that uses close-talk microphone recordings to estimate training targets lets a speech enhancement model adapt to real far-field data, cutting CER to 29.80% on MISP2023.

  3. MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses

    eess.AS 2025-07 reject novelty 5.0 of 10

    MMW combines a Mamba-based Mix Block, a Frame Diarization Mamba layer, and multi-scale GRPO to reduce side-talk interference in Whisper ASR, reporting WER as low as 3.71%.

  4. Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge

    cs.SD 2025-05 conditional novelty 5.0 of 10

    Training a neural speech enhancer on pseudo labels derived from aligned close-talk recordings improves far-field meeting ASR, reaching second place in the MISP-Meeting Challenge.

  5. The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026

    eess.AS 2026-07 conditional novelty 4.0 of 10

    A cascaded smart-glasses TSA-ASR system with a dominant-speaker overlap fallback achieved 7.10% tcpCER on two-person dialogues and 34.04% on multi-party meetings, ranking second on the meeting track.

Pith tools