Pith. sign in

REVIEW 5 cited by

CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.09249 v2 pith:FQLJN7XH submitted 2020-04-20 cs.SD cs.CLeess.AS

CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings

classification cs.SD cs.CLeess.AS
keywords speechrecognitionchallengemultispeakerchime-6trackunsegmentedchime
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Following the success of the 1st, 2nd, 3rd, 4th and 5th CHiME challenges we organize the 6th CHiME Speech Separation and Recognition Challenge (CHiME-6). The new challenge revisits the previous CHiME-5 challenge and further considers the problem of distant multi-microphone conversational speech diarization and recognition in everyday home environments. Speech material is the same as the previous CHiME-5 recordings except for accurate array synchronization. The material was elicited using a dinner party scenario with efforts taken to capture data that is representative of natural conversational speech. This paper provides a baseline description of the CHiME-6 challenge for both segmented multispeaker speech recognition (Track 1) and unsegmented multispeaker speech recognition (Track 2). Of note, Track 2 is the first challenge activity in the community to tackle an unsegmented multispeaker speech recognition scenario with a complete set of reproducible open source baselines providing speech enhancement, speaker diarization, and speech recognition modules.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

    cs.SD 2026-07 conditional novelty 6.0

    Proxy-supervised joint fine-tuning of a BSRNN separator with ASR, speaker-similarity, VAD and DNSMOS losses on a new 71k real-conversation corpus yields the best SIM and timing F1 on REAL-T.

  2. ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching

    eess.AS 2025-07 conditional novelty 6.0

    ZipVoice-Dialog is a flow-matching non-autoregressive model for zero-shot spoken dialogue generation that uses curriculum learning and speaker-turn embeddings, paired with a new 6.8k-hour OpenDialog dataset, and repor...

  3. Balancing ASR and diarization in end-to-end LLMs for multi-talker speech recognition

    eess.AS 2026-06 unverdicted novelty 4.0

    LLM-based multi-talker ASR with dual-encoder, feature interleaving, length-aware speaker loss, and adaptive ASR threshold achieves 18% and 24% relative gains over baselines on AliMeeting and Aishell4.

  4. SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription

    eess.AS 2026-06 unverdicted novelty 4.0

    SoulX-Transcriber is a unified LLM framework for end-to-end multi-speaker transcription using two-stage training (speaker-aware pre-training then supervised fine-tuning) that reports strong results on AliMeeting, AISH...

  5. Spatial Speech Perception Systems: A Survey of Sound Source Localization, Directional Enhancement, and Speech Recognition

    eess.AS 2026-07 unverdicted novelty 2.0

    A survey of spatial speech perception systems covering sound source localization, directional enhancement, and automatic speech recognition methods and their integration.