Pith. sign in

REVIEW 6 cited by

The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.13734 v2 pith:G7SXG5M4 submitted 2023-06-23 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords challengechimechallengeschime-7dasrdevicesdiarizationdifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The CHiME challenges have played a significant role in the development and evaluation of robust automatic speech recognition (ASR) systems. We introduce the CHiME-7 distant ASR (DASR) task, within the 7th CHiME challenge. This task comprises joint ASR and diarization in far-field settings with multiple, and possibly heterogeneous, recording devices. Different from previous challenges, we evaluate systems on 3 diverse scenarios: CHiME-6, DiPCo, and Mixer 6. The goal is for participants to devise a single system that can generalize across different array geometries and use cases with no a-priori information. Another departure from earlier CHiME iterations is that participants are allowed to use open-source pre-trained models and datasets. In this paper, we describe the challenge design, motivation, and fundamental research questions in detail. We also present the baseline system, which is fully array-topology agnostic and features multi-channel diarization, channel selection, guided source separation and a robust ASR model that leverages self-supervised speech representations (SSLR).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models

    eess.AS 2025-06 conditional novelty 6.0 of 10

    An LLM conditioned on speaker embeddings and utterance time boundaries jointly transcribes and timestamps overlapping multi-speaker speech.

  2. SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A pseudo-supervised method that uses close-talk microphone recordings to estimate training targets lets a speech enhancement model adapt to real far-field data, cutting CER to 29.80% on MISP2023.

  3. M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

    eess.AS 2025-06 reject novelty 5.0 of 10

    Release of M3SD, a 770+ hour pseudo-labeled multi-scenario, multi-language audio-visual speaker diarization dataset, built from YouTube and Bilibili videos without manual annotation.

  4. Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge

    cs.SD 2025-05 conditional novelty 5.0 of 10

    Training a neural speech enhancer on pseudo labels derived from aligned close-talk recordings improves far-field meeting ASR, reaching second place in the MISP-Meeting Challenge.

  5. The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge

    eess.AS 2025-07 conditional novelty 4.0 of 10

    A challenge system combining speaker diarization, speaker embeddings, and a Qwen2.5 LLM adapter architecture reports 18.08% tcpWER on multilingual multi-speaker ASR, far below the 60.39% baseline.

  6. Exploring Speaker Diarization with Mixture of Experts

    cs.SD 2025-06 conditional novelty 4.0 of 10

    A speaker diarization system that adds a shared-and-soft mixture-of-experts layer to a memory-augmented sequence-to-sequence model reports lower error rates on CHiME-6, DiPCo, and Mixer 6, with some DIHARD-III claims ...

Pith tools