Pith. sign in

REVIEW 1 cited by

An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.03697 v2 pith:SF3ZAGKA submitted 2024-01-08 cs.SD eess.AS

classification cs.SDeess.AS
keywords approachchallengeextractionspeechaudio-quality-basedmispmulti-strategyspeaker
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper describes our audio-quality-based multi-strategy approach for the audio-visual target speaker extraction (AVTSE) task in the Multi-modal Information based Speech Processing (MISP) 2023 Challenge. Specifically, our approach adopts different extraction strategies based on the audio quality, striking a balance between interference removal and speech preservation, which benifits the back-end automatic speech recognition (ASR) systems. Experiments show that our approach achieves a character error rate (CER) of 24.2% and 33.2% on the Dev and Eval set, respectively, obtaining the second place in the challenge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge

    cs.SD 2025-05 conditional novelty 4.0 of 10

    A hybrid diarization and ASR system with a CER-supervised bridging module achieved the best results in two MISP 2025 tracks.

Pith tools