DQ-Data2vec adds online K-means quantizers with cluster counts matched to language and phoneme counts to decouple these features during multilingual speech pre-training, improving phoneme and word error rates on CommonVoice.
XLS-R: Self- supervised cross-lingual speech representation learning at scale,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
DQ-Data2vec adds online K-means quantizers with cluster counts matched to language and phoneme counts to decouple these features during multilingual speech pre-training, improving phoneme and word error rates on CommonVoice.