pith. machine review for the scientific record. sign in

Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model benchmark

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

years

2026 1 2025 1

verdicts

UNVERDICTED 2

representative citing papers

Kimi-Audio Technical Report

eess.AS · 2025-04-25 · unverdicted · novelty 5.0

Kimi-Audio is an open-source audio foundation model that achieves state-of-the-art results on speech recognition, audio understanding, question answering, and conversation after pre-training on more than 13 million hours of speech, sound, and music data.

citing papers explorer

Showing 2 of 2 citing papers.

  • VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cs.CL · 2026-05-07 · unverdicted · none · ref 116

    VITA-QinYu is the first expressive end-to-end spoken language model supporting role-playing and singing alongside conversation, trained on 15.8K hours of data and outperforming prior models on expressiveness and conversational benchmarks.

  • Kimi-Audio Technical Report eess.AS · 2025-04-25 · unverdicted · none · ref 50

    Kimi-Audio is an open-source audio foundation model that achieves state-of-the-art results on speech recognition, audio understanding, question answering, and conversation after pre-training on more than 13 million hours of speech, sound, and music data.