ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

Alberto Abad; Alexander Polok; Carlos Carvalho; Chenda Li; Chyi-Jiunn Lin; Da-Hee Yang; Francisco Teixeira; Jiatong Shi; Jinchuan Tian; Masao Someki

arxiv: 2606.21854 · v1 · pith:VIH5QVCTnew · submitted 2026-06-20 · 📡 eess.AS · cs.SD

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

Masao Someki , Alexander Polok , Carlos Carvalho , Chyi-Jiunn Lin , Da-Hee Yang , Jiatong Shi , Jinchuan Tian , Nelson Enrique Yalta Soplin

show 9 more authors

Samuele Cornell Siddhant Arora Francisco Teixeira Wei Wang William Chen Alberto Abad Chenda Li Shinji Watanabe Wangyou Zhang

This is my paper

classification 📡 eess.AS cs.SD

keywords espnet3trainingdatasetemphexperimentsresearchspeechaudio

0 comments

read the original abstract

Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort to support such experiments. We present ESPnet3, a speech and audio research framework built on a modular system architecture with configuration-driven dataset composition and unified Python-based workflows. ESPnet3 introduces a DataOrganizer abstraction for flexible dataset integration and dataset sharding for memory-efficient large-scale training, while allowing recipe-specific logic through lightweight stage overrides. In OWSM pre-training experiments, ESPnet3 reduces per-epoch training time by \emph{21.1 minutes} compared to ESPnet2 and achieves \emph{>80\% GPU utilization} in multi-node training. Fine-tuning experiments show that new models and datasets can be integrated with around \emph{46 lines of additional code}. ESPnet3 will be publicly released with model checkpoints and training logs.

This paper has not been read by Pith yet.

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

discussion (0)