Pith. sign in

REVIEW 14 cited by

BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.04658 v2 pith:SOE42A6Z submitted 2022-06-09 cs.SD cs.CLcs.LGeess.AS

classification cs.SDcs.CLcs.LGeess.AS
keywords audiobigvganvariousvocoderenvironmentshigh-fidelitylarge-scalemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite recent progress in generative adversarial network (GAN)-based vocoders, where the model generates raw waveform conditioned on acoustic features, it is challenging to synthesize high-fidelity audio for numerous speakers across various recording environments. In this work, we present BigVGAN, a universal vocoder that generalizes well for various out-of-distribution scenarios without fine-tuning. We introduce periodic activation function and anti-aliased representation into the GAN generator, which brings the desired inductive bias for audio synthesis and significantly improves audio quality. In addition, we train our GAN vocoder at the largest scale up to 112M parameters, which is unprecedented in the literature. We identify and address the failure modes in large-scale GAN training for audio, while maintaining high-fidelity output without over-regularization. Our BigVGAN, trained only on clean speech (LibriTTS), achieves the state-of-the-art performance for various zero-shot (out-of-distribution) conditions, including unseen speakers, languages, recording environments, singing voices, music, and instrumental audio. We release our code and model at: https://github.com/NVIDIA/BigVGAN

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors

    eess.AS 2026-07 conditional novelty 6.0 of 10

    On 1,168 professional voice actors, a misidentification floor in speaker embeddings survives calibration, normalization, and discriminative re-ranking, and the same floor makes fixed-threshold voice-clone attribution ...

  2. What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Underrepresented gender in training suffers higher deepfake-detection error; WavLM gaps stay large under balance, and all post-hoc calibrations leave the EER gap fixed at 1.317 pp.

  3. RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    Hybrid two-stage diffusion transformer architecture for instruction-guided audio editing via rectified flow that performs joint attention at low resolution then alternates joint and cross-attention at high resolution ...

  4. Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training

    cs.SD 2026-03 conditional novelty 6.0 of 10

    Jointly training the watermark embedder/detector with the source separator enables ~1% bit-error-rate recovery of per-stem watermarks after mixing and separation, where independent training yields 15–35%.

  5. OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

    cs.SD 2026-02 conditional novelty 6.0 of 10

    A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.

  6. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  7. Hearing Hands: Generating Sounds from Physical Interactions in 3D Scenes

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A rectified flow model conditioned on 3D hand trajectories and rendered scene video generates realistic hand-scene interaction sounds, with a human study finding near-chance discrimination (47% misclassified).

  8. SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns

    eess.AS 2026-03 conditional novelty 5.5 of 10

    SEMamba++ combines Frequency GLP (FAN-based global-periodic + local conv) with multi-resolution parallel TFDP and learnable softplus mapping to outperform GSR baselines on VCTK, URGENT and AATC while remaining efficient.

  9. Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Faster IndexTTS-2 compiles all components of IndexTTS-2 into TensorRT engines, cutting end-to-end latency by up to 3.6x while adding streaming and batching.

  10. NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference

    eess.AS 2025-08 conditional novelty 5.0 of 10

    NanoCodec achieves competitive speech quality at 12.5 frames per second and 0.6-1.78 kbps, with a causal decoder for low-latency speech LLM inference.

  11. Traceable TTS: Toward Watermark-Free TTS with Strong Traceability

    eess.AS 2025-07 reject novelty 5.0 of 10

    A joint training loop makes an F5-TTS model produce audio that a paired wav2vec 2.0/LCNN discriminator can recognize, enabling watermark-free attribution; however, the reported generalization gain is not isolated from...

  12. SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A reference-guided diffusion model with a low-frequency upsampling module achieves marginal quality improvements over prior SVS baselines on Opencpop, with significant caveats about statistical significance and reprod...

  13. FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching

    eess.AS 2025-05 reject novelty 5.0 of 10

    FlowSE applies rectified flow matching with a DiT backbone to speech enhancement, reporting better DNSMOS and WER results and a much lower real-time factor than diffusion baselines.

  14. EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

    cs.SD 2025-05 reject novelty 4.0 of 10

    EZ-VC combines discrete units from a multilingual self-supervised encoder (Xeus) with an F5-TTS flow-matching decoder to achieve zero-shot any-to-any voice conversion, without text labels or multiple disentangling encoders.

Pith tools