REVIEW 21 cited by
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language models (LLM) have demonstrated the capability to handle a variety of generative tasks. This paper presents the UniAudio system, which, unlike prior task-specific approaches, leverages LLM techniques to generate multiple types of audio (including speech, sounds, music, and singing) with given input conditions. UniAudio 1) first tokenizes all types of target audio along with other condition modalities, 2) concatenates source-target pair as a single sequence, and 3) performs next-token prediction using LLM. Also, a multi-scale Transformer model is proposed to handle the overly long sequences caused by the residual vector quantization based neural codec in tokenization. Training of UniAudio is scaled up to 165K hours of audio and 1B parameters, based on all generative tasks, aiming to obtain sufficient prior knowledge not only in the intrinsic properties of audio but also the inter-relationship between audio and other modalities. Therefore, the trained UniAudio model has the potential to become a foundation model for universal audio generation: it shows strong capability in all trained tasks and can seamlessly support new audio generation tasks after simple fine-tuning. Experiments demonstrate that UniAudio achieves state-of-the-art or at least competitive results on most of the 11 tasks. Demo and code are released at https://github.com/yangdongchao/UniAudio
Forward citations
Cited by 21 Pith papers
-
Simultaneous Speech-to-Speech Translation Without Aligned Data
Hibiki-Zero performs simultaneous speech-to-speech translation without word-level aligned data, using sentence-level supervision plus GRPO reinforcement learning with BLEU-based process rewards, and reports state-of-t...
-
Scaling Laws of Motion Forecasting and Planning -- Technical Report
Motion forecasting models improve with compute as a power law, with optimal model size growing 1.5x faster than dataset size, and closed-loop driving failures also decreasing with scale.
-
ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models
Hierarchical multi-prompt representation generation plus generalized flow matching yields high-quality single-stage waveform diffusion from 12.5 Hz latents and efficient LDM TTS.
-
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
Fine-grained Soundscape Control for Augmented Hearing
Aurchestra enables real-time, on-device per-class sound extraction and volume control for up to five simultaneous sound classes on hearables.
-
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
Semantic-token infilling plus a frozen flow-matching decoder and a TTS-based GRPO reward produces more intelligible and natural text-based speech edits than prior AR and NAR baselines.
-
iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation
iMontage repurposes a pretrained video diffusion model to generate coherent yet highly dynamic image sets from arbitrary numbers of input images.
-
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.
-
Ego-centric Predictive Model Conditioned on Hand Trajectories
Ego-PM predicts future hand trajectories and then uses them to condition latent diffusion video generation, jointly outputting actions and future frames in egocentric and robotic scenes.
-
Your Spending Needs Attention: Modeling Financial Habits with Transformers
A causal transformer pre-trained with next-token prediction on tokenized bank transactions, fused end-to-end with tabular features, lifts recommendation test AUC by 1.25% relative over a LightGBM baseline at Nubank.
-
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
A single multimodal language model can generate intelligible speech and synchronized background audio jointly from a silent video, transcript, and reference voice, outperforming separately concatenated speech and audi...
-
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
A causal audio language model with continuous-valued tokens and masked next-token prediction matches diffusion-based text-to-audio quality with smaller, streamable models.
-
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
SemAlignVC strips source-speaker timbre by aligning a speech semantic encoder to BERT text embeddings, then resynthesizes the content conditioned only on a target voice reference.
-
Qwen-Audio-3.0-Gen-Preview Technical Report
One diffusion-transformer system with a shared audio codec generates standalone audio, multi-speaker dialogue, and time-structured mixed scenes, with strongest measured advantages in speaker similarity, cross-turn con...
-
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models
Introduces CCG-CFG with inconsistency-based dynamic scales and hard-sample mining distillation to boost emotional alignment in auto-regressive TTS, reporting up to 12% absolute gains in emotion recognition accuracy.
-
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations
QTTS models speech as sequences from a multi-codebook RVQ audio codec whose first codebook is trained with ASR supervision, aiming for higher-fidelity TTS than single-codebook systems.
-
Exploring State-Space-Model based Language Model in Music Generation
SiMBA, a Mamba-based decoder, converges faster and matches or slightly trails a Transformer baseline in a single-codebook text-to-music system under limited compute.
-
BoSS: Beyond-Semantic Speech
Current spoken-language models perform poorly on a new five-task evaluation of beyond-semantic speech signals, including dialect, emotion, age, and non-verbal cues.
-
Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration
The paper proposes a systematic classification and two roadmaps for using foundation models (LLMs and wireless foundation models) to design Synesthesia of Machines systems for 6G, with preliminary case-study evidence ...
-
Small Data Explainer -- The impact of small data methods in everyday life
A review and explainer that frames small data methods through the recurring challenges of similarity, transfer, and uncertainty and maps them to application areas and technical approaches.
Discussion (0). Sign in to comment.