Pith. sign in

REVIEW 19 cited by

Stable Audio Open

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.14358 v2 pith:A7REK5XU submitted 2024-07-19 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords modelsmodelopentext-to-audioaccessibleacrossallowingarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model's performance is competitive with the state-of-the-art across various metrics. Notably, the reported FDopenl3 results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Doppelganger: Sound Effects and Their Synthetic Twins

    cs.SD 2026-07 accept novelty 6.5 of 10

    Instance-pair training matches synthetic sound-effect twins to their real sources on unseen events (~80% R@1), while class supervision degrades below the frozen baseline and the mapping stays generator-specific.

  2. On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A single mean shift in the latent space of several neural codecs achieves competitive music bandwidth extension on some metrics, implying a largely linear structure.

  3. FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving

    cs.DC 2026-07 conditional novelty 6.0 of 10

    FlashDiff reduces diffusion serving latency by 30–97% and raises throughput 1.2–2.2× by adaptively skipping refinement of latent regions that no longer need it.

  4. Qwen-Music Technical Report

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Qwen-Music generates high-fidelity vocal songs via 25 Hz semantic tokens, Melody-CoT planning, and DiT rendering, claiming SOTA on 13/16 metrics and expert preference over proprietary systems.

  5. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  6. An Empirical Analysis of Task-Induced Encoder Bias in Fr\'echet Audio Distance

    eess.AS 2026-02 conditional novelty 6.0 of 10

    No single tested audio encoder catches all quality issues: reconstruction-trained encoders detect signal degradation, speech-trained encoders detect temporal order, and classification-trained encoders detect semantic ...

  7. SemanticAudio: Audio Generation and Editing in Semantic Space

    eess.AS 2026-01 conditional novelty 6.0 of 10

    SemanticAudio improves text-to-audio alignment by generating a compact semantic plan first with a Flow Matching planner and then rendering acoustic latents from that plan, and it performs training-free audio editing b...

  8. Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music

    cs.SD 2025-08 conditional novelty 6.0 of 10

    Amadeus generates symbolic music by autoregressively predicting note-level latents and decoding their attributes in parallel with a masked discrete diffusion model, yielding faster and more controllable generation tha...

  9. JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment

    cs.SD 2025-07 conditional novelty 6.0 of 10

    JAM is a 530M-parameter flow-matching song generator that adds word- and phoneme-level timing control and duration control, achieving strong lyric fidelity and musicality scores when ground-truth timings are provided.

  10. WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling

    cs.SD 2025-07 conditional novelty 6.0 of 10

    WildFX generates multi-track audio datasets by rendering real DAW effect graphs with commercial plugins inside Docker, and demonstrates the pipeline on blind mixing-graph estimation.

  11. MusGO: A Community-Driven Framework For Assessing Openness in Music-Generative AI

    cs.SD 2025-07 conditional novelty 6.0 of 10

    MusGO is a community-refined framework with 13 openness categories, applied to 16 music-generative models to produce a public openness leaderboard.

  12. AI-Generated Song Detection via Lyrics Transcripts

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Transcribing audio with Whisper and classifying the transcript with LLM2Vec detects AI-generated songs from audio alone, nearly matching clean-lyrics accuracy and beating audio-based detectors under perturbations and ...

  13. Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

    cs.SD 2025-06 conditional novelty 6.0 of 10

    OSSL is the first self-hosted, mood-annotated video-music dataset, and a video adapter on MusicGen-Medium improves film music generation over text-only baselines.

  14. In-the-wild Audio Spatialization with Flexible Text-guided Localization

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A text-guided latent diffusion model converts monaural audio into binaural audio whose perceived directions and distances follow user-specified text prompts.

  15. EgoZero: Robot Learning from Smart Glasses

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Robot policies trained only on egocentric human videos from smart glasses transfer zero-shot to a Franka gripper, with 70% success across 7 manipulation tasks.

  16. RenderBox: Expressive Performance Rendering with Text Control

    eess.AS 2025-02 conditional novelty 6.0 of 10

    RenderBox is a text-and-score conditioned diffusion model that renders expressive, controllable audio performances across piano, guitar, saxophone, violin, and orchestral instruments.

  17. KVAE: Family of Tokenizers for Multimodal Generative Models

    cs.CV 2026-08 conditional novelty 5.0 of 10

    KVAE introduces image, video, and full-band audio tokenizers whose reconstruction and downstream generation quality is competitive with, and often better than, current open-source tokenizers in head-to-head tests.

  18. Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Boomerang sampling, applied to a pretrained music diffusion model, creates audio variations that improve beat tracking when training data is scarce and can change instruments via text prompts.

  19. SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A cascaded pipeline of audio compression, latent diffusion extraction, and generative correction achieves state-of-the-art target speech extraction quality and intelligibility on Libri2Mix and out-of-domain data.

Pith tools