Pith. sign in

REVIEW 3 cited by

LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08203 v1 pith:C6KVMETO submitted 2024-06-12 eess.AS cs.SD

classification eess.AScs.SD
keywords audiogenerationmodelsdiffusionlatentmodelstepsflow
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, the application of diffusion models has facilitated the significant development of speech and audio generation. Nevertheless, the quality of samples generated by diffusion models still needs improvement. And the effectiveness of the method is accompanied by the extensive number of sampling steps, leading to an extended synthesis time necessary for generating high-quality audio. Previous Text-to-Audio (TTA) methods mostly used diffusion models in the latent space for audio generation. In this paper, we explore the integration of the Flow Matching (FM) model into the audio latent space for audio generation. The FM is an alternative simulation-free method that trains continuous normalization flows (CNF) based on regressing vector fields. We demonstrate that our model significantly enhances the quality of generated audio samples, achieving better performance than prior models. Moreover, it reduces the number of inference steps to ten steps almost without sacrificing performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Room Impulse Response Generation Conditioned on Acoustic Parameters

    cs.SD 2025-07 conditional novelty 5.0 of 10

    MaskGIT conditioned on acoustic parameters, operating on Descript Audio Codec tokens, generates room impulse responses that outperform StoRIR and FastRIR in objective and MUSHRA evaluations.

  2. AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion

    cs.SD 2025-05 conditional novelty 5.0 of 10

    AudioTurbo fine-tunes a diffusion model on deterministic noise-audio pairs created by the pretrained Auffusion model, achieving strong text-to-audio results in 10 inference steps.

  3. SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A purely algebraic interval-splitting consistency objective trains few-step generative models without JVP computations and recovers MeanFlow's differential identity as a special limit.

Pith tools