REVIEW 3 major objections 4 minor 1 cited by
Sound Scene Synthesis at the DCASE 2024 Challenge
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Text-to-sound-scene systems still trail expert recordings by 36 percent, and FAD tracks human judgment closely enough to serve as a proxy.
desk verdict Honest, useful challenge report with clean evaluation protocol; the FAD-human correlation is real but rests on 5 points and a single engineered reference, and the paper itself flags the small sample. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Fréchet Audio Distance computed over PANN-Wavegram-Logmel embeddings, paired with a perceptual score built from Foreground Fit, Background Fit, and Audio Quality ratings: $\text{Perceptual Score} = (2FF + BF + AQ)/4$. FAD compares the mean ($\mu$) and covariance ($\Sigma$) of the embedding distributions of the reference and generated audio sets, and the specific embedding was chosen because prior work showed that FAD's agreement with human perception depends on the embedding. The reference side of both measures is a 250-caption evaluation set recorded by a single sound engineer, and the weighted perceptual formula gives foreground accuracy the largest say in the final ranking.
What would settle it
Have a second sound engineer record a fresh reference set for the same 250 captions, recompute FAD for the four submitted systems against both reference sets, and compare with the existing human ratings; if the FAD-based ranking or the 36 percent gap shifts while human ratings stay stable, the single-reference assumption is the source of the instability.
Extended reading notes
Core claim
The paper's central claim is that a standardized framework combining FAD computed on PANN-Wavegram-Logmel embeddings with ratings from 14 expert listeners gives a stable, interpretable comparison of sound scene synthesis systems. Using that framework, the best submitted system reaches an average perceptual score of 5.832 against the sound engineer reference's 8.793, a gap of about 36 percent, while FAD values range from 35.985 for the best system to 53.728 for the lowest-ranked one. The objective and subjective measures agree strongly on foreground fit ($r=0.94$) and background fit ($r=0.94$) and less strongly on overall audio quality ($r=0.77$), which the paper treats as useful but weak evidence because only five systems were compared.
Load-bearing premise
The evaluation assumes that a single set of reference recordings made by one sound engineer is the correct target for each caption, so a system is judged by how close it comes to that one distribution.
Editorial extensions
If this is right
- FAD can serve as an inexpensive screen for future sound scene synthesis comparisons, since it tracks human foreground and background fit at $r=0.94$ across the systems tested.
- The 36 percent gap between the best system and the reference quantifies the headroom remaining for generative sound models, giving later work a concrete improvement target.
- The 'Foreground with Background in the background' caption structure and the separate FF, BF, and AQ ratings allow future evaluations to diagnose failure by scene layer rather than by overall quality alone.
- The organizers' accounting of roughly 120 hours of expert effort plus platform and compute costs explains why the task was discontinued, implying that sustainable generative-audio benchmarks will need cheaper reference and rating protocols.
- The drop from 32 submissions in the 2023 edition to 4 in 2024, paired with the removal of training-data constraints, suggests that evaluator overhead and reliance on large pre-existing models shape participation as much as synthesis skill.
Reading between the lines
- The strong FAD-human correlation suggests a testable two-stage benchmark design: use FAD to pre-screen many systems, then spend limited human rating effort only on the top FAD candidates.
- The single-reference design is the fragile part of the framework, because for open-ended captions many acoustic realizations can legitimately fit the same text and would be penalized for departing from one engineer's choices.
- The 36 percent gap might shrink or grow if the reference set were expanded to multiple engineers' recordings, and checking that sensitivity would clarify whether the gap reflects model weakness or reference idiosyncrasy.
- The difficulty of sustaining annual human evaluation points toward automated or semi-automated proxies, but the paper's own $r=0.77$ correlation on audio quality warns that FAD alone would misrank systems when overall quality is the criterion.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports on DCASE 2024 Challenge Task 7, a text-to-sound generation task for sound scenes. The authors describe a standardized evaluation framework combining an objective metric, Fréchet Audio Distance (FAD) with PANN-Wavegram-Logmel embeddings, and human perceptual ratings on three scales (Foreground Fit, Background Fit, Audio Quality), aggregated into a weighted Perceptual Score. The dataset consists of 310 audio-caption pairs (60 development, 250 evaluation) created by a sound engineer, and the task constrains outputs to 4-second mono clips without music or intelligible speech. Four submitted systems and an AudioLDM baseline were evaluated. The main reported findings are a substantial gap between the sound-engineer reference and the best submitted system, and strong correlations between FAD and the subjective metrics (0.94, 0.94, 0.77). The paper also discusses the decision to discontinue the task in 2025 due to cost and shifting research scope.
Significance. If the evaluation framework is valid, it provides a reusable protocol for comparing text-to-audio models in a constrained sound-scene setting. The manuscript has concrete strengths: it specifies the prompt structure, dataset sizes, rater blinding, self-rating removal, inter-rater agreement (Cronbach's alpha = 0.959), and it releases official evaluation software. The authors also explicitly acknowledge that the small number of systems limits the strength of the FAD-human correlation evidence. However, the paper's central claim that FAD is validated as a perceptual quality measure is threatened by the use of a single-engineer reference set and by the embedding-selection history in the authors' prior work. The paper also contains a numerical inconsistency in the headline performance-gap figure. These issues are load-bearing and require revision before the framework can be considered established.
major comments (3)
- [§3.1, §4.1, Eq. (1)] The FAD reference set in Eq. (1) is built from audio created by a single sound engineer, apparently one recording per prompt (§3.1). For open-ended text-to-audio generation, many acoustically different recordings can legitimately satisfy the same caption, so FAD(r,g) penalizes any valid output that is far in embedding space from that engineer's specific rendition. This threatens the paper's interpretation of FAD as a perceptual quality measure and could change system rankings if a different engineer's recordings were used as the reference. The authors should test the stability of the reported correlations and rankings by constructing reference sets from multiple independent engineers (or by using a larger set of references per prompt) and should state this limitation explicitly in the paper.
- [§5.2, Table 1] The headline claim of a "substantial 36% performance gap" between the reference (8.793) and the best submitted system (5.832) is not supported by the numbers in Table 1: the relative gap is (8.793 − 5.832)/8.793 ≈ 33.7%, not 36%. Please either correct the percentage or explain the calculation; this number appears in the Abstract and Section 5.2 and is reported as a key result.
- [§5.2, §4.1] The reported FAD-human correlations (0.94, 0.94, 0.77) are computed on only five systems (four submissions plus the baseline). Moreover, the PANN-Wavegram-Logmel embedding used for FAD was chosen in the authors' prior work [8] specifically to maximize correlation with human perception. Because that prior selection is not independent of the present validation, the correlations should be interpreted with caution. The paper acknowledges the small sample but does not discuss this selection issue; please add a discussion of this limitation and, ideally, provide confidence intervals or a leave-one-system-out analysis to assess robustness.
minor comments (4)
- [§1] In the Introduction, "motivated by the recent advances generative models" is missing "in" after "advances"; it should read "recent advances in generative models."
- [§3.2] In the sentence about Room Tone 1, there is a stray space before the period and "sounds" may be intended to be part of the category name; please rephrase for clarity (e.g., "Room Tone 1 (labeled as 'Nothing') sounds").
- [§5.1] The paper states that 24 evaluation captions were used for subjective rating but does not specify whether the FAD scores in Table 1 and Figure 1 were computed on those 24 captions or on the full 250-caption evaluation set; please clarify this to ensure the correlation analysis is interpretable.
- [§5.2] The phrase "weak evidence" for the FAD-human correlation is slightly misleading: with n=5, a correlation of 0.94 is large in magnitude but statistically fragile. Consider reporting exact p-values, confidence intervals, or a permutation test to make the strength of the evidence precise.
Circularity Check
No significant circularity: the challenge evaluation is an independent comparison of FAD against human ratings, not a fitted prediction.
full rationale
This paper reports the results of a DCASE challenge; it contains no derivation of a predicted quantity from fitted inputs. The objective metric, FAD, is defined in Eq. (1) using an external pretrained embedding, and the reference audio sets are fixed by the challenge design (Sec. 3.1, Sec. 4.1). The human perceptual scores are collected independently by a panel of raters (Sec. 4.2), and the reported correlations (Sec. 5.2) compare two separately measured quantities. The only near-circular element is that the FAD embedding was selected in the authors' prior work [8] to maximize FAD-human correlation; however, the challenge data and ratings used in Sec. 5.2 are not shown to be the same data used for that selection, so the correlation is an out-of-sample check rather than a fitted result. The paper itself cautions that the correlation is weak evidence due to the small number of systems (Sec. 5.2). The reference-set assumption flagged by the skeptic—one sound engineer's recordings as target—is a validity limitation for open-ended text-to-audio, not a circularity, because FAD is not defined in terms of the human ratings and no result is equivalent to its own input. No load-bearing self-citation chain or uniqueness theorem is invoked. Therefore the paper is self-contained as an evaluation report and receives a score of 0.
Assumptions & free parameters
free parameters (2)
- Perceptual score weights =
2, 1, 1 for foreground fit, background fit, audio quality
- FAD embedding model =
PANN-Wavegram-Logmel
assumptions (4)
- domain assumption Human ratings on 0-10 scales for foreground fit, background fit, and audio quality are valid ground truth for sound scene synthesis quality.
- domain assumption FAD with PANN-Wavegram-Logmel embeddings is a valid objective proxy for perceptual quality.
- domain assumption The 24 evaluation captions selected for subjective rating are representative of the 250-caption evaluation set.
- domain assumption The reference audio clips produced by one sound engineer are the correct target for each text prompt.
Cite this review
Pith. "Pith review of Sound Scene Synthesis at the DCASE 2024 Challenge." pith.science (2026). https://pith.science/paper/AIMB6Q4L
@misc{pith2026250108587,
author = {Pith},
title = {Pith review of: Sound Scene Synthesis at the DCASE 2024 Challenge},
year = {2026},
howpublished = {\url{https://pith.science/paper/AIMB6Q4L}},
note = {Machine review of arXiv:2501.08587}
}
read the original abstract
This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. Recent advances in sound synthesis and generative models have enabled the creation of realistic and diverse audio content. We introduce a standardized evaluation framework for comparing different sound scene synthesis systems, incorporating both objective and subjective metrics. The challenge attracted four submissions, which are evaluated using the Fr\'echet Audio Distance (FAD) and human perceptual ratings. Our analysis reveals significant insights into the current capabilities and limitations of sound scene synthesis systems, while also highlighting areas for future improvement in this rapidly evolving field.
Figures
Forward citations
Cited by 1 Pith paper
-
SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision
An agentic pipeline that plans, retrieves/generates, and deterministically renders multi-event soundscapes, and shows those structured outputs improve audio-language model reasoning over real-only data.
Reference graph
Works this paper leans on
-
[8]
ACKNOWLEDGEMENTS We thank all the raters who did the subjective evaluation: Xie Zhi- Dong, Li XinYu, Liu HaiCheng, Zou XiaoYan, Sun Yu, Hae Chun Chung, Jae Hoon Jung, Yi Yuan, Haohe Liu, Xubo Liu, Mark D. Plumbley, Wenwu Wang, Sagnik Ghosh, Gaurav Verma, Sid- dharath Narayan Shakya, Shubham Sharma, Shivesh Singh, Urszula Oszczapinska, Paige Brady, Angjeli...
-
[1]
INTRODUCTION This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. The challenge is motivated by the recent advances generative models for the creation of realistic and diverse audio con- tent, as proposed in [1] and following the last year’s version [2]
work page 2024
-
[2]
This is a more flexible setup than the category- based generation used in the last year [2]
PROBLEM AND TASK DEFINITION We defined the challenge as a text-to-sound generation task, where systems must generate realistic environmental audio based on tex- tual descriptions. This is a more flexible setup than the category- based generation used in the last year [2]. Each prompt follows the following structure: ”Foregroundwith Background in the backg...
-
[3]
Sound Scene Synthesis at the DCASE 2024 Challenge
DATASET AND BASELINE 3.1. Dataset Creation The challenge dataset contains 310 audio-captions in total, with 60 samples designated for development and 250 for evaluation. All audio content was carefully designed by a sound engineer to match specific prompts, ensuring high-quality and consistent sound scenes. The audio samples were sourced from Freesound.or...
work page Pith review arXiv 2025
-
[4]
EV ALUATION METHODOLOGY 4.1. Objective Evaluation We employed the Fr ´echet Audio Distance (FAD) [6] with PANN- Wavegram-Logmel [7] embeddings as our primary objective metric. The embedding was chosen to maximize the correlation between the FAD score and the human perception [8]. The FAD computation is defined as: FAD(r, g) = ∥µr − µg∥2 + Tr(Σr + Σg − 2 p...
-
[5]
System Performance Table 1 summarizes the evaluation results
RESULTS 5.1. System Performance Table 1 summarizes the evaluation results. The evaluation process encompassed four submitted systems [9, 10, 11, 12] assessed by a 1Room tone is a recorded sound with no specific sound event and used to capture natural noise of a recording environment. 2https://freesound.org/ 3https://sound-effects.bbcrewind.co.uk/search 4h...
work page 2024
-
[6]
First, the generative aspect of organizing this challenge has been costly and labor intensive
DISCONTINUATION OF THE TASK It is worth mentioning why the organizers decided not to continue the DCASE challenge in 2025 despite the successful challenges in 2023 and 2024. First, the generative aspect of organizing this challenge has been costly and labor intensive. In this year’s challenge, it took (a) about 40 hours to create and refine the evaluation...
work page 2025
-
[7]
CONCLUSION The DCASE 2024 Challenge Task 7 has provided valuable insights into the current state of sound scene synthesis while highlighting several crucial areas for future development. While the submit- ted systems demonstrated promising capabilities, the significant gap between synthetic and reference audio quality indicates substantial room for improv...
work page 2024
Show all 43 references
-
[9]
A proposal for foley sound synthesis challenge,
K. Choi, S. Oh, M. Kang, and B. McFee, “A proposal for foley sound synthesis challenge,”arXiv preprint arXiv:2207.10760, 2022
2022 arXiv
-
[10]
Foley sound synthesis at the dcase 2023 challenge,
K. Choi, J. Im, L. Heller, B. Mcfee, K. Imoto, Y . Okamoto, M. Lagrange, and S. Takamichi, “Foley sound synthesis at the dcase 2023 challenge,” in 2023 Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2023) , 2023
2023
-
[11]
AudioLDM: Text-to-audio generation with latent diffusion models,
H. Liu, Z. Chen, Y . Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” arXiv preprint arXiv:2301.12503, 2023
2023 arXiv
-
[12]
Audiocaps: Gen- erating captions for audios in the wild,
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Gen- erating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the As- sociation for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Paper...
2019
-
[13]
Audio set: An ontology and human-labeled dataset for audio events,
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2017, p...
2017
-
[14]
Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms.” in INTERSPEECH, 2019
2019
-
[15]
Panns: Large-scale pretrained audio neural net- works for audio pattern recognition,
Q. Kong, Y . Cao, T. Iqbal, Y . Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural net- works for audio pattern recognition,” IEEE/ACM Transac- tions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
-
[16]
Correlation of fr´echet audio dis- tance with human perception of environmental audio is em- bedding dependent,
M. Tailleur, J. Lee, M. Lagrange, K. Choi, L. M. Heller, K. Imoto, and Y . Okamoto, “Correlation of fr´echet audio dis- tance with human perception of environmental audio is em- bedding dependent,” in 2024 32nd European Signal Process- ing Conference (EUSIPCO). IEEE, 2024
2024
-
[17]
Sound scene synthesis with audioldm and tango2 for dcase 2024 task7,
X. ZhiDong, L. XinYu, L. HaiCheng, Z. XiaoYan, and S. Yu, “Sound scene synthesis with audioldm and tango2 for dcase 2024 task7,” Samsung Research China-Nanjing, Nanjing, China, Tech. Rep., July 2024
2024
-
[18]
Sound scene synthesis based on gan using contrastive learning and effective time-frequency swap cross attention mechanism,
H. C. Chung and J. H. Jung, “Sound scene synthesis based on gan using contrastive learning and effective time-frequency swap cross attention mechanism,” KT Corporation, Seoul, Re- public of Korea, Tech. Rep., July 2024
2024
-
[19]
Dif- fusion based sound scene synthesis for dcase challenge 2024 task 7,
Y . Yuan, H. Liu, X. Liu, M. D. Plumbley, and W. Wang, “Dif- fusion based sound scene synthesis for dcase challenge 2024 task 7,” University of Surrey, Guildford, United Kingdom, Tech. Rep., July 2024
2024
-
[20]
Sound scene synthesis based on fine-tuned latent diffusion model for dcase challenge 2024 task 7,
S. Ghosh, G. Verma, S. N. Shakya, S. Sharma, and S. Singh, “Sound scene synthesis based on fine-tuned latent diffusion model for dcase challenge 2024 task 7,” Indian Institute of Technology Mandi, Kamand, Mandi, India, Tech. Rep., July 2024
2024
-
[21]
Challenge on sound scene synthesis: Evaluating text-to-audio generation,
J. Lee, M. Tailleur, L. M. Heller, K. Choi, M. Lagrange, B. McFee, K. Imoto, and Y . Okamoto, “Challenge on sound scene synthesis: Evaluating text-to-audio generation,” in Au- dio Imagination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Generation
2024
-
[22]
T-foley: A controllable waveform-domain diffusion model for temporal-event-guided foley sound synthesis,
Y . Chung, J. Lee, and J. Nam, “T-foley: A controllable waveform-domain diffusion model for temporal-event-guided foley sound synthesis,” in ICASSP 2024-2024 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 6820–6824
2024
-
[23]
Mambafoley: Foley sound generation using selec- tive state-space models,
M. F. Colombo, F. Ronchini, L. Comanducci, and F. An- tonacci, “Mambafoley: Foley sound generation using selec- tive state-space models,” arXiv preprint arXiv:2409.09162 , 2024
2024 arXiv
-
[24]
Audio generation with multiple conditional diffu- sion model,
Z. Guo, J. Mao, R. Tao, L. Yan, K. Ouchi, H. Liu, and X. Wang, “Audio generation with multiple conditional diffu- sion model,” in Proceedings of the AAAI Conference on Arti- ficial Intelligence, vol. 38, no. 16, 2024, pp. 18 153–18 161
2024
-
[25]
Picoaudio: En- abling precise timestamp and frequency controllability of audio events in text-to-audio generation,
Z. Xie, X. Xu, Z. Wu, and M. Wu, “Picoaudio: En- abling precise timestamp and frequency controllability of audio events in text-to-audio generation,” arXiv preprint arXiv:2407.02869, 2024
2024 arXiv
-
[26]
Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,
H. Liu, Y . Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y . Wang, W. Wang, Y . Wang, and M. D. Plumbley, “Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024
2024
-
[27]
Auffusion: Leveraging the power of diffusion and large language models for text-to- audio generation,
J. Xue, Y . Deng, Y . Gao, and Y . Li, “Auffusion: Leveraging the power of diffusion and large language models for text-to- audio generation,” arXiv preprint arXiv:2401.01044, 2024
2024 arXiv
-
[28]
Ezaudio: Enhancing text-to-audio gen- eration with efficient diffusion transformer,
J. Hai, Y . Xu, H. Zhang, C. Li, H. Wang, M. Elhi- lali, and D. Yu, “Ezaudio: Enhancing text-to-audio gen- eration with efficient diffusion transformer,” arXiv preprint arXiv:2409.10819, 2024
2024 arXiv
-
[29]
Fugatto 1: Foundational generative audio transformer opus 1,
Anonymous, “Fugatto 1: Foundational generative audio transformer opus 1,” in Submitted to The Thirteenth International Conference on Learning Representations, 2024, under review. [Online]. Available: https://openreview.net/ forum?id=B2Fqu7Y2cd
2024
-
[30]
Stable audio open,
Z. Evans, J. D. Parker, C. Carr, Z. Zukowski, J. Tay- lor, and J. Pons, “Stable audio open,” arXiv preprint arXiv:2407.14358, 2024
2024 arXiv
-
[31]
Improving text- to-audio models with synthetic captions,
Z. Kong, S.-g. Lee, D. Ghosal, N. Majumder, A. Mehrish, R. Valle, S. Poria, and B. Catanzaro, “Improving text- to-audio models with synthetic captions,” arXiv preprint arXiv:2406.15487, 2024
2024 arXiv
-
[32]
Syncfusion: Multi- modal onset-synchronized video-to-audio foley synthesis,
M. Comunit `a, R. F. Gramaccioni, E. Postolache, E. Rodol `a, D. Comminiello, and J. D. Reiss, “Syncfusion: Multi- modal onset-synchronized video-to-audio foley synthesis,” in ICASSP 2024-2024 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP) ...
2024
-
[33]
Sonicvisionlm: Play- ing sound with vision language models,
Z. Xie, S. Yu, Q. He, and M. Li, “Sonicvisionlm: Play- ing sound with vision language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26 866–26 875
2024
-
[34]
Video-foley: Two-stage video-to-sound generation via temporal event condition for fo- ley sound,
J. Lee, J. Im, D. Kim, and J. Nam, “Video-foley: Two-stage video-to-sound generation via temporal event condition for fo- ley sound,” arXiv preprint arXiv:2408.11915, 2024
2024
-
[35]
Read, watch and scream! sound generation from text and video,
Y . Jeong, Y . Kim, S. Chun, and J. Lee, “Read, watch and scream! sound generation from text and video,”arXiv preprint arXiv:2407.05551, 2024
2024 arXiv
-
[36]
Movie gen: A cast of media foundation models,
A. Polyak, A. Zohar, A. Brown, A. Tjandra, A. Sinha, A. Lee, A. Vyas, B. Shi, C.-Y . Ma, C.-Y . Chuang, et al. , “Movie gen: A cast of media foundation models,” arXiv preprint arXiv:2410.13720, 2024
2024 arXiv
-
[37]
Video-guided foley sound generation with multimodal controls,
Z. Chen, P. Seetharaman, B. Russell, O. Nieto, D. Bour- gin, A. Owens, and J. Salamon, “Video-guided foley sound generation with multimodal controls,” arXiv preprint arXiv:2411.17698, 2024
2024 arXiv
-
[38]
Vintage: Joint video and text conditioning for holistic audio generation,
S. S. Kushwaha and Y . Tian, “Vintage: Joint video and text conditioning for holistic audio generation,” arXiv preprint arXiv:2412.10768, 2024
2024 arXiv
-
[39]
Frieren: Efficient video-to-audio generation with rectified flow matching,
Y . Wang, W. Guo, R. Huang, J. Huang, Z. Wang, F. You, R. Li, and Z. Zhao, “Frieren: Efficient video-to-audio generation with rectified flow matching,” arXiv preprint arXiv:2406.00320, 2024
2024 arXiv
-
[40]
Masked gener- ative video-to-audio transformers with enhanced synchronic- ity,
S. Pascual, C. Yeh, I. Tsiamas, and J. Serr `a, “Masked gener- ative video-to-audio transformers with enhanced synchronic- ity,” inEuropean Conference on Computer Vision. Springer, 2025, pp. 247–264
2025
-
[41]
Temporally aligned audio for video with autoregression,
I. Viertola, V . Iashin, and E. Rahtu, “Temporally aligned audio for video with autoregression,” arXiv preprint arXiv:2409.13689, 2024
2024 arXiv
-
[42]
Gotta hear them all: Sound source aware vision to audio generation,
W. Guo, H. Wang, W. Cai, and J. Ma, “Gotta hear them all: Sound source aware vision to audio generation,” arXiv preprint arXiv:2411.15447, 2024
2024 arXiv
-
[43]
Taming multimodal joint train- ing for high-quality video-to-audio synthesis,
H. K. Cheng, M. Ishii, A. Hayakawa, T. Shibuya, A. Schwing, and Y . Mitsufuji, “Taming multimodal joint train- ing for high-quality video-to-audio synthesis,” arXiv preprint arXiv:2412.15322, 2024
2024 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.