REVIEW 3 major objections 5 minor 1 cited by
Generative Data Augmentation Challenge: Synthesis of Room Acoustics for Speaker Distance Estimation
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Simulated room acoustics can train speaker distance estimation to about 20 cm error.
desk verdict A useful challenge benchmark undercut by a one-room validation of the simulator that feeds both tasks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Treble wave-based room acoustics simulation built on the Discontinuous Galerkin method, which generates dense grids of monaural and Ambisonic RIRs for ten furnished rooms, plus ten GWA hybrid-simulation rooms. The paper pairs this with an oracle SDE model, a convolutional recurrent neural network with attention and GRUs, trained on the simulated RIRs convolved with VCTK speech; the oracle performance serves as the upper bound that participant generative systems are expected to approach. The challenge protocol itself, with two sparse-enrollment scenarios, is what turns the simulation into a test of generative data augmentation.
What would settle it
Measure real RIRs in several of the other simulated Treble room layouts and compare T20, EDF, and DRR against the simulation; if errors substantially exceed the Room_0 values, the claim that the simulated dataset matches real-world data would be falsified. Alternatively, train an SDE model on real measured RIRs and compare its test error to the oracle's 0.21 m to see if the simulated training data is as informative as real data.
Extended reading notes
Core claim
The paper establishes that the challenge dataset is legitimate and useful. First, it validates the Treble wave-based simulation against a real measured room (Room_0), reporting broadband T20 mean absolute percentage error of 13.5%, EDF MSE of 22 dB, and DRR MSE of 14.5 dB; the authors interpret these low errors as evidence that the simulated rooms match real-world acoustics. Second, it shows that an SDE model trained from scratch on the full simulated dataset reaches 0.208 m and 0.209 m MAE in the two scenarios, versus 1.65 m and 1.69 m for a baseline model trained only on measured RIRs from a different dataset, indicating that the simulated RIRs carry enough acoustic detail to support accurate distance estimation.
Load-bearing premise
The paper assumes that the Treble simulation is accurate enough to serve as ground truth for all twenty rooms, but this accuracy is validated on only one measured room with 20 RIRs.
Editorial extensions
If this is right
- If simulated RIRs are trustworthy ground truth, generative models need only produce RIRs that match simulation accuracy to be useful for SDE training.
- A participant system that reconstructs RIRs close to the simulation should push SDE error down from the 1.65 m baseline toward the 0.21 m oracle.
- The two enrollment scenarios show that sparse center or corner measurements may be enough to generate room-wide training data, lowering the cost of dataset construction.
- The fixed SDE architecture means any improvement in Task 2 is attributable to data quality, not model changes, making the challenge a direct test of generative RIR fidelity.
Reading between the lines
- The 20 cm oracle error may approach what is achievable from single-channel audio given physical limits such as source-receiver placement uncertainty and reverberation variability; participant systems are unlikely to beat it by much.
- The validation on one control room leaves open the possibility that simulation accuracy varies by room type; a direct comparison of measured and simulated RIRs in a bathroom or meeting room would test whether the single-room validation transfers.
- A natural extension is to ask whether a model fine-tuned on generated RIRs generalizes to real rooms outside the challenge set, which would test the broader claim that simulation-based augmentation replaces measurement rather than just imitating it.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This ICASSP workshop challenge paper proposes a generative data augmentation benchmark for room acoustics, with two tasks: Task 1 asks participants to generate RIRs at unseen source-receiver locations from sparse enrollment data, and Task 2 asks them to use their generated RIRs to train a fixed speaker-distance-estimation (SDE) model. The dataset combines 10 simulated rooms from Treble's wave-based solver, 10 rooms from the GWA dataset, and one measured control room (Room_0) with corresponding simulations. The authors report objective metrics for Task 1 (T20 MAPE, EDF MSE, DRR MSE) and SDE results for Task 2, including an oracle model trained on the full challenge dataset and a C4DM-trained baseline. The central claim is that the simulated rooms match real-world data, supported by a comparison of Treble simulations against 20 measured RIRs in Room_0 (Table III), and that the oracle SDE results (Table IV) show the dataset supports accurate distance estimation.
Significance. If the challenge dataset is accepted as realistic, this paper provides a useful, reproducible benchmark for generative RIR synthesis and its downstream utility. The two-task design is sensible, the use of a fixed SDE architecture is a good way to isolate data quality, and the release of dataset, evaluation code, and baselines is a concrete contribution. The metrics chosen (T20, EDF, DRR) are standard in room acoustics, and the SDE evaluation with an open-source baseline is externally grounded. However, the significance of the paper is directly tied to the validity of the simulator-as-ground-truth assumption. The evidence for that assumption is currently thin: one measured room, no error bars, and no validation of the GWA rooms. The oracle SDE result, while encouraging, only demonstrates that the simulated data are internally consistent and learnable, not that they are faithful to any real acoustic space. The paper is therefore a solid challenge description whose central legitimacy claim requires additional support or more careful framing.
major comments (3)
- [Section IV.A, Table III] The claim that 'the low errors indicate that the rooms simulated with Treble software in our challenge dataset matches real-world data' is load-bearing, but it rests on a single comparison: 20 measured RIRs from Room_0 against their simulated counterparts. No error bars, confidence intervals, or per-position breakdowns are provided, and the room dimensions in the dataset range from 1.75 m to 6.3 m with varying furniture and materials, so one room cannot establish accuracy across the other nine Treble rooms. The Table III numbers themselves include a 13.5% broadband T20 MAPE, a 14.5 dB DRR MSE, and a 41.1 dB DRR MSE at 4 kHz, which are not trivially 'low' without a comparison baseline. Since simulated RIRs are the ground truth for Task 1 and the training/test source for Task 2, this validation gap propagates into every downstream evaluation. Please either validate the simulator on a larger set of measured rooms, provide per-room error bars and a baseline error for reference, or substantially weaken the real-world fidelity claim.
- [Section V.A, Table IV] The oracle SDE results are presented as evidence that the challenge dataset supports accurate distance estimation, but they are not independent evidence of simulator realism. Both the oracle training data and the Task 2 test set are generated by convolving speech with RIRs from the same simulation pipeline, so low MAE on held-out simulated positions demonstrates internal consistency and task learnability, not fidelity to real rooms. The comparison with the C4DM baseline is also confounded by training-domain shift and by the distance distribution mismatch, as the authors themselves acknowledge in Section V.B. To support the realism claim, the SDE evaluation should include at least a small set of measured RIRs from a real room (e.g., the Room_0 recordings) as an external test set, or the paper should explicitly state that the oracle result only shows learnability within the simulated domain.
- [Section III.A] The ten GWA rooms are included in the challenge dataset, but no validation evidence is provided for them. They are generated by a different simulation method (hybrid FDTD and ray tracing) than the Treble rooms, so the Room_0 validation does not transfer automatically. Please state explicitly that the GWA rooms are unvalidated, or provide a comparable accuracy assessment. This matters because Task 1 participants are required to generate RIRs for all twenty rooms, and the evaluation treats all simulated RIRs as ground truth.
minor comments (5)
- [Section V.B] In the sentence 'ensuring the training process utilizes most of the the dataset', 'the the' should be 'the'.
- [Table III caption] There is an extra space before the period in the caption: 'Quantitative Error of the Treble Simulated vs. Measured Room .'
- [Reference [16]] The title contains a typo: 'discontinous' should be 'discontinuous'.
- [Section III.A] The term 'monoaural' appears in several places and should be 'monaural'.
- [Section IV.A] The sentence 'We set Treble's wave-based simulations as the upper bound of the performance of generative RIR systems' is clear in intent, but it should be clarified that this is an upper bound under the assumption that the Treble simulator is accurate; otherwise it risks circularity.
Circularity Check
No significant circularity; the simulator validation uses external measured data, and the oracle SDE results are an internal upper bound, not evidence of real-world fidelity.
full rationale
The paper's derivation chain is not circular. The central external claim—that Treble's wave-based simulator produces realistic RIRs—rests on a direct comparison against measured RIRs in Room_0 (Table III), which is independent of the challenge dataset and is not fitted to any prediction. The oracle SDE results in Table IV are explicitly presented as an upper bound for participants and demonstrate internal consistency of the simulated dataset; the paper does not use the low oracle MAE as evidence that the simulated rooms match real-world data. The metrics in Eqs. (1)-(4) are standard objective measures and are not constructed from the quantities they evaluate. The generalization from the single validated Room_0 to the other simulated rooms is an extrapolation and a limitation, but it is not a circular reduction: the statement that the simulated rooms 'match real-world data' is an inductive claim, not an identity forced by definition. Reference [14] is a vendor validation page, and several authors are affiliated with the vendor, but it is not load-bearing because Table III supplies direct external measured evidence. No fitted parameter is renamed as a prediction, and no self-citation chain forces the paper's conclusions.
Assumptions & free parameters
assumptions (5)
- domain assumption Simulated RIRs from the Treble wave-based solver (DG method) are accurate enough to serve as ground truth for the challenge.
- domain assumption The GWA simulated rooms are reliable ground truth.
- domain assumption Convolving RIRs with VCTK speech yields realistic reverberant speech for training and evaluating SDE.
- domain assumption The fixed baseline SDE model [19] is a suitable evaluation instrument; using it unmodified isolates data quality effects.
- domain assumption The two enrollment scenarios (center-to-corner and corner-to-corner) represent realistic sparse measurement settings.
Cite this review
Pith. "Pith review of Generative Data Augmentation Challenge: Synthesis of Room Acoustics for Speaker Distance Estimation." pith.science (2026). https://pith.science/paper/WLQM7452
@misc{pith2026250113250,
author = {Pith},
title = {Pith review of: Generative Data Augmentation Challenge: Synthesis of Room Acoustics for Speaker Distance Estimation},
year = {2026},
howpublished = {\url{https://pith.science/paper/WLQM7452}},
note = {Machine review of arXiv:2501.13250}
}
read the original abstract
This paper describes the synthesis of the room acoustics challenge as a part of the generative data augmentation workshop at ICASSP 2025. The challenge defines a unique generative task that is designed to improve the quantity and diversity of the room impulse responses dataset so that it can be used for spatially sensitive downstream tasks: speaker distance estimation. The challenge identifies the technical difficulty in measuring or simulating many rooms' acoustic characteristics precisely. As a solution, it proposes generative data augmentation as an alternative that can potentially be used to improve various downstream tasks. The challenge website, dataset, and evaluation code are available at https://sites.google.com/view/genda2025.
Figures
Forward citations
Cited by 1 Pith paper
-
DiffusionRIR: Room Impulse Response Interpolation using Diffusion Models
A diffusion inpainting model reconstructs missing room impulse responses of microphone arrays from measured neighbors, outperforming spline interpolation in simulated rooms.
Reference graph
Works this paper leans on
-
[1]
A binaural room impulse response database for the evaluation of dereverberation algorithms,
M. Jeub, M. Schafer, and P. Vary, “A binaural room impulse response database for the evaluation of dereverberation algorithms,” in 2009 16th International Conference on Digital Signal Processing , pp. 1–5, IEEE, 2009
work page 2009
-
[2]
Estimation of room acoustic parameters: The ACE challenge,
J. Eaton, N. D. Gaubitch, A. H. Moore, and P. A. Naylor, “Estimation of room acoustic parameters: The ACE challenge,” IEEE/ACM Trans- actions on Audio, Speech, and Language Processing , vol. 24, no. 10, pp. 1681–1693, 2016
work page 2016
-
[3]
Statistics of natural reverberation enable perceptual separation of sound and space,
J. Traer and J. H. McDermott, “Statistics of natural reverberation enable perceptual separation of sound and space,” Proceedings of the National Academy of Sciences , vol. 113, no. 48, pp. E7856–E7865, 2016
work page 2016
-
[4]
Soundspaces: Audio-visual navigation in 3d environments,
C. Chen, U. Jain, C. Schissler, S. V . A. Gari, Z. Al-Halah, V . K. Ithapu, P. Robinson, and K. Grauman, “Soundspaces: Audio-visual navigation in 3d environments,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16, pp. 17–36, Springer, 2020
work page 2020
-
[5]
S. Koyama, T. Nishida, K. Kimura, T. Abe, N. Ueno, and J. Brunnström, “MeshRIR: A dataset of room impulse responses on meshed grid points for evaluating sound field analysis and synthesis methods,” in 2021 IEEE workshop on applications of signal processing to audio and acoustics (WASPAA), pp. 1–5, IEEE, 2021
work page 2021
-
[6]
GW A: A large high-quality acoustic dataset for audio processing,
Z. Tang, R. Aralikatti, A. J. Ratnarajah, and D. Manocha, “GW A: A large high-quality acoustic dataset for audio processing,” in ACM SIGGRAPH 2022 Conference Proceedings , pp. 1–9, 2022
work page 2022
-
[7]
G. Götz, S. J. Schlecht, and V . Pulkki, “A dataset of higher-order ambisonic room impulse responses and 3d models measured in a room with varying furniture,” in 2021 Immersive and 3D Audio: from Architecture to Automotive (I3DA) , pp. 1–8, IEEE, 2021
work page 2021
-
[8]
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo,et al., “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4015–4026, 2023
work page 2023
Show all 21 references
-
[9]
CV AE-GAN: Fine-Grained Image Generation through Asymmetric Training,
J. Bao, D. Chen, F. Wen, H. Li, and G. Hua, “CV AE-GAN: Fine-Grained Image Generation through Asymmetric Training,” in IEEE International Conference on Computer Vision (ICCV) , pp. 2764–2773, 2017
2017
-
[10]
Generative Hierarchical Features from Synthesizing Images,
Y . Xu, Y . Shen, J. Zhu, C. Yang, and B. Zhou, “Generative Hierarchical Features from Synthesizing Images,” in IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR) , pp. 4430–4430, 2021
2021
-
[11]
The potential of neural speech synthesis-based data augmentation for personalized speech en- hancement,
A. Kuznetsova, A. Sivaraman, and M. Kim, “The potential of neural speech synthesis-based data augmentation for personalized speech en- hancement,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1–5, IEEE, 2023
2023
-
[12]
Speech recognition with augmented synthesized speech,
A. Rosenberg, Y . Zhang, B. Ramabhadran, Y . Jia, P. Moreno, Y . Wu, and Z. Wu, “Speech recognition with augmented synthesized speech,” in 2019 IEEE automatic speech recognition and understanding workshop (ASRU), pp. 996–1002, IEEE, 2019
2019
-
[13]
Data augmentation approaches in natural language processing: A survey,
B. Li, Y . Hou, and W. Che, “Data augmentation approaches in natural language processing: A survey,” Ai Open , vol. 3, pp. 71–90, 2022
2022
-
[14]
Treble technologies
“Treble technologies.” https://docs.treble.tech/validation
-
[15]
Time-domain room acoustic simulations with extended-reacting porous absorbers using the discontinuous Galerkin method,
F. Pind, C.-H. Jeong, A. P. Engsig-Karup, J. S. Hesthaven, and J. Strømann-Andersen, “Time-domain room acoustic simulations with extended-reacting porous absorbers using the discontinuous Galerkin method,” The Journal of the Acoustical Society of America , vol. 148, no. 5, pp....
2020
-
[16]
Massively parallel nodal discontinous Galerkin finite element method simulator for room acoustics,
A. Melander, E. Strøm, F. Pind, A. P. Engsig-Karup, C.-H. Jeong, T. Warburton, N. Chalmers, and J. S. Hesthaven, “Massively parallel nodal discontinous Galerkin finite element method simulator for room acoustics,” The International Journal of High Performance Computing Applica...
2024
-
[17]
3382-2, Acoustics—Measurement of room acoustic parameters— Part 2: Reverberation time in ordinary rooms
I. 3382-2, Acoustics—Measurement of room acoustic parameters— Part 2: Reverberation time in ordinary rooms . International Organization for Standardization, 2008
2008
-
[18]
New method of measuring reverberation time,
M. R. Schroeder, “New method of measuring reverberation time,” The Journal of the Acoustical Society of America , vol. 37, no. 6_Supplement, pp. 1187–1188, 1965
1965
-
[19]
Speaker distance estimation in enclosures from single-channel audio,
M. Neri, A. Politis, D. Krause, M. Carli, and T. Virtanen, “Speaker distance estimation in enclosures from single-channel audio,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024
2024
-
[20]
Database of omnidirectional and B-format room impulse responses,
R. Stewart and M. Sandler, “Database of omnidirectional and B-format room impulse responses,” in 2010 IEEE International Conference on Acoustics, Speech and Signal Processing , pp. 165–168, IEEE, 2010
2010
-
[21]
CSTR VCTK Corpus: English multi-speaker corpus for CSTR V oice Cloning Toolkit (version 0.92),
J. Yamagishi, C. Veaux, and K. MacDonald, “CSTR VCTK Corpus: English multi-speaker corpus for CSTR V oice Cloning Toolkit (version 0.92),” Nov 2019
2019
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.