REVIEW 3 major objections 5 minor 47 references
VR-PTOLEMAIC: A Virtual Environment for the Perceptual Testing of Spatial Audio Algorithms
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read VR-PTOLEMAIC claims a virtual-reality MUSHRA platform using measured room impulse responses at 25 positions supports spatial audio evaluation and yields perceptual feedback comparable to traditional setups.
desk verdict A solid, buildable VR-MUSHRA system paper whose headline comparability claim outruns a validation with no non-VR baseline and no inferential statistics. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the measured second-order Ambisonic room impulse response paired with a head-tracked binaural decoder. Each of the 25 virtual seats maps to one measured A-RIR; the audio engine convolves the selected anechoic sample with either that measured response or a reconstructed response, producing a multichannel Ambisonic stream. A real-time binaural decoder rotates the stream according to the listener's head orientation and delivers it over closed headphones, so the listener's head movement becomes part of the evaluation. This chain ties every MUSHRA rating to a specific room position and a specific head orientation, which is what makes the subjective test spatially anchored.
What would settle it
Run the same listeners and the same stimulus set through two administrations: the VR platform as described, and a conventional non-interactive binaural version with fixed head position and identical headphone playback. If the ranking between the hidden reference, the low-pass anchor, and the two reconstruction algorithms changes across administrations, or if hidden-reference scores differ by more than the internal consistency of repeated VR trials, then the claimed comparability to traditional setups is refuted.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that the MUSHRA protocol—a multi-stimulus test with a hidden reference and an anchor—survives translation into an interactive VR room. The authors build the test on measured second-order Ambisonic room impulse responses at 25 chair positions, so the reference and hidden reference are the actual acoustics of a real seminar room, while the conditions under test are reconstructed responses generated by different sound field reconstruction algorithms. The validation results show consistent separation among conditions on all four rating attributes, and participant questionnaire responses were generally positive, with only mild discomfort reported from prolonged headset wear. The paper therefore concludes that the platform effectively supports the evaluation of spatial audio algorithms and offers perceptual feedback comparable to traditional setups.
Load-bearing premise
The load-bearing premise is that head-tracked binaural decoding of measured second-order Ambisonic room impulse responses preserves the same perceptual quality differences among algorithms as a real listening room, so VR-based MUSHRA ratings are comparable to traditional listening-test ratings; the paper presents no direct non-VR comparison to confirm that equivalence.
Editorial extensions
If this is right
- If the central claim holds, spatial audio algorithms can be compared perceptually at 25 different room positions using one measured room impulse response acquisition, without repositioning loudspeakers or microphones between trials.
- The built-in tracking data lets evaluators see where listeners moved and how long they stayed at each seat, adding a behavioral correlate to quality scores such as localizability.
- Because the audio pipeline runs in real time on an all-in-one headset and a laptop, standardized spatial audio listening tests could be run outside an acoustically treated listening room.
- The clear separation between hidden reference, anchor, and reconstructed stimuli in the reported MUSHRA results supports using the platform for future comparisons of sound field reconstruction algorithms.
Reading between the lines
- The paper's comparison to traditional setups is qualitative; a stricter claim would need a within-subjects equivalence test between VR and non-VR administration of the same stimulus set, and the paper does not report one.
- Because head orientation is tracked, the platform could be extended to evaluate orientation-dependent attributes such as externalization and directional fidelity, which a fixed-headphone MUSHRA cannot capture.
- The 25 discrete seats suggest a natural next step of interpolating between measured responses to evaluate moving sources or walk-through auralization, which the current system does not implement.
- The planned open release would let other groups swap in their own measured responses, turning the system into a generic perceptual testbed for any Ambisonic capture; that generality is left implicit.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents VR-PTOLEMAIC, a virtual-reality system for perceptual testing of spatial audio algorithms. The platform couples a Unity-based VR application with a Max-based audio processor via OSC: users can move among 25 predefined listening positions of a reconstructed seminar room, select stimuli through a virtual MUSHRA interface, and hear measured or reconstructed second-order Ambisonic room impulse responses encoded binaurally with head-tracked SPARTA ambiBIN decoding. The system also logs head position, rotation, and teleportation behaviour. The authors validate the platform with a listening test in which 15 participants (11 after MUSHRA screening) rated four stimuli against a reference across four attributes (basic audio quality, localizability, spatial quality, timbral quality), and they report generally positive usability feedback and exploratory behavioural tracking results. The paper concludes that the platform effectively supports spatial audio evaluation and offers perceptual feedback comparable to traditional setups.
Significance. If the comparability claim were supported, the paper would make a useful practical contribution: a concrete, buildable VR implementation of MUSHRA with measured room impulse responses at 25 positions, real-time head-tracked binaural rendering, hidden-reference and anchor conditions, and behavioural tracking. The system description is clear enough to be reproducible, and the use of the HOMULA-RIR dataset anchors the work in real measurements. However, the validation as reported does not establish the central conclusion of equivalence with traditional testing: there is no non-VR baseline condition, no inferential statistical analysis, and the usability evidence is qualitative. The tool is promising, but the evidence presented is preliminary and the comparability claim outstrips the data.
major comments (3)
- [Abstract and §5 (Conclusion)] The claim that the platform offers 'perceptual feedback comparable to traditional setups' is not supported by the validation described in §3. The study contains no non-VR or real-room baseline condition, so there is no evidence that MUSHRA ratings obtained in the VR environment approximate ratings from a conventional listening test. The observed ability to distinguish the low-pass anchor and the hidden reference is compatible with VR-specific distortions introduced by the non-individualized SPARTA ambiBIN decoding, head-tracked playback, and free navigation among 25 positions. Either add a comparative baseline condition (e.g., the same MUSHRA test with fixed orientation and non-head-tracked binaural or loudspeaker reproduction) or remove and explicitly qualify the comparability claim.
- [§3, §4, Fig. 5] The manuscript asserts that participants 'were able to distinguish between the different reconstruction methods with a high degree of consistency', but no inferential statistics are provided. Figure 5 shows aggregated MUSHRA ratings without confidence intervals, error bars, effect sizes, or significance tests, and the number of valid participants after screening is only 11. Without a per-participant or per-item statistical analysis, the discrimination claim is not quantitatively established. Report condition means and confidence intervals and, if appropriate, a within-subjects test or effect-size measure.
- [§4 (Results and Discussion)] The usability and immersivity conclusions rest on an informal post-session questionnaire ('Most users described the system as intuitive...'). No rating scales, no item-level results, no quantification of agreement, and no analysis of the reported mild discomfort are provided. With 11 or 15 participants, this evidence supports only anecdotal impressions, not the general statement that the system 'effectively supports the assessment of spatial audio algorithms'. The claims should be scaled back to what the qualitative data can sustain, or the questionnaire should be presented as a formal usability instrument with defined scales.
minor comments (5)
- [Header] The convention header reads '22th – 26th June 2025'; this should be '22nd – 26th June 2025'.
- [§2, device name] The text says 'Meta 3 Quest 2 VR-headset'; the correct product name is Meta Quest 2.
- [§3.1, anchor description] The low-pass anchor is described as 'Low audio quality solution Slp' with a cutoff frequency of3.5 kHz; add a space after 'of' and use a non-breaking space in '3.5 kHz'.
- [§3.1, attribute scales] The Localizability attribute is described with the scale 'More difficult - Easier', which is ambiguous: clarify whether higher ratings correspond to easier or more difficult localizability, and ensure the plot in Fig. 5 uses the same polarity.
- [References] Several references are missing publisher or venue details (e.g., [12], [35]), and the formatting of reference [20] appears inconsistent. Please normalize the bibliography.
Circularity Check
No circularity identified: the platform's usability claim is an empirical result, not a derivation that reduces to its own inputs.
full rationale
The paper describes VR-PTOLEMAIC, a VR-based MUSHRA testing platform, and validates it through a listening experiment. There is no equation, fitted parameter, or model-derived prediction in the paper; the central claim is that the platform 'effectively supports the evaluation of spatial audio algorithms, offering perceptual feedback comparable to traditional setups' (Sec. 5). This claim is supported by an empirical usability study, not by a derivation from assumptions. The validation compares several SFR algorithms, including some from the authors' prior work ([9], [30]), but the paper explicitly states that the SFR discrimination finding is 'not central to this study's validation' (Sec. 4). The platform's usability is assessed via hidden-reference screening, user teleportation tracking, and subjective questionnaire feedback, none of which is equivalent by construction to the claim being made. Self-citations to the HOMULA-RIR dataset [30] and parametric virtual miking [9] are to published, externally reproducible resources (a measured-RIR dataset and a peer-reviewed method) used as test components, not as a load-bearing justification for the platform's validity. The conclusion about comparability to traditional setups is under-evidenced because no non-VR baseline condition was run, but this is a correctness/validity limitation, not circular reasoning. The manuscript contains no circular definition, no fitted input renamed as prediction, and no uniqueness assertion imported from the authors' prior work. Thus, no significant circularity is present.
Assumptions & free parameters
assumptions (4)
- domain assumption Binaural rendering of second-order Ambisonic RIRs with head tracking accurately represents the spatial audio quality of the real room.
- domain assumption The HOMULA-RIR dataset's measured A-RIRs at the 25 chair positions are accurate and representative for the sound field evaluation.
- domain assumption MUSHRA methodology transfers unmodified to a VR environment.
- domain assumption Participants' prior experience in music and no hearing impairments make them suitable assessors.
Cite this review
Pith. "Pith review of VR-PTOLEMAIC: A Virtual Environment for the Perceptual Testing of Spatial Audio Algorithms." pith.science (2026). https://pith.science/paper/M5YPSUZX
@misc{pith2026250800501,
author = {Pith},
title = {Pith review of: VR-PTOLEMAIC: A Virtual Environment for the Perceptual Testing of Spatial Audio Algorithms},
year = {2026},
howpublished = {\url{https://pith.science/paper/M5YPSUZX}},
note = {Machine review of arXiv:2508.00501}
}
read the original abstract
The perceptual evaluation of spatial audio algorithms is an important step in the development of immersive audio applications, as it ensures that synthesized sound fields meet quality standards in terms of listening experience, spatial perception and auditory realism. To support these evaluations, virtual reality can offer a powerful platform by providing immersive and interactive testing environments. In this paper, we present VR-PTOLEMAIC, a virtual reality evaluation system designed for assessing spatial audio algorithms. The system implements the MUSHRA (MUlti-Stimulus test with Hidden Reference and Anchor) evaluation methodology into a virtual environment. In particular, users can position themselves in each of the 25 simulated listening positions of a virtually recreated seminar room and evaluate simulated acoustic responses with respect to the actually recorded second-order ambisonic room impulse responses, all convolved with various source signals. We evaluated the usability of the proposed framework through an extensive testing campaign in which assessors were asked to compare the reconstruction capabilities of various sound field reconstruction algorithms. Results show that the VR platform effectively supports the assessment of spatial audio algorithms, with generally positive feedback on user experience and immersivity.
Figures
Reference graph
Works this paper leans on
-
[1]
INTRODUCTION Spatial audio is a trending research field focused on un- derstanding, recreating, and optimizing 3D auditory envi- ronments [1], providing listeners with an immersive and interactive listening experience. Applications range from virtual and augmented reality [2] to gaming [3], telecon- ferencing and remote concerts [4], where precise sound l...
work page Pith review arXiv 2025
-
[2]
EV ALUA TION SYSTEM Perceptual tests for spatial audio in VR require a mul- timodal environment that involves visual content, audio content and interactivity in order to maintain in the asses- sor a sense of presence and immersivity during the eval- uation procedure. The following section presents a VR- based evaluation system for assessing spatial audio ...
work page 2025
-
[3]
SYSTEM V ALIDA TION To validate the proposed VR-based evaluation framework, we used SFR as a test scenario and conducted listening tests within the VR environment. SFR is a fundamen- tal task in spatial audio, focused on estimating the pres- sure field in regions where direct measurements cannot be obtained. Capturing the acoustic field over a broad spati...
work page 2025
-
[4]
RESULTS AND DISCUSSION The proposed VR-based system was evaluated consider- ing its ability to facilitate subjective spatial audio assess- ment and its impact on user interaction. Fifteen partici- pants (13 male, 2 female) took part in the listening tests, with an average age of 26.9 years (SD = 4.0). All but one held a university degree, and all had prio...
work page 2025
-
[5]
CONCLUSION AND FUTURE WORK We introduced a VR-based system for evaluating spa- tial audio reproduction, integrating real acoustic measure- ments into an immersive and interactive environment. The system was assessed through structured listening tests, which included a comparative evaluation of SFR algo- rithms, and focused on perceptual attributes such as...
-
[6]
Telecommunications of the Future
ACKNOWLEDGMENTS This work has been funded by ”REPERTORIUM project. Grant agreement number 101095065. Horizon Europe. Cluster II. Culture, Creativity and Inclusive Society. Call HORIZON-CL2-2022-HERITAGE-01-02.”. This work was partially supported by the European Union – Next Generation EU under the Italian National Recovery and Resilience Plan (NRRP), Miss...
work page 2022
-
[7]
Microphone array processing for paramet- ric spatial audio techniques,
A. Politis, “Microphone array processing for paramet- ric spatial audio techniques,” 2016
work page 2016
-
[8]
Computa- tional and optimization design in geometric acous- tics,
A. Bassuet, D. Rife, and L. Dellatorre, “Computa- tional and optimization design in geometric acous- tics,” Building Acoustics , vol. 21, no. 1, pp. 75–85, 2014
work page 2014
Show all 47 references
-
[9]
Audio aug- mented reality: A systematic review of technologies, applications, and future research directions,
J. Yang, A. Barde, and M. Billinghurst, “Audio aug- mented reality: A systematic review of technologies, applications, and future research directions,” journal of the audio engineering society , vol. 70, no. 10, pp. 788–809, 2022
2022
-
[10]
3d sound spatialization with game engines: the virtual acous- tics performance of a game engine and a middleware for interactive audio design,
H. B. Fırat, L. Maffei, and M. Masullo, “3d sound spatialization with game engines: the virtual acous- tics performance of a game engine and a middleware for interactive audio design,” Virtual Reality, vol. 26, no. 2, pp. 539–558, 2022
2022
-
[11]
A virtual symphony orchestra for studies on concert hall acoustics,
J. P ¨atynen, “A virtual symphony orchestra for studies on concert hall acoustics,” 2011
2011
-
[12]
Spatial audio signal pro- cessing for binaural reproduction of recorded acoustic scenes–review and challenges,
B. Rafaely, V . Tourbabin, E. Habets, Z. Ben-Hur, H. Lee, H. Gamper, L. Arbel, L. Birnie, T. Abhaya- pala, and P. Samarasinghe, “Spatial audio signal pro- cessing for binaural reproduction of recorded acoustic scenes–review and challenges,” Acta Acustica, vol. 6, p. 47, 2022
2022
-
[13]
An overview of machine learning and other data- based methods for spatial audio capture, processing, and reproduction,
M. Cobos, J. Ahrens, K. Kowalczyk, and A. Poli- tis, “An overview of machine learning and other data- based methods for spatial audio capture, processing, and reproduction,” vol. 2022, no. 1, p. 10
2022
-
[14]
A comparative analysis of the directional sound radi- ation of historical violins,
M. Pezzoli, A. Canclini, F. Antonacci, and A. Sarti, “A comparative analysis of the directional sound radi- ation of historical violins,” The Journal of the Acous- tical Society of America, vol. 152, no. 1, pp. 354–367, 2022
2022
-
[15]
Soundfield reconstruction in reverber- ant rooms based on compressive sensing and image- source models of early reflections,
S. Damiano, F. Borra, A. Bernardini, F. Antonacci, and A. Sarti, “Soundfield reconstruction in reverber- ant rooms based on compressive sensing and image- source models of early reflections,” in 2021 IEEE Workshop on Applications of Signal Processing to Au- dio and Acoustics (...
2021
-
[16]
A parametric approach to virtual miking for sources of arbitrary directivity,
M. Pezzoli, F. Borra, F. Antonacci, S. Tubaro, and A. Sarti, “A parametric approach to virtual miking for sources of arbitrary directivity,” vol. 28, pp. 2333–
-
[17]
but also on subjective experiments that address how listeners actually experience the audio, assessing percep- tual attributes such as localization accuracy, spatial clarity, immersion, and externalization. The standard approach for subjective sound qual- ity evaluation is bas...
-
[18]
Reconstruction of sound field through diffusion models,
F. Miotello, L. Comanducci, M. Pezzoli, A. Bernar- dini, F. Antonacci, and A. Sarti, “Reconstruction of sound field through diffusion models,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 1476–1480. ISSN: 2379-190X
2024
-
[19]
A zero-shot physics-informed dictionary learning ap- proach for sound field reconstruction,
S. Damiano, F. Miotello, M. Pezzoli, A. Bernardini, F. Antonacci, A. Sarti, and T. Van Waterschoot, “A zero-shot physics-informed dictionary learning ap- proach for sound field reconstruction,” in ICASSP 2025-2025 IEEE International Conference on Acous- tics, Speech and Signal...
2025
-
[20]
Physics- informed neural network for volumetric sound field reconstruction of speech signals,
M. Olivieri, X. Karakonstantis, M. Pezzoli, F. An- tonacci, A. Sarti, and E. Fernandez-Grande, “Physics- informed neural network for volumetric sound field reconstruction of speech signals,” vol. 2024, no. 1
2024
-
[21]
Generative models for sound field reconstruction,
E. Fernandez-Grande, X. Karakonstantis, D. Caviedes-Nozal, and P. Gerstoft, “Generative models for sound field reconstruction,” The Journal of the Acoustical Society of America , vol. 153, no. 2, pp. 1179–1190, 2023
2023
-
[22]
Zacharov, Sensory evaluation of sound
N. Zacharov, Sensory evaluation of sound. CRC Press, 2018
2018
-
[23]
Towards hrtf personalization using denoising diffusion models,
J. C. Albarrac ´ın S´anchez, L. Comanducci, M. Pezzoli, and F. Antonacci, “Towards hrtf personalization using denoising diffusion models,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1–5, 2025
2025
-
[24]
Joint sam- pling theory and subjective investigation of plane- wave and spherical harmonics formulations for bin- aural reproduction,
Z. Ben-Hur, J. Sheaffer, and B. Rafaely, “Joint sam- pling theory and subjective investigation of plane- wave and spherical harmonics formulations for bin- aural reproduction,” Applied Acoustics , vol. 134, pp. 138–144, 2018
2018
-
[25]
Subjective and objective evaluations of a scattered sound field in a scale model opera house,
J. K. Ryu and J. Y . Jeon, “Subjective and objective evaluations of a scattered sound field in a scale model opera house,” The Journal of the Acoustical Society of America, vol. 124, no. 3, pp. 1538–1549, 2008
2008
-
[26]
Bs. 1116-3,“,
R. ITU-R, “Bs. 1116-3,“,” Methods for the sub- jective assesment of small impairments in audio systems, ” International Telecommunication Union- Radiocommunication Sector, 2015
2015
-
[27]
ITU-R, “1534-3,” Method for the subjective assess- ment of intermediate quality level of audio systems , 2015
B. ITU-R, “1534-3,” Method for the subjective assess- ment of intermediate quality level of audio systems , 2015. 11th Convention of the European Acoustics Association M´alaga, Spain • 22th – 26th June 2025 •
2015
-
[28]
Itu-t recommendation p. 800. methods for ob- jective and subjective assessment of quality,
T. Itu, “Itu-t recommendation p. 800. methods for ob- jective and subjective assessment of quality,” 1996
1996
-
[29]
Method for the subjective assessment of intermediate quality level of audio systems,
B. Series, “Method for the subjective assessment of intermediate quality level of audio systems,” Inter- national Telecommunication Union Radiocommuni- cation Assembly, vol. 2, 2014
2014
-
[30]
HOMULA-RIR: A room impulse response dataset for teleconferencing and spatial audio applications ac- quired through higher-order microphones and uniform linear microphone arrays,
F. Miotello, P. Ostan, M. Pezzoli, L. Coman- ducci, A. Bernardini, F. Antonacci, and A. Sarti, “HOMULA-RIR: A room impulse response dataset for teleconferencing and spatial audio applications ac- quired through higher-order microphones and uniform linear microphone arrays,” in...
-
[31]
Scaling sound quality using models for paired-comparison and ranking data,
F. Wickelmaier, N. Umbach, K. Sergin, and S. Choisel, “Scaling sound quality using models for paired-comparison and ranking data,” in Conference Paper , DAGA 2012 Congress 38th German Annual Conference on Acoustics, 2012
2012
-
[32]
Direct and indirect listen- ing test methods—a discussion based on audio-visual spatial coherence experiments,
C. Pike and H. Stenzel, “Direct and indirect listen- ing test methods—a discussion based on audio-visual spatial coherence experiments,” in Audio Engineering Society Convention 143 , Audio Engineering Society, 2017
2017
-
[33]
Free-field study on auditory localiza- tion and discrimination performance in older adults,
C. Freigang, K. Schmiedchen, I. Nitsche, and R. R ¨ubsamen, “Free-field study on auditory localiza- tion and discrimination performance in older adults,” vol. 232, no. 4, pp. 1157–1172
-
[34]
On the use of subjective hrtf evaluations for creating global percep- tual similarity metrics of assessors and assessees.,
A. Andreopoulou and B. F. Katz, “On the use of subjective hrtf evaluations for creating global percep- tual similarity metrics of assessors and assessees.,” in ICAD, pp. 13–20, 2015
2015
-
[35]
As- sessing ambisonics sound source localization by means of virtual reality and gamification tools,
E. Medina, R. Viveros-Mu ˜noz, and F. Otondo, “As- sessing ambisonics sound source localization by means of virtual reality and gamification tools,” Ap- plied Sciences, vol. 14, no. 17, p. 7986, 2024
2024
-
[36]
Audio quality evaluation in virtual reality: multiple stimulus ranking with behavior tracking,
O. Rummukainen, T. Robotham, S. J. Schlecht, A. Plinge, J. Herre, and E. A. Habels, “Audio quality evaluation in virtual reality: multiple stimulus ranking with behavior tracking,” inAudio Engineering Society Conference: 2018 AES International Conference on Audio for Virtual a...
2018
-
[37]
A flexible software tool for perceptual evaluation of audio mate- rial and vr environments,
S. Gorzynski, N. Kaplanis, and S. Bech, “A flexible software tool for perceptual evaluation of audio mate- rial and vr environments,” in Audio Engineering Soci- ety Convention 149, Audio Engineering Society, 2020
2020
-
[38]
Am- plitude matching for multizone sound field control,
T. Abe, S. Koyama, N. Ueno, and H. Saruwatari, “Am- plitude matching for multizone sound field control,” IEEE/ACM Transactions on Audio, Speech, and Lan- guage Processing, vol. 31, pp. 656–669, 2022
2022
-
[39]
Object-based six-degrees-of-freedom ren- dering of sound scenes captured with multiple am- bisonic receivers,
L. McCormack, A. Politis, T. McKenzie, C. Hold, and V . Pulkki, “Object-based six-degrees-of-freedom ren- dering of sound scenes captured with multiple am- bisonic receivers,” vol. 70, no. 5, pp. 355–372. Pub- lisher: Audio Engineering Society
-
[40]
Re- construction of reverberant sound fields over large spatial domains,
A. Figueroa-Duran and E. Fernandez-Grande, “Re- construction of reverberant sound fields over large spatial domains,” The Journal of the Acoustical So- ciety of America, vol. 157, no. 1, pp. 180–190, 2025
2025
-
[41]
The hisstools impulse response toolbox: Convolution for the masses,
A. Harker and P. A. Tremblay, “The hisstools impulse response toolbox: Convolution for the masses,” in Proc. of the international computer music conference , pp. 148–155, The International Computer Music As- sociation, 2012
2012
-
[42]
Sparta & compass: Real-time implementations of linear and parametric spatial audio reproduction and processing methods,
L. McCormack and A. Politis, “Sparta & compass: Real-time implementations of linear and parametric spatial audio reproduction and processing methods,” in AES International Conference on Immersive and Interactive Audio, Audio Engineering Society, 2019
2019
-
[43]
Damiano, F
S. Damiano, F. Miotello, M. Pezzoli, A. Bernardini, F. Antonacci, A. Sarti, and T. Waterschoot, A Zero- Shot Physics-Informed Dictionary Learning Approach for Sound Field Reconstruction
-
[44]
EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation,
J. Richter, Y .-C. Wu, S. Krenn, S. Welker, B. Lay, S. Watanabe, A. Richard, and T. Gerkmann, “EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation,” in Inter- speech 2024, pp. 4873–4877, ISCA
2024
-
[45]
Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, in- sights, and applications,
B. Li, X. Liu, K. Dinesh, Z. Duan, and G. Sharma, “Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, in- sights, and applications,” vol. 21, no. 2, pp. 522–535
-
[47]
A spatial audio qual- ity inventory (saqi),
A. Lindau, V . Erbes, S. Lepa, H.-J. Maempel, F. Brinkman, and S. Weinzierl, “A spatial audio qual- ity inventory (saqi),” Acta Acustica united with Acus- tica, vol. 100, no. 5, pp. 984–994, 2014
2014
-
[2348]
Conference Name: IEEE/ACM Transactions on Audio, Speech, and Language Processing
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.