Pith. sign in

REVIEW 3 major objections 5 minor 47 references

VR-PTOLEMAIC: A Virtual Environment for the Perceptual Testing of Spatial Audio Algorithms

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read VR-PTOLEMAIC claims a virtual-reality MUSHRA platform using measured room impulse responses at 25 positions supports spatial audio evaluation and yields perceptual feedback comparable to traditional setups.

desk verdict A solid, buildable VR-MUSHRA system paper whose headline comparability claim outruns a validation with no non-VR baseline and no inferential statistics. read the letter →

arxiv 2508.00501 v1 pith:M5YPSUZX submitted 2025-08-01 eess.AS cs.SD

classification eess.AScs.SD
keywords virtualacousticsMUSHRArealityperceptualevaluationspatialaudiosoundfieldreconstructionbinauralrenderingbehavioraltracking
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that a virtual-reality implementation of the MUSHRA listening test can serve as a practical platform for perceptual evaluation of spatial audio algorithms. The platform recreates a seminar room in VR, constrains listening to 25 measured positions, and renders audio by convolving anechoic source samples with measured or reconstructed second-order Ambisonic room impulse responses, decoded binaurally with head tracking. In a validation session, 15 listeners rated four sound field reconstruction conditions (hidden reference, low-pass anchor, and two reconstruction methods) on basic audio quality, localizability, spatial quality, and timbral quality, and their ratings separated the conditions consistently. On this evidence the authors conclude that the platform supports spatial audio assessment and provides perceptual feedback comparable to traditional setups. If true, standardized spatial audio evaluation could move out of the treated listening room into a lightweight, headset-based environment.

What carries the argument

The load-bearing object is the measured second-order Ambisonic room impulse response paired with a head-tracked binaural decoder. Each of the 25 virtual seats maps to one measured A-RIR; the audio engine convolves the selected anechoic sample with either that measured response or a reconstructed response, producing a multichannel Ambisonic stream. A real-time binaural decoder rotates the stream according to the listener's head orientation and delivers it over closed headphones, so the listener's head movement becomes part of the evaluation. This chain ties every MUSHRA rating to a specific room position and a specific head orientation, which is what makes the subjective test spatially anchored.

What would settle it

Run the same listeners and the same stimulus set through two administrations: the VR platform as described, and a conventional non-interactive binaural version with fixed head position and identical headphone playback. If the ranking between the hidden reference, the low-pass anchor, and the two reconstruction algorithms changes across administrations, or if hidden-reference scores differ by more than the internal consistency of repeated VR trials, then the claimed comparability to traditional setups is refuted.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that the MUSHRA protocol—a multi-stimulus test with a hidden reference and an anchor—survives translation into an interactive VR room. The authors build the test on measured second-order Ambisonic room impulse responses at 25 chair positions, so the reference and hidden reference are the actual acoustics of a real seminar room, while the conditions under test are reconstructed responses generated by different sound field reconstruction algorithms. The validation results show consistent separation among conditions on all four rating attributes, and participant questionnaire responses were generally positive, with only mild discomfort reported from prolonged headset wear. The paper therefore concludes that the platform effectively supports the evaluation of spatial audio algorithms and offers perceptual feedback comparable to traditional setups.

Load-bearing premise

The load-bearing premise is that head-tracked binaural decoding of measured second-order Ambisonic room impulse responses preserves the same perceptual quality differences among algorithms as a real listening room, so VR-based MUSHRA ratings are comparable to traditional listening-test ratings; the paper presents no direct non-VR comparison to confirm that equivalence.

Editorial extensions

If this is right

  • If the central claim holds, spatial audio algorithms can be compared perceptually at 25 different room positions using one measured room impulse response acquisition, without repositioning loudspeakers or microphones between trials.
  • The built-in tracking data lets evaluators see where listeners moved and how long they stayed at each seat, adding a behavioral correlate to quality scores such as localizability.
  • Because the audio pipeline runs in real time on an all-in-one headset and a laptop, standardized spatial audio listening tests could be run outside an acoustically treated listening room.
  • The clear separation between hidden reference, anchor, and reconstructed stimuli in the reported MUSHRA results supports using the platform for future comparisons of sound field reconstruction algorithms.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's comparison to traditional setups is qualitative; a stricter claim would need a within-subjects equivalence test between VR and non-VR administration of the same stimulus set, and the paper does not report one.
  • Because head orientation is tracked, the platform could be extended to evaluate orientation-dependent attributes such as externalization and directional fidelity, which a fixed-headphone MUSHRA cannot capture.
  • The 25 discrete seats suggest a natural next step of interpolating between measured responses to evaluate moving sources or walk-through auralization, which the current system does not implement.
  • The planned open release would let other groups swap in their own measured responses, turning the system into a generic perceptual testbed for any Ambisonic capture; that generality is left implicit.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents VR-PTOLEMAIC, a virtual-reality system for perceptual testing of spatial audio algorithms. The platform couples a Unity-based VR application with a Max-based audio processor via OSC: users can move among 25 predefined listening positions of a reconstructed seminar room, select stimuli through a virtual MUSHRA interface, and hear measured or reconstructed second-order Ambisonic room impulse responses encoded binaurally with head-tracked SPARTA ambiBIN decoding. The system also logs head position, rotation, and teleportation behaviour. The authors validate the platform with a listening test in which 15 participants (11 after MUSHRA screening) rated four stimuli against a reference across four attributes (basic audio quality, localizability, spatial quality, timbral quality), and they report generally positive usability feedback and exploratory behavioural tracking results. The paper concludes that the platform effectively supports spatial audio evaluation and offers perceptual feedback comparable to traditional setups.

Significance. If the comparability claim were supported, the paper would make a useful practical contribution: a concrete, buildable VR implementation of MUSHRA with measured room impulse responses at 25 positions, real-time head-tracked binaural rendering, hidden-reference and anchor conditions, and behavioural tracking. The system description is clear enough to be reproducible, and the use of the HOMULA-RIR dataset anchors the work in real measurements. However, the validation as reported does not establish the central conclusion of equivalence with traditional testing: there is no non-VR baseline condition, no inferential statistical analysis, and the usability evidence is qualitative. The tool is promising, but the evidence presented is preliminary and the comparability claim outstrips the data.

major comments (3)
  1. [Abstract and §5 (Conclusion)] The claim that the platform offers 'perceptual feedback comparable to traditional setups' is not supported by the validation described in §3. The study contains no non-VR or real-room baseline condition, so there is no evidence that MUSHRA ratings obtained in the VR environment approximate ratings from a conventional listening test. The observed ability to distinguish the low-pass anchor and the hidden reference is compatible with VR-specific distortions introduced by the non-individualized SPARTA ambiBIN decoding, head-tracked playback, and free navigation among 25 positions. Either add a comparative baseline condition (e.g., the same MUSHRA test with fixed orientation and non-head-tracked binaural or loudspeaker reproduction) or remove and explicitly qualify the comparability claim.
  2. [§3, §4, Fig. 5] The manuscript asserts that participants 'were able to distinguish between the different reconstruction methods with a high degree of consistency', but no inferential statistics are provided. Figure 5 shows aggregated MUSHRA ratings without confidence intervals, error bars, effect sizes, or significance tests, and the number of valid participants after screening is only 11. Without a per-participant or per-item statistical analysis, the discrimination claim is not quantitatively established. Report condition means and confidence intervals and, if appropriate, a within-subjects test or effect-size measure.
  3. [§4 (Results and Discussion)] The usability and immersivity conclusions rest on an informal post-session questionnaire ('Most users described the system as intuitive...'). No rating scales, no item-level results, no quantification of agreement, and no analysis of the reported mild discomfort are provided. With 11 or 15 participants, this evidence supports only anecdotal impressions, not the general statement that the system 'effectively supports the assessment of spatial audio algorithms'. The claims should be scaled back to what the qualitative data can sustain, or the questionnaire should be presented as a formal usability instrument with defined scales.
minor comments (5)
  1. [Header] The convention header reads '22th – 26th June 2025'; this should be '22nd – 26th June 2025'.
  2. [§2, device name] The text says 'Meta 3 Quest 2 VR-headset'; the correct product name is Meta Quest 2.
  3. [§3.1, anchor description] The low-pass anchor is described as 'Low audio quality solution Slp' with a cutoff frequency of3.5 kHz; add a space after 'of' and use a non-breaking space in '3.5 kHz'.
  4. [§3.1, attribute scales] The Localizability attribute is described with the scale 'More difficult - Easier', which is ambiguous: clarify whether higher ratings correspond to easier or more difficult localizability, and ensure the plot in Fig. 5 uses the same polarity.
  5. [References] Several references are missing publisher or venue details (e.g., [12], [35]), and the formatting of reference [20] appears inconsistent. Please normalize the bibliography.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identified: the platform's usability claim is an empirical result, not a derivation that reduces to its own inputs.

full rationale

The paper describes VR-PTOLEMAIC, a VR-based MUSHRA testing platform, and validates it through a listening experiment. There is no equation, fitted parameter, or model-derived prediction in the paper; the central claim is that the platform 'effectively supports the evaluation of spatial audio algorithms, offering perceptual feedback comparable to traditional setups' (Sec. 5). This claim is supported by an empirical usability study, not by a derivation from assumptions. The validation compares several SFR algorithms, including some from the authors' prior work ([9], [30]), but the paper explicitly states that the SFR discrimination finding is 'not central to this study's validation' (Sec. 4). The platform's usability is assessed via hidden-reference screening, user teleportation tracking, and subjective questionnaire feedback, none of which is equivalent by construction to the claim being made. Self-citations to the HOMULA-RIR dataset [30] and parametric virtual miking [9] are to published, externally reproducible resources (a measured-RIR dataset and a peer-reviewed method) used as test components, not as a load-bearing justification for the platform's validity. The conclusion about comparability to traditional setups is under-evidenced because no non-VR baseline condition was run, but this is a correctness/validity limitation, not circular reasoning. The manuscript contains no circular definition, no fitted input renamed as prediction, and no uniqueness assertion imported from the authors' prior work. Thus, no significant circularity is present.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no free parameters or new entities. Its central claim rests on domain assumptions about the fidelity of binaural ambisonic rendering, the accuracy of the HOMULA-RIR dataset, and the transferability of MUSHRA to VR, none of which are independently validated in the paper.

assumptions (4)
  • domain assumption Binaural rendering of second-order Ambisonic RIRs with head tracking accurately represents the spatial audio quality of the real room.
    The entire evaluation relies on the fidelity of the SPARTA ambiBIN decoder and head tracking; introduced in Sec. 2.3 and never validated.
  • domain assumption The HOMULA-RIR dataset's measured A-RIRs at the 25 chair positions are accurate and representative for the sound field evaluation.
    The dataset is the ground truth for comparing reconstruction algorithms; the paper takes its accuracy for granted.
  • domain assumption MUSHRA methodology transfers unmodified to a VR environment.
    The paper implements the MUSHRA UI in VR but does not test whether the VR presentation (visual scene, head tracking, teleportation) changes how participants rate audio compared to a standard lab setup.
  • domain assumption Participants' prior experience in music and no hearing impairments make them suitable assessors.
    Assessor recruitment follows common practice, but the paper gives no formal screening tests beyond self-report.

how reviews work

0 comments
Cite this review

Pith. "Pith review of VR-PTOLEMAIC: A Virtual Environment for the Perceptual Testing of Spatial Audio Algorithms." pith.science (2026). https://pith.science/paper/M5YPSUZX

@misc{pith2026250800501,
  author       = {Pith},
  title        = {Pith review of: VR-PTOLEMAIC: A Virtual Environment for the Perceptual Testing of Spatial Audio Algorithms},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/M5YPSUZX}},
  note         = {Machine review of arXiv:2508.00501}
}
read the original abstract

The perceptual evaluation of spatial audio algorithms is an important step in the development of immersive audio applications, as it ensures that synthesized sound fields meet quality standards in terms of listening experience, spatial perception and auditory realism. To support these evaluations, virtual reality can offer a powerful platform by providing immersive and interactive testing environments. In this paper, we present VR-PTOLEMAIC, a virtual reality evaluation system designed for assessing spatial audio algorithms. The system implements the MUSHRA (MUlti-Stimulus test with Hidden Reference and Anchor) evaluation methodology into a virtual environment. In particular, users can position themselves in each of the 25 simulated listening positions of a virtually recreated seminar room and evaluate simulated acoustic responses with respect to the actually recorded second-order ambisonic room impulse responses, all convolved with various source signals. We evaluated the usability of the proposed framework through an extensive testing campaign in which assessors were asked to compare the reconstruction capabilities of various sound field reconstruction algorithms. Results show that the VR platform effectively supports the assessment of spatial audio algorithms, with generally positive feedback on user experience and immersivity.

Figures

Figures reproduced from arXiv: 2508.00501 by the authors.

Figure 4
Figure 4. All the measured and reconstructed A-RIRs are [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

47 extracted references · 47 canonical work pages

  1. [1]

    INTRODUCTION Spatial audio is a trending research field focused on un- derstanding, recreating, and optimizing 3D auditory envi- ronments [1], providing listeners with an immersive and interactive listening experience. Applications range from virtual and augmented reality [2] to gaming [3], telecon- ferencing and remote concerts [4], where precise sound l...

  2. [2]

    The following section presents a VR- based evaluation system for assessing spatial audio per- ception

    EV ALUA TION SYSTEM Perceptual tests for spatial audio in VR require a mul- timodal environment that involves visual content, audio content and interactivity in order to maintain in the asses- sor a sense of presence and immersivity during the eval- uation procedure. The following section presents a VR- based evaluation system for assessing spatial audio ...

  3. [3]

    SFR is a fundamen- tal task in spatial audio, focused on estimating the pres- sure field in regions where direct measurements cannot be obtained

    SYSTEM V ALIDA TION To validate the proposed VR-based evaluation framework, we used SFR as a test scenario and conducted listening tests within the VR environment. SFR is a fundamen- tal task in spatial audio, focused on estimating the pres- sure field in regions where direct measurements cannot be obtained. Capturing the acoustic field over a broad spati...

  4. [4]

    Fifteen partici- pants (13 male, 2 female) took part in the listening tests, with an average age of 26.9 years (SD = 4.0)

    RESULTS AND DISCUSSION The proposed VR-based system was evaluated consider- ing its ability to facilitate subjective spatial audio assess- ment and its impact on user interaction. Fifteen partici- pants (13 male, 2 female) took part in the listening tests, with an average age of 26.9 years (SD = 4.0). All but one held a university degree, and all had prio...

  5. [5]

    CONCLUSION AND FUTURE WORK We introduced a VR-based system for evaluating spa- tial audio reproduction, integrating real acoustic measure- ments into an immersive and interactive environment. The system was assessed through structured listening tests, which included a comparative evaluation of SFR algo- rithms, and focused on perceptual attributes such as...

  6. [6]

    Telecommunications of the Future

    ACKNOWLEDGMENTS This work has been funded by ”REPERTORIUM project. Grant agreement number 101095065. Horizon Europe. Cluster II. Culture, Creativity and Inclusive Society. Call HORIZON-CL2-2022-HERITAGE-01-02.”. This work was partially supported by the European Union – Next Generation EU under the Italian National Recovery and Resilience Plan (NRRP), Miss...

  7. [7]

    Microphone array processing for paramet- ric spatial audio techniques,

    A. Politis, “Microphone array processing for paramet- ric spatial audio techniques,” 2016

  8. [8]

    Computa- tional and optimization design in geometric acous- tics,

    A. Bassuet, D. Rife, and L. Dellatorre, “Computa- tional and optimization design in geometric acous- tics,” Building Acoustics , vol. 21, no. 1, pp. 75–85, 2014

Show all 47 references
  1. [9]

    Audio aug- mented reality: A systematic review of technologies, applications, and future research directions,

    J. Yang, A. Barde, and M. Billinghurst, “Audio aug- mented reality: A systematic review of technologies, applications, and future research directions,” journal of the audio engineering society , vol. 70, no. 10, pp. 788–809, 2022

  2. [10]

    3d sound spatialization with game engines: the virtual acous- tics performance of a game engine and a middleware for interactive audio design,

    H. B. Fırat, L. Maffei, and M. Masullo, “3d sound spatialization with game engines: the virtual acous- tics performance of a game engine and a middleware for interactive audio design,” Virtual Reality, vol. 26, no. 2, pp. 539–558, 2022

  3. [11]

    A virtual symphony orchestra for studies on concert hall acoustics,

    J. P ¨atynen, “A virtual symphony orchestra for studies on concert hall acoustics,” 2011

  4. [12]

    Spatial audio signal pro- cessing for binaural reproduction of recorded acoustic scenes–review and challenges,

    B. Rafaely, V . Tourbabin, E. Habets, Z. Ben-Hur, H. Lee, H. Gamper, L. Arbel, L. Birnie, T. Abhaya- pala, and P. Samarasinghe, “Spatial audio signal pro- cessing for binaural reproduction of recorded acoustic scenes–review and challenges,” Acta Acustica, vol. 6, p. 47, 2022

  5. [13]

    An overview of machine learning and other data- based methods for spatial audio capture, processing, and reproduction,

    M. Cobos, J. Ahrens, K. Kowalczyk, and A. Poli- tis, “An overview of machine learning and other data- based methods for spatial audio capture, processing, and reproduction,” vol. 2022, no. 1, p. 10

  6. [14]

    A comparative analysis of the directional sound radi- ation of historical violins,

    M. Pezzoli, A. Canclini, F. Antonacci, and A. Sarti, “A comparative analysis of the directional sound radi- ation of historical violins,” The Journal of the Acous- tical Society of America, vol. 152, no. 1, pp. 354–367, 2022

  7. [15]

    Soundfield reconstruction in reverber- ant rooms based on compressive sensing and image- source models of early reflections,

    S. Damiano, F. Borra, A. Bernardini, F. Antonacci, and A. Sarti, “Soundfield reconstruction in reverber- ant rooms based on compressive sensing and image- source models of early reflections,” in 2021 IEEE Workshop on Applications of Signal Processing to Au- dio and Acoustics (...

  8. [16]

    A parametric approach to virtual miking for sources of arbitrary directivity,

    M. Pezzoli, F. Borra, F. Antonacci, S. Tubaro, and A. Sarti, “A parametric approach to virtual miking for sources of arbitrary directivity,” vol. 28, pp. 2333–

  9. [17]

    but also on subjective experiments that address how listeners actually experience the audio, assessing percep- tual attributes such as localization accuracy, spatial clarity, immersion, and externalization. The standard approach for subjective sound qual- ity evaluation is bas...

  10. [18]

    Reconstruction of sound field through diffusion models,

    F. Miotello, L. Comanducci, M. Pezzoli, A. Bernar- dini, F. Antonacci, and A. Sarti, “Reconstruction of sound field through diffusion models,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 1476–1480. ISSN: 2379-190X

  11. [19]

    A zero-shot physics-informed dictionary learning ap- proach for sound field reconstruction,

    S. Damiano, F. Miotello, M. Pezzoli, A. Bernardini, F. Antonacci, A. Sarti, and T. Van Waterschoot, “A zero-shot physics-informed dictionary learning ap- proach for sound field reconstruction,” in ICASSP 2025-2025 IEEE International Conference on Acous- tics, Speech and Signal...

  12. [20]

    Physics- informed neural network for volumetric sound field reconstruction of speech signals,

    M. Olivieri, X. Karakonstantis, M. Pezzoli, F. An- tonacci, A. Sarti, and E. Fernandez-Grande, “Physics- informed neural network for volumetric sound field reconstruction of speech signals,” vol. 2024, no. 1

  13. [21]

    Generative models for sound field reconstruction,

    E. Fernandez-Grande, X. Karakonstantis, D. Caviedes-Nozal, and P. Gerstoft, “Generative models for sound field reconstruction,” The Journal of the Acoustical Society of America , vol. 153, no. 2, pp. 1179–1190, 2023

  14. [22]

    Zacharov, Sensory evaluation of sound

    N. Zacharov, Sensory evaluation of sound. CRC Press, 2018

  15. [23]

    Towards hrtf personalization using denoising diffusion models,

    J. C. Albarrac ´ın S´anchez, L. Comanducci, M. Pezzoli, and F. Antonacci, “Towards hrtf personalization using denoising diffusion models,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1–5, 2025

  16. [24]

    Joint sam- pling theory and subjective investigation of plane- wave and spherical harmonics formulations for bin- aural reproduction,

    Z. Ben-Hur, J. Sheaffer, and B. Rafaely, “Joint sam- pling theory and subjective investigation of plane- wave and spherical harmonics formulations for bin- aural reproduction,” Applied Acoustics , vol. 134, pp. 138–144, 2018

  17. [25]

    Subjective and objective evaluations of a scattered sound field in a scale model opera house,

    J. K. Ryu and J. Y . Jeon, “Subjective and objective evaluations of a scattered sound field in a scale model opera house,” The Journal of the Acoustical Society of America, vol. 124, no. 3, pp. 1538–1549, 2008

  18. [26]

    Bs. 1116-3,“,

    R. ITU-R, “Bs. 1116-3,“,” Methods for the sub- jective assesment of small impairments in audio systems, ” International Telecommunication Union- Radiocommunication Sector, 2015

  19. [27]

    ITU-R, “1534-3,” Method for the subjective assess- ment of intermediate quality level of audio systems , 2015

    B. ITU-R, “1534-3,” Method for the subjective assess- ment of intermediate quality level of audio systems , 2015. 11th Convention of the European Acoustics Association M´alaga, Spain • 22th – 26th June 2025 •

  20. [28]

    Itu-t recommendation p. 800. methods for ob- jective and subjective assessment of quality,

    T. Itu, “Itu-t recommendation p. 800. methods for ob- jective and subjective assessment of quality,” 1996

  21. [29]

    Method for the subjective assessment of intermediate quality level of audio systems,

    B. Series, “Method for the subjective assessment of intermediate quality level of audio systems,” Inter- national Telecommunication Union Radiocommuni- cation Assembly, vol. 2, 2014

  22. [30]

    HOMULA-RIR: A room impulse response dataset for teleconferencing and spatial audio applications ac- quired through higher-order microphones and uniform linear microphone arrays,

    F. Miotello, P. Ostan, M. Pezzoli, L. Coman- ducci, A. Bernardini, F. Antonacci, and A. Sarti, “HOMULA-RIR: A room impulse response dataset for teleconferencing and spatial audio applications ac- quired through higher-order microphones and uniform linear microphone arrays,” in...

  23. [31]

    Scaling sound quality using models for paired-comparison and ranking data,

    F. Wickelmaier, N. Umbach, K. Sergin, and S. Choisel, “Scaling sound quality using models for paired-comparison and ranking data,” in Conference Paper , DAGA 2012 Congress 38th German Annual Conference on Acoustics, 2012

  24. [32]

    Direct and indirect listen- ing test methods—a discussion based on audio-visual spatial coherence experiments,

    C. Pike and H. Stenzel, “Direct and indirect listen- ing test methods—a discussion based on audio-visual spatial coherence experiments,” in Audio Engineering Society Convention 143 , Audio Engineering Society, 2017

  25. [33]

    Free-field study on auditory localiza- tion and discrimination performance in older adults,

    C. Freigang, K. Schmiedchen, I. Nitsche, and R. R ¨ubsamen, “Free-field study on auditory localiza- tion and discrimination performance in older adults,” vol. 232, no. 4, pp. 1157–1172

  26. [34]

    On the use of subjective hrtf evaluations for creating global percep- tual similarity metrics of assessors and assessees.,

    A. Andreopoulou and B. F. Katz, “On the use of subjective hrtf evaluations for creating global percep- tual similarity metrics of assessors and assessees.,” in ICAD, pp. 13–20, 2015

  27. [35]

    As- sessing ambisonics sound source localization by means of virtual reality and gamification tools,

    E. Medina, R. Viveros-Mu ˜noz, and F. Otondo, “As- sessing ambisonics sound source localization by means of virtual reality and gamification tools,” Ap- plied Sciences, vol. 14, no. 17, p. 7986, 2024

  28. [36]

    Audio quality evaluation in virtual reality: multiple stimulus ranking with behavior tracking,

    O. Rummukainen, T. Robotham, S. J. Schlecht, A. Plinge, J. Herre, and E. A. Habels, “Audio quality evaluation in virtual reality: multiple stimulus ranking with behavior tracking,” inAudio Engineering Society Conference: 2018 AES International Conference on Audio for Virtual a...

  29. [37]

    A flexible software tool for perceptual evaluation of audio mate- rial and vr environments,

    S. Gorzynski, N. Kaplanis, and S. Bech, “A flexible software tool for perceptual evaluation of audio mate- rial and vr environments,” in Audio Engineering Soci- ety Convention 149, Audio Engineering Society, 2020

  30. [38]

    Am- plitude matching for multizone sound field control,

    T. Abe, S. Koyama, N. Ueno, and H. Saruwatari, “Am- plitude matching for multizone sound field control,” IEEE/ACM Transactions on Audio, Speech, and Lan- guage Processing, vol. 31, pp. 656–669, 2022

  31. [39]

    Object-based six-degrees-of-freedom ren- dering of sound scenes captured with multiple am- bisonic receivers,

    L. McCormack, A. Politis, T. McKenzie, C. Hold, and V . Pulkki, “Object-based six-degrees-of-freedom ren- dering of sound scenes captured with multiple am- bisonic receivers,” vol. 70, no. 5, pp. 355–372. Pub- lisher: Audio Engineering Society

  32. [40]

    Re- construction of reverberant sound fields over large spatial domains,

    A. Figueroa-Duran and E. Fernandez-Grande, “Re- construction of reverberant sound fields over large spatial domains,” The Journal of the Acoustical So- ciety of America, vol. 157, no. 1, pp. 180–190, 2025

  33. [41]

    The hisstools impulse response toolbox: Convolution for the masses,

    A. Harker and P. A. Tremblay, “The hisstools impulse response toolbox: Convolution for the masses,” in Proc. of the international computer music conference , pp. 148–155, The International Computer Music As- sociation, 2012

  34. [42]

    Sparta & compass: Real-time implementations of linear and parametric spatial audio reproduction and processing methods,

    L. McCormack and A. Politis, “Sparta & compass: Real-time implementations of linear and parametric spatial audio reproduction and processing methods,” in AES International Conference on Immersive and Interactive Audio, Audio Engineering Society, 2019

  35. [43]

    Damiano, F

    S. Damiano, F. Miotello, M. Pezzoli, A. Bernardini, F. Antonacci, A. Sarti, and T. Waterschoot, A Zero- Shot Physics-Informed Dictionary Learning Approach for Sound Field Reconstruction

  36. [44]

    EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation,

    J. Richter, Y .-C. Wu, S. Krenn, S. Welker, B. Lay, S. Watanabe, A. Richard, and T. Gerkmann, “EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation,” in Inter- speech 2024, pp. 4873–4877, ISCA

  37. [45]

    Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, in- sights, and applications,

    B. Li, X. Liu, K. Dinesh, Z. Duan, and G. Sharma, “Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, in- sights, and applications,” vol. 21, no. 2, pp. 522–535

  38. [47]

    A spatial audio qual- ity inventory (saqi),

    A. Lindau, V . Erbes, S. Lepa, H.-J. Maempel, F. Brinkman, and S. Weinzierl, “A spatial audio qual- ity inventory (saqi),” Acta Acustica united with Acus- tica, vol. 100, no. 5, pp. 984–994, 2014

  39. [2348]

    Conference Name: IEEE/ACM Transactions on Audio, Speech, and Language Processing

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.