{"id":"c0a3e078-7fe0-486a-87f5-2c6fbbd1d4e9","arxiv_id":"2508.00501","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"VR-PTOLEMAIC implements MUSHRA listening tests in virtual reality using measured room impulse responses, and a 15-participant study reports generally positive user feedback.","lead":"This paper presents VR-PTOLEMAIC, a virtual reality system for running MUSHRA listening tests on spatial audio algorithms inside a simulated seminar room. It combines real measured room acoustics with head-tracked headphone playback to make subjective audio testing more interactive.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Sec. 5 claim that the platform offers perceptual feedback comparable to traditional setups is not supported because the validation has no non-VR baseline condition; observed discrimination and usability are compatible with VR-specific bias.","rationale":"I agree with the reader's weakest_assumption: the unvalidated bridge between VR binaural rendering and traditional MUSHRA outcomes is what the conclusion depends on. The system itself is coherent: the OSC/Unity/Max pipeline is described, the HOMULA-RIR measurements provide a real acoustic basis, and the hidden-reference screening is a meaningful internal control. However, none of those elements compares VR ratings to a non-VR baseline. Independent statistical support (error bars, tests) is also absent, so the central comparative claim remains conjectural. This does not warrant rejection; the contribution is a usable, documented tool, and the authors disclose a planned open-source release. A conditional verdict requiring a matched non-VR validation is the right level. My proposed experiment directly targets the premise, and would either substantiate or force a downgrade of the comparability wording.","tokens_in":8672,"tokens_out":6140,"duration_ms":67462,"concrete_test":"Run a matched-subjects MUSHRA comparison with the same five stimuli, four attributes, and 25 positions in two conditions: (A) the full VR-PTOLEMAIC system, and (B) a conventional non-VR binaural MUSHRA interface using the same SPARTA ambiBIN renderer, same headphones/amp, fixed head orientation, and position selection via mouse or keyboard. Use at least 20 screened listeners per condition in a counterbalanced design. Test per-attribute mean equivalence with two one-sided t-tests at a ±5-point bound on the 0–100 MUSHRA scale and compare stimulus rank order across conditions. If equivalence and rank agreement hold, the Sec. 5 comparability claim is supported; otherwise the claim should be weakened to 'usable VR test environment' and the comparability statement removed or qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, stated in the Abstract and reiterated in Sec. 5 ('offering perceptual feedback comparable to traditional setups'), requires that MUSHRA ratings obtained in VR approximate ratings that would be obtained in a conventional listening test. The validation in Sec. 3 does not include any non-VR or real-room condition. It shows only that, after applying the MUSHRA screening rule, 11 assessors could distinguish four SFR variants in the VR environment and reported generally positive usability. That evidence is equally consistent with VR-specific distortions: SPARTA ambiBIN decoding with a generic (likely non-individualized) HRTF, head-tracked playback, and free choice among 25 positions could shift spatial or timbral judgments relative to a fixed-orientation binaural or loudspeaker setup. Moreover, Sec. 4 reports no inferential statistics or effect sizes, so even the internal discrimination finding is asserted from the aggregated plots rather than tested. The buildable system and hidden-reference screening are real strengths, but the comparability conclusion is a load-bearing empirical claim that the current experiment cannot settle.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents VR-PTOLEMAIC, a virtual-reality system for perceptual testing of spatial audio algorithms. The platform couples a Unity-based VR application with a Max-based audio processor via OSC: users can move among 25 predefined listening positions of a reconstructed seminar room, select stimuli through a virtual MUSHRA interface, and hear measured or reconstructed second-order Ambisonic room impulse responses encoded binaurally with head-tracked SPARTA ambiBIN decoding. The system also logs head position, rotation, and teleportation behaviour. The authors validate the platform with a listening test in which 15 participants (11 after MUSHRA screening) rated four stimuli against a reference across four attributes (basic audio quality, localizability, spatial quality, timbral quality), and they report generally positive usability feedback and exploratory behavioural tracking results. The paper concludes that the platform effectively supports spatial audio evaluation and offers perceptual feedback comparable to traditional setups.","tokens_in":8843,"tokens_out":2337,"duration_ms":26447,"significance":"If the comparability claim were supported, the paper would make a useful practical contribution: a concrete, buildable VR implementation of MUSHRA with measured room impulse responses at 25 positions, real-time head-tracked binaural rendering, hidden-reference and anchor conditions, and behavioural tracking. The system description is clear enough to be reproducible, and the use of the HOMULA-RIR dataset anchors the work in real measurements. However, the validation as reported does not establish the central conclusion of equivalence with traditional testing: there is no non-VR baseline condition, no inferential statistical analysis, and the usability evidence is qualitative. The tool is promising, but the evidence presented is preliminary and the comparability claim outstrips the data.","major_comments":[{"comment":"The claim that the platform offers 'perceptual feedback comparable to traditional setups' is not supported by the validation described in §3. The study contains no non-VR or real-room baseline condition, so there is no evidence that MUSHRA ratings obtained in the VR environment approximate ratings from a conventional listening test. The observed ability to distinguish the low-pass anchor and the hidden reference is compatible with VR-specific distortions introduced by the non-individualized SPARTA ambiBIN decoding, head-tracked playback, and free navigation among 25 positions. Either add a comparative baseline condition (e.g., the same MUSHRA test with fixed orientation and non-head-tracked binaural or loudspeaker reproduction) or remove and explicitly qualify the comparability claim.","section":"Abstract and §5 (Conclusion)"},{"comment":"The manuscript asserts that participants 'were able to distinguish between the different reconstruction methods with a high degree of consistency', but no inferential statistics are provided. Figure 5 shows aggregated MUSHRA ratings without confidence intervals, error bars, effect sizes, or significance tests, and the number of valid participants after screening is only 11. Without a per-participant or per-item statistical analysis, the discrimination claim is not quantitatively established. Report condition means and confidence intervals and, if appropriate, a within-subjects test or effect-size measure.","section":"§3, §4, Fig. 5"},{"comment":"The usability and immersivity conclusions rest on an informal post-session questionnaire ('Most users described the system as intuitive...'). No rating scales, no item-level results, no quantification of agreement, and no analysis of the reported mild discomfort are provided. With 11 or 15 participants, this evidence supports only anecdotal impressions, not the general statement that the system 'effectively supports the assessment of spatial audio algorithms'. The claims should be scaled back to what the qualitative data can sustain, or the questionnaire should be presented as a formal usability instrument with defined scales.","section":"§4 (Results and Discussion)"}],"minor_comments":[{"comment":"The convention header reads '22th – 26th June 2025'; this should be '22nd – 26th June 2025'.","section":"Header"},{"comment":"The text says 'Meta 3 Quest 2 VR-headset'; the correct product name is Meta Quest 2.","section":"§2, device name"},{"comment":"The low-pass anchor is described as 'Low audio quality solution Slp' with a cutoff frequency of3.5 kHz; add a space after 'of' and use a non-breaking space in '3.5 kHz'.","section":"§3.1, anchor description"},{"comment":"The Localizability attribute is described with the scale 'More difficult - Easier', which is ambiguous: clarify whether higher ratings correspond to easier or more difficult localizability, and ensure the plot in Fig. 5 uses the same polarity.","section":"§3.1, attribute scales"},{"comment":"Several references are missing publisher or venue details (e.g., [12], [35]), and the formatting of reference [20] appears inconsistent. Please normalize the bibliography.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a potentially useful open tool, but the central empirical claim of equivalence with traditional listening tests is not backed by the presented data. The authors should either add a baseline comparison or clearly reframe the contribution as a system description with preliminary usability data. I also note that the validation uses the authors' own SFR algorithms and dataset; while not circular for the usability claim, the SFR discrimination results should be interpreted as a demonstration of the platform rather than as a general validation of spatial audio quality assessment."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper does what it says: a working VR implementation of MUSHRA with measured second-order Ambisonic RIRs at 25 positions, head-tracked binaural playback via SPARTA, and behavioral tracking. Prior work had VR-based MUSHRA-like tools, but this one integrates a specific real-room dataset and teleportation navigation, which is a concrete, useful addition. The system description is clear and comprehensible, and the authors are honest that the SFR evaluation results are not the main point.\n\nWhere the paper gets into trouble is the conclusion. The abstract and Section 5 claim the platform \"effectively supports\" evaluation and offers \"perceptual feedback comparable to traditional setups.\" The validation only shows that 11 screened listeners could rank four SFR variants in VR and that usability questionnaires were positive. There is no non-VR baseline, no inferential statistics or error bars in Fig. 5, and no control for VR-specific bias (e.g., generic HRTF, head-tracking, free position choice). Those results are consistent with the platform being usable, but they cannot support comparability with traditional listening tests. The discrimination itself is asserted from aggregated plots, not tested.\n\nMinor points: the sample is small and homogeneous (13 male, music-experienced, mostly university-educated), and the code is only planned for release. These are not disqualifying for a systems paper, but they should be acknowledged.\n\nOverall, the contribution is incremental but genuinely useful for the spatial audio evaluation community. With a tempered conclusion and an explicit statement of the comparability limitation, it would be a fine conference paper. As is, it passes the bar for peer review but would need revision before acceptance.\n\nI'd bring it to a reading group if anyone is working on VR audio evaluation.\n\nBest.","headline":"A solid, buildable VR-MUSHRA system paper whose headline comparability claim outruns a validation with no non-VR baseline and no inferential statistics.","tokens_in":9414,"tokens_out":2193,"would_cite":false,"duration_ms":22251,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"VR-PTOLEMAIC claims a virtual-reality MUSHRA platform using measured room impulse responses at 25 positions supports spatial audio evaluation and yields perceptual feedback comparable to traditional setups.","keywords":["virtual acoustics","MUSHRA","virtual reality","perceptual evaluation","spatial audio","sound field reconstruction","binaural rendering","behavioral tracking"],"falsifier":"Run the same listeners and the same stimulus set through two administrations: the VR platform as described, and a conventional non-interactive binaural version with fixed head position and identical headphone playback. If the ranking between the hidden reference, the low-pass anchor, and the two reconstruction algorithms changes across administrations, or if hidden-reference scores differ by more than the internal consistency of repeated VR trials, then the claimed comparability to traditional setups is refuted.","tokens_in":8467,"feed_emoji":"🎧","tokens_out":9051,"duration_ms":81261,"temperature":0.7,"pith_summary":"The paper claims that a virtual-reality implementation of the MUSHRA listening test can serve as a practical platform for perceptual evaluation of spatial audio algorithms. The platform recreates a seminar room in VR, constrains listening to 25 measured positions, and renders audio by convolving anechoic source samples with measured or reconstructed second-order Ambisonic room impulse responses, decoded binaurally with head tracking. In a validation session, 15 listeners rated four sound field reconstruction conditions (hidden reference, low-pass anchor, and two reconstruction methods) on basic audio quality, localizability, spatial quality, and timbral quality, and their ratings separated the conditions consistently. On this evidence the authors conclude that the platform supports spatial audio assessment and provides perceptual feedback comparable to traditional setups. If true, standardized spatial audio evaluation could move out of the treated listening room into a lightweight, headset-based environment.","feed_headline":"VR room runs spatial audio tests at 25 seats","feed_subtitle":"MUSHRA-style VR evaluation with measured room acoustics separated reconstruction algorithms across four quality attributes.","key_machinery":"The load-bearing object is the measured second-order Ambisonic room impulse response paired with a head-tracked binaural decoder. Each of the 25 virtual seats maps to one measured A-RIR; the audio engine convolves the selected anechoic sample with either that measured response or a reconstructed response, producing a multichannel Ambisonic stream. A real-time binaural decoder rotates the stream according to the listener's head orientation and delivers it over closed headphones, so the listener's head movement becomes part of the evaluation. This chain ties every MUSHRA rating to a specific room position and a specific head orientation, which is what makes the subjective test spatially anchored.","core_discovery":"On the paper's own terms, the discovery is that the MUSHRA protocol—a multi-stimulus test with a hidden reference and an anchor—survives translation into an interactive VR room. The authors build the test on measured second-order Ambisonic room impulse responses at 25 chair positions, so the reference and hidden reference are the actual acoustics of a real seminar room, while the conditions under test are reconstructed responses generated by different sound field reconstruction algorithms. The validation results show consistent separation among conditions on all four rating attributes, and participant questionnaire responses were generally positive, with only mild discomfort reported from prolonged headset wear. The paper therefore concludes that the platform effectively supports the evaluation of spatial audio algorithms and offers perceptual feedback comparable to traditional setups.","pith_inferences":["The paper's comparison to traditional setups is qualitative; a stricter claim would need a within-subjects equivalence test between VR and non-VR administration of the same stimulus set, and the paper does not report one.","Because head orientation is tracked, the platform could be extended to evaluate orientation-dependent attributes such as externalization and directional fidelity, which a fixed-headphone MUSHRA cannot capture.","The 25 discrete seats suggest a natural next step of interpolating between measured responses to evaluate moving sources or walk-through auralization, which the current system does not implement.","The planned open release would let other groups swap in their own measured responses, turning the system into a generic perceptual testbed for any Ambisonic capture; that generality is left implicit."],"forward_implications":["If the central claim holds, spatial audio algorithms can be compared perceptually at 25 different room positions using one measured room impulse response acquisition, without repositioning loudspeakers or microphones between trials.","The built-in tracking data lets evaluators see where listeners moved and how long they stayed at each seat, adding a behavioral correlate to quality scores such as localizability.","Because the audio pipeline runs in real time on an all-in-one headset and a laptop, standardized spatial audio listening tests could be run outside an acoustically treated listening room.","The clear separation between hidden reference, anchor, and reconstructed stimuli in the reported MUSHRA results supports using the platform for future comparisons of sound field reconstruction algorithms."],"supporting_citations":[{"why":"defines the MUSHRA protocol with hidden reference and anchor that the virtual interface implements and the reliability screening follows.","marker":"[21]"},{"why":"supplies the measured second-order Ambisonic room impulse responses at 25 positions used as the explicit reference and hidden reference.","marker":"[30]"},{"why":"provides the real-time multichannel convolution that renders each anechoic sample through the selected response.","marker":"[33]"},{"why":"provides the real-time binaural decoder with head rotation that delivers the audio to headphones.","marker":"[34]"},{"why":"one of the two sound field reconstruction algorithms rated by listeners, representing a parametric virtual-miking approach.","marker":"[9]"},{"why":"the other rated reconstruction algorithm, an amplitude-matching method used as a comparative condition.","marker":"[38]"},{"why":"prior work combining multiple-stimulus ranking with behavior tracking in VR, which motivates the platform's tracking component.","marker":"[28]"}],"fun_headline_variants":["VR room tests spatial audio with hidden references at 25 positions","MUSHRA-style audio evaluation now lives in a virtual venue","25 VR seats grade spatial audio reconstruction algorithms fairly","Virtual hall assesses sound field algorithms against measured acoustics","Interactive VR arena hears out spatial audio algorithms at 25 spots"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that head-tracked binaural decoding of measured second-order Ambisonic room impulse responses preserves the same perceptual quality differences among algorithms as a real listening room, so VR-based MUSHRA ratings are comparable to traditional listening-test ratings; the paper presents no direct non-VR comparison to confirm that equivalence.","fun_headline_variants_meta":{"raw":{"variants":["VR room tests spatial audio with hidden references at 25 positions","MUSHRA-style audio evaluation now lives in a virtual venue","25 VR seats grade spatial audio reconstruction algorithms fairly","Virtual hall assesses sound field algorithms against measured acoustics","Interactive VR arena hears out spatial audio algorithms at 25 spots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000166,"raw_usage":{"total_tokens":1223,"prompt_tokens":885,"completion_tokens":338,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":501,"completion_tokens_details":{"reasoning_tokens":257}},"tokens_in":501,"tokens_out":338,"duration_ms":4202,"temperature":1.0,"reasoning_tokens":257,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:05:53.717151+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same listeners and the same stimulus set through two administrations: the VR platform as described, and a conventional non-interactive binaural version with fixed head position and identical headphone playback. If the ranking between the hidden reference, the low-pass anchor, and the two reconstruction algorithms changes across administrations, or if hidden-reference scores differ by more than the internal consistency of repeated VR trials, then the claimed comparability to traditional setups is refuted.","supporting_citations":[{"cited_title":"Generative models for sound field reconstruction,","cited_arxiv_id":null,"evidence_quote":"defines the MUSHRA protocol with hidden reference and anchor that the virtual interface implements and the reliability screening follows."},{"cited_title":"HOMULA-RIR: A room impulse response dataset for teleconferencing and spatial audio applications ac- quired through higher-order microphones and uniform linear microphone arrays,","cited_arxiv_id":null,"evidence_quote":"supplies the measured second-order Ambisonic room impulse responses at 25 positions used as the explicit reference and hidden reference."},{"cited_title":"Free-field study on auditory localiza- tion and discrimination performance in older adults,","cited_arxiv_id":null,"evidence_quote":"provides the real-time multichannel convolution that renders each anechoic sample through the selected response."},{"cited_title":"On the use of subjective hrtf evaluations for creating global percep- tual similarity metrics of assessors and assessees.,","cited_arxiv_id":null,"evidence_quote":"provides the real-time binaural decoder with head rotation that delivers the audio to headphones."},{"cited_title":"Audio aug- mented reality: A systematic review of technologies, applications, and future research directions,","cited_arxiv_id":null,"evidence_quote":"one of the two sound field reconstruction algorithms rated by listeners, representing a parametric virtual-miking approach."},{"cited_title":"Am- plitude matching for multizone sound field control,","cited_arxiv_id":null,"evidence_quote":"the other rated reconstruction algorithm, an amplitude-matching method used as a comparative condition."},{"cited_title":"Itu-t recommendation p. 800. methods for ob- jective and subjective assessment of quality,","cited_arxiv_id":null,"evidence_quote":"prior work combining multiple-stimulus ranking with behavior tracking in VR, which motivates the platform's tracking component."}],"review_version":1}