{"id":"404060af-eb8c-445a-8551-2258977751d9","arxiv_id":"2506.10331","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new UGC omnidirectional audio-visual quality dataset with MOS and a CIQNet-VGGish-transformer baseline that outperforms four video-only methods on this dataset.","lead":"This paper builds a dataset of 300 user-generated 360-degree videos with synchronized audio, human quality ratings, and head movement data. It also proposes a baseline audio-visual quality model that combines video and audio features and reports the best scores on this dataset.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SOTA claim is confounded: CIQNet baseline row matches the authors' own no-audio ablation, baselines train at 30x lower LR, and a single split with no error bars supports a ~0.02 SROCC margin.","rationale":"I read this as a dataset-plus-baseline paper. The dataset construction is described in detail and appears to follow ITU-T procedures; I do not see a soundness problem there. The central assertion with correctness risk is the SOTA result in Table 2. The reader's weakest assumption already targets comparison fairness. My stress-test agrees and adds two concrete aggravating details: the CIQNet row and 'Ours (Without Audio)' row are identical in all four metrics, which suggests the baseline was not independently reproduced; and the learning-rate gap is 30x, with only one split and no error bars, so the reported advantage is within noise. A controlled rerun would settle whether the SOTA claim lands. This does not change the reader's conditional verdict; it strengthens the conditions under which the paper should be accepted.","tokens_in":7379,"tokens_out":6314,"duration_ms":76345,"concrete_test":"Retrain CIQNet, ProVQA, DOVER, FastVQA, and the proposed model under one shared protocol: same optimizer and learning rate (test both 1e-3 and 3e-5), same pretrained initializations, and the same 10 random 80/20 splits of the proposed dataset, reporting mean plus/minus std SROCC/PLCC and a paired significance test. Also run a scene-disjoint split with no same scene or camera in train and test to check for content leakage. If the proposed model does not win consistently, the SOTA claim in Sec. IV-B should be softened to a competitive-performance claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the proposed model 'achieves SOTA performance' (Sec. IV-B) rests entirely on Table 2, and that table is assembled under a confounded protocol. First, the CIQNet baseline row (SROCC 0.8045, PLCC 0.8254, KROCC 0.6186, RMSE 0.7198) is numerically identical to the authors' own 'Ours (Without Audio)' ablation row. This indicates that the CIQNet comparison is effectively the authors' video-only branch, not an independently reproduced baseline, so part of the reported gain may reflect different training conditions rather than the proposed audio-visual design. Second, Sec. IV-A.4 states the proposed model uses Adam at lr=1e-3 while all compared methods are fine-tuned at lr=3e-5, a 30x difference with no sensitivity analysis. Third, all numbers come from one random 80/20 split (60 test videos), with no error bars or significance tests; the SROCC margin over ProVQA (0.8245 vs 0.8081) and CIQNet (0.8045) is about 0.02, easily within split-to-split variability. These issues do not invalidate the dataset, but they do mean the SOTA statement is not yet supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a new dataset for audio-visual quality assessment of user-generated omnidirectional videos (UGC-ODV), containing 300 sequences captured with two consumer-grade 360-degree cameras, covering 10 scene types, with subjective MOS scores and head-movement data collected from 136 subjects following ITU-T BT.500/P.910 procedures. The authors also propose a no-reference audio-visual quality assessment baseline model consisting of a CIQNet-based video feature extractor, a VGGish-based audio feature extractor, and a transformer-based fusion module. The paper claims that this is the first work on UGC-ODV audio-visual quality assessment and that the proposed model achieves state-of-the-art performance on the proposed dataset.","tokens_in":7644,"tokens_out":2672,"duration_ms":30955,"significance":"The dataset contribution is potentially valuable: it addresses a genuine gap, since prior omnidirectional video quality datasets predominantly focus on PGC content and ignore audio. The subjective experiment is described in reasonable detail, follows established ITU recommendations, and includes outlier screening and SSQ-based dizziness filtering, lending credibility to the MOS ground truth. However, the model's claimed SOTA performance is not currently supported by the reported experiments. The evaluation uses a single random split, unequal training protocols, and a suspiciously identical baseline row in Table 2. If the dataset is released and the evaluation protocol is corrected, the paper could make a useful contribution to the community. The architectural novelty is modest—the main components are borrowed from CIQNet and VGGish—but the audio-visual fusion and the UGC-ODV application are of interest.","major_comments":[{"comment":"The proposed model is trained with Adam at learning rate 1e-3, while DOVER, FastVQA, CIQNet, and ProVQA are fine-tuned with learning rate 3e-5. No sensitivity analysis is reported, so the claimed SOTA margin could be an artifact of unequal training settings. The authors should fine-tune all methods under the same protocol or at least report results across a grid of learning rates and demonstrate that the ranking remains stable.","section":"IV-A.4 and Table 2"},{"comment":"The CIQNet row (SROCC 0.8045, PLCC 0.8254, KROCC 0.6186, RMSE 0.7198) is numerically identical to the 'Ours (Without Audio)' ablation row. This suggests that the CIQNet comparison is not an independently reproduced baseline but effectively the authors' own video-only branch. Please clarify how the CIQNet result was obtained; if it is an independent reproduction, explain why the numbers coincide exactly. Otherwise, the comparison must be re-run and the SOTA claim re-evaluated.","section":"Table 2"},{"comment":"All results are based on a single random 80/20 split of 300 videos, yielding 60 test videos, with no error bars, confidence intervals, or significance tests. The SROCC gap between the proposed model (0.8245) and ProVQA (0.8081) is only about 0.02, which is plausibly within split-to-split variability. The authors should report k-fold cross-validation or multiple random splits with standard deviations, and include a significance test such as paired bootstrap or Wilcoxon signed-rank.","section":"IV-A.1 and IV-B"},{"comment":"The manuscript contains no statement about releasing the dataset, MOS scores, head movement data, or code, and no cross-dataset evaluation is performed. Since the dataset is the primary claimed contribution, the authors should specify its availability and provide a clear data-release plan; without this, the community cannot independently verify the MOS values or reproduce the model results.","section":"General (Dataset and Code Availability)"}],"minor_comments":[{"comment":"The text contains typographical and formatting inconsistencies: '12,06 million' in Section II should read '12.06 million'; the abstract renders 'A VQA' with a spurious space; and the phrase 'audio-visual' is sometimes hyphenated and sometimes not. Please standardize.","section":"Abstract and Section II"},{"comment":"The header row of Table 1 is not fully clear: the 'HM/EM/MOS' column mixes three separate data types, and the D-SA V360 row leaves the MOS entry blank without explanation. Please separate the columns or add a footnote clarifying missing entries.","section":"Table 1"},{"comment":"The list of compared methods mentions DOVER, FastVQA, CIQNet, and ProVQA, but the ablation rows 'Ours (Without Audio)', 'Ours (Cat)', and 'Ours (Add)' are not introduced in that list. Please describe the ablation variants explicitly in the text.","section":"Section IV-A.3"},{"comment":"The sentence about head movement data says that valid subjective ratings imply valid head movement data, but this assumption is not justified. At minimum, please report the correlation between head movement data quality and subjective ratings or cite prior work supporting this decision.","section":"Section II-C"}],"recommendation":"major_revision","confidential_remarks":"The dataset contribution is potentially strong, but the empirical evaluation of the model needs substantial rework. I would encourage the editor to require the authors to disclose data and code availability, and to rerun the comparisons under a fair and statistically sound protocol before considering acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the dataset is the real contribution, and it looks like a useful one; the model is a stock combination of CIQNet + VGGish + transformer fusion, and the paper's SOTA claim is not currently supported by the experiments.\n\nThe dataset itself is the thing to look at. 300 user-generated omnidirectional clips, 10 scene types, 4K to 8K, captured with Insta360 cameras, with MOS from 136 subjects after BT.500-style outlier screening and head movement data. That fills a real gap: the prior ODV quality datasets either have no audio (VQA-ODV, VOD-VQA, VRVQW) or are not quality-assessment datasets (D-SAV360). The subjective protocol is described in enough detail to be believable, and the MOS confidence intervals are shown. For the immersive QoE and AVQA community, this is a genuinely new resource that could catalyze UGC-360 quality research.\n\nThe model side is weaker. The baseline is a direct combination of existing components, which is acceptable for a dataset paper, but the evaluation has real problems. The CIQNet baseline row in Table 2 is numerically identical to the authors' own \"Ours (Without Audio)\" ablation, so it is not an independent comparison. The proposed model is trained at lr=1e-3 while all compared methods are fine-tuned at lr=3e-5, a 30x difference with no sensitivity analysis. Everything rests on a single random 80/20 split with no error bars; the SROCC margin over ProVQA is about 0.02, well within split-to-split variability. So the SOTA statement is not yet supported. The dataset and code are not released, which also makes the numbers hard to verify.\n\nNone of this invalidates the dataset. The paper needs a major revision: fair comparison protocol (same training conditions, multiple splits or significance tests), independent reproductions of baselines, and ideally release of the dataset and MOS. As it stands, the dataset deserves serious referee time; the model evaluation needs to be redone.","headline":"A genuinely new UGC-ODV audio-visual dataset with a credible subjective experiment, but the SOTA claim for the baseline model is undercut by a confounded comparison protocol.","tokens_in":8194,"tokens_out":2712,"would_cite":false,"duration_ms":31259,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces a 300-clip user-generated omnidirectional video dataset with audio, subjective Mean Opinion Scores, and head-movement traces, along with a no-reference audio-visual quality model that reaches SROCC 0.8245 and PLCC…","keywords":["user-generated omnidirectional video","audio-visual quality assessment","no-reference quality assessment","omnidirectional video dataset","subjective quality scores","head movement data","multimodal fusion","mean opinion score"],"falsifier":"Re-run the four baselines under the same training protocol as the proposed model, with the same optimizer, learning rate, and 80/20 split; if CIQNet or ProVQA then reaches or exceeds SROCC 0.8245, the claimed state-of-the-art result and the attribution of the gain to audio-visual fusion would be falsified.","tokens_in":7164,"feed_emoji":"🎥","tokens_out":9997,"duration_ms":106813,"temperature":0.7,"pith_summary":"The paper constructs the first audio-visual quality assessment dataset for user-generated omnidirectional video: 300 clips captured by five people with consumer 360-degree cameras across ten scene types, with natural video and audio distortions, human Mean Opinion Scores, and head-movement recordings. The central claim is that audio-visual quality of such content is predictable by a no-reference model that combines video features, audio features, and a transformer-based fusion module, and that this model outperforms existing video-only quality measures. A fair-minded reader would care because everyday 360-degree content is growing quickly while existing quality assessment work has concentrated on professionally produced video and has mostly ignored the audio channel.","feed_headline":"First audio-visual quality dataset for amateur 360° video","feed_subtitle":"Amateur 360° clips pair sight and sound; a new no-reference model predicts perceived quality with SROCC 0.8245.","key_machinery":"The load-bearing mechanism is the paired dataset plus a three-module baseline. The visual branch partitions each equirectangular-projected frame into latitude sub-regions, trains per-region encoders, and applies a backdoor-adjustment weighting scheme to reduce the influence of dimensional confounders on quality features; the audio branch converts the soundtrack into a Mel spectrogram and extracts features with VGGish; and a transformer block with self-attention and cross-attention fuses the two feature streams before a quality regression head. The dataset supplies the subjective ground truth, including 5,026 retained valid opinion scores and more than 12 million head-movement entries, that makes training and comparison possible.","core_discovery":"On its own terms, the paper's finding is that user-generated omnidirectional video quality is better assessed jointly from sight and sound than from video alone. The proposed baseline uses a causal-intervention visual branch, a VGGish-based audio branch, and a transformer fusion module; on the new 300-video dataset it achieves SROCC 0.8245 and PLCC 0.8590, while the same model without the audio branch scores 0.8045 and 0.8254, and replacing the fusion transformer with additive or concatenative merging lowers performance further. The dataset is presented as the first to address audio-visual quality assessment specifically for user-generated omnidirectional content.","pith_inferences":["A natural next test, left implicit in the paper, is cross-dataset generalization: fine-tune the model on these 300 clips and evaluate it on professionally captured omnidirectional audio-visual content to see whether the audio-visual advantage persists outside user-generated conditions.","The 120 Hz head-movement traces could support viewport-dependent or gaze-contingent quality models, a direction the paper records data for but does not exploit.","Because part of the dataset carries four-channel audio, a future audio branch could treat spatial audio directly instead of collapsing the soundtrack to a monaural Mel representation."],"forward_implications":["Audio contributes measurably to perceived omnidirectional video quality, since dropping the audio branch lowers SROCC from 0.8245 to 0.8045.","The choice of fusion method matters: concatenative and additive fusion both underperform the transformer fusion module in the reported results.","Existing 2D user-generated video quality methods transfer poorly to omnidirectional content, with the best 2D baseline reaching only 0.7974 SROCC on this dataset.","The dataset provides a shared benchmark with MOS, head movement, and multi-resolution content on which future no-reference audio-visual quality models can be trained and compared."],"supporting_citations":[{"why":"Defines the causal-intervention visual feature extraction backbone used by the proposed model and is the strongest ODV-only baseline compared.","marker":"[9]"},{"why":"Supplies the pretraining data for the omnidirectional image quality feature extractor used in the visual branch.","marker":"[20]"},{"why":"Supplies the VGGish network that the audio branch uses to turn Mel spectrograms into audio quality features.","marker":"[17]"},{"why":"Provides the pretraining corpus that initializes the audio feature extraction network.","marker":"[21]"},{"why":"Provides the screening and Mean Opinion Score calculation procedure that turns raw ratings into ground-truth labels.","marker":"[11]"},{"why":"Guides the design of the subjective test environment and protocol for the omnidirectional audio-visual sequences.","marker":"[10]"},{"why":"Serves as a state-of-the-art omnidirectional no-reference video quality baseline that is compared and that degrades on user-generated content.","marker":"[8]"},{"why":"Serves as a 2D user-generated video quality baseline that highlights the need for spherical awareness in the comparison.","marker":"[15]"},{"why":"Appears in the dataset comparison as an existing omnidirectional audio-visual dataset without MOS, supporting the novelty claim for the proposed dataset.","marker":"[6]"}],"fun_headline_variants":["Audio boosts quality scores for amateur 360° video","Sight and sound beat video alone for 360° video quality","First dataset linking audio-visual quality in 360° video","User-generated 360° video quality: audio matters","New baseline for audio-visual quality in 360° video"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that this model is state of the art assumes the comparison against prior methods is fair, yet the proposed model is trained with a higher learning rate than all compared baselines and no sensitivity analysis is reported.","fun_headline_variants_meta":{"raw":{"variants":["Audio boosts quality scores for amateur 360° video","Sight and sound beat video alone for 360° video quality","First dataset linking audio-visual quality in 360° video","User-generated 360° video quality: audio matters","New baseline for audio-visual quality in 360° video"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000378,"raw_usage":{"total_tokens":1969,"prompt_tokens":865,"completion_tokens":1104,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":1021}},"tokens_in":481,"tokens_out":1104,"duration_ms":11335,"temperature":1.0,"reasoning_tokens":1021,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:29:19.197034+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the four baselines under the same training protocol as the proposed model, with the same optimizer, learning rate, and 80/20 split; if CIQNet or ProVQA then reaches or exceeds SROCC 0.8245, the claimed state-of-the-art result and the attribution of the gain to audio-visual fusion would be falsified.","supporting_citations":[{"cited_title":"Omnidirectional video quality assessment with causal intervention,","cited_arxiv_id":null,"evidence_quote":"Defines the causal-intervention visual feature extraction backbone used by the proposed model and is the strongest ODV-only baseline compared."},{"cited_title":"Spatial attention-based non-reference perceptual quality prediction network for omnidirectional images,","cited_arxiv_id":null,"evidence_quote":"Supplies the pretraining data for the omnidirectional image quality feature extractor used in the visual branch."},{"cited_title":"The robust feature extraction of audio signal by using vggish model,","cited_arxiv_id":null,"evidence_quote":"Supplies the VGGish network that the audio branch uses to turn Mel spectrograms into audio quality features."},{"cited_title":"ITU, 2012","cited_arxiv_id":null,"evidence_quote":"Provides the screening and Mean Opinion Score calculation procedure that turns raw ratings into ground-truth labels."},{"cited_title":"ITU, 2008","cited_arxiv_id":null,"evidence_quote":"Guides the design of the subjective test environment and protocol for the omnidirectional audio-visual sequences."},{"cited_title":"Blind vqa on 360 video via progressively learning from pixels, frames, and video,","cited_arxiv_id":null,"evidence_quote":"Serves as a state-of-the-art omnidirectional no-reference video quality baseline that is compared and that degrades on user-generated content."},{"cited_title":"Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,","cited_arxiv_id":null,"evidence_quote":"Serves as a 2D user-generated video quality baseline that highlights the need for spherical awareness in the comparison."},{"cited_title":"D-sav360: A dataset of gaze scanpaths on 360 ambisonic videos,","cited_arxiv_id":null,"evidence_quote":"Appears in the dataset comparison as an existing omnidirectional audio-visual dataset without MOS, supporting the novelty claim for the proposed dataset."}],"review_version":1}