{"id":"28d77f75-ea23-4cbf-b7b4-13d1e39c8c4d","arxiv_id":"2506.14803","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new 360-degree video dataset and a recurrent attention-based super-resolution model, S3PO, are introduced and shown to improve PSNR on omnidirectional video benchmarks over several prior VSR models.","lead":"This paper builds a new benchmark dataset of 360-degree videos and a deep learning model called S3PO for 4x super-resolution of equirectangular projection frames. The authors report that S3PO beats most existing video super-resolution models on 360-degree test sets, though the comparison uses baselines that were not retrained on the new dataset.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Evaluation protocol confounds architecture with domain adaptation: baselines are not retrained on 360VDS, and reported margins are small and untested for significance.","rationale":"I read the paper as an empirical claim: a new dataset plus a recurrent 360-specific architecture and distortion-aware loss yields state-of-the-art 360° VSR. The architecture description and ablations are credible; the ablation study (Tables V-VII) is a genuine strength, and the WSS-L1 loss is a reasonable contribution. The load-bearing point is external validity of the comparison, not internal inconsistency. The paper's own ablation shows domain adaptation contributes only about 0.05 dB PSNR to S3PO (Table VII), which is smaller than the reported margins, but that does not control for how much more the conventional baselines could gain from being fine-tuned on 360VDS; their published weights were trained for conventional video, and they may suffer a larger domain shift than an architecture already designed for ERP input. The small margins and ties, together with no statistical testing, mean the tables as presented cannot distinguish architectural superiority from adaptation advantage or run-to-run noise. This matches the reader's weakest assumption. The conditional verdict is appropriate; if the proposed retraining/significance check were run and the margins persisted, the claim would be substantially strengthened.","tokens_in":20612,"tokens_out":6975,"duration_ms":66118,"concrete_test":"Fine-tune every reported baseline (BasicVSR, RBPN, EDVR, RSDN, TGA) on the 360VDS training split under the same protocol as S3PO: initialize from the authors' published weights (or from their conventional training), use the same degradation (BI or BD), optimizer, learning-rate schedule, and number of epochs, and evaluate on the same 45-clip test set. Use each baseline's standard loss, and also run a sensitivity check with WSS-L1. Report per-clip paired differences and 95% bootstrap confidence intervals for PSNR, SSIM, WS-PSNR, and WS-SSIM. Separately, inspect source-video identifiers to confirm that no source video contributes clips to both training and test splits. If any retrained baseline matches or exceeds S3PO on any metric, or if the confidence intervals include zero for the reported margins, the state-of-the-art claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, that S3PO is state-of-the-art for 360° VSR, is supported only by an uncontrolled comparison. In Sec. V-A1, S3PO is initialized on Vimeo90K and then fine-tuned on the 360VDS training split; in Sec. V-B, Table I compares against conventional VSR baselines evaluated with their published pretrained weights, and the caption explicitly says those baselines 'use the original degradation, as presented by the corresponding authors.' The model with target-domain fine-tuning therefore enjoys a domain-adaptation advantage that is not available to the baselines. The reported margins are small (BI PSNR: 27.26 vs 27.14 for RBPN and 27.05 for BasicVSR; BD PSNR: 27.51 vs 27.32 for RSDN; MiG PSNR: 31.16 vs 30.89 for BasicVSR), and some metrics are effectively ties (SSIM 0.8227 in BI with BasicVSR; WS-SSIM 0.8453 vs 0.8452 in MiG). No confidence intervals, per-clip paired tests, or repeated-seed evaluations are provided for any table. A second, compounding gap is that the 360VDS train/test split is described as a random split of 590 clips (Sec. IV-B), with no statement that clips from the same source video are kept in one split; if source-video overlap exists, fine-tuning on 360VDS could have seen near-duplicate test content. Under these conditions the tables do not establish the headline claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces 360VDS, a new dataset of 590 equirectangular video clips for 360° video super-resolution, and proposes a model called S3PO built on recurrent propagation with a sliding-window feature extractor, spatial/channel attention, dual-duct residual refinement, and a weighted spherically smooth L1 (WSS-L1) loss. S3PO is first trained on Vimeo90K and then fine-tuned on 360VDS, and it is evaluated on the 360VDS test set, the MiG Panorama test set, and a new 360UHD high-resolution subset, under both bicubic (BI) and blur (BD) degradations. The paper reports that S3PO outperforms existing conventional and 360°-specific VSR models on most metrics, and it includes ablations of the feature extractor, attention, loss function, recurrent residue, mutual exchange, and domain adaptation.","tokens_in":20891,"tokens_out":5698,"duration_ms":55621,"significance":"The dataset contribution is potentially valuable: 360VDS is larger and more diverse than the existing MiG Panorama benchmark, and the paper's analysis of spatial and temporal complexity supports this. The architectural study is also useful, particularly the ablation evidence that the recurrent design without explicit alignment can work on ERP frames, and the claim that the WSS-L1 loss helps is tested against a Smooth-L1 baseline. The paper states that code and weights will be released, which would aid reproducibility. However, the headline claim of state-of-the-art performance rests on an evaluation protocol that gives S3PO a domain-adaptation advantage over baselines that are not retrained on the target 360VDS domain; the reported margins are small, and no significance testing is provided. As a result, the empirical superiority claim is not presently established, although the underlying architecture and dataset are credible contributions that could be supported by a fairer comparison.","major_comments":[{"comment":"The central comparison is confounded: S3PO is initialized on Vimeo90K and then fine-tuned on the 360VDS training split (Section V-A1), while the conventional VSR baselines are evaluated with their original published settings, as indicated by the Table I caption stating that these models 'use the original degradation, as presented by the corresponding authors.' This means the evaluation conflates architectural merit with domain adaptation. The reported advantages are small—BI PSNR 27.26 dB vs 27.14 dB for RBPN, BD PSNR 27.51 dB vs 27.32 dB for RSDN, and an exact SSIM tie with BasicVSR under BI (0.8227)—and no confidence intervals, per-clip paired tests, or repeated-seed evaluations are provided. To support the claim of state-of-the-art performance, the authors should retrain or fine-tune the baselines on the 360VDS training split and report per-clip paired significance tests (e.g., Wilcoxon signed-rank or paired t-test with confidence intervals). Without such a control, the headline conclusion is not established.","section":"Section V-B, Tables I–III"},{"comment":"The dataset construction creates a potential train/test contamination risk. The 590 clips are derived from 301 source videos, and the paper states that the 590 clips are 'split randomly' into 45 test and 545 training clips. If clips from the same source video appear on both sides of the split, then a model fine-tuned on the training side may have seen near-duplicate content at test time, which would inflate its measured performance. The authors should either split at the source-video level (ensuring no source video contributes clips to both train and test) or provide a provenance mapping showing that all test clips come from source videos not represented in the training set. This is load-bearing for every quantitative claim made on 360VDS.","section":"Section IV-B"},{"comment":"The high-resolution evaluation in Table III compares S3PO only against BasicVSR, rather than against the full set of baselines used in Tables I and II. Since the abstract claims superiority over 'most state-of-the-art' conventional and 360°-specific models, restricting the 360UHD comparison to a single baseline is insufficient to support that claim at high resolutions. Additionally, the 360UHD clips are randomly selected from the same 45-clip 360VDS test set, so the comparison is on a subset of the already small test set and inherits the split-contamination concern raised above. The authors should either compare against additional baselines on 360UHD or temper the claim to be specific to the 360VDS and MiG test sets.","section":"Section V-B, Table III"}],"minor_comments":[{"comment":"The text contains several typos: '360VSD' instead of '360VDS' in Section V-A2, 'aixs' instead of 'axis' in the Conclusion, and 'BLOOD' instead of 'BOLD' in the Table III caption.","section":"General"},{"comment":"The definition 'CA Feat(·) = Conv3×3(ReLU(Conv3×3(·))' is missing a closing parenthesis for the inner Conv3×3. Please correct the notation to 'CA Feat(·) = Conv3×3(ReLU(Conv3×3(·)))'.","section":"Eqn. (2)"},{"comment":"The cyclic treatment variant is referred to as 'S3PO-cyc' in Table IV but as 'S3PO-cyclic' in the text; please use a single consistent name.","section":"Section V-D, Table IV"},{"comment":"The paper describes S3PO as applicable in both online and offline settings, but the 360° Feature Extractor uses the future frame Ft+1 (Eqn. 1), which means the model requires a one-frame look-ahead and is not suitable for true online processing in its current form. Please clarify this point in the applicability discussion.","section":"Section III-C"}],"recommendation":"major_revision","confidential_remarks":"The central empirical claim is not supported by the current evaluation protocol because of the domain-adaptation asymmetry and the absence of significance testing. These issues are fixable within the scope of the manuscript by retraining/fine-tuning baselines on 360VDS and reporting paired statistics, and by splitting the dataset at source-video level. I would also ask the editor to verify that the dataset release plan (code and weights) is honored, as that is a key part of the paper's value."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contributions here are the 360VDS dataset and the S3PO architecture's 360-specific pieces. The dataset (590 clips, diverse SI/TI, plus an 8-clip UHD test set) is a useful asset for a field that has very few benchmarks. The model itself is a sensible hybrid: recurrent propagation without optical flow, a 360 feature extractor with CBAM-style attention, dual-duct residual blocks, and a latitude-weighted Smooth L1 loss. The ablation study is genuinely informative — each component gives a measurable bump, and the attention maps do seem to track ERP distortion. The writing is clear and the authors are transparent about their training protocol.\n\nThe problem is that the headline claim — S3PO outperforms state-of-the-art VSR models on 360 video — is not actually established by the experiments. In Table I, S3PO is initialized on Vimeo90K and then fine-tuned on the 360VDS training split, while every baseline is evaluated with its published pretrained weights (trained on REDS, Vimeo90K, etc.). The caption states this explicitly. That gives S3PO a domain-adaptation advantage that the baselines do not have. The reported margins are tiny: 0.12–0.21 dB PSNR on BI, and some metrics are ties (SSIM 0.8227 with BasicVSR). No confidence intervals, no per-clip paired tests, no repeated seeds. The MiG Panorama results are cleaner in the sense that S3PO was not fine-tuned on that dataset, but the same confound remains because baselines still come from the conventional domain. The 360UHD comparison pits S3PO against only BasicVSR and carries the same issue.\n\nThere is also a potential data-leakage problem in the dataset split. The paper says 590 clips were randomly split into 45 test and 545 train, but does not say whether clips derived from the same source video were kept together. If a clip from one source video ends up in both train and test, the evaluation is inflated. That needs to be addressed or at least clarified.\n\nThe WSS-L1 loss is a simple cosine-weighted Smooth L1; that is fine as a minor contribution, not a major one. The architecture leans on the authors' prior R2D2 work, which is cited and acknowledged, so no circularity issue.\n\nThis paper is for people working on 360/omnidirectional video processing and VSR benchmarks. It deserves a serious referee because the dataset alone is a service to the community and the architecture is a reasonable baseline. But the comparison must be redone: retrain or at least fine-tune baselines on 360VDS, enforce source-video separation in the split, and report significance. I would accept it for peer review with those expectations.","headline":"The dataset and 360-specific components are worth attention, but the SOTA claim rests on an uneven comparison: S3PO is fine-tuned on the target domain while baselines use off-the-shelf weights, and the margins are small.","tokens_in":21495,"tokens_out":2417,"would_cite":false,"duration_ms":25582,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A distortion-aware recurrent network, S3PO, is claimed to outperform conventional video super-resolution models on 360-degree footage by accounting for equirectangular distortion and horizontal cyclicity.","keywords":["360-degree video super-resolution","equirectangular projection","spherical distortion","recurrent neural network","attention mechanism","weighted loss function","omnidirectional video dataset","video super-resolution"],"falsifier":"Retrain or fine-tune the top two conventional baselines, such as BasicVSR and RSDN, on the 360VDS training set under the same bicubic and blur degradations used for S3PO, then evaluate on the 45-clip test set with repeated runs and confidence intervals; if their PSNR or WS-PSNR reaches or exceeds S3PO's values of 27.26 dB BI and 27.51 dB BD, or if the reported margins over frozen baselines disappear under matched training, the paper's superiority claim would be falsified.","tokens_in":20385,"feed_emoji":"🎥","tokens_out":7443,"duration_ms":66327,"temperature":0.7,"pith_summary":"The paper tries to establish that equirectangular 360-degree video has its own super-resolution problem: conventional VSR models transfer with decent results, but they miss the latitude-dependent distortion and horizontal seam continuity of the ERP format. To support this, it builds a 590-clip benchmark, 360VDS, and proposes S3PO, a recurrent network that drops optical-flow alignment, extracts joint features from three unaligned panorama frames with spatial and channel attention, and trains with a latitude-weighted smooth-L1 loss. On 360VDS and the MiG Panorama test set, the paper reports S3PO as the best model on PSNR, SSIM, WS-PSNR, and WS-SSIM under both bicubic and blur degradations, including against the prior 360-degree-specific model. A domain-adaptation step, pretraining on conventional VSR data and then fine-tuning on 360VDS, is presented with ablations showing each spherical-aware component contributes to the gain.","feed_headline":"S3PO beats 360° video super-resolution baselines without optical flow","feed_subtitle":"A recurrent model with a latitude-weighted loss lifts PSNR on 360° video, backed by a new 590-clip benchmark.","key_machinery":"The load-bearing mechanism is the Weighted Spherically Smooth-L1 (WSS-L1) loss, whose per-pixel weight is $\\psi_{i,j} = \\cos\\left(\\frac{(i + 0.5 - \\mathrm{height}/2)\\pi}{\\mathrm{height}}\\right)$, so training emphasis falls on the equatorial band where ERP distortion is smallest and viewer attention is highest. Around it sits the 360-degree Feature Extractor, which takes three consecutive unaligned ERP frames, builds a joint feature map with shared convolutions, correlates each neighbour with the target frame, and applies spatial and channel attention; these local features are then fused with a recurrent hidden state and the previous super-resolved output, refined in ten dual-duct residual blocks with mutual information exchange, and upsampled by pixel shuffle. Replacing explicit optical-flow alignment with this correlated-feature extraction is what lets the model handle large motion and cyclic motion across the ERP seam.","core_discovery":"On the paper's own terms, the central discovery is that a VSR model built around ERP geometry can outperform both conventional video super-resolution models and existing 360-degree image and video super-resolution approaches on 360-degree content. Table I reports S3PO at 27.26 dB PSNR under bicubic degradation and 27.51 dB under blur degradation on 360VDS, ahead of the best conventional baselines, RBPN at 27.14 dB and RSDN at 27.32 dB, with the same ordering on SSIM, WS-PSNR, and WS-SSIM. Table II reports the same claim on the four-clip MiG Panorama test set, where S3PO reaches 30.42 dB WS-PSNR versus 30.22 dB for BasicVSR. The paper also finds that conventional recurrent models are surprisingly usable on ERP frames, but that their performance is capped because they ignore spherical distortion and the continuous left-right boundary. S3PO's edge is attributed to the combination of an attention-based 360-degree feature extractor, recurrent global fusion that avoids explicit alignment, dual-duct residual refinement with mutual information exchange, and the WSS-L1 loss that weights equatorial pixels more heavily than polar pixels.","pith_inferences":["A matched-domain comparison, in which top conventional baselines are retrained or fine-tuned on the 360VDS training set under the same BI and BD degradations, could narrow the reported margins, which are often below 0.2 dB PSNR; the paper compares against baselines used only with their published pretrained weights.","The cosine latitude weight addresses vertical distortion but not the left-right seam at training time; the small gain from the S3PO-cyc variant suggests that seam-aware padding, spherical convolutions, or rotation-invariant training could push further.","Because S3PO avoids alignment and uses only three frames, its runtime and parameter profile could translate into bandwidth savings for tile-based adaptive 360-degree streaming, although the paper does not test streaming or quality-of-experience outcomes.","The benchmark's 480x360 training resolution is far below 4K deployment resolutions; the 360UHD results suggest generalization, but fine-tuning at higher resolution on a larger high-resolution 360-degree corpus is a natural next test."],"forward_implications":["A single model can super-resolve ERP 360-degree video by a factor of four without optical-flow alignment, using only three input frames, so it is applicable to both online and offline processing settings.","The new 360VDS benchmark provides 590 clips with varied spatial and temporal complexity, plus an eight-clip 360UHD high-resolution subset ranging from HD to 4K, giving the community a more diverse test bed than the four-clip MiG Panorama set.","The latitude-weighted WSS-L1 loss is proposed as a standard training objective for ERP-based 360-degree video enhancement, not just as an evaluation metric.","Conventional VSR models remain usable on 360-degree video but are improved upon by 360-degree-specific modelling, and this gap also appears on high-resolution 360UHD clips where S3PO beats BasicVSR on every clip and metric.","A cyclic-treatment variant that stitches the frame edges before feature extraction yields a small PSNR gain, confirming horizontal cyclicity as a distinct and separable source of error in ERP super-resolution."],"supporting_citations":[{"why":"Provides the RBPN recurrent back-projection baseline that S3PO is compared against and outperforms on 360VDS under bicubic degradation.","marker":"[7]"},{"why":"Provides the RSDN recurrent structure-detail baseline, the strongest blur-degradation competitor on 360VDS.","marker":"[9]"},{"why":"Provides the BasicVSR baseline, the strongest conventional recurrent model used across 360VDS, MiG Panorama, and 360UHD comparisons.","marker":"[10]"},{"why":"Defines the prior 360-degree video super-resolution work SMFN and supplies the MiG Panorama test set used for the external benchmark.","marker":"[15]"},{"why":"Defines the weighted-to-spherically-uniform SSIM metric used to evaluate quality on the sphere.","marker":"[27]"},{"why":"Supplies the SpyNet optical-flow model used in the ablation to replace the 360-degree feature extractor, showing that conventional alignment hurts ERP super-resolution.","marker":"[28]"},{"why":"Supplies the spatial and channel attention module adapted inside the 360-degree feature extractor.","marker":"[33]"},{"why":"Supplies the Vimeo90K conventional video dataset used for initial training before fine-tuning on 360VDS, enabling the domain-adaptation claim.","marker":"[43]"},{"why":"Defines the WS-PSNR distortion-aware metric and the ITU-style latitude weight map that motivates the WSS-L1 loss.","marker":"[51]"}],"fun_headline_variants":["New 360° video super-res model beats baselines without optical flow","S3PO: Recurrent 360° video super-resolution beats optical flow methods","Attention-based 360° super-res lifts PSNR, no alignment needed","New dataset and model push 360° video SR beyond conventional VSR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central comparison assumes that evaluating baselines only with their published pretrained weights and original degradation is a fair representation of their performance ceiling on 360-degree video, while S3PO is fine-tuned on the 360VDS training domain; no control retraining or significance testing is provided.","fun_headline_variants_meta":{"raw":{"variants":["New 360° video super-res model beats baselines without optical flow","S3PO: Recurrent 360° video super-resolution beats optical flow methods","Attention-based 360° super-res lifts PSNR, no alignment needed","New dataset and model push 360° video SR beyond conventional VSR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000596,"raw_usage":{"total_tokens":2862,"prompt_tokens":1090,"completion_tokens":1772,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":706,"completion_tokens_details":{"reasoning_tokens":1699}},"tokens_in":706,"tokens_out":1772,"duration_ms":11568,"temperature":1.0,"reasoning_tokens":1699,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:22:11.408235+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain or fine-tune the top two conventional baselines, such as BasicVSR and RSDN, on the 360VDS training set under the same bicubic and blur degradations used for S3PO, then evaluate on the 45-clip test set with repeated runs and confidence intervals; if their PSNR or WS-PSNR reaches or exceeds S3PO's values of 27.26 dB BI and 27.51 dB BD, or if the reported margins over frozen baselines disappear under matched training, the paper's superiority claim would be falsified.","supporting_citations":[{"cited_title":"Recurrent back-projection network for video super-resolution,","cited_arxiv_id":null,"evidence_quote":"Provides the RBPN recurrent back-projection baseline that S3PO is compared against and outperforms on 360VDS under bicubic degradation."},{"cited_title":"Video super- resolution with recurrent structure-detail network,","cited_arxiv_id":null,"evidence_quote":"Provides the RSDN recurrent structure-detail baseline, the strongest blur-degradation competitor on 360VDS."},{"cited_title":"Basicvsr: The search for essential components in video super-resolution and beyond,","cited_arxiv_id":null,"evidence_quote":"Provides the BasicVSR baseline, the strongest conventional recurrent model used across 360VDS, MiG Panorama, and 360UHD comparisons."},{"cited_title":"Weighted-to-spherically- uniform ssim objective quality evaluation for panoramic video,","cited_arxiv_id":null,"evidence_quote":"Defines the weighted-to-spherically-uniform SSIM metric used to evaluate quality on the sphere."},{"cited_title":"Optical flow estimation using a spatial pyramid network,","cited_arxiv_id":null,"evidence_quote":"Supplies the SpyNet optical-flow model used in the ablation to replace the 360-degree feature extractor, showing that conventional alignment hurts ERP super-resolution."},{"cited_title":"Ahg8: Ws-psnr for 360 video objective quality evaluation,","cited_arxiv_id":null,"evidence_quote":"Defines the WS-PSNR distortion-aware metric and the ITU-style latitude weight map that motivates the WSS-L1 loss."}],"review_version":1}