{"id":"62d7d66d-3e5f-4f59-a0e3-8564faf0ca2d","arxiv_id":"2412.03552","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Imagine360 generates 360-degree equirectangular videos from ordinary perspective video anchors using a dual-branch diffusion model with antipodal attention and elevation-aware handling.","lead":"Imagine360 turns an ordinary perspective video into a full 360-degree spherical video by filling in everything outside the camera's view. The method combines a panorama branch and a perspective branch of a video diffusion model, plus attention to opposite (antipodal) pixels, to generate coherent motion across the full sphere.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'first perspective-to-360° video generation framework' claim is contradicted by the paper's own cited related work, VidPanos [31], which takes perspective videos and generates panoramic videos; the omission also weakens the superiority claim.","rationale":"The reader's weakest_assumption focused on the elevation estimator, which is a legitimate robustness concern but is explicitly acknowledged as a limitation in the paper. The present concern—that the paper's 'first' claim is contradicted by its own cited prior work—directly targets the novelty component of the central claim and is easily falsifiable by reading [31]. Even if Imagine360 is a strong technical contribution, an incorrect novelty claim and a missing comparison against the most relevant existing method materially undermine the stated contribution. The reader already flagged the evaluation as not independently verifiable; this concern adds a concrete factual dimension that can be settled by a single literature check. The CONDITIONAL verdict remains appropriate, but the conditions should explicitly include correcting the 'first' claim and adding a comparison with VidPanos if it meets the task definition.","tokens_in":15542,"tokens_out":18378,"duration_ms":186280,"concrete_test":"Retrieve the VidPanos paper (Ma et al., SIGGRAPH Asia 2024, cited as [31]) and verify whether its input is a perspective video (e.g., a casual panning video) and its output is an equirectangular 360°×180° video or other VR-viewable panoramic video. If so, the 'first perspective-to-360° video generation framework' sentence in the abstract and Sec. 5 is contradicted, and the main comparison (Tab. 1) should be re-run to include VidPanos under the same input conditions, or the paper should explicitly justify why VidPanos is not directly comparable.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract and Sec. 5 assert that Imagine360 is 'the first perspective-to-360° video generation framework.' Yet Sec. 2.3 cites VidPanos [31] (Ma et al., SIGGRAPH Asia 2024) as one of the few works on panorama video generation, without describing it or comparing against it. VidPanos takes a casually captured perspective panning video as input and produces a panoramic video that can be explored interactively, which is exactly the perspective-to-360° task defined in Sec. 1. If so, the 'first' component of the central claim is factually inaccurate. Furthermore, because VidPanos is a direct competitor, its absence from Tab. 1 and the qualitative comparisons means the 'superior graphics quality and motion coherence among SOTA' claim is not established against a method that targets the same input condition. The related-work text groups VidPanos with 360DVD and 4K4DGen and asserts that such approaches 'struggle to bridge the distribution gap between panoramic and perspective videos,' but no evidence or experiment supports that this critique applies to VidPanos, which specifically addresses perspective-to-panorama generation. This is load-bearing because the novelty claim is a binary, checkable assertion, and the performance claim lacks a head-to-head evaluation against the most relevant prior work.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Imagine360, a perspective-to-360° video generation framework that takes a narrow-FOV perspective video as anchor and generates a full equirectangular 360° video. The method builds on a dual-branch diffusion design (a panorama branch and a perspective branch) with cross-domain spherical attention augmented by an antipodal mask, plus elevation-aware data sampling and inference-time elevation estimation. The authors fine-tune spatial LoRA layers and motion modules on a combination of WEB360 and a newly collected dataset of 8,630 YouTube 360° videos. Quantitative comparisons on 100 test cases use Vbench metrics and Q-Align VQA, plus a human study, and ablations cover the dual-branch design, antipodal mask, elevation-aware designs, and fine-tuning strategy.","tokens_in":15879,"tokens_out":3646,"duration_ms":38113,"significance":"If the claims hold, Imagine360 would be a useful step toward user-friendly 360° video creation, because it uses only a perspective video as input rather than requiring text prompts plus optical flow or high-quality panoramic images. The dual-branch structure with antipodal attention is a reasonable way to inject global spherical constraints into a pretrained video diffusion model, and the elevation-aware mask handling addresses a real practical issue. The paper also contributes a nontrivial dataset and reports ablations for the main design choices. However, the central novelty claim and the superiority claim are weakened by the omission of VidPanos, a directly relevant prior method, and by evaluation concerns: the test set is undescribed, the baselines receive different input modalities, and no error bars or significance tests are given. The architecture itself appears defensible, but the load-bearing claims need substantial revision and supporting evidence.","major_comments":[{"comment":"The abstract, Sec. 1, and Sec. 5 repeatedly claim 'the first perspective-to-360° video generation framework.' Yet Sec. 2.3 cites VidPanos [31], which takes casually captured panning perspective videos and produces panoramic videos that can be explored interactively. That is the same input condition and output modality as the task defined in Sec. 1. The paper neither describes VidPanos nor compares against it, and the blanket statement that such methods 'struggle to bridge the distribution gap between panoramic and perspective videos' is made without evidence specific to VidPanos. Because the novelty claim is a binary assertion and the performance claim needs a head-to-head comparison against the most relevant prior method, this is a load-bearing omission.","section":"Sec. 2.3 and Sec. 5"},{"comment":"The comparison in Table 1 does not establish 'superior graphics quality and motion coherence among state-of-the-art 360° video generation methods' because the baselines are given different input modalities: 360DVD receives text and optical flow, AnimateDiff+LatentLab360 receives only the first frame, and Follow-Your-Canvas receives a masked perspective canvas. Only Imagine360 receives the full perspective video. The relative gains could partly reflect the additional conditioning information rather than the proposed architecture. The paper should either compare against methods that accept the same input condition (e.g., VidPanos) or restrict the claim to methods with similar input settings.","section":"Sec. 4.1 and Table 1"},{"comment":"The quantitative evaluation is performed on 'a total of 100 test cases' with no description of where these cases come from, how they are selected, or whether they overlap with the 8,630 YouTube videos collected in Sec. 3.5. If the test videos are drawn from the same distribution as the training data, the reported numbers would be inflated. Additionally, no error bars, confidence intervals, or significance tests are reported; several differences are small (e.g., Motion Smoothness 0.9806 vs. 0.9771, Subject Consistency 0.9710 vs. 0.9629). The paper should describe the test set construction, verify no overlap with training data, and report variance across multiple runs or statistical significance.","section":"Sec. 4.2"},{"comment":"The claim that 'without loss of generality, we can set the azimuthal angle to zero by assuming the viewer is rotating with the camera' is not a valid simplification for general perspective video inputs. It fails for input clips that pan, which are common in casually captured videos. The method as described can adjust for pitch changes but cannot adapt to changing yaw. Since the introduction and Sec. 3.4 describe the goal as handling 'diverse perspective video inputs' and 'general video inputs,' the paper should state clearly that panning videos are out of scope, or extend the method to handle yaw variation.","section":"Sec. 3.4"},{"comment":"The pipeline's robustness depends on an off-the-shelf single-image elevation estimator (PerspectiveFields) followed by LinearRegression smoothing. The authors acknowledge in the Limitations paragraph that inaccurate elevation estimation can produce artifacts, and Table D shows that smoothing is important. However, no analysis is given of how estimation errors translate into mask errors or how large the resulting artifacts can be across the 100 test cases. Since the elevation-aware design is one of the three key contributions, the paper should quantify the estimator's accuracy on the test set or provide a sensitivity analysis around the estimated pitch angles.","section":"Sec. 3.4 and Limitations"}],"minor_comments":[{"comment":"The sentence 'Since we are the first perspective-to-360° video generation framework, it's infeasible to find a method that has the exactly same input condition as ours' repeats the novelty claim that is contradicted by VidPanos [31] and should be revised.","section":"Sec. 4.1"},{"comment":"The notation for the elevation angle is inconsistent: the text uses both `phi` (pitch) and `theta` (yaw) in the camera pose `(FOV, theta_{1:T}, phi_{1:T})`, but later the mask positional encoding concatenates a sinusoidal embedding for `theta`; clarify which angle is being encoded.","section":"Sec. 3.1"},{"comment":"The references [54] and [55] appear to be the same paper (Taming Stable Diffusion for Text to 360 Panorama Image Generation); this duplication should be fixed.","section":"Sec. 2.2 and Sec. 2.3"},{"comment":"Figure 2 is dense and the text labels 'Spatial Layer Spatial LoRA Motion Module' are difficult to parse; a clearer legend or separate subfigures would improve readability.","section":"Fig. 2"},{"comment":"The phrase 'we can set the azimuthal angle to zero by assuming the viewer is rotating with the camera' conflates a coordinate-system choice with an assumption about the input video; the paper should separate these.","section":"Sec. 3.4"},{"comment":"The human evaluation in Table 3 reports average rankings but does not state how many videos were judged per method or whether the evaluators were blinded to method identity; this information should be added.","section":"Sec. 4.2 and Table 3"},{"comment":"When retraining 360DVD on the authors' data, the paper does not state how many training steps or which hyperparameters were used; a fair comparison of the data contribution requires this detail.","section":"Sec. B.2 and Table C"},{"comment":"The comparison to SIG-SS [18] should specify that this method was designed for a different task (spherical image generation from a few normal-FOV images) and that the comparison is therefore indirect.","section":"Sec. A and Table A"},{"comment":"The typo 'GPUss' should be corrected to 'GPUs'.","section":"Sec. 4.1"}],"recommendation":"major_revision","confidential_remarks":"The 'first perspective-to-360° video generation framework' claim is likely to be the main point of contention with the editors and reviewers. VidPanos (SIGGRAPH Asia 2024) is a directly relevant prior work that should have been compared or at least carefully distinguished. The evaluation also needs strengthening: the test set must be described and checked against the training collection, and the baseline comparisons should be made fair or the claims appropriately scoped. These issues are fixable within a revision, hence major_revision rather than reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I think the stress-test note is on target. The paper claims to be the first perspective-to-360° video generation framework, but it cites VidPanos [31], which takes a casual panning perspective video and generates a full 360° panoramic video. That is the same task. The authors mention VidPanos in Sec. 2.3 but do not describe it, compare against it, or include it in Tab. 1. Then Sec. 4.1 claims it's infeasible to find a method with the same input condition. That's a checkable binary claim, and it's wrong. The omission also weakens the 'superior to state-of-the-art' claim, because the closest competitor is not evaluated.\n\nWhat the paper does well: the dual-branch video denoising with panorama and perspective branches is a reasonable adaptation of PanFusion, and the antipodal mask for reversed motion is a new idea. The elevation-aware training/inference is also thoughtful, and the ablations in Tabs. 2 and B show each component helps. The extended YouTube dataset is a useful resource. The authors are also honest about the elevation estimation weakness in the limitations.\n\nThe soft spots beyond the VidPanos issue: the yaw assumption in Sec. 3.4. Setting the azimuth to zero 'without loss of generality' only holds for non-panning videos; in-the-wild videos often pan, and that's exactly the case VidPanos targets. The evaluation is also less convincing than it looks: the 100-case test set is not described or sourced, so it may overlap with the YouTube training data; there are no error bars; and the baselines receive different input types (360DVD gets text and optical flow, Follow-Your-Canvas gets a masked canvas, AnimateDiff gets a first frame). The qualitative claims are plausible, but the numbers don't nail down the ordering.\n\nBottom line: this is solid engineering with a few genuinely new components, but the novelty framing is inaccurate and the evaluation omits a direct competitor. A serious referee should require a position on VidPanos and a comparison (where feasible), plus better test-set documentation. I'd send it to review; the damage is reparable. For a reading group, it's a maybe—useful to see the dual-branch design, but pair it with VidPanos to keep the first-claim in perspective.","headline":"Solid dual-branch 360 video generator, but the 'first perspective-to-360' claim is contradicted by VidPanos, which the paper cites but never compares against.","tokens_in":16376,"tokens_out":4913,"would_cite":false,"duration_ms":47926,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Imagine360 claims to be the first framework to lift an ordinary perspective video into a full 360-degree equirectangular video, and reports that it beats existing text- and image-guided 360° video generators on visual quality and motion…","keywords":["360-degree video generation","perspective-to-360 video","video outpainting","diffusion models","antipodal attention","elevation-aware masking","equirectangular projection","panorama video"],"falsifier":"Take a perspective video with a measured, sharply nonlinear pitch trajectory (for example, a drone that tilts down 30° and pauses), run the inference pipeline, and compare the projected anchor region against the actual frames; if the anchor drifts or geometry distorts wherever the estimated pitch deviates from the measured trajectory, then the elevation-aware inference fails on general inputs. A second decisive check is to feed the same 40-frame clip with yaw set to a nonzero constant; if the output no longer matches the input view or the antipodal motion reverses incorrectly, the zero-yaw assumption is violated for panning footage.","tokens_in":15375,"feed_emoji":"🌐","tokens_out":8002,"duration_ms":72011,"temperature":0.7,"pith_summary":"Imagine360 sets out to establish that an ordinary narrow-field-of-view perspective video, the kind a phone records, can be turned into a complete 360-degree equirectangular video in which the unseen surroundings are visually consistent and move coherently with the input clip. The paper's central claim is that this is achievable with a dual-branch video diffusion model that denoises the panorama and perspective domains side by side, an antipodal attention mask that enforces the reversed motion seen between opposite hemispheres, and elevation-aware masking that adapts to changing camera pitch. If the claim holds, immersive spherical video creation no longer requires panoramic cameras, panoramic optical flow, or high-quality panorama images as inputs; everyday footage becomes a sufficient anchor. The paper reports that its model outperforms existing 360° video generation methods on frame quality, motion smoothness, subject consistency, and human-rated structure plausibility.","feed_headline":"A single perspective clip now grows into a full 360° video","feed_subtitle":"Imagine360 fills in the unseen hemisphere with antipodal-aware, coherent motion from a narrow-FOV anchor.","key_machinery":"The central mechanism is the dual-branch video denoising U-Net with cross-domain spherical attention. A panorama branch denoises the full equirectangular latent, while a perspective branch projects that same latent into 20 perspective views through icosahedron modeling and denoises them with a pre-trained perspective video prior; both branches are built from space-time disentangled video diffusion U-Nets, with spatial LoRA layers and a fine-tuned panorama motion module. At the end of each block, the two branches exchange information through attention whose mask highlights two kinds of correspondences: the spherical mask for pixels directly mapped between panorama and perspective domains, and the antipodal mask for each pixel's opposite-hemisphere counterpart, which extends the receptive field across the sphere and encodes the reversed-motion relationship. Circular padding keeps the left and right edges of the equirectangular frame continuous, and elevation-aware mask positional encodings condition the model on the estimated pitch trajectory.","core_discovery":"The paper claims to be the first perspective-to-360° video generation framework, and the design goal is to make the generated sphere follow the anchor video's appearance and motion while plausibly inventing everything outside the anchor's field of view. The discovery is that the strong generative prior of perspective video diffusion survives the transfer to the spherical domain if the model denoises a global panorama latent and local perspective views in parallel and couples them through cross-domain attention at every block. The antipodal mask sharpens that coupling by explicitly linking each pixel with its counterpart on the opposite hemisphere, so the model learns that forward motion in the viewing direction implies backward motion in the opposite direction. The elevation-aware designs let the mask geometry follow the input's pitch trajectory, so the anchor can sit higher or lower in the sphere across frames. The paper's experiments, including human evaluation, report the best graphics quality, motion smoothness, subject consistency, and panorama structure plausibility among the compared methods.","pith_inferences":["Because the antipodal mask is a generic architectural prior for spherical coordinates, it could be inserted into other equirectangular generative models, such as text-to-360° video generators, to improve long-range motion consistency without retraining from scratch.","A testable extension is to feed the same anchor video at different yaw offsets; if the 'viewer rotates with the camera' assumption holds, the generated panorama should be equivalent up to a horizontal shift, and the antipodal motion should reverse accordingly.","The method's reliance on a single-image pitch estimator implies that a video-native or temporally aware elevation estimator would likely reduce artifacts on handheld footage with rapid, nonlinear tilts—an improvement the paper itself flags as future work.","The antipodal-motion prior could transfer to other spherical vision tasks, such as 360° video inpainting or depth completion, wherever opposite-hemisphere consistency is physically required."],"forward_implications":["Perspective video anchors can replace text prompts or panoramic optical flow as the control signal for 360° video generation, opening the task to ordinary recorded footage.","Fine-tuning only spatial LoRA layers plus the panorama-branch motion module, while freezing the perspective motion module, is sufficient to learn spherical motion from a few thousand panorama videos.","Explicit antipodal attention yields visible reversed motion: when the camera advances in the anchor view, objects in the opposite view recede.","Elevation-aware masking and linear smoothing of the pitch estimate reduce geometry artifacts in videos with changing camera tilt.","The same trained pipeline performs panorama image outpainting, producing more style-coherent results than dedicated image-based panorama outpainting methods."],"supporting_citations":[{"why":"Supplies the dual-branch panorama–perspective design that the framework extends from image generation to video.","marker":"[54]"},{"why":"Provides the video outpainting baseline, the anchor-video feature extraction via a query-based Transformer, and the motion-module initialization.","marker":"[6]"},{"why":"Provides the space-time disentangled video diffusion U-Net structure and motion module that both branches build on.","marker":"[16]"},{"why":"Estimates per-frame elevation at inference; the elevation-aware pipeline and its acknowledged failure mode both rest on this estimator.","marker":"[25]"},{"why":"Is the main text-guided 360° video baseline and the source of the public panorama video dataset used in training.","marker":"[46]"},{"why":"Is the image-guided 360° video baseline that animates a panoramic image, one of the comparison methods.","marker":"[30]"},{"why":"Supplies the video quality, motion smoothness, and subject consistency metrics used in evaluation.","marker":"[24]"},{"why":"Supplies the learned overall video quality assessment score.","marker":"[47]"},{"why":"Is used to segment collected web videos into shots during dataset construction.","marker":"[39]"},{"why":"Generates the text captions for the collected web videos used as conditioning during training.","marker":"[9]"}],"fun_headline_variants":["Perspective clip becomes full 360° video with antipodal mask","First framework to lift single view into coherent 360° video","Imagine360: One anchor view expands to entire scene sphere","Dual-branch diffusion turns perspective clips into 360° video","Elevation-aware model invents unseen hemisphere from one clip"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the off-the-shelf single-image pitch estimator, after linear regression smoothing, yields a correct enough elevation trajectory to place the anchor video on the spherical canvas; if the elevation estimate is off, the generated 360° video contains artifacts, as the paper itself notes.","fun_headline_variants_meta":{"raw":{"variants":["Perspective clip becomes full 360° video with antipodal mask","First framework to lift single view into coherent 360° video","Imagine360: One anchor view expands to entire scene sphere","Dual-branch diffusion turns perspective clips into 360° video","Elevation-aware model invents unseen hemisphere from one clip"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000396,"raw_usage":{"total_tokens":2101,"prompt_tokens":996,"completion_tokens":1105,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":1019}},"tokens_in":612,"tokens_out":1105,"duration_ms":8814,"temperature":1.0,"reasoning_tokens":1019,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:16:15.954164+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a perspective video with a measured, sharply nonlinear pitch trajectory (for example, a drone that tilts down 30° and pauses), run the inference pipeline, and compare the projected anchor region against the actual frames; if the anchor drifts or geometry distorts wherever the estimated pitch deviates from the measured trajectory, then the elevation-aware inference fails on general inputs. A second decisive check is to feed the same 40-frame clip with yaw set to a nonzero constant; if the output no longer matches the input view or the antipodal motion reverses incorrectly, the zero-yaw assumption is violated for panning footage.","supporting_citations":[{"cited_title":"Taming stable diffusion for text to 360 panorama image generation","cited_arxiv_id":null,"evidence_quote":"Supplies the dual-branch panorama–perspective design that the framework extends from image generation to video."},{"cited_title":"Perspective fields for single image cam- era calibration","cited_arxiv_id":null,"evidence_quote":"Estimates per-frame elevation at inference; the elevation-aware pipeline and its acknowledged failure mode both rest on this estimator."},{"cited_title":"360dvd: Controllable panorama video generation with 360-degree video diffusion model","cited_arxiv_id":null,"evidence_quote":"Is the main text-guided 360° video baseline and the source of the public panorama video dataset used in training."},{"cited_title":"Vbench: Com- prehensive benchmark suite for video generative models","cited_arxiv_id":null,"evidence_quote":"Supplies the video quality, motion smoothness, and subject consistency metrics used in evaluation."},{"cited_title":"Q-align: Teaching lmms for visual scoring via discrete text-defined levels","cited_arxiv_id":null,"evidence_quote":"Supplies the learned overall video quality assessment score."}],"review_version":1}