{"id":"48a51ba1-356e-4a5c-ba38-166aee694d3e","arxiv_id":"2412.07770","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A diffusion model trained on 1 million 360-degree videos synthesizes novel views with camera translation and enables 3D reconstruction from a single image.","lead":"Researchers collected over one million 360-degree YouTube videos and trained a diffusion model, ODIN, that generates new camera views of real-world scenes from a single image. The model can move the camera through a scene, and the authors demonstrate 3D reconstruction from these generated views.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"360-1M reconstruction gains are partly circular because Dust3R is used in the data pipeline, the pseudo-ground truth, and the final reconstruction; non-circular scene-geometry evidence is missing.","rationale":"The reader's weakest-assumption analysis identifies Dust3R's dual role in data curation and evaluation as the central risk, and my review converges on the same point. This is the single most load-bearing concern because the abstract's scene-level geometric claim ('infer the geometry and layout of the scene') is supported almost entirely by the 360-1M reconstruction benchmark, while the other benchmarks either measure view synthesis quality rather than geometry (MipNeRF360, DTU) or address isolated objects (GSO). The circularity is not merely a theoretical worry: the data pipeline actively selects pairs that Dust3R scores highly, so ODIN is trained to be compatible with Dust3R, and the evaluation then uses the same model to produce both the reference and the reconstruction. An independent geometry pipeline would settle whether ODIN's output is genuinely geometrically consistent or merely Dust3R-friendly. I do not see a stronger objection. The dataset and the scalable correspondence-search pipeline are real contributions, the NVS results on standard benchmarks provide some non-circular evidence of improved image quality, and the qualitative demonstrations are suggestive. Those positives are exactly why the appropriate verdict is conditional rather than rejection: the central claim is plausible but not yet independently verified. I therefore leave the reader's CONDITIONAL verdict unchanged.","tokens_in":13527,"tokens_out":2668,"duration_ms":29324,"concrete_test":"Re-run the Table 4 evaluation on the held-out 360-1M set with COLMAP replacing Dust3R in both stages: build the pseudo-ground-truth point cloud from the ground-truth source frames with COLMAP SfM/MVS, and reconstruct the ODIN-generated trajectory with COLMAP instead of Dust3R. If ODIN's Chamfer-Distance/IoU advantage over Zero-1-to-3 persists under this independent geometry pipeline, the circularity concern is resolved; if the advantage shrinks or reverses, the current 360-1M numbers overstate ODIN's geometric accuracy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that ODIN can infer the geometry and layout of real-world scenes from a single image—is supported quantitatively mainly by the 360-1M reconstruction results (Table 4 in Appendix C and Section 6.3). That evaluation is partially circular. Dust3R is used in three connected places: (1) to find frame correspondences and estimate relative poses when building the training data (Section 3.1); (2) to build the pseudo-ground-truth point cloud from all ground-truth views of each held-out 360-1M scene (Section 6.3); and (3) to reconstruct a 3D scene from ODIN's generated images (Section 5.3). Because training pairs were filtered by Dust3R confidence, ODIN is explicitly trained to produce images that Dust3R can register. It is therefore plausible that the reported Chamfer-Distance and IoU improvements reflect agreement with Dust3R's particular inductive biases rather than true geometric accuracy. The non-circular evidence is weaker: DTU shows only a small improvement and is object-centric; MipNeRF360 measures image quality, not geometry; and GSO demonstrates object-level reconstruction comparable to Zero-1-to-3, not free-camera scene geometry. Thus the scene-level geometric claim rests on an evaluation loop that the paper does not break.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces 360-1M, a dataset of over one million 360-degree videos from YouTube, together with a scalable pipeline for extracting multi-view frame correspondences using Dust3R-based pose estimation and confidence filtering, graph-based correspondence propagation, and metric scale anchoring via monocular depth. The authors train ODIN, a latent diffusion model for novel view synthesis conditioned on relative rotation and translation, with a motion-masking loss to handle dynamic content. They report improved LPIPS on DTU and MipNeRF360, improved Chamfer distance and IoU on Google Scanned Objects, and improved Chamfer distance and IoU on a held-out 360-1M set, claiming that ODIN can generate free-camera views of real-world scenes and enable 3D reconstruction from a single image.","tokens_in":13912,"tokens_out":6179,"duration_ms":53121,"significance":"If the claims hold, this is a significant contribution: 360-1M is the largest real-world multi-view dataset to date, the correspondence-search pipeline is practical and scalable, and ODIN demonstrates a new capability of long-range free-camera view synthesis from a single image. The motion-masking technique is a simple and useful idea for training on in-the-wild video. However, the quantitative evidence for the scene-geometry claim is currently undermined by an evaluation loop involving Dust3R and by missing statistical rigor. The dataset and code release, if provided, would be valuable resources.","major_comments":[{"comment":"The 360-1M reconstruction evaluation is circular. Dust3R is used (i) in Section 3.1 to select training correspondences by mean confidence threshold tau=4, (ii) in Section 6.3 to build pseudo-ground-truth point clouds from ground-truth views, and (iii) in Section 5.3 to reconstruct geometry from ODIN-generated images. Because ODIN is trained on pairs filtered by Dust3R confidence, it is incentivized to produce images that Dust3R can register; the reported Chamfer Distance and IoU then measure agreement with Dust3R's inductive biases rather than independent geometric accuracy. The authors should break this loop, for example by evaluating with an independent SfM/MVS pipeline such as COLMAP or against datasets with ground-truth 3D scans such as ScanNet or Matterport3D, and should report results with the reconstruction method fixed across all compared methods.","section":"Section 6.3 / Table 4 (Appendix C)"},{"comment":"There is a mismatch between the text and the table: Section 6.3 states 'We compare with ZeroNVS for scene reconstruction on a held-out set of 360-1M (Table 4 in Appendix),' but Table 4 is headed 'Comparison with Zero 1-to-3.' If the baseline is actually Zero-1-to-3, the comparison is not meaningful for scene-level reconstruction because that model is object-centric; if the baseline is ZeroNVS, the header must be corrected. This must be resolved before the scene-reconstruction claim can be assessed.","section":"Section 6.3 / Table 4"},{"comment":"The non-circular evidence for the central scene-geometry claim is thin. Table 1 (DTU) shows only a 0.002 LPIPS improvement over ZeroNVS on an object-centric benchmark, Table 2 (MipNeRF360) reports image-quality metrics rather than geometry, and Table 3 (GSO) is object-level and shows only a small Chamfer improvement over Zero-1-to-3. The only scene-level geometric evaluation (Table 4) is the circular one discussed above. The paper would be substantially stronger if it added a scene-level geometric evaluation that does not use Dust3R at any stage, or if it explicitly qualified the claim to exclude scene geometry.","section":"Section 6.2 / Tables 1-3"},{"comment":"No error bars, confidence intervals, or significance tests are reported anywhere in the experimental section. Given the small margins (e.g., LPIPS 0.380 vs 0.378 in Table 1; Chamfer distance 0.0717 vs 0.0697 in Table 3), the reader cannot determine whether the improvements are statistically meaningful. The authors should report variances over evaluation scenes or runs, or at least per-scene results.","section":"Tables 1-4"}],"minor_comments":[{"comment":"The abstract and Section 1 use 'Odin' while the rest of the paper uses 'ODIN'; please standardize the capitalization.","section":"Abstract / throughout"},{"comment":"Equation (2) uses epsilon_theta but the text defines the denoiser as f_theta; please align the notation.","section":"Section 5.2, Eq. (2)"},{"comment":"The paper uses 'Mip-NeRF 360' in Section 6.1 and 'MipNeRF360' in Table 2; please use one consistent name.","section":"Section 6.1 / Table 2"},{"comment":"Table 4 is referenced as evaluating 360-1M, but the table caption does not state the evaluation set; add a clear caption that identifies the dataset and the baseline.","section":"Appendix C, Table 4"},{"comment":"The paper reports an average video length of 6.3 minutes while Figure 5 shows a long-tail distribution; clarify whether the mean is computed over all videos or only over videos that yielded correspondences.","section":"Section 4.2"},{"comment":"The paper states that code, models, and dataset will be open-sourced, but no code or data access is provided in the submission; for reproducibility, please include a link or an appendix with dataset metadata details.","section":"Section 1 / Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The circularity concern raised in the stress-test note is accurate and is the main obstacle to acceptance. The 360-1M reconstruction evaluation uses Dust3R in the training-data pipeline, in the pseudo-ground-truth construction, and in the reconstruction from generated images, so the reported gains may reflect agreement with Dust3R rather than true geometric accuracy. I recommend requiring a non-circular scene-geometry evaluation, correction of the Table 4 baseline inconsistency, and the addition of statistical significance measures before the paper can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real news here is the data. 360-1M—a million 360-degree videos with 363 million frame correspondences—is a genuinely useful resource, and the graph-based propagation trick to find long-range pairs is clever. The scale calibration with Depth Anything is also sensible. If the authors release the dataset and code, that alone is worth a citation. The model, ODIN, is a fairly straightforward extension of Zero-1-to-3 with translation conditioning and a motion-masking loss, but the motion masking is a practical solution to training on dynamic video without manual filtering. The qualitative results show something real: ODIN can produce plausible free-translation views of scenes, which ZeroNVS cannot do.\n\nNow the soft spots, and they are not minor. The strongest quantitative claim—that ODIN infers scene geometry—is supported mainly by the 360-1M reconstruction benchmark in Table 4. That evaluation is partly circular. Dust3R is used in three connected places: to find training correspondences and filter them by confidence, to build the pseudo-ground-truth point cloud from all ground-truth views, and to reconstruct geometry from ODIN's generated images. Since ODIN is explicitly trained to produce images that Dust3R can register, the reported Chamfer-Distance and IoU improvements may reflect agreement with Dust3R's biases rather than true geometric accuracy. The paper never breaks this loop, for example by checking against COLMAP or multi-view stereo on a subset. The non-circular evidence is much thinner: the DTU improvement is 0.378 vs 0.380 LPIPS, MipNeRF360 measures image quality rather than geometry, and GSO shows only object-level parity with Zero-1-to-3. The paper also gives no error bars and compares against a single baseline on the 360-1M reconstruction task.\n\nThe authors are upfront about some limitations, like not modeling dynamic elements, which is fair. But the central claim needs stronger support. A referee should ask for a non-circular geometry evaluation, more baselines, and ideally an ablation separating the value of the data from the architecture changes.\n\nWho is this for? Anyone building large-scale multi-view datasets, and researchers in single-image novel view synthesis. The dataset is the contribution; the model is a promising but not yet proven step. It deserves serious peer review, with the evaluation concerns taken seriously.","headline":"The 360-1M dataset and correspondence pipeline are a genuine contribution, but the headline scene-geometry result rests on an evaluation loop that uses Dust3R for both the pseudo-ground truth and the reconstruction.","tokens_in":14371,"tokens_out":1823,"would_cite":true,"duration_ms":50546,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Training a diffusion model on over a million 360° videos lets it synthesize free-camera views of real scenes from a single image and reconstruct their 3D geometry.","keywords":["360-degree video dataset","novel view synthesis","diffusion models","3D reconstruction from a single image","multi-view correspondences","camera pose estimation","motion masking","scene generation"],"falsifier":"On a held-out set of 360° scenes, recover camera trajectories with an independent metric-scale system such as LiDAR-equipped scanning or COLMAP with known calibration, generate ODIN views from one frame, reconstruct them with Dust3R, and measure the alignment error between the reconstructed point cloud and the independent geometry; a large misalignment alongside small Dust3R reconstruction error would falsify the claim that ODIN's views are geometrically accurate.","tokens_in":13322,"feed_emoji":"🎥","tokens_out":5714,"duration_ms":50456,"temperature":0.7,"pith_summary":"The paper sets out to show that 360° video, mined at scale, can supply the multi-view training data that real-world 3D generation has been missing, and that a diffusion model trained on it can turn a single photograph into new views taken from freely moved camera positions. To that end the authors build 360-1M, a dataset of over a million 360° videos with hundreds of millions of frame correspondences and relative poses. Their model ODIN learns to generate those views and, because the views are geometrically consistent, the scene can then be reconstructed in 3D. The claim matters because prior generative models were effectively limited to rotating around a single object, whereas ODIN can move through a scene, which is what robotics, AR, and graphics applications need.","feed_headline":"Million 360° videos let one photo become a 3D scene","feed_subtitle":"Trained on 360-1M, ODIN generates new views from a single photo and reconstructs the scene.","key_machinery":"The load-bearing machinery is a scalable correspondence-mining pipeline for 360° video. Frames are sampled at one per second, each equirectangular frame is projected into four views at 90° yaw increments, and pairs within a 20-frame window are fed to Dust3R, which returns relative poses and confidence maps; a mean-confidence threshold filters out non-overlapping pairs. A graph-propagation step then links frames that share a common correspondent, recovering long-range pairs without exhaustive search, and the dimensionless Dust3R poses are anchored to metric scale by fitting a scale factor against monocular depth from Depth Anything. The generative side is a latent diffusion U-Net whose conditioning includes rotation and translation and whose output is multiplied by a learned motion mask with an auxiliary loss that keeps the mask from collapsing to zero; at inference, views are sampled along a smooth trajectory and the resulting image set is fed back through Dust3R to build the 3D scene.","core_discovery":"ODIN, a latent diffusion model conditioned on both camera rotation and translation, is trained on 360-1M and is argued to be the first model that can reasonably synthesize real-world 3D scenes and reconstruct their geometry from a single input image with free camera movement. On the DTU and Mip-NeRF 360 novel-view-synthesis benchmarks it improves LPIPS over prior single-image methods without fine-tuning, and on Google Scanned Objects and a held-out 360-1M split it improves Chamfer distance and volumetric IoU for 3D reconstruction. The enabling observation is that a 360° video contains, in principle, many views of the same content from different positions: by rotating the equirectangular projection of nearby frames, one can align them to overlapping views and recover their relative pose.","pith_inferences":["The correspondence-mining recipe should transfer to any video source with wide fields of view or camera motion, not only 360° footage, potentially enlarging the pool of real-world multi-view data.","Because Dust3R is used both to build the pseudo-ground truth and to reconstruct ODIN's generations, some of the reported geometric gains could reflect ODIN learning to produce images that Dust3R finds easy to align; an independent geometric check would separate true geometry from this feedback loop.","The metric-scale anchoring inherits the bias of monocular depth estimation, so applications like robotics that need accurate absolute scale will likely require additional calibration.","Extending motion masking from a soft filter to explicit modeling of moving objects would turn the static-scene assumption into a full 4D generator, which the paper itself identifies as the next step."],"forward_implications":["A single image of a real scene becomes enough to generate a sequence of views that supports 3D reconstruction, without per-scene optimization or known camera poses.","Novel-view-synthesis models can move the camera through an environment rather than only rotating around a central point, extending generative 3D from objects to scenes.","The motion-masking loss allows training on in-the-wild, partially dynamic video, removing the need to manually filter or curate static scenes.","ODIN improves LPIPS on Mip-NeRF 360 and Chamfer distance and IoU on Google Scanned Objects and a held-out 360-1M split relative to ZeroNVS and Zero-1-to-3.","The released 360-1M dataset, with 363 million correspondences and poses, provides a resource for other multi-view and 3D learning tasks."],"supporting_citations":[{"why":"Supplies the relative pose estimation and confidence filtering used to mine correspondences from 360° video and to reconstruct scenes from ODIN's generated views.","marker":"[52]"},{"why":"Provides the latent diffusion architecture and conditioning setup that ODIN builds on, and serves as the main object-centric comparison baseline.","marker":"[24]"},{"why":"The scene-level novel-view-synthesis baseline trained on real multi-view datasets; used for comparison on DTU, Mip-NeRF 360, and held-out 360-1M reconstruction.","marker":"[39]"},{"why":"Monocular depth estimator whose predictions anchor the dimensionless Dust3R pointmaps to metric scale in the dataset pipeline.","marker":"[60]"},{"why":"Defines the Mip-NeRF 360 real-scene benchmark used for novel view synthesis evaluation.","marker":"[4]"},{"why":"Defines the DTU multi-view benchmark used for novel view synthesis evaluation.","marker":"[1]"},{"why":"Google Scanned Objects, the object dataset used for 3D reconstruction evaluation via Chamfer distance and IoU.","marker":"[13]"}],"fun_headline_variants":["One photo, a million 360° videos, and a full 3D scene","Turning a single image into a 3D world with 1M 360 videos","From stills to scenes: Odin learns 3D from 360° video","Million 360° videos teach AI to imagine 3D scenes from one photo","Odins eye: single images become 3D scenes via 360° video training"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation of 3D reconstruction quality assumes Dust3R's pose and pointmap estimates are accurate enough to serve as ground truth, even though the same model is used to find training correspondences and to reconstruct ODIN's output.","fun_headline_variants_meta":{"raw":{"variants":["One photo, a million 360° videos, and a full 3D scene","Turning a single image into a 3D world with 1M 360 videos","From stills to scenes: Odin learns 3D from 360° video","Million 360° videos teach AI to imagine 3D scenes from one photo","Odins eye: single images become 3D scenes via 360° video training"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000608,"raw_usage":{"total_tokens":2853,"prompt_tokens":986,"completion_tokens":1867,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":1756}},"tokens_in":602,"tokens_out":1867,"duration_ms":11548,"temperature":1.0,"reasoning_tokens":1756,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:29:42.415797+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a held-out set of 360° scenes, recover camera trajectories with an independent metric-scale system such as LiDAR-equipped scanning or COLMAP with known calibration, generate ODIN views from one frame, reconstruct them with Dust3R, and measure the alignment error between the reconstructed point cloud and the independent geometry; a large misalignment alongside small Dust3R reconstruction error would falsify the claim that ODIN's views are geometrically accurate.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the latent diffusion architecture and conditioning setup that ODIN builds on, and serves as the main object-centric comparison baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Mip-NeRF 360 real-scene benchmark used for novel view synthesis evaluation."},{"cited_title":"Aanæs, R","cited_arxiv_id":null,"evidence_quote":"Defines the DTU multi-view benchmark used for novel view synthesis evaluation."},{"cited_title":"Downs, A","cited_arxiv_id":null,"evidence_quote":"Google Scanned Objects, the object dataset used for 3D reconstruction evaluation via Chamfer distance and IoU."}],"review_version":1}