{"id":"7bb47b70-a2ce-4634-88e2-a163ebb7bd5a","arxiv_id":"2505.02388","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"MetaScenes converts 706 ScanNet scenes into simulatable 3D replicas with 15,366 objects and ranked candidate assets, and introduces Scan2Sim for automated asset replacement.","lead":"The authors built MetaScenes, a dataset of 706 simulated 3D rooms converted from real-world scans, with 15,366 objects across 831 categories and human-ranked replacement assets, plus a model that automates picking good replacements. The value is in offering a scalable path from real scans to interactive 3D scenes for robot training and sim-to-real transfer.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The quantitative support for MetaScenes and Scan2Sim relies on the same human annotation procedure it aims to automate; without inter-annotator agreement or independent pose/quality ground truth, reported gains may reflect annotation conventions rather than true replica fidelity.","rationale":"The paper's central claim—that real-world scans can be converted into large, high-quality, simulatable training scenes without artist-driven design—requires that the human ranking/placement annotations be reliable and that the quality evidence be independent of those annotations. The reader's weakest assumption correctly identifies ranking reliability as a key risk. My read extends that concern: the pose-alignment ground truth on ScanNet++ is generated by the same annotation procedure, and the Sec. 3.3 CD comparison has no error bars and mixes image-to-3D candidates conditioned on the target image, so the evaluation loop is partly closed. This is an internal independence problem, not a disagreement with field consensus. The dataset's scale, category diversity, and physics-optimization machinery are credible contributions, and the existence of the resource is not in doubt. But the headline quantitative claims—Top-1 asset selection, pose-alignment gains, and the CD advantage over Scan2CAD—are only as strong as the annotation procedure, whose reliability is stated but not measured. A re-annotation study would settle whether these gains survive independent labels. The reader's CONDITIONAL verdict remains appropriate; no change is needed, but the condition should explicitly include independent re-annotation and noise-floor reporting.","tokens_in":24593,"tokens_out":6207,"duration_ms":86280,"concrete_test":"Have an independent annotation team, blind to the original labels, re-rank asset candidates and re-place objects for a random sample of ~500 objects across ~50 MetaScenes scenes plus the 10 ScanNet++ scenes used in Table 3. Compute rank correlation / Krippendorff's alpha for rankings and the displacement between the two pose placements. Re-run Scan2Sim's evaluation on the subset where the two teams agree; if Top-1 or pose CD changes materially, or if raw agreement is below a pre-registered threshold, the reported gains are not established. As a secondary probe, repeat Table 2 with image-to-3D generated candidates removed to check for reconstruction leakage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. 3.2 builds the dataset around human rankings of replacement assets; those rankings are the training signal for Scan2Sim and the reference for every asset-selection metric in Table 2. The only reliability check reported (Supp. A.2) is a 10% per-batch QC pass with a 98% threshold; no inter-annotator agreement, no rank correlation, and no noise floor is given. If the rankings are noisy or biased, Top-1 of 28.4%, the CD values, and the comparison to Scan2CAD all inherit that bias.\n\nThe same issue affects the independent-looking pose-alignment evaluation: Sec. 4.1 states that the ScanNet++ ground truth 'is annotated following the same procedure in Sec. 3.2.' The pose-alignment gains on ScanNet++ (CD 0.21 vs ACDC 0.26) can therefore indicate agreement with the annotators' placement conventions (center alignment, longest-side scaling, 30-degree increments) rather than true geometric fidelity to the scan.\n\nSec. 3.3's quality comparison (0.25 vs 0.35) is reported without error bars, normalization, or a description of how CD is computed, and the candidate pool includes image-to-3D reconstructions (TripoSR, InstantMesh, Michelangelo) conditioned on the target object's own image, so low CD can partly reflect overfitting to the target appearance. The central claim that real scans can be converted to high-quality training scenes without artist-driven design rests on this self-referential evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MetaScenes, a large-scale simulatable 3D scene dataset built by replacing objects in 706 real-world ScanNet scans with 15,366 simulatable assets spanning 831 fine-grained categories, with at least six candidate assets per object (98,423 unique 3D assets in total). The construction pipeline combines room-layout estimation, foundation-model-driven asset curation (text-to-3D, image-to-3D, and retrieval), human ranking and placement annotation, and physics-based optimization. The paper also proposes Scan2Sim, a multimodal alignment model trained on the ranking annotations to automate asset retrieval and pose alignment, and two downstream benchmarks: Micro-Scene Synthesis for small-object layouts and cross-domain vision-and-language navigation (VLN). The central claims are that MetaScenes offers a scalable alternative to artist-driven scene creation, that Scan2Sim outperforms existing baselines on asset selection and pose alignment, and that training on MetaScenes improves agent generalization and sim-to-real transfer.","tokens_in":24902,"tokens_out":4697,"duration_ms":50896,"significance":"If the claims hold, MetaScenes is a substantial resource: it provides a large real-to-sim dataset with per-object candidate pools, human preference rankings, physical attributes, spatial relations, and downstream benchmarks, and it tackles a genuinely important scalability problem in embodied AI. The paper deserves credit for combining dataset construction with a concrete baseline model, two evaluation tasks, and a real-robot deployment, and for attempting quantitative quality analysis against Scan2CAD. However, the load-bearing quality evidence is currently thinner than the claims require. The human ranking annotations are both the training signal for Scan2Sim and the reference for all asset-selection metrics, yet no inter-annotator reliability is reported; the ScanNet++ pose ground truth is produced by the same annotation procedure; and the Chamfer-distance comparisons lack variance and metric details. These are fixable with additional analysis and experiments, but they currently leave the core quality claims under-supported.","major_comments":[{"comment":"The ranking annotations are the training signal for Scan2Sim (Eq. 2) and the reference for every asset-selection metric in Table 2, but the only reliability check reported is a 10% per-batch QC pass with a 98% threshold. No inter-annotator agreement, rank correlation, or noise floor is reported. If the rankings are noisy or systematically biased, the Top-1 accuracy of 28.4%, the CD/ECD/IoU/color-histogram comparisons, and the conclusion that Scan2Sim outperforms baselines all inherit that noise. Please report inter-annotator agreement (e.g., Kendall's W or Krippendorff's alpha on a subset of objects annotated by multiple annotators) and a sensitivity analysis of Table 2 under label noise, such as training with corrupted rankings or evaluating on a consensus-only subset.","section":"Sec. 3.2, Supp. A.2, Eq. (2), Table 2"},{"comment":"The ScanNet++ pose ground truth 'is annotated following the same procedure in Sec. 3.2', meaning the same annotation interface and placement conventions (center alignment, longest-side scaling, and 30-degree rotation increments) were used. The pose-alignment gains on ScanNet++ (CD 0.21 vs ACDC 0.26) may therefore indicate agreement with the annotators' placement conventions rather than true geometric fidelity to the scans. Please evaluate on independent pose ground truth, for example ICP-refined alignments or manually verified absolute poses with finer rotation resolution, and report absolute pose errors rather than only differences from the annotation convention.","section":"Sec. 4.1, Table 3"},{"comment":"The Chamfer-distance quality comparison (0.25 vs 0.35) is reported without error bars, normalization, or a precise definition of the metric, including which point clouds are compared, whether the distance is symmetric or one-sided, and what units are used. Additionally, the candidate pool includes image-to-3D reconstructions (TripoSR, InstantMesh, Michelangelo) conditioned on the target object's own image, so low CD may partly reflect appearance overfitting to the target view rather than true replica fidelity. Please specify the CD computation protocol, report per-category statistics and variances, and ablate retrieval-only vs generation-inclusive candidate pools to separate these effects.","section":"Sec. 3.3"},{"comment":"The VLN results are reported as single runs without multiple seeds or variance, and the Heldout Scenes differences are small (SR 52.64 vs 51.21 for ProcTHOR, with the combined dataset at 51.36). The Heldout Domains evaluation uses only 10 ScanNet++ scenes. These results are load-bearing for the claim that MetaScenes improves agent generalization, but without variance estimates and significance testing the observed gains may be within noise. Please report means and standard errors over at least three seeds and a paired statistical test (e.g., bootstrap over trajectories or scenes).","section":"Sec. 4.3, Table 5"}],"minor_comments":[{"comment":"The modality notation 'I+TØI', 'T→P', etc., is used without definition; please define the arrow notation in the text or table caption.","section":"Sec. 4.1, Table 2"},{"comment":"The losses in Eqs. (2) and (3) use vector-valued scores q but do not explicitly show the softmax or candidate dimension; clarify the indexing and whether σ is softmax over the L candidates.","section":"Sec. 3.4, Eqs. (1)-(3)"},{"comment":"There is a typo: 'ProcPHOR' should be 'ProcTHOR'.","section":"Sec. 4.3, para. 2"},{"comment":"The phrase 'similarity score' is used for a Chamfer distance where lower is better; please call it a distance or clarify the sign convention.","section":"Sec. 3.3"},{"comment":"The labels 'A/B/C' in Fig. 1 for scene-level randomization and object-level augmentation are not explained in the caption; please annotate them.","section":"Fig. 1, Sec. 3.2"},{"comment":"The asterisk in 'Shape-E*' and 'Michelangelo*' is not defined in the main text; define it as texture optimization in the caption.","section":"Fig. A5, Supp. A.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong candidate if the authors can supply the missing reliability evidence and de-circularize the quality evaluations. The lack of inter-annotator agreement and the self-referential ScanNet++ pose ground truth are the main risks; both are addressable within the paper's scope. I would also encourage an explicit limitations section, since the current manuscript reports no error bars or negative results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one for the dataset, not the methods. The genuinely new piece is the ranked-candidate annotation layer: 15,366 objects across 831 categories in 706 ScanNet-derived simulatable rooms, each with 10 human-ranked replacement assets, 98,423 total, assembled from text/image-to-3D generation and Objaverse retrieval. Scan2CAD, R3DS, and ACDC all give you a single best match; they don't give you a ranked candidate pool with replacement rationale. The Micro-Scene Synthesis benchmark for small-object layouts is a real gap-filler, and the VLN transfer result (about 5.3% SR over ProcTHOR on held-out ScanNet++ scenes) is the kind of claim that matters for embodied AI. The construction is also clearly serious: MCMC physics optimization, verified scene graphs, quality-control checks on annotation batches.\n\nThat said, the evaluation is thinner than the resource, and every soft spot is fixable. The pose-alignment gains on ScanNet++ (CD 0.21 vs ACDC 0.26) are weakened because the ScanNet++ ground truth was annotated using the same procedure as the training data - the comparison partly measures agreement with the annotators' placement conventions (center alignment, longest-side scaling, 30-degree rotation increments) rather than fidelity to the scan. The Sec. 3.3 quality comparison (0.25 vs 0.35) has no metric definition, normalization, or error bars, and since the candidate pool includes image-to-3D reconstructions conditioned on the target object's own image, low CD can partly reflect appearance overfitting. Table 5 is the strongest empirical claim, but there are no seeds or variance reported; a 5% gap on 10 held-out scenes could shift. And the supervision itself - the human rankings - deserves an inter-annotator agreement number; the 10% QC pass at a 98% threshold is an audit, not a reliability statistic. Minor note: the real-world AGV deployment is qualitative only, and the physics attributes come from GPT-4V estimates rather than measurement, which is fine for training data but not a physical ground truth.\n\nNone of this kills the central claim. The dataset's existence is credible, and its value for embodied AI training is plausible. But the headline claim of automated, high-quality replacement inherits the evaluation's weaknesses, so the paper should not be accepted as-is. For a reader: anyone building real-to-sim training environments gets something here. Send it to serious review - it merits referee time - with the requirement that the authors release the dataset link and code, report seeded evaluation, define the CD metric, and clarify the ScanNet++ annotation protocol. Conditional accept after revision, not a desk reject.","headline":"A substantial real-to-sim dataset resource with genuinely new ranked-candidate annotation; the evaluation claims need a disclosure-and-rigor pass before acceptance.","tokens_in":25487,"tokens_out":4559,"would_cite":true,"duration_ms":54485,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that human preference rankings of candidate 3D assets let a learned model turn real-world scans into large interactive training scenes without artist-driven design.","keywords":["3D scene dataset","real-to-sim","asset retrieval","multi-modal alignment","embodied AI","scene synthesis","vision-and-language navigation","ScanNet"],"falsifier":"Have several independent annotators rank the same candidate asset sets for a sample of objects and measure rank agreement; if agreement is low, or if a Scan2Sim model trained on one set of rankings predicts another annotator's choices no better than a random baseline, then the claim that these rankings teach a reliable automated asset-selection model would be refuted.","tokens_in":24400,"feed_emoji":"🏠","tokens_out":5656,"duration_ms":64669,"temperature":0.7,"pith_summary":"MetaScenes claims that a large, simulation-ready 3D scene dataset can be built from real-world scans by replacing every scanned object with a high-quality simulatable asset, and that the replacement process can be automated well enough to remove the usual reliance on artist-designed scenes. The dataset contains 15,366 objects across 831 fine-grained categories in 706 rooms reconstructed from ScanNet, with at least six candidate assets per object. The paper's key move is to have human annotators rank those candidates, turning subjective replacement quality into training labels for Scan2Sim, a multimodal model that selects the best asset from images, text, and point clouds. Two benchmarks, micro-scene synthesis for small-object layouts and cross-domain vision-and-language navigation, are used to show that training on these scenes transfers to unseen rooms and to a real robot. If true, embodied AI could scale training environments directly from everyday scans rather than from manual 3D design.","feed_headline":"15,366 scanned objects get automatic simulatable doubles","feed_subtitle":"Human-ranked asset swaps let a learned model rebuild real rooms as interactive training arenas.","key_machinery":"The load-bearing mechanism is Scan2Sim, a multi-modal contrastive retrieval model that aligns an object's image and text description with candidate 3D point clouds, trained with a cross-modal matching loss against human ranking annotations. It is supported by a construction pipeline that uses SAM and GPT-4V to caption scanned objects, text-to-3D retrieval and generation to build candidate assets, and a physics-based MCMC optimization that adjusts placements to remove collisions and floating objects. The rankings are what convert subjective replacement quality into a learnable objective, so the model and the benchmarks both depend on them.","core_discovery":"The central claim is that human preference rankings of replacement assets provide the supervision needed to learn automated replica creation from real-world scans. On the paper's evidence, Scan2Sim trained on these rankings selects the best asset at 28.4% Top-1 accuracy on the MetaScenes test set, outperforming baselines including GPT-4V at 16.5% and ULIP-2 at 13.1%, and the resulting replicas are closer to the original scans than Scan2CAD's, with a Chamfer distance of 0.25 versus 0.35. The same scenes, after physics-based optimization, support a navigation agent that improves held-out-domain success rate by 5.34 percentage points over training on procedurally generated ProcTHOR scenes. The paper presents the dataset itself, the Scan2Sim pipeline, and the two benchmarks as a package: the dataset is the evidence, the ranking annotations are the ground truth, and the benchmarks are the demonstration that the replicas are useful for embodied agents.","pith_inferences":["If ranking noise is the main bottleneck, then replacing or augmenting human rankings with a learned preference model, or with pairwise comparisons from a larger annotator pool, could push Scan2Sim-style selection well beyond the reported 28.4% Top-1 accuracy.","The VLN result that navigation to small items is a weak point suggests the same dataset could be used to probe whether object-goal navigation failures concentrate on small objects, and whether manipulation policies trained in these scenes inherit that benefit.","Because the candidate pool is built partly from generative models, the pipeline's ceiling is tied to generator quality; as image-to-3D generation improves, the same annotation pipeline should yield higher-fidelity replicas without redesign.","A per-category breakdown of Chamfer distance versus Scan2CAD would clarify whether the reported accuracy gain comes uniformly from all objects or mostly from small items where Scan2CAD has no equivalents."],"forward_implications":["New real-world scans, such as the ScanNet++ scenes tested in the paper, can be converted into simulatable replicas without per-scene artist work.","Training on MetaScenes improves object-goal navigation on unseen scenes and domains compared with training on procedurally generated scenes alone.","The Micro-Scene Synthesis benchmark makes small-object layout generation a measurable task, giving manipulation research a data source it previously lacked.","The ranking annotations provide a reusable evaluation target for any future automated asset-selection system.","The resulting scenes are physically optimized and interactable, so they can be dropped directly into embodied-agent simulators."],"supporting_citations":[{"why":"Supplies all 706 real-world ScanNet scenes and object point clouds that MetaScenes converts into replicas.","marker":"[9]"},{"why":"Provides the large Objaverse asset pool used for text-to-3D retrieval of candidate replacements.","marker":"[14]"},{"why":"Scan2CAD is the prior ScanNet-to-CAD conversion baseline that MetaScenes compares against for diversity and accuracy.","marker":"[1]"},{"why":"ULIP-2 supplies the frozen image and text encoders and the pretrained point-cloud encoder used to build Scan2Sim.","marker":"[99]"},{"why":"GPT-4V generates the object captions and physical attributes used to create candidate assets and scene metadata.","marker":"[103]"},{"why":"ACDC serves as the baseline method for automated asset selection and pose alignment in replica creation.","marker":"[10]"},{"why":"SPOC is the navigation model trained and evaluated in the vision-and-language navigation domain-transfer benchmark.","marker":"[18]"},{"why":"PhyScene is one of the physics-guided scene synthesis baselines evaluated on the Micro-Scene Synthesis task.","marker":"[101]"},{"why":"ScanNet++ provides the held-out real-world domains used to test cross-domain transfer of the replica-creation pipeline.","marker":"[104]"}],"fun_headline_variants":["Automated replica creation from 15k scanned objects","Scan2Sim uses human-ranked swaps to build virtual twins","AI learns to swap assets for realistic 3D scenes","15k scans become interactive training arenas automatically","Replica pipeline beats GPT-4V on asset choice"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that human annotators rank replacement assets consistently and correctly; the paper reports quality checks on only 10% of batches and gives no inter-annotator reliability numbers, so if those rankings are noisy or biased, the retrieval model, the Chamfer-distance comparison, and the benchmark conclusions all inherit that noise.","fun_headline_variants_meta":{"raw":{"variants":["Automated replica creation from 15k scanned objects","Scan2Sim uses human-ranked swaps to build virtual twins","AI learns to swap assets for realistic 3D scenes","15k scans become interactive training arenas automatically","Replica pipeline beats GPT-4V on asset choice"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000712,"raw_usage":{"total_tokens":3212,"prompt_tokens":964,"completion_tokens":2248,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":580,"completion_tokens_details":{"reasoning_tokens":2170}},"tokens_in":580,"tokens_out":2248,"duration_ms":18248,"temperature":1.0,"reasoning_tokens":2170,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:52:21.729952+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have several independent annotators rank the same candidate asset sets for a sample of objects and measure rank agreement; if agreement is low, or if a Scan2Sim model trained on one set of rankings predicts another annotator's choices no better than a random baseline, then the claim that these rankings teach a reliable automated asset-selection model would be refuted.","supporting_citations":[{"cited_title":"Ulip-2: Towards scalable multimodal pre-training for 3d understanding","cited_arxiv_id":null,"evidence_quote":"ULIP-2 supplies the frozen image and text encoders and the pretrained point-cloud encoder used to build Scan2Sim."},{"cited_title":"Physcene: Physically interactable 3d scene synthesis for embodied ai","cited_arxiv_id":null,"evidence_quote":"PhyScene is one of the physics-guided scene synthesis baselines evaluated on the Micro-Scene Synthesis task."},{"cited_title":"Scannet++: A high-fidelity dataset of 3d indoor scenes","cited_arxiv_id":null,"evidence_quote":"ScanNet++ provides the held-out real-world domains used to test cross-domain transfer of the replica-creation pipeline."}],"review_version":1}