{"id":"bd5f6d67-8d26-44c5-a7eb-471f62a89484","arxiv_id":"2508.11898","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A multi-camera robot policy that builds a bird's-eye-view representation with deformable attention reportedly improves in-distribution, out-of-distribution, and few-shot manipulation performance by 11%, 17%, and 84% over baselines.","lead":"OmniD is a robot-control method that merges images from several cameras into one top-down view, and it is claimed to generalize better when cameras or backgrounds change. The reason to read it is that robots that only work in their training lab are a central barrier to real-world deployment, and the paper reports especially large gains in few-shot settings.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Full text is an unrelated cond-mat paper (Cr2Ge2Se3Te3); the abstract's OmniD claims are entirely unsupported by the provided manuscript.","rationale":"The reader's verdict is REJECT with low confidence; I agree with rejection but locate the load-bearing problem at an earlier point than the reader's weakest_assumption. The reader identifies possible benchmark artifacts as the weakest assumption. That is a relevant secondary concern, but before any benchmark validity question can even be asked, the submitted full text must actually describe OmniD and its experiments. It does not: the body is an unrelated condensed-matter paper. Thus the most load-bearing concern is the absence of evidence for every component of the central claim. This is not an internal inconsistency of a technical argument but a complete mismatch. I would treat this as a compilation/upload artifact risk; if the correct full text is retrieved, the verdict could change. Under the provided manuscript, keeping the reader's REJECT verdict is appropriate.","tokens_in":5562,"tokens_out":3506,"duration_ms":40269,"concrete_test":"Use the arXiv API to fetch the official metadata and PDF for 2508.11898 and 2508.11899. Check (a) whether the title/abstract for 2508.11898 is the OmniD abstract and not the Cr2Ge2Se3Te3 abstract; (b) whether the downloaded PDF full text contains the OmniD method sections (BEV representation, deformable attention OFG, diffusion policy) and the benchmark tables with the claimed percentages, or instead contains the Cr2Ge2Se3Te3 calculations. If the PDF does not contain OmniD material, the concern lands and the abstract's quantitative claims have no in-scope supporting evidence. If a correct OmniD full text exists, re-review the technical claims using that text.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a deformable-attention BEV diffusion policy ('OmniD') beats baselines by 11%, 17%, and 84% in in-distribution, OOD, and few-shot settings. For this claim to hold, the paper must present the method and the experiments. The supplied full text is a first-principles study of Cr2Ge2Se3Te3 (arXiv:2508.11899v1) with no mention of OmniD, BEV, diffusion policy, robotics, or any benchmark. No architecture, no training details, no baseline definitions, no dataset, no evaluation protocol, no results. Therefore the evidence base required for the central claim is absent. This is not a subtle flaw in a derivation; it is a complete mismatch between the claimed contribution and the supporting text. The only way the claim could still be true is if the correct OmniD manuscript exists and this is an upload/compilation artifact, but as submitted the paper cannot be verified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission is internally inconsistent. The abstract and title describe OmniD, a multi-view robot visuomotor diffusion policy with a bird's-eye-view (BEV) representation and a deformable-attention Omni-Feature Generator (OFG), and claim average improvements of 11%, 17%, and 84% over the best baseline in in-distribution, out-of-distribution, and few-shot experiments. The full text, however, is a condensed-matter first-principles study of the monolayer Cr2Ge2Se3Te3, with no robotics content, no BEV representation, no diffusion policy, no OFG, and no experimental evaluation. The central claim of the paper is therefore completely unsupported by the supplied manuscript.","tokens_in":5778,"tokens_out":1944,"duration_ms":24056,"significance":"If the abstract's claims were substantiated by a proper method section and benchmark evaluation, the proposed approach would be of interest to the robot-learning community: multi-view BEV fusion for better out-of-distribution generalization and few-shot manipulation is an active and important problem, and a public benchmark and training-code release would be a useful contribution. However, as submitted, the manuscript contains no method description, no evaluation protocol, no baseline definitions, and no results. The significance of the claimed contributions cannot be assessed because the evidence base is absent.","major_comments":[{"comment":"The central claim—that OmniD achieves 11%, 17%, and 84% average improvement over the best baseline—is stated in the abstract but nowhere supported by the body. The supplied full text is arXiv:2508.11899v1, a DFT study of Cr2Ge2Se3Te3, and contains no mention of OmniD, BEV, diffusion policy, robot manipulation, or any benchmark. This is not a missing detail but a complete absence of the method and evidence required to evaluate the paper.","section":"Abstract and full text"},{"comment":"Even if the correct body had been uploaded, the abstract's quantitative claims are presented without the experimental context needed to interpret them: no task suite, no demonstration counts, no camera-perturbation protocol, no baseline list, no standard deviations, and no ablations. The 84% few-shot figure is particularly under-determined. Without these details the claimed improvements are not verifiable, and the out-of-distribution and few-shot gains could be benchmark artifacts rather than evidence of generalization.","section":"Abstract (experimental claims)"},{"comment":"The architecture that is the paper's contribution—the Omni-Feature Generator (OFG), deformable attention, image-to-BEV synthesis, and the diffusion policy—is not defined anywhere in the submitted manuscript. The reader cannot determine what is proposed, how it differs from prior BEV visuomotor policies, or what hyperparameters, input representations, or training objectives are used. The manuscript is internally inconsistent: the title and abstract promise a robotics paper, while the body delivers an unrelated materials-science paper. This cannot be fixed by local revision; the manuscript would need to be replaced.","section":"Full text, Sections I–III"}],"minor_comments":[{"comment":"The GitHub link (https://github.com/1mather/omnid.git) is listed but no repository contents, license, or usage instructions are described in the manuscript; as submitted, the link is unverifiable.","section":"Abstract"},{"comment":"The full text uses a different arXiv identifier (2508.11899) and has a different title and author list from the abstract. This should be corrected in any resubmission to avoid a mismatch between the declared and actual content.","section":"Full text, headers"}],"recommendation":"reject","confidential_remarks":"For the editor: the mismatch between the abstract and the full text is so complete that this looks like the wrong file was uploaded rather than a substantive flaw in the intended robotic-policy paper. However, as submitted, the manuscript cannot be reviewed as a robotics paper because none of the claimed method or results are present. If the authors submit the correct full text, it should be treated as a new submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, the abstract describes a robot manipulation policy (OmniD) with a deformable-attention BEV fusion and big gains (11/17/84%); the full text is a first-principles study of a Janus monolayer Cr2Ge2Se3Te3. Different field, different authors, different claims. Second, that mismatch is the whole story: whatever merit the materials work has, it has zero connection to the claimed robotics result, and the robotics result has zero supporting evidence here.\n\nWhat is good: the abstract is readable and the architecture is plausible. Combining deformable attention with BEV for visuomotor diffusion policies is a reasonable extension of things that work in driving; the OFG idea (selectively attend to task-relevant features, suppress view-specific noise) is sensible. The claimed gains, if reproducible, would matter for generalization and few-shot learning in manipulation. Also, the abstract points to a GitHub repo, which suggests the authors intend to ship code and a benchmark.\n\nSoft spots, in proportion. The absence is total. No method section, no task suite, no baseline definitions, no standard deviations, no ablations, no camera perturbation protocol, no demonstration counts. The 84% few-shot number is especially under-determined: few-shot gains of that size live or die on whether the evaluation tasks share structure with training and whether the baseline is tuned fairly. None of that can be checked. The materials paper cannot rescue any of this; it is not the same manuscript. I would not try to extract insight from the cond-mat text about the robotics claim — there is no bridge.\n\nThe charitable interpretation is an upload/compilation mixup. The arXiv watermark in the full text is 2508.11899, while the abstract is from 2508.11898, suggesting the wrong PDF was attached. If the real OmniD paper exists, a proper review could change the picture substantially. But based on what is actually in front of us, there is nothing coherent to referee. This is not a case where I can identify a good paper with a weak section; it is a case where the evidence base is missing.\n\nBottom line: who is this for? A reader interested in the materials paper might get value from that text, but not from this submission. The robotics community gets a list of claims and no method. I wouldn't send this to peer review as submitted, and I wouldn't cite it. If the authors resubmit with the correct full text, then it deserves a serious referee. As is: desk reject.","headline":"The uploaded text is a cond-mat paper; the robotics abstract is unsupported, so this submission is unverifiable as is.","tokens_in":6301,"tokens_out":1683,"would_cite":false,"duration_ms":18589,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A multi-view robot policy that fuses camera feeds into a single bird's-eye map claims large gains over single-view baselines in new scenes.","keywords":["visuomotor policy","bird's-eye view","multi-view fusion","deformable attention","out-of-distribution generalization","few-shot imitation","diffusion policy","robot manipulation"],"falsifier":"Train OmniD and the strongest baseline on the same demonstrations, then evaluate on an OOD suite where the scenery is changed but the camera viewpoints are kept exactly the same. If OmniD's advantage over the baseline mostly disappears, the gains are about background robustness rather than viewpoint generalization. Alternatively, ablate the deformable attention by replacing OFG with a conv-based BEV encoder of equal parameter count; if the 17% OOD gap does not collapse, deformable attention is not the load-bearing component.","tokens_in":5459,"feed_emoji":"🤖","tokens_out":4200,"duration_ms":40261,"temperature":0.7,"pith_summary":"The paper claims that visuomotor policies overfit to fixed camera positions and backgrounds, and proposes a fix: fuse multiple camera views into a single bird's-eye-view (BEV) representation before generating actions. The fusion is guided by a deformable attention mechanism that selects task-relevant features and suppresses view-specific noise. On benchmarks, the method reports average improvements of 11% over the best baseline in-distribution, 17% out-of-distribution, and 84% in few-shot settings. If these numbers hold, the BEV representation is not just a convenience but a genuine generalization mechanism.","feed_headline":"One bird's-eye map beats single-view robot policies","feed_subtitle":"Deformable attention filters noisy views, giving 17% better out-of-distribution and 84% better few-shot results.","key_machinery":"The Omni-Feature Generator (OFG), a deformable attention module that projects multi-view image features onto a controllable bird's-eye-view grid. Deformable attention lets each BEV query sample a sparse set of relevant pixels across views rather than attending to the full image, which is what suppresses background and view-specific noise while keeping task-relevant geometry.","core_discovery":"OmniD builds a unified BEV representation from multiple image observations and feeds it into a diffusion policy. The Omni-Feature Generator uses deformable attention to sample only task-relevant image features, constructing a 3D-aware grid that is invariant to individual camera viewpoints and backgrounds. The central claim is that this representation directly attacks the two failure modes named in the paper: overfitting to training-time camera positions and poor multi-view fusion. The reported result is that OmniD outperforms the best baseline by 11% in-distribution, 17% out-of-distribution, and 84% in few-shot demonstrations.","pith_inferences":["The abstract does not describe the evaluation protocol, so the reported percentages should be read as claims about particular benchmarks; a natural next test is to vary camera intrinsics and lighting systematically while holding the task fixed.","A testable consequence the paper leaves implicit: if OFG is truly discarding view-specific noise, then the policy's performance should degrade gracefully as camera positions drift, rather than show a sudden cliff; measuring the performance-vs-displacement curve would directly probe the mechanism.","The BEV idea could transfer to other visuomotor settings such as mobile manipulation or navigation, where a ground-plane representation is equally natural; the paper does not claim this, but the mechanism is not task-specific."],"forward_implications":["If the reported gains generalize, switching visuomotor policies from image-space fusion to BEV fusion would reduce performance drops when robots are deployed in new rooms or with cameras at different positions.","The 84% few-shot improvement suggests that the BEV representation drastically lowers the number of demonstrations needed to learn a new task, which matters for real-world data collection cost.","The deformable attention design implies that policies can scale to many cameras without quadratic attention cost, since each BEV query attends to a small subset of features.","The method can be combined with existing diffusion policy heads, meaning it is a representation-level improvement that does not require a new action decoder."],"supporting_citations":[],"fun_headline_variants":["Bird's-eye view policy beats camera-specific overfitting","Multi-view robot policy generalizes via BEV representation","OmniD fuses views into one map for 84% better few-shot","Deformable attention builds robust BEV for robot control","Single BEV map boosts robot policy generalization by 17%"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The reported 17% out-of-distribution and 84% few-shot gains assume the benchmark tasks used to measure them really shift camera positions and task structures the way the motivation describes, and that the 'best baseline' is a fair, well-tuned comparison.","fun_headline_variants_meta":{"raw":{"variants":["Bird's-eye view policy beats camera-specific overfitting","Multi-view robot policy generalizes via BEV representation","OmniD fuses views into one map for 84% better few-shot","Deformable attention builds robust BEV for robot control","Single BEV map boosts robot policy generalization by 17%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1251,"prompt_tokens":693,"completion_tokens":558,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":473}},"tokens_in":437,"tokens_out":558,"duration_ms":5717,"temperature":1.0,"reasoning_tokens":473,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:41:45.662683+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train OmniD and the strongest baseline on the same demonstrations, then evaluate on an OOD suite where the scenery is changed but the camera viewpoints are kept exactly the same. If OmniD's advantage over the baseline mostly disappears, the gains are about background robustness rather than viewpoint generalization. Alternatively, ablate the deformable attention by replacing OFG with a conv-based BEV encoder of equal parameter count; if the 17% OOD gap does not collapse, deformable attention is not the load-bearing component.","supporting_citations":[],"review_version":1}