{"id":"601d701e-e202-484f-9fcb-a40a19397392","arxiv_id":"2411.14770","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"AMR is an end-to-end vision-based local navigation system that reaches a desired relative pose to an object with centimeter-level precision, using a reference image plus mask and multi-modal sensing.","lead":"AMR is a navigation system that lets a robot move to a precise spot relative to an object, such as one meter away and facing it, using only a photo of the object and a requested pose. It is trained in simulation, transfers to real robots, and reaches errors of a few centimeters, which matters for docking, inspection, and manipulation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'front side' goal specification (Sec. II) is ambiguous for objects without a dominant side and for oblique reference images, so the claim of reaching 'any object' is not supported for those cases.","rationale":"I agree with the reader's weakest assumption: the Sec. II front-side parametrization is the least secure link in the argument for 'any object'. The paper itself documents the failure mode in Fig. 7c(4), and the real-kitchen sink/cup errors (14.8 cm and 8.7 cm) provide supporting evidence. My stress-test pass did not find a separate, stronger objection. The quantitative evidence otherwise is substantial: large-scale photorealistic simulation data, 2,000 closed-loop evaluation tasks, a real-robot deployment with 17/18 completed runs for the hist=4 model, ablation studies, and a self-identified limitation. These support a conditional acceptance rather than unconditional acceptance: AMR appears to deliver centimeter-level precision for objects that have a dominant side and when the reference image is not too oblique, but the universal 'any object' wording in the abstract overstates the demonstrated scope. Because the reader's verdict already conditions on precisely this assumption, my concern does not move the verdict; it reinforces the conditional framing.","tokens_in":11808,"tokens_out":4295,"duration_ms":50649,"concrete_test":"In the Isaac Sim test scenes, select a set of rotationally symmetric or side-ambiguous objects (e.g., cylinders, round trash cans, vases) not seen in training. For each object, render a reference image from a 45° azimuth and issue the goal C = (front, d = 0.5 m, θ = 0°). Run AMR from 20 different starting poses per object and record the distribution of final robot azimuth relative to the object. If the final side selected is not consistently the intended front (e.g., the modal side appears in fewer than 70% of runs, or multiple sides appear with comparable frequency), the parametrization is ambiguous in practice and the 'any object' claim fails. Also repeat the same protocol with a reference image taken from a frontal viewpoint; if the success rate rises sharply, the ambiguity is specifically in the goal specification, not in the navigation policy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on Sec. II's goal parametrization, which defines the object's front as 'the most visible side' in the reference image and then expresses the desired approach side S relative to that front. This is well-defined only when the object has four dominant, distinct sides and the reference image is approximately frontal. For a cylindrical or symmetric object, or for a reference image taken at an oblique angle such as 45°, two or more sides are equally visible, so the same C = {S, d, θ} does not specify a unique physical goal pose; the robot literally cannot know which side is intended. The paper itself reports this failure in Fig. 7c(4): 'a robot may go to the wrong side when there is no dominantly visible side.' The real-kitchen sink and cup results are consistent with this fragility (14.8 cm and 8.7 cm mean distance errors, versus 1.8-3.1 cm for the other objects). Because the headline 3 cm / 1° statistics are computed only over completed runs and only over objects for which the convention is reasonably unambiguous, they do not substantiate the universal 'any object' claim. This is a scope overclaim rather than an internal inconsistency: the system may be very reliable for near-frontal, dominant-side objects, but the paper's title and abstract assert more.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents AMR (Aim My Robot), an end-to-end learned local navigation system that, given a reference image of a target object with a mask and a relative pose specification C={S,d,θ}, outputs base waypoints and camera tilt commands to position the robot at the desired pose with claimed centimeter-level precision. The system is trained in Isaac Sim on 500k trajectories from HSSD scenes using behavior cloning and DAgger, with a transformer that combines MAE-encoded RGB, depth-as-positional-encoding, LiDAR tokens, and autoregressive waypoint decoding. The evaluation covers 2,000 tasks in 5 held-out HSSD scenes with 166 unseen objects, plus real-kitchen trials on 6 objects; the paper reports median distance errors of roughly 2–4 cm and angular errors under about 2° in most conditions, with a 95.9% completion rate when the object is initially visible. Qualitative demonstrations include closing a fridge drawer and a forklift pallet-loading task.","tokens_in":12069,"tokens_out":4720,"duration_ms":47135,"significance":"If the reported precision holds, AMR is a significant step beyond conventional 1-m-goal navigation: it is map-free, CAD-free, trained entirely in simulation, and explicitly conditions on a target side, distance, and angle rather than just a point or image goal. The evaluation is unusually thorough for a systems paper, with large-scale photorealistic training and testing, held-out unseen objects, a meaningful ablation study, and real-robot deployment with downstream manipulation demonstrations. The central caveat is that the universal 'any object' claim is not supported for objects without a dominant side or for oblique reference images, and the paper's own failure analysis and real-world sink/cup results illustrate this limitation.","major_comments":[{"comment":"The goal parametrization defines the approach side S ∈ {front, back, left, right} relative to the 'most visible side' in the reference image, assuming 'common objects have 4 dominant sides.' For cylindrical or symmetric objects, or when the reference image is taken from an oblique angle such as 45°, two or more sides are equally visible and the same specification C={S,d,θ} does not determine a unique physical goal pose. This is acknowledged as a failure mode in Fig. 7c(4) ('a robot may go to the wrong side when there is no dominantly visible side'), and the real-kitchen cup and sink results in Table III (8.7 cm and 14.8 cm mean distance errors, versus 1.8–3.1 cm for the other objects) are consistent with this fragility. The abstract's 'reach any object' claim is therefore overstated; the paper should restrict the scope to objects with a well-defined dominant side or require a near-frontal reference image, and say so explicitly.","section":"§II, Fig. 2c"},{"comment":"The headline median errors (3 cm, 1°) are computed only over completed runs; Fig. 6 reports completion rates of 95.9% (visible) and 90.4% (invisible). The expected error over all attempted runs is therefore higher than the reported conditional median, and incomplete runs may have arbitrarily large final pose error. The paper already includes all runs in the ablation plot (Fig. 8, caption: 'We consider both complete and incomplete runs'), so the main quantitative evaluation should do the same, or the abstract and Section V should explicitly qualify the errors as completion-conditional.","section":"§V, 'We consider a run complete...'"},{"comment":"The real-world claim of 'little degradation' rests on only 3 runs per object, with hand-measured ground truth, and the sink and cup exceptions (14.8 cm and 8.7 cm distance errors) are mentioned only in the table caption. These exceptions should be discussed in the main text, because they are the largest deviations and directly bear on the generality of the method. With n=3, reporting only the mean also obscures the run-to-run spread; per-run values or a scatter plot should be provided so the reader can assess consistency.","section":"§V-C, Table III"}],"minor_comments":[{"comment":"There is a typo: 'navigate to objects with precisely' should read 'navigate to objects with precision.'","section":"§I, first paragraph"},{"comment":"The symbol C is used both for the goal condition (Section II) and for the classifier output in Section III-B3 ('C(·)'); this notational collision should be resolved to avoid confusion.","section":"§III-B2"},{"comment":"The DAgger procedure is described in one sentence in Section IV ('identified failures... and used DAgger to augment the dataset') with no details on the number of iterations, the failure criteria, or how the expert labels the augmented states; this hampers reproducibility of the training pipeline.","section":"§III-A, Trajectory generation"},{"comment":"The text states that the error distribution contains final pose error for only the completed runs, while the Fig. 8 caption says 'We consider both complete and incomplete runs'; clarify which protocol applies in each figure and table.","section":"§V, first paragraph and Fig. 8 caption"},{"comment":"The final sentence of Related Work is incomplete: 'As of today, the most general and capable mobile manipulation systems' ends without a predicate. Finish the sentence or remove it.","section":"§VI, last sentence"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically sound and the experimental effort is substantial; the main weakness is the gap between the title/abstract and the demonstrated scope. The 'any object' claim is contradicted by the paper's own failure analysis and by the real-world sink/cup errors, so the manuscript needs a careful rewriting of the claims and a brief discussion of the ambiguous-goal limitation. This is fixable without new experiments; I would not reject over it, but the current phrasing would mislead readers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious look. AMR gives you a concrete way to specify a precise relative pose to an object—reference image plus mask plus side, distance, angle—and then actually reaches it. The simulation scale is real (500k trajectories, 2000 test tasks), the real-kitchen results show the thing works on a physical robot, and the ablations convince me that LiDAR, depth, and the footprint token each earn their place. The residual autoregressive decoding is a nice trick for keeping precision while using discrete action tokens. I believe the central claim: for objects with a well-defined front side, and when the reference image is roughly frontal, the system lands within a few centimeters and a couple of degrees. That alone is a useful contribution beyond the 1-meter-radius success criterion that most prior instance-goal navigation uses.\n\nThe soft spots are real but not fatal. First, the abstract says \"any object,\" but the goal parametrization depends on the object having four dominant sides and the reference image being unambiguous. The paper's own Fig. 7c(4) shows a failure when there is no dominantly visible side (e.g., looking from 45°), and cylindrical or symmetric objects are problematic by construction. That is a scope overclaim, not an internal inconsistency, and the paper does acknowledge it. Second, the headline 3 cm / 1° numbers come from completed runs only; the completion rate is reported separately, so precision and reliability are somewhat conflated. There are no error bars or variance figures, which matters more because the real-world table is based on three runs per object. Third, training samples distances in [0.1, 0.5] m, while the real experiments use d = 1 m—outside the training distribution. It apparently still works, but that gap deserves a sentence of explanation. Fourth, there is a small numeric inconsistency: the text says 160 unseen objects, while Table I lists 166.\n\nI would not let these issues block peer review. The system is novel, the sim-to-real evidence is genuine, and the failure modes are honestly discussed. I would ask the authors to either soften the \"any object\" language or extend the goal representation to handle view ambiguity, and to report error distributions for all runs, not just those that completed within 1 m.\n\nWho is this for? Robotics people working on last-mile navigation, docking, and mobile manipulation. It deserves a serious referee.","headline":"A solid, well-engineered precision navigation system whose 'any object' claim outruns its goal parametrization.","tokens_in":12608,"tokens_out":1665,"would_cite":false,"duration_ms":18811,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A vision-based navigation system claims centimeter-level precision to any object in a room, without maps or CAD models.","keywords":["aim-my-robot","precision local navigation","object-centric navigation","sim-to-real transfer","RGB-D","LiDAR","transformer policy","imitation learning"],"falsifier":"Take a cylindrical or spherical object with no planar sides, render a reference image from a 45-degree angle, request a 'front' approach at 1.0 m, and measure the robot's final pose across many runs: if the robot consistently chooses the wrong side or the median distance error exceeds the claimed 3 cm, the dominant-front parametrization fails.","tokens_in":1678,"feed_emoji":"🤖","tokens_out":7440,"duration_ms":80630,"temperature":0.7,"pith_summary":"The paper introduces Aim-My-Robot (AMR), a local navigation system that claims to position a robot to a specified side, distance, and angle of any nearby object with centimeter-level accuracy. It uses only a masked reference image of the target object and streams of RGB-D and LiDAR data, requiring neither a metric map nor an object 3D model. The authors report a median final error of 3 cm and 1 degree on unseen objects in simulation, with little degradation when deployed on a real robot in a kitchen. If this holds, it would close the gap between standard navigation success (within 1 m) and the precision needed for docking, inspection, and manipulation.","feed_headline":"Robot navigates to the centimeter without maps or object models","feed_subtitle":"A simulation-trained system positions a robot at a specified side, distance, and angle of any nearby object.","key_machinery":"The key machinery is a goal parametrization plus a transformer-based policy: the target object is specified by a masked reference image (the mask identifies the instance), and the desired pose is given as C = {S, d, theta} where S is one of four sides relative to the most visible side in the reference image, d is the approach distance, and theta is the approach angle. The policy encodes RGB and depth as tokens (depth is injected as a 3D positional embedding for the RGB patches), encodes LiDAR as directional-bin tokens, fuses them with the goal and robot footprint in a multi-modal context encoder, and decodes a base trajectory autoregressively using multi-token classification with residual predictions to preserve precision after discretization. The data pipeline generates precise demonstrations with AIT* planning in Isaac Sim using the HSSD scene dataset, and training is augmented with DAgger. The architecture is designed to track the target object with a learned mask decoder and to reason about robot size, which the ablations show reduces collision rates.","core_discovery":"The central claim is that an end-to-end learned policy can map multi-modal observations and an object-centric goal specification directly to precise base trajectories and camera tilt commands, eliminating the need for maps and object 3D models. The goal is specified by a reference image with a target mask and a relative pose parameterized by approach side (front, back, left, right), distance (0.1-1.0 m), and angle (0, plus or minus 15, plus or minus 30 degrees). The model is trained entirely in simulation on 500k trajectories across 54 photorealistic scenes, and achieves a median error of 3 cm and 1 degree on 2000 navigation tasks in unseen environments, including objects never seen during training. In a real kitchen, the robot completed most runs with distance errors between 1.8 and 3.1 cm for four of six objects, and it succeeded in closing a fridge drawer and aligning a forklift to a pallet for open-loop insertion.","pith_inferences":["If this level of precision generalizes, the 1-meter success radius used in many navigation benchmarks becomes an outdated metric, and a new benchmark should measure endpoint pose error in centimeters.","The dominant-front parametrization could be extended to a full 6-DoF goal specification, such as a desired camera view direction, which would handle objects without a dominant side like cylinders or spheres.","The depth-as-positional-embedding design for RGB-D fusion is a transferable architectural idea that could reduce token counts in other vision-language-action models.","The recipe of photorealistic simulation plus model-based planner demonstrations may generalize to other precise behaviors beyond navigation, such as docking and alignment tasks, but this is not claimed by the paper."],"forward_implications":["AMR can be interfaced with high-level planners that output a mask and pose parameters, making it a modular precision layer for task planning systems.","The robot can perform downstream manipulation using its body as an end effector, demonstrated by closing a fridge drawer in open loop after reaching the goal.","The system adapts to different robot kinematics: it runs on an omnidirectional base and, after fine-tuning with about 500 demonstrations, on a simulated forklift with Ackermann steering.","The robot can navigate to objects that are initially out of view, with a 90.4% completion rate in simulation, which exceeds what a classical pose-estimation baseline can do since that baseline requires initial visibility.","The centimeter-level endpoint accuracy makes open-loop final actions feasible, such as driving the forklift forward so the fork fully inserts into a pallet."],"supporting_citations":[{"why":"Supplies the HSSD dataset of diverse room layouts and 10,000+ objects used for photorealistic training and evaluation.","marker":"[21]"},{"why":"Provides the AIT* planner with a Reeds-Shepp state space that generates precise, collision-free demonstration trajectories in simulation.","marker":"[22]"},{"why":"Supplies the frozen Masked Autoencoder backbone used to tokenize RGB images and the reference image.","marker":"[23]"},{"why":"Provides the Isaac Sim simulator used to render photorealistic images and generate 500k training trajectories.","marker":"[24]"},{"why":"The DAgger algorithm augments the dataset by collecting corrective demonstrations from the model's own failures in simulation.","marker":"[26]"},{"why":"FoundationPose serves as the classical pose-estimation baseline that AMR must beat; the comparison defines the precision claim.","marker":"[13]"},{"why":"Behavior Transformers inspire the multi-token classification with residual predictions used to decode precise waypoints.","marker":"[25]"}],"fun_headline_variants":["Robot hits 3 cm accuracy without maps or object models","Sim-to-real robot navigates to any object at centimeter precision","No maps, no 3D models: Robot reaches exact pose on any object","Precision navigation to centimeters, trained in sim, works in real","Aim My Robot: navigate to any object with centimeter accuracy"],"cache_read_input_tokens":14720,"weakest_assumption_plain":"The goal specification assumes the target object has a dominant front side defined as the most visible side in the reference image, so the requested approach side and angle are unambiguous; for objects without a dominant side, such as cylinders, or when the reference image is oblique, the robot may aim at the wrong side.","fun_headline_variants_meta":{"raw":{"variants":["Robot hits 3 cm accuracy without maps or object models","Sim-to-real robot navigates to any object at centimeter precision","No maps, no 3D models: Robot reaches exact pose on any object","Precision navigation to centimeters, trained in sim, works in real","Aim My Robot: navigate to any object with centimeter accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00074,"raw_usage":{"total_tokens":3263,"prompt_tokens":864,"completion_tokens":2399,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":2309}},"tokens_in":480,"tokens_out":2399,"duration_ms":17345,"temperature":1.0,"reasoning_tokens":2309,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:54:59.138019+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a cylindrical or spherical object with no planar sides, render a reference image from a 45-degree angle, request a 'front' approach at 1.0 m, and measure the robot's final pose across many runs: if the robot consistently chooses the wrong side or the median distance error exceeds the claimed 3 cm, the dominant-front parametrization fails.","supporting_citations":[{"cited_title":"Masked autoencoders are scalable vision learners,","cited_arxiv_id":null,"evidence_quote":"Supplies the frozen Masked Autoencoder backbone used to tokenize RGB images and the reference image."},{"cited_title":"Nvidia isaac sim,","cited_arxiv_id":null,"evidence_quote":"Provides the Isaac Sim simulator used to render photorealistic images and generate 500k training trajectories."},{"cited_title":"A reduction of imitation learning and structured prediction to no-regret online learning,","cited_arxiv_id":null,"evidence_quote":"The DAgger algorithm augments the dataset by collecting corrective demonstrations from the model's own failures in simulation."},{"cited_title":"Behavior transformers: Cloning k modes with one stone,","cited_arxiv_id":null,"evidence_quote":"Behavior Transformers inspire the multi-token classification with residual predictions used to decode precise waypoints."}],"review_version":1}