{"id":"45e275e3-0ee7-48a6-a7cf-61ffa5ddd185","arxiv_id":"2411.12503","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"The paper describes the tasks, simulation, evaluation metrics, reference policy, and prizes of a 2025 robot manipulation challenge combining touch and vision.","lead":"This paper announces the ManiSkill-ViTac 2025 robotics competition, in which teams train robots for contact-rich tasks like peg insertion and lock opening using touch and vision, across three separate tracks. It matters as a proposed standard benchmark for comparing vision-tactile manipulation algorithms and sensor designs.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of a small sim-to-real gap (Section I) is asserted without any quantitative comparison between the SAPIEN/IPC tactile simulation, the Zhang depth simulation, and the real GelSight/RealSense platform; absent this validation, simulation rankings may not transfer to the real tasks.","rationale":"The reader's UNVERDICTED verdict is appropriate. This is a challenge announcement rather than a research preprint, so it does not need a full scientific study to be useful, but it does make an explicit, correctness-relevant empirical claim: that the simulation environment closely matches real-world performance. That claim is load-bearing because the challenge's value hinges on sim-to-real transfer; without it, Stage 1 simulation rankings and Track 3 simulated sensor-design evaluations cannot be assumed to reflect real-world behavior. I agree with the reader that simulator fidelity is the weakest assumption. Independent support does exist in the form of clear challenge organization, a well-specified action/substep protocol, and prior published methods [15,17], but those prior citations do not contain the specific validation for this platform. The internal algebra of Equations 2-13 appears coherent, and no mathematical contradiction was found. The concern is not that the claim contradicts established consensus; it is that the claim is entirely unsupported by reported evidence. A single controlled transfer experiment with the reference baseline would largely settle the issue, which is why the verdict should remain UNCHANGED rather than moving to ACCEPT or REJECT.","tokens_in":12485,"tokens_out":2854,"duration_ms":28828,"concrete_test":"Run the provided TD3 baseline (Section V) on the Peg Insertion and Lock Opening tasks: train in the simulation, then deploy the identical policy on the real Section III-C platform without retraining. Record success rate over at least 100 episodes in each environment, and at matched states compare marker-flow displacement fields and depth maps between sim and real. If the success-rate gap exceeds 10 percentage points, or the normalized marker and depth error exceeds a pre-set tolerance, the 'small sim-to-real gap' claim in Section I is unsupported and should be weakened or removed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's main scientific premise is that the simulator closely matches the real platform (Section I, Key Features), so policies trained in simulation will transfer and challenge rankings will be meaningful. This premise enters all three tracks: Track 1 and Track 2 are scored in simulation in Stage 1, and Track 3's sensor designs are verified in simulation only. Yet the paper provides no measured comparison of simulated versus real sensor outputs or task outcomes. Section III-B describes the FEM/IPC marker-displacement model and Zhang's active-stereo depth simulation, and Section III-C describes the real hardware, but no plot, table, or success-rate comparison connects the two. Figure 1 even shows the simulated scene contains only silicone parts and target objects, omitting the gripper and F/T sensor used in the real platform, which could alter contact mechanics. The cited prior protocol [17] may address sim-to-real issues, but the results are not reported here. Because the challenge's central deliverable is a meaningful sim-to-real benchmark, the absence of any validation is a load-bearing gap: if the sim-to-real gap is actually large, the Stage 1 simulation rankings and Track 3 design evaluations would not predict real-world performance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces the ManiSkill-ViTac 2025 challenge for contact-rich manipulation with vision and tactile sensing. It proposes three tracks: tactile-only manipulation, tactile-vision fusion, and tactile sensor structure design. The platform description covers a SAPIEN-based simulation with IPC/FEM tactile sensing and Zhang's active-stereo depth simulation, plus a real platform with GelSight Mini sensors, a RealSense D415, a parallel gripper, and positioning stages. The paper specifies observations, actions, reward functions, and submission formats for each track, and it claims that the simulation closely matches the real platform with a small sim-to-real gap, making simulation rankings meaningful for real-world performance.","tokens_in":12771,"tokens_out":2936,"duration_ms":31617,"significance":"The challenge has the potential to be a useful community resource: it provides a concrete multi-track evaluation protocol, a physically grounded simulation pipeline, and a baseline TD3 method. If the sim-to-real claim were quantitatively validated, the standardized metrics and the shared hardware platform would give the field a much-needed common testbed for vision-tactile manipulation. However, the paper's central premise, the small sim-to-real gap, is asserted rather than demonstrated, and the action/reward specification contains internal inconsistencies. These issues must be addressed before the proposed evaluations can be considered reproducible or the rankings interpretable.","major_comments":[{"comment":"The claim of a 'small sim-to-real gap' is the load-bearing premise for the meaningfulness of Stage 1 simulation rankings and of Track 3's simulation-only sensor-design evaluation, but the manuscript contains no quantitative comparison of simulated versus real tactile signals, depth images, or task success rates. Figure 1's caption further notes that the simulated scene includes only the silicone parts and target objects, omitting the gripper and F/T sensor used in the real platform; the effect of this omission on contact mechanics is not discussed. Please add quantitative sim-to-real validation (for example, marker-displacement error, depth error, and policy success rates in simulation and on the real platform) or explicitly reframe the claim as 'simulation-only evaluation' and state what conclusions can be drawn from the rankings.","section":"Section I and Section III-B/III-C"},{"comment":"The Lock Opening task is specified inconsistently. Section IV-A-2-b defines actions as displacement increments in x, y, and z only, and Algorithm 2 applies only linear velocities. Yet Section V-A-1 states that the Track 1 actor outputs a three-dimensional action vector [ax, ay, a_theta] relative to the peg, and both the success criterion ('Key and Lock Alignment') and the reward function (Eqs. 18-21) include an angular error e_theta. Either the action space must include a rotation component or the angular terms must be removed; as written, participants cannot determine the intended control loop for this task.","section":"Sections IV-A-2-b, IV-A-3-b, V-A-1, and Eqs. (18)-(21)"},{"comment":"Several constants that determine the behavior of the environment and the reward are named but never specified: max_action, v_max, omega_max, x_max, y_max, theta_max, the z-axis step size and step limit, the error thresholds tau_xy, tau_theta, tau_xyz, the per-step penalty P, and the lock-opening error scaling factor 500 in Eq. (19). Without these values, the challenge protocol is not reproducible and the reward-shaping behavior cannot be assessed. Please state all constants explicitly or link to a released configuration file.","section":"Sections IV-A-2, IV-A-3, V-A-2, and Algorithms 1-2"}],"minor_comments":[{"comment":"Please clarify the notation in Eq. (2): state explicitly that v is a linear velocity vector and that v*Delta_t is the displacement, and define omega as a scalar rotation rate with axis d. The label '4 views in simulation' in Figure 1 is unexplained and should be expanded.","section":"Eq. (2) and Fig. 1"},{"comment":"The Track 3 evaluation of 'innovativeness of the sensor design' is not defined beyond a general reference to structural design and marker distribution. A scoring rubric or at least a list of criteria would be needed for a fair competition and for reproducibility of the evaluation.","section":"Section IV-C-3"},{"comment":"The schedule entries '2025-TBD' for the winner announcement and award ceremony are uninformative; please provide concrete dates or state that they will be announced after the challenge launches.","section":"Section VI-A"}],"recommendation":"major_revision","confidential_remarks":"This is a challenge-proposal paper rather than a validated study. The main risk is the unsupported sim-to-real claim; that is a load-bearing correctness issue and should be fixed with quantitative evidence or by narrowing the claim. The Lock Opening action/reward inconsistency and the unspecified constants are also blocking issues for reproducibility. The paper could become suitable for publication after these points are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this is a challenge announcement, not a research result, and that is fine for what it is. The 2025 edition adds Track 3 (sensor structure design) to the 2024 format, and the paper gives enough implementation detail—IPC/FEM elastomer simulation, marker flow, Zhang's active stereo, a TD3 baseline, and per-track reward and submission rules—that a team could enter the competition from the text alone. That level of specification is genuinely useful, and the authors get credit for it.\n\nThe soft spot is the load-bearing sim-to-real claim. Section I lists 'small sim-to-real gap' as a key feature, and the whole challenge design depends on simulation rankings meaning something for the real hardware. The paper offers no quantitative evidence: no comparison of simulated vs real tactile or depth readings, no task success numbers. The cited prior protocol [17] may contain that evidence, but it is not summarized here. Figure 1 even shows the simulated scene without the gripper and F/T sensor used in the real platform, which could change contact mechanics. A reader cannot verify the central premise.\n\nThere are also internal inconsistencies. Track 1's Lock Opening action spec uses only x/y/z translations, but the reward and evaluation include an angular error e_theta. Several thresholds and constants (P, tau_xy, tau_xyz, vmax, step limits) are named but never given values. The 'surface point consistency' success criterion compares tactile markers to their initial positions, which is a strange proxy for task success and is not justified. These are fixable, but as standalone text they are gaps.\n\nThe reader's UNVERDICTED verdict makes sense: there is no new result to check. I am less bothered by the absence of results—this is a call for participation—but the sim-to-real assertion is a factual claim, not an omission, and it needs support.\n\nI would not send this to full archival peer review in its current form. It is a workshop-level challenge description. If the authors add a validation section (even a summary of published results from [17]) and clean up the Lock Opening inconsistencies, it could become a legitimate benchmark paper for RAL. Until then, treat it as a useful pointer for anyone considering the challenge.","headline":"Useful challenge spec, but the central sim-to-real claim is asserted without evidence; treat as a workshop call, not a benchmark paper.","tokens_in":13347,"tokens_out":3683,"would_cite":false,"duration_ms":35319,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The ManiSkill-ViTac 2025 challenge aims to establish the first open, standardized benchmark that fairly compares tactile-only and vision-tactile policies for contact-rich manipulation, with a simulation platform intended to transfer to an…","keywords":["robotic manipulation","tactile sensing","vision-tactile fusion","reinforcement learning","simulation benchmark","sim-to-real transfer","sensor design","contact-rich manipulation"],"falsifier":"Run the provided reference TD3 policy on the Track 1 peg insertion task in simulation and record its success rate, then run the identical policy on the described real platform under the identical evaluation criteria; if the real success rate falls far below the simulated one (or the marker-flow images differ visibly under matched contact), the claim of a small sim-to-real gap is falsified.","tokens_in":12273,"feed_emoji":"🤖","tokens_out":5741,"duration_ms":53045,"temperature":0.7,"pith_summary":"This paper presents the ManiSkill-ViTac 2025 challenge, an open competition built around a claim: contact-rich manipulation skills are best learned with both vision and touch, and the community lacks a benchmark where different approaches can be compared fairly. The challenge provides three independent tracks—tactile-only manipulation (peg insertion and lock opening), vision-tactile fusion insertion, and tactile sensor structure design—all sharing the same evaluation metrics, simulation environment, and real hardware platform. The paper's central assertion is that the simulator reproduces the real GelSight Mini and RealSense sensors closely enough that policies trained in simulation transfer to the identical real platform with a small sim-to-real gap. If that assertion holds, teams without physical robots can develop competitive policies, and the challenge's rankings would genuinely reflect real-world manipulation skill. The intended significance is a community standard for multi-modal, contact-rich manipulation research.","feed_headline":"Three-track challenge benchmarks robots that learn by touch and sight","feed_subtitle":"A shared simulator and real robot platform let researchers compare tactile, fusion, and sensor-design policies fairly.","key_machinery":"The load-bearing object is the simulation stack built on the SAPIEN environment plus the Incremental Potential Contact (IPC) method, which discretizes the GelSight Mini elastomer as a tetrahedral finite-element mesh and solves deformation without surface intersections or inversions. Marker points on the mesh surface are carried by barycentric interpolation and rendered into camera-frame pixel flows, producing the tactile observation. For vision, depth is synthesized by physics-grounded active stereo sensor simulation, which renders stereo infrared images and computes depth the way a RealSense D415 would. These two modules jointly generate the marker-flow and depth/point-cloud observations that feed the TD3-based reference policy, while the real platform is designed to mirror the same sensing geometry. The machinery converts raw physical contact and depth into the standardized observation spaces used by all three tracks.","core_discovery":"On its own terms, the paper establishes a unified platform for vision-tactile manipulation research rather than a single algorithmic result. Its discovery claim is that a carefully constructed simulation—using Incremental Potential Contact (IPC) finite-element deformation of the tactile elastomer and physics-grounded active stereo depth simulation—can stand in for the real GelSight Mini and RealSense D415 sensors closely enough that reinforcement-learned policies transfer across the gap. The paper operationalizes this through standardized metrics (excessive-error checks, step limits, insertion depth, and surface-point consistency) applied identically in simulation and on a real platform composed of a translation and rotary stage, dual GelSight Mini sensors, a Robotiq Hand-E gripper, and the D415 camera. Three tracks are derived from this platform: tactile-only manipulation, tactile-vision fusion, and iterative design of the tactile elastomer geometry.","pith_inferences":["Beyond the paper: the small sim-to-real gap is asserted without quantitative sensor-level or task-level comparison here; the challenge's Stage 2 real evaluations will be the first real test of whether simulated rankings predict physical performance.","Beyond the paper: the same platform could support sim-to-real transfer-learning studies—such as domain randomization or fine-tuning on real data—rather than only zero-shot transfer, since simulator and real setup share geometry and observation spaces.","Beyond the paper: the sensor-design track could produce geometries that are optimal only within the IPC-based simulator; validating those designs on the real GelSight Mini would be needed before claiming they improve real-world touch.","Beyond the paper: if the benchmark succeeds, it may generalize the vision-only manipulation benchmark concept to contact-rich manipulation, giving tactile sensing a first-class evaluation community."],"forward_implications":["Simulation-trained policies will transfer to the identical real hardware platform, so teams without access to robots can still compete meaningfully.","The same evaluation metrics will permit direct, fair comparison of tactile-only, fusion, and sensor-design approaches in one benchmark.","Researchers can iterate on tactile sensor geometry in simulation via Track 3, testing custom silicone designs before physical fabrication.","The reference TD3 policy with a shared tactile encoder provides a baseline that later entrants can measure against.","If the challenge attracts broad participation, it will establish a standard evaluation protocol for contact-rich multi-modal manipulation."],"supporting_citations":[{"why":"ManiSkill2, a vision-based manipulation benchmark, represents the existing class of benchmarks that lack tactile sensing and motivates the challenge's gap claim.","marker":"[2]"},{"why":"Provides the SAPIEN simulation environment that hosts the challenge's physics and rendering.","marker":"[13]"},{"why":"Supplies the Incremental Potential Contact method that solves elastomer deformation without intersections or inversions.","marker":"[14]"},{"why":"Gives the physics-grounded active stereo depth simulation used to emulate the RealSense D415's depth output.","marker":"[15]"},{"why":"The Segment Anything Model segments peg and hole point clouds in the real-platform depth pipeline.","marker":"[16]"},{"why":"Supplies the general-purpose sim2real protocol and marker-based tactile reference method that the baseline policy builds on.","marker":"[17]"},{"why":"The Twin-Delayed DDPG algorithm used as the reference reinforcement-learning policy for Tracks 1 and 2.","marker":"[18]"}],"fun_headline_variants":["Robots learn by touch and vision in ManiSkill-ViTac 2025","Three tracks merge tactile and visual skill learning","Tactile-vision fusion challenge: 3 tracks for robot manipulation","Benchmarking touch-sight robot skills across 3 tracks","Touch and sight: robot skill challenge expands to 3 tracks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire structure depends on the simulator reproducing the real GelSight Mini and RealSense D415 closely enough for a policy trained in simulation to work on the real platform; the paper asserts this small sim-to-real gap but shows no quantitative comparison of sensor readings or task success.","fun_headline_variants_meta":{"raw":{"variants":["Robots learn by touch and vision in ManiSkill-ViTac 2025","Three tracks merge tactile and visual skill learning","Tactile-vision fusion challenge: 3 tracks for robot manipulation","Benchmarking touch-sight robot skills across 3 tracks","Touch and sight: robot skill challenge expands to 3 tracks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000455,"raw_usage":{"total_tokens":2229,"prompt_tokens":832,"completion_tokens":1397,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":448,"completion_tokens_details":{"reasoning_tokens":1307}},"tokens_in":448,"tokens_out":1397,"duration_ms":10807,"temperature":1.0,"reasoning_tokens":1307,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:27:06.276280+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the provided reference TD3 policy on the Track 1 peg insertion task in simulation and record its success rate, then run the identical policy on the described real platform under the identical evaluation criteria; if the real success rate falls far below the simulated one (or the marker-flow images differ visibly under matched contact), the claim of a small sim-to-real gap is falsified.","supporting_citations":[{"cited_title":"ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills,","cited_arxiv_id":null,"evidence_quote":"ManiSkill2, a vision-based manipulation benchmark, represents the existing class of benchmarks that lack tactile sensing and motivates the challenge's gap claim."},{"cited_title":"Sapien: A simulated part-based interactive environment,","cited_arxiv_id":null,"evidence_quote":"Provides the SAPIEN simulation environment that hosts the challenge's physics and rendering."},{"cited_title":"Incremental potential contact: intersection-and inversion-free, large-deformation dynamics","cited_arxiv_id":null,"evidence_quote":"Supplies the Incremental Potential Contact method that solves elastomer deformation without intersections or inversions."},{"cited_title":"Close the optical sensing domain gap by physics-grounded active stereo sensor simulation,","cited_arxiv_id":null,"evidence_quote":"Gives the physics-grounded active stereo depth simulation used to emulate the RealSense D415's depth output."},{"cited_title":"General- purpose sim2real protocol for learning contact-rich manipulation with marker-based visuotactile sensors,","cited_arxiv_id":null,"evidence_quote":"Supplies the general-purpose sim2real protocol and marker-based tactile reference method that the baseline policy builds on."},{"cited_title":"Twin-delayed ddpg: A deep reinforcement learning technique to model a continuous movement of an intelligent robot agent,","cited_arxiv_id":null,"evidence_quote":"The Twin-Delayed DDPG algorithm used as the reference reinforcement-learning policy for Tracks 1 and 2."}],"review_version":1}