{"id":"f5bb647f-12c4-4ee2-971b-c0f577e729ca","arxiv_id":"2607.22530","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"An action-conditioned visuo-tactile world model generates synthetic camera-plus-touch rollouts that, mixed with real demonstrations, improve downstream contact-rich manipulation policies.","lead":"ViTacWorld is a world model that predicts both camera views and fingertip tactile images from robot actions, so it can generate imagined contact-rich manipulation videos. Adding these imagined videos to real demonstrations improved real-robot success rates on plugging, peeling, and insertion tasks in the paper's experiments.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported policy gain is not isolated from manual filtering or added data volume: no control adds equivalent unfiltered or real replay rollouts, so Table 1's improvements may not be caused by the learned world-model dynamics.","rationale":"The reader's primary concern is also mine: the manual filter plus missing data-volume baseline makes the mechanism attribution untestable from Table 1. I did not find a more fundamental internal inconsistency. The vision/tactile generation ablations in Table 2 are real evidence that pretraining helps prediction quality, but prediction quality is not the same as proving that downstream gains come from the generator. The limitations section's open admission supports rather than resolves the concern. Because the paper still provides genuine real-robot results and a plausible architecture, a conditional verdict remains appropriate; the proposed control experiments should be prerequisites, not optional additions. The 10-trial statistical weakness is secondary; it would be partly addressed by the multi-seed evaluation in the concrete test. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":14681,"tokens_out":6190,"duration_ms":64831,"concrete_test":"Run a matched-size augmentation ablation on the four tasks with three training sets: (1) expert demos + 200 unfiltered ViTacWorld rollouts; (2) expert demos + 200 real-robot rollouts sampled from the same rollout policy (or action-noised replays of expert demos) and filtered by the same manual criterion; (3) expert demos + 200 ViTacWorld rollouts filtered as in the paper. Keep training steps, batch size, and policy architecture identical; use at least 3 seeds and ≥20 trials per task. If (2) or (1) matches or exceeds (3)'s average success, the improvement in Table 1 is attributable to data volume or curation rather than to the learned world-model dynamics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"ViTacWorld's central claim—that action-conditioned visuo-tactile world models generate policy-improvement contact-rich data—is supported by Table 1, where averaging over downstream policies improves from 42.5% to 67.5% when expert demonstrations are augmented with ViTacWorld rollouts. The load-bearing assumption is that the improvement is due to the world model's learned dynamics. That assumption is not isolated. Section 3.4 states that generated rollouts are filtered 'according to task success and visual-tactile plausibility,' and Section 6 admits the selection 'still relies partly on manual inspection.' The paper provides no baseline that adds an equivalent number of (a) unfiltered ViTacWorld rollouts, (b) real-robot replay/random rollouts, or (c) additional expert demonstrations matched for dataset size. Consequently, the observed gains could be produced by the extra data volume alone or by the human curator's exclusion of implausible trajectories, rather than by the learned model. Since the generation pipeline's scalability is part of the contribution, leaving the curation step human-in-the-loop weakens both the mechanism and the 'scalable' claim. A controlled baseline is therefore necessary before attributing Table 1's improvement to the world model. The 10-trial/no-seed evaluation makes this worse, but the missing control is the decisive confound.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ViTacWorld, an action-conditioned visuo-tactile world model that generates temporally aligned visual and tactile rollouts conditioned on robot action chunks. The model is initialized from a Cosmos-Predict2.5 video prior, pretrained on large-scale public visuo-tactile data (OmniViTac) and task-aligned Isaac Sim data, and then fine-tuned on real expert demonstrations and policy rollouts from a Franka Panda robot. The authors claim that augmenting expert demonstrations with ViTacWorld-generated, filtered 'dream' rollouts improves downstream tactile policies, and that the same model can serve as a lightweight policy evaluator by predicting task success from imagined rollouts. Experiments are conducted on four contact-rich manipulation tasks (charger plugging, cucumber peeling, U-block insertion, cuboid insertion) with 10 real-robot trials per condition. The main quantitative claim is an average success-rate improvement from 42.5% (expert-only) to 67.5% for the π0.5 + tactile policy when round-1 ViTacWorld rollouts are added, with a further improvement to 80.0% after round-2 rollouts. The paper also reports generation-quality ablations (PSNR/SSIM/LPIPS) in Table 2.","tokens_in":14980,"tokens_out":3922,"duration_ms":45189,"significance":"If the central claim is validated, ViTacWorld would be a useful contribution to contact-rich manipulation: it provides a concrete recipe for leveraging public tactile datasets and simulation to scale visuo-tactile data, and it extends action-conditioned video world models to include tactile streams as a first-class output. The paper's strengths include the use of a large pretraining corpus, task-aligned simulation, real-robot evaluation, matched initial-state evaluation in the appendix, and an honest limitations statement. The main empirical claim, however, is not yet isolated from confounds: the dream data are manually filtered, no control condition adds equivalent unfiltered or non-world-model data, and all success rates rest on 10 trials without confidence intervals or significance testing. Because the paper's contribution is specifically about the world model as a scalable data generator, these confounds directly affect the validity of the central conclusion.","major_comments":[{"comment":"The central claim that ViTacWorld's learned dynamics cause the policy improvement in Table 1 is not isolated from two confounds: added data volume and manual curation. The dream data are filtered 'according to task success and visual-tactile plausibility' (Sec. 3.4), and Sec. 6 admits that selection 'still relies partly on manual inspection.' The paper provides no control condition that adds an equivalent number of (a) unfiltered ViTacWorld rollouts, (b) real-robot replay or random rollouts, or (c) additional expert demonstrations matched for dataset size. Without such a baseline, the observed 42.5% to 67.5% average improvement could be produced by the extra training data per se or by the human excluding implausible trajectories, rather than by the learned visuo-tactile dynamics. Since scalability is part of the contribution, a human-in-the-loop curation step weakens the mechanism. Pleas","section":"Sec. 3.4, Sec. 4.2, Sec. 6"},{"comment":"All real-robot success rates are based on 10 trials per task with no confidence intervals, no repeated seeds, and no significance testing. Several reported improvements are one-trial changes; for example, U-Block π0.5 stays 60 in both conditions and Cuboid π0.5 stays 40, so the 'consistently improves' statement is not statistically supported. A difference of one trial (e.g., U-Block ACT+tactile from 30 to 40) is well within chance for n=10. Please report binomial confidence intervals, run more trials or seeds, or apply a paired significance test across the matched initial states; otherwise the headline improvement may reflect noise.","section":"Sec. 4.1, Table 1"},{"comment":"The policy-evaluation contribution is also under-validated. Success labels for imagined rollouts are assigned by the authors' judgment of whether the trajectory 'completes the task with plausible visual-tactile interaction' (Sec. A.1). Table 4 reports 90% agreement on a single task (U-Block), and Table 3 gives only per-task aggregate gaps without confidence intervals or an automated, pre-registered success criterion. This makes it difficult to distinguish a genuine evaluation signal from the human already knowing the real outcome. Please define an objective success-detection rule for generated rollouts, report agreement across all tasks and initial conditions, and compare against a chance-level or image-similarity baseline.","section":"Sec. A.1, Table 3, Table 4"}],"minor_comments":[{"comment":"The 'first framework' claim should be qualified. Refs. [15,16,17,31,32,33] already describe visuo-tactile world or world-action models that predict future tactile observations. The novelty appears to be specifically the use of such a model as a data generator for downstream policy learning, not visuo-tactile trajectory generation per se. Please phrase the claim to avoid overstatement.","section":"Abstract, Sec. 2"},{"comment":"'Visual-tactile plausibility' is never operationally defined. Please specify what criteria are used (e.g., contact consistency, end-effector motion, object permanence) and whether the filter is applied by the authors, a learned model, or a combination. This matters because the contribution depends on what the filter removes.","section":"Sec. 3.4"},{"comment":"The OmniViTac-to-simulation sampling ratio of approximately 2:1 is presented without sensitivity analysis. Since simulated data are a key pretraining ingredient, a brief study varying this ratio (or at least reporting the loss curves) would strengthen the claim that the chosen balance is not arbitrary.","section":"Appendix C.2"},{"comment":"The view-wise PSNR/SSIM/LPIPS numbers are reported without error bars, and the tactile-stream differences between variants are very small (e.g., PSNR 35.127 vs 35.225). Please report standard deviations over validation clips and note that pixel-level metrics may be dominated by static or non-contact frames.","section":"Table 2"},{"comment":"Minor formatting issues: 'V AE' in Sec. 3.2, 'VT-W AM' in Sec. 2, and the reference author list for [52] ('P. Intelligence') are nonstandard. Please normalize the author lists and spacing.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important scaling problem and the experimental setup is nontrivial, but the central causal claim—that the world model's learned dynamics improve downstream policies—is currently confounded by manual filtering and added data volume. The missing control is the decisive issue, not the 10-trial statistics alone. If the authors add the recommended ablation (equivalent unfiltered or real-rollout data) and show that ViTacWorld's selection retains the gains, I would support acceptance. The 'first' novelty claim should also be checked against the cited visuo-tactile world-model papers before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: ViTacWorld is a reasonable bet. Adapting an action-conditioned video prior (Cosmos) to also emit an image-like tactile stream, then using it to generate training data for downstream tactile policies, attacks a real bottleneck. The training recipe — pretrain on public tactile data plus task-aligned simulation, then finetune on real demonstrations and policy rollouts — is sensible, and the paper is honest that the dream-data selection still relies partly on manual inspection. There is a real-robot evaluation, which is more than many world-model papers do.\n\nWhat is actually new is the specific system: visual plus tactile streams in one action-conditioned generator, trained with this two-stage recipe, used as an offline data generator. The related work is thorough and actually lists prior visuo-tactile world models, so the \"first framework\" claim is overstated, but the emphasis on generating data for downstream policies is a distinct twist.\n\nNow the soft spots. The main result — policy improvement from the generated rollouts — does not have a clean control. The pipeline filters rollouts manually by \"task success and visual-tactile plausibility,\" so we don't know whether the gain comes from the learned world-model dynamics, from just having more data, or from the human curator throwing away bad trajectories. A baseline adding an equivalent number of unfiltered generated rollouts, or real-robot replay/random rollouts, would isolate the generator. That is the load-bearing gap. The 10-trial evaluation, no seeds, no intervals, makes one-trial differences (e.g., U-Block 70 to 80) effectively noise. The generation quality metrics in Table 2 are fine but secondary. The policy evaluation section is a nice exploratory addition, though the \"predicted success\" is itself judged by humans, so it inherits the same problem.\n\nOverall: the central direction is plausible and the paper is worth engaging with, but the strong claim — that the world model itself generates policy-improvement contact-rich data — is not yet supported. A serious referee should push for the control baselines and statistical rigor.\n\nI'd bring it to a reading group to discuss what a proper dream-data control looks like, but I wouldn't cite the result yet. It deserves peer review, not a desk reject.\n\nBest.","headline":"Plausible and timely idea, but the headline result isn't isolated from manual data curation and extra data volume; the stats are too thin to support the strongest claim.","tokens_in":15516,"tokens_out":1896,"would_cite":false,"duration_ms":19693,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A world model that generates visuo-tactile rollouts improves contact-rich policies from 42.5% to 67.5% average success.","keywords":["world model","tactile sensing","contact-rich manipulation","policy learning","data augmentation","diffusion transformer","robot manipulation","visuo-tactile data"],"falsifier":"Train the same downstream policies on an equal number of real policy rollouts or random replays instead of ViTacWorld-generated dreams; if success rates match or exceed 67.5%, the central claim is false. Also rerun the real-robot evaluations with many more trials and multiple random seeds; if the 42.5-to-67.5 gap shrinks or reverses, the difference may be noise.","tokens_in":1301,"feed_emoji":"🖐️","tokens_out":4331,"duration_ms":66886,"temperature":0.7,"pith_summary":"The paper argues that an action-conditioned world model can learn to generate future frames for both cameras and tactile sensors from robot actions. These generated rollouts, when filtered and added to real demonstrations, improve real-robot policies on four contact-rich tasks. The authors also claim their model is the first to use a world model for visuo-tactile-action trajectory generation and policy evaluation. If correct, the approach offers a scalable way to obtain contact-rich training data without additional real-robot teleoperation.","feed_headline":"Imagined touch data lifts robot success from 42.5 to 67.5 percent","feed_subtitle":"A world model generating vision-and-touch rollouts trains better contact-rich policies; a second round reaches 80 percent.","key_machinery":"The core object is a stream-aware diffusion transformer that treats the tactile sensor as an additional image view alongside a main camera and wrist camera. Each stream is encoded to latent tokens; stream identity embeddings condition generation, in-stream self-attention preserves modality integrity, and cross-view attention exchanges contact information between tactile and visual tokens. The model inherits a pretrained action-conditioned video prior, is pretrained on real and simulated visuo-tactile trajectories, then finetuned on real demonstrations and policy rollouts. This design lets the model generate temporally aligned visual and tactile frames from a shared action sequence.","core_discovery":"The central claim is that action-conditioned visuo-tactile world models can serve as data generators for downstream tactile policies. The model, after pretraining on large real and simulated visuo-tactile trajectories and finetuning on real policy rollouts, predicts aligned visual and tactile observations conditioned on robot actions. The paper shows that augmenting expert demonstrations with 200 selected generated successful rollouts improves average real-robot success from 42.5% to 67.5% for a tactile policy, and that a second round of generated data reaches 80%. The model also provides a lightweight policy evaluation by predicting success or failure from the same initial conditions, match","pith_inferences":["The paper leaves open whether the manual filter for selecting successful generated rollouts is the main driver of improvement; an automated or randomized-dream baseline could separate the generator's contribution from the filter's.","The transfer to a different image-based tactile sensor described in the appendix suggests the method may generalize across sensor hardware, but only a small fine-tuning experiment demonstrates it; more systematic sensor-agnostic training could strengthen this.","If the 'smaller sim-to-real gap for tactile signals' holds, this recipe could extend to other contact-rich domains such as dexterous hands or soft manipulation, where tactile simulators are less mature."],"forward_implications":["Augmenting expert data with ViTacWorld rollouts improves success for a tactile policy from 42.5% to 67.5% across four tasks; a second round of generated data reaches 80% average success.","Gains appear for both vision-only and tactile policies, suggesting the generated data carries general task-level supervision.","The same world model can predict policy outcomes from the same initial states, offering a pre-deployment evaluation signal that matched real results on 9 of 10 U-block trials.","Pretraining with large real and task-aligned simulated data improves generative quality, with the largest perceptual gains on the tactile stream.","The approach can be iterated: using an improved policy to generate a new round of dream data leads to further improvement, at least until tasks approach saturation."],"fun_headline_variants":["ViTacWorld: imagined touch data lifts robot success to 67.5%","Synthetic rollouts from visuo-tactile world model boost policy to 67.5%","World model generates touch data, raising contact-rich robot success","Action-conditioned world model improves robot policies with generated touch","From 42.5% to 67.5%: ViTacWorld's generated rollouts help robots"],"cache_read_input_tokens":16768,"weakest_assumption_plain":"The measured gains rely on the assumption that the improvement comes from the world model's learned dynamics, but the selected generated rollouts are filtered by human judgment of success and plausibility (Sec. 3.4) and success rates are measured over only 10 trials per task with no error bars, so the effect could partly come from the filter or from trial noise.","fun_headline_variants_meta":{"raw":{"variants":["ViTacWorld: imagined touch data lifts robot success to 67.5%","Synthetic rollouts from visuo-tactile world model boost policy to 67.5%","World model generates touch data, raising contact-rich robot success","Action-conditioned world model improves robot policies with generated touch","From 42.5% to 67.5%: ViTacWorld's generated rollouts help robots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000638,"raw_usage":{"total_tokens":2811,"prompt_tokens":817,"completion_tokens":1994,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":1886}},"tokens_in":561,"tokens_out":1994,"duration_ms":13506,"temperature":1.0,"reasoning_tokens":1886,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T04:23:46.244059+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same downstream policies on an equal number of real policy rollouts or random replays instead of ViTacWorld-generated dreams; if success rates match or exceed 67.5%, the central claim is false. Also rerun the real-robot evaluations with many more trials and multiple random seeds; if the 42.5-to-67.5 gap shrinks or reverses, the difference may be noise.","supporting_citations":[],"review_version":1}