Pith. sign in

REVIEW 5 cited by

This&That: Language-Gesture Controlled Video Generation for Robot Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.05530 v2 pith:HNKD5BV6 submitted 2024-07-08 cs.RO cs.AIcs.CV

classification cs.ROcs.AIcs.CV
keywords videoplanninginstructionsrobottaskbehaviorcloningcomplex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Clear, interpretable instructions are invaluable when attempting any complex task. Good instructions help to clarify the task and even anticipate the steps needed to solve it. In this work, we propose a robot learning framework for communicating, planning, and executing a wide range of tasks, dubbed This&That. This&That solves general tasks by leveraging video generative models, which, through training on internet-scale data, contain rich physical and semantic context. In this work, we tackle three fundamental challenges in video-based planning: 1) unambiguous task communication with simple human instructions, 2) controllable video generation that respects user intent, and 3) translating visual plans into robot actions. This&That uses language-gesture conditioning to generate video predictions, as a succinct and unambiguous alternative to existing language-only methods, especially in complex and uncertain environments. These video predictions are then fed into a behavior cloning architecture dubbed Diffusion Video to Action (DiVA), which outperforms prior state-of-the-art behavior cloning and video-based planning methods by substantial margins.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A dual-stream diffusion VLA that predicts actions and future observations in separate streams with shared attention beats VLA/world-model baselines on simulated and real robot tasks.

  2. Ego-centric Predictive Model Conditioned on Hand Trajectories

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Ego-PM predicts future hand trajectories and then uses them to condition latent diffusion video generation, jointly outputting actions and future frames in egocentric and robotic scenes.

  3. PR-Aware Automated Unit Test Generation: Challenges and Opportunities

    cs.SE 2026-05 unverdicted novelty 5.0 of 10

    EvoSuite produced at least one fail-to-pass test for 36% of PRs versus 13% for GPT-4o, but both tools generated no meaningful change-capturing tests for 64% of the PRs evaluated.

  4. World Action Models: The Next Frontier in Embodied AI

    cs.RO 2026-05 unverdicted novelty 4.0 of 10

    The paper introduces World Action Models as a new paradigm unifying predictive world modeling with action generation in embodied foundation models and provides a taxonomy of existing approaches.

  5. World Action Models: A Survey

    cs.RO 2026-06 unverdicted novelty 3.0 of 10

    A survey that clarifies boundaries and organizes World Action Models by generation requirements and predictive substrates, identifying a trend toward generating less of the future.

Pith tools