Pith. sign in

REVIEW 22 cited by

ALOHA Unleashed: A Simple Recipe for Robot Dexterity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13126 v1 pith:JE7PP4FU submitted 2024-10-17 cs.RO

classification cs.RO
keywords learningchallengingrecipetasksalohademonstrateimitationmanipulation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work has shown promising results for learning end-to-end robot policies using imitation learning. In this work we address the question of how far can we push imitation learning for challenging dexterous manipulation tasks. We show that a simple recipe of large scale data collection on the ALOHA 2 platform, combined with expressive models such as Diffusion Policies, can be effective in learning challenging bimanual manipulation tasks involving deformable objects and complex contact rich dynamics. We demonstrate our recipe on 5 challenging real-world and 3 simulated tasks and demonstrate improved performance over state-of-the-art baselines. The project website and videos can be found at aloha-unleashed.github.io.

Discussion (0). Sign in to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design

    cs.RO 2026-07 conditional novelty 7.0 of 10

    A single diffusion transformer trains on tokenized robot bodies and motions to generate and optimize robot designs for unseen rewards and trajectories, outpacing evolutionary search in speed and often in reward.

  2. Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Action chunking in robotic behavioral cloning works mainly because it acts as a delayed-prediction policy and an implicit ensemble, not because of temporal consistency or horizon reduction.

  3. It's Not Just More Demos: Counterfactual Action Sensitivity Coverage for Data-Efficient Robust Robot Imitation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Pairing clean and nuisance observations to measure action drift lets CFNBC select 20–30 counterfactual repair examples that outperform matched random selection for robust imitation.

  4. $\pi\mathbf{R}^2$: Reactive Real-time Flow Policies

    cs.RO 2026-07 conditional novelty 6.0 of 10

    πR² makes flow-matching VLA policies reactive by splitting conditioning into fresh proprioception and stale vision-language features and using a one-step-per-call staircase noise schedule, reaching ~25 Hz closed-loop ...

  5. TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    TempoVLA learns a single VLA policy with controllable execution speed via variable-speed trajectory augmentation and explicit speed conditioning.

  6. EquiBim: Learning Symmetry-Equivariant Policy for Bimanual Manipulation

    cs.RO 2026-03 conditional novelty 6.0 of 10

    Adding a loss that enforces left-right equivariance between observations and actions improves average bimanual imitation policy success by +2.7 to +9.5 points across four observation/action settings.

  7. Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models

    cs.LG 2025-10 conditional novelty 6.0 of 10

    FPO fine-tunes flow-matching vision-language-action policies with a PPO-style objective that replaces intractable policy ratios with per-sample conditional flow-matching loss differences, reaching 87.2% average succes...

  8. Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A robot policy trained on one real demonstration plus AI-generated 3D views succeeds from novel initial poses, including opposite-side starts, across six real manipulation tasks.

  9. Reactive In-Air Clothing Manipulation with Confidence-Aware Dense Correspondence and Visuotactile Affordance

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A dual-arm robot folds and hangs crumpled shirts in mid-air using confidence-aware visual correspondences and touch-supervised grasp affordance.

  10. Constraint-Preserving Data Generation for Visuomotor Policy Learning

    cs.RO 2025-08 conditional novelty 6.0 of 10

    CP-Gen uses keypoint-trajectory constraints to turn a single expert demonstration into many geometry- and pose-varied robot demos, and policies trained on them transfer zero-shot to the real world.

  11. Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

    cs.RO 2025-07 conditional novelty 6.0 of 10

    Gaze-guided foveated patch tokenization reduces ViT tokens by 94%, accelerates training 7x and inference 3x, and improves robustness to distractors in bimanual manipulation policies.

  12. CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Causal Diffusion Policy adds historical action conditioning and attention cache sharing to diffusion-based robot policies, improving success rates on most tested manipulation tasks under degraded observations.

  13. DemoSpeedup: Accelerating Visuomotor Policies via Entropy-Guided Demonstration Acceleration

    cs.RO 2025-06 conditional novelty 6.0 of 10

    DemoSpeedup accelerates visuomotor policies by downsampling high-entropy segments of demonstrations, achieving roughly 2x faster execution with maintained or improved success rates.

  14. RTFF: Random-to-Target Fabric Flattening Policy using Dual-Arm Manipulator

    cs.RO 2025-10 conditional novelty 5.0 of 10

    A hybrid imitation-learning and visual-servoing policy anchored on a template mesh flattens randomly wrinkled fabric and aligns it to an arbitrary flat target on a real dual-arm robot.

  15. RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction

    cs.RO 2025-09 conditional novelty 5.0 of 10

    Robot policies trained on human interventions that rewind to a familiar state and then correct the mistake achieve higher long-horizon success and better data efficiency than imitation on full demonstrations alone.

  16. Improving Generalization Ability of Robotic Imitation Learning by Resolving Causal Confusion in Observations

    cs.RO 2025-07 conditional novelty 5.0 of 10

    Causal-ACT masks task-irrelevant image features and lifts out-of-distribution transfer success from 0.23 to 0.82 in a simulated ALOHA cube transfer task.

  17. SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training

    cs.RO 2025-07 conditional novelty 5.0 of 10

    Simulation-pretrained policies, with digital-twin demos for critic bootstrapping and action proposals, cut real-world RL training time while reaching near-perfect success on three manipulation tasks.

  18. Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Knowledge insulation blocks gradients from a continuous action expert into a VLM backbone while training with discrete action tokens, yielding faster training, better language following, and strong real-robot results.

  19. A Survey: Learning Embodied Intelligence from Physical Simulators and World Models

    cs.RO 2025-07 conditional novelty 4.0 of 10

    Embodied intelligence learning is reviewed through the complementary lenses of physical simulators and world models, with a proposed IR-L0 to IR-L4 robot capability taxonomy.

  20. Haptic-Informed ACT with a Soft Gripper and Recovery-Informed Training for Pseudo Oocyte Manipulation

    cs.RO 2025-06 conditional novelty 4.0 of 10

    A haptic-augmented ACT controller with a soft gripper reaches 80% success on a pseudo-oocyte pick-and-place task, versus 50% for standard ACT.

  21. Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision

    cs.CV 2025-06 accept novelty 3.0 of 10

    A comprehensive review of cross-view video understanding that uses both first-person and third-person cameras, organized into a three-direction taxonomy with a dataset catalog and future research gaps.

  22. Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research

    cs.RO 2025-06 accept novelty 1.0 of 10

    A perspective article reviews the state of using foundation models for laboratory automation and proposes a roadmap for fully autonomous experiments.

Pith tools