Pith. sign in

REVIEW 14 cited by

BAKU: An Efficient Transformer for Multi-Task Policy Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07539 v2 pith:RHSOEBKK submitted 2024-06-11 cs.RO

classification cs.RO
keywords bakulearningtasksactiondatademonstrationsefficientimprovement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training generalist agents capable of solving diverse tasks is challenging, often requiring large datasets of expert demonstrations. This is particularly problematic in robotics, where each data point requires physical execution of actions in the real world. Thus, there is a pressing need for architectures that can effectively leverage the available training data. In this work, we present BAKU, a simple transformer architecture that enables efficient learning of multi-task robot policies. BAKU builds upon recent advancements in offline imitation learning and meticulously combines observation trunks, action chunking, multi-sensory observations, and action heads to substantially improve upon prior work. Our experiments on 129 simulated tasks across LIBERO, Meta-World suite, and the Deepmind Control suite exhibit an overall 18% absolute improvement over RT-1 and MT-ACT, with a 36% improvement on the harder LIBERO benchmark. On 30 real-world manipulation tasks, given an average of just 17 demonstrations per task, BAKU achieves a 91% success rate. Videos of the robot are best viewed at https://baku-robot.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Feel the Force: Contact-Driven Learning from Humans

    cs.RO 2025-06 conditional novelty 7.0 of 10

    FeelTheForce trains a robot policy on human tactile demonstrations, predicting desired contact forces and using a PD controller to track them on the robot gripper, achieving 77% success across five force-sensitive tasks.

  2. Patch Policy: Efficient Embodied Control via Dense Visual Representations

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Patch Policy shows that frozen dense ViT patch tokens, consumed through a block-causal attention mask, let lightweight robot policies beat pooled-feature policies and even a fine-tuned 7B vision-language-action model.

  3. TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    TempoVLA learns a single VLA policy with controllable execution speed via variable-speed trajectory augmentation and explicit speed conditioning.

  4. Nautilus: From One Prompt to Plug-and-Play Robot Learning

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    NAUTILUS is a prompt-driven harness that automates plug-and-play adapters, typed contracts, and validation for policies, benchmarks, and robots in learning research.

  5. Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

    cs.RO 2026-01 conditional novelty 6.0 of 10

    Slot-based object-centric visual representations, especially with robot-video pretraining, improve out-of-distribution generalization of robotic manipulation policies compared to global and dense pre-trained features.

  6. FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A compact 950-million-parameter robot policy trained in about 200 GPU-hours matches or beats multi-billion-parameter baselines on most manipulation benchmarks, including a new best score on CALVIN ABC.

  7. AMPLIFY: Actionless Motion Priors for Robot Learning from Videos

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A three-stage pipeline that turns keypoint tracks into discrete motion tokens, predicts them from action-free video, and decodes them into actions yields large few-shot and zero-shot policy improvements in robot manipulation.

  8. Touch begins where vision ends: Generalizable policies for contact-rich manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A localize-then-execute policy that combines vision-language reaching, semantic background augmentation, and residual reinforcement learning with tactile sensing reaches about 90% success on millimeter-precision manip...

  9. SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    SwitchVLA trains a vision-language-action policy to handle mid-execution instruction changes by conditioning on contact state and a three-way behavior mode, using only existing single-task demonstrations.

  10. EgoZero: Robot Learning from Smart Glasses

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Robot policies trained only on egocentric human videos from smart glasses transfer zero-shot to a Franka gripper, with 70% success across 7 manipulation tasks.

  11. Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Knowledge insulation blocks gradients from a continuous action expert into a VLM backbone while training with discrete action tokens, yielding faster training, better language following, and strong real-robot results.

  12. ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Fine-tuning a vision-language-action robot model on teacher-generated reasoning rationales raises average simulated manipulation success by up to 8.6 percentage points over the SpatialVLA baseline.

  13. Bootstrapping Imitation Learning for Long-horizon Manipulation via Hierarchical Data Collection Space

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Breaking long manipulation tasks into atomic subtasks and collecting demonstrations from varied starting poses improves imitation learning success using fewer demonstration frames.

  14. Is an object-centric representation beneficial for robotic manipulation ?

    cs.AI 2025-06 reject novelty 4.0 of 10

    Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to un...

Pith tools