Pith. sign in

REVIEW 13 cited by

On Bringing Robots Home

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.16098 v1 pith:5GTQ4ERM submitted 2023-11-27 cs.RO cs.AIcs.CVcs.LG

classification cs.ROcs.AIcs.CVcs.LG
keywords dobb-ehomehomesmachinesminutesrobottaskbeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Throughout history, we have successfully integrated various machines into our homes. Dishwashers, laundry machines, stand mixers, and robot vacuums are a few recent examples. However, these machines excel at performing only a single task effectively. The concept of a "generalist machine" in homes - a domestic assistant that can adapt and learn from our needs, all while remaining cost-effective - has long been a goal in robotics that has been steadily pursued for decades. In this work, we initiate a large-scale effort towards this goal by introducing Dobb-E, an affordable yet versatile general-purpose system for learning robotic manipulation within household settings. Dobb-E can learn a new task with only five minutes of a user showing it how to do it, thanks to a demonstration collection tool ("The Stick") we built out of cheap parts and iPhones. We use the Stick to collect 13 hours of data in 22 homes of New York City, and train Home Pretrained Representations (HPR). Then, in a novel home environment, with five minutes of demonstrations and fifteen minutes of adapting the HPR model, we show that Dobb-E can reliably solve the task on the Stretch, a mobile robot readily available on the market. Across roughly 30 days of experimentation in homes of New York City and surrounding areas, we test our system in 10 homes, with a total of 109 tasks in different environments, and finally achieve a success rate of 81%. Beyond success percentages, our experiments reveal a plethora of unique challenges absent or ignored in lab robotics. These range from effects of strong shadows, to variable demonstration quality by non-expert users. With the hope of accelerating research on home robots, and eventually seeing robot butlers in every home, we open-source Dobb-E software stack and models, our data, and our hardware designs at https://dobb-e.com

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Eval-Actions: Fine-Grained Execution Quality Evaluation for Robotic Manipulation

    cs.RO 2026-01 conditional novelty 7.0 of 10

    A new benchmark and multimodal evaluator for scoring robotic manipulation execution quality (smoothness, safety, efficiency) and detecting whether a trajectory came from a policy or teleoperation.

  2. Feel the Force: Contact-Driven Learning from Humans

    cs.RO 2025-06 conditional novelty 7.0 of 10

    FeelTheForce trains a robot policy on human tactile demonstrations, predicting desired contact forces and using a PD controller to track them on the robot gripper, achieving 77% success across five force-sensitive tasks.

  3. Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Pretraining a VLA model on 18,561 hours of robot-synthesized egocentric human video mixed with robot data improves out-of-distribution manipulation success in simulation and on a real dual-arm robot.

  4. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  5. XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

    cs.RO 2025-11 unverdicted novelty 6.0 of 10

    XR-1 introduces Unified Vision-Motion Codes learned by dual-branch VQ-VAE and applies them in a three-stage training pipeline to outperform prior VLA models on 120+ real-world manipulation tasks across six robot embodiments.

  6. UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies

    cs.RO 2025-10 conditional novelty 6.0 of 10

    Embodiment-Aware Diffusion Policy steers a UMI-trained diffusion policy with controller tracking-cost gradients at inference time, improving aerial manipulation success in simulation and real flights.

  7. FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A compact 950-million-parameter robot policy trained in about 200 GPU-hours matches or beats multi-billion-parameter baselines on most manipulation benchmarks, including a new best score on CALVIN ABC.

  8. HannesImitation: Grasping with the Hannes Prosthetic Hand via Imitation Learning

    cs.RO 2025-08 conditional novelty 6.0 of 10

    An imitation-learning diffusion policy controlling wrist and finger movements from an eye-in-hand camera grasps diverse objects with the Hannes prosthetic hand in three real-world scenarios.

  9. Tactile Beyond Pixels: Multisensory Touch Representations for Robot Manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Sparsh-X is a transformer trained on about one million unlabeled touch interactions that fuses image, audio, motion, and pressure into representations that boost downstream robot manipulation performance over tactile-...

  10. eFlesh: Highly customizable Magnetic Touch Sensing using Cut-Cell Microstructures

    cs.RO 2025-06 conditional novelty 6.0 of 10

    eFlesh is a customizable 3D-printed magnetic tactile sensor that localizes contact to 0.5 mm, estimates force within 0.27 N, and boosts precise robot manipulation success to 91%.

  11. LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

    cs.RO 2026-07 conditional novelty 5.0 of 10

    LeapBot-WA shows robot policies can be trained with latent world-model predictions instead of pixel video generation, hitting state-of-the-art for predictive action models and staying competitive with generative WAMs.

  12. Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Knowledge insulation blocks gradients from a continuous action expert into a VLM backbone while training with discrete action tokens, yielding faster training, better language following, and strong real-robot results.

  13. Spatial-Temporal Aware Visuomotor Diffusion Policy Learning

    cs.RO 2025-07 conditional novelty 3.0 of 10

    A diffusion-based visuomotor policy gains 3D and 4D scene awareness from a dynamic Gaussian world model, improving simulated and real robot manipulation success rates.

Pith tools