Pith. sign in

REVIEW 3 cited by

Instruction-driven history-aware policies for robotic manipulations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.04899 v3 pith:HPMFGVY4 submitted 2022-09-11 cs.RO cs.AIcs.CLcs.CVcs.LG

classification cs.ROcs.AIcs.CLcs.CVcs.LG
keywords tasksapproachinstructionsmanipulationaddresschallengingenvironmentsgeneralization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In human environments, robots are expected to accomplish a variety of manipulation tasks given simple natural language instructions. Yet, robotic manipulation is extremely challenging as it requires fine-grained motor control, long-term memory as well as generalization to previously unseen tasks and environments. To address these challenges, we propose a unified transformer-based approach that takes into account multiple inputs. In particular, our transformer architecture integrates (i) natural language instructions and (ii) multi-view scene observations while (iii) keeping track of the full history of observations and actions. Such an approach enables learning dependencies between history and instructions and improves manipulation precision using multiple views. We evaluate our method on the challenging RLBench benchmark and on a real-world robot. Notably, our approach scales to 74 diverse RLBench tasks and outperforms the state of the art. We also address instruction-conditioned tasks and demonstrate excellent generalization to previously unseen variations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training and Evaluating Diffusion Policies with Long Context Lengths

    cs.RO 2026-06 conditional novelty 6.0 of 10

    Naive long-context Diffusion Policies succeed with UNet+Cross-Attention and sufficient data; variable-history training cuts sample complexity in the low-data regime.

  2. SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation

    cs.RO 2025-01 conditional novelty 6.0 of 10

    SAM2Act reports 86.8% average success across 18 RLBench tasks, and the memory variant SAM2Act+ reaches 94.3% on the new MemoryBench tasks.

  3. Graph-Based Operator Learning from Limited Data on Irregular Domains

    cs.LG 2025-05 reject novelty 5.0 of 10

    GOLA combines attention-based graph message passing with a learnable Fourier encoder and reports lower relative L2 error than GKN on four 2D PDE benchmarks, especially with few training samples.

Pith tools