Pith. sign in

REVIEW 4 cited by

LoHoRavens: A Long-Horizon Language-Conditioned Benchmark for Robotic Tabletop Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.12020 v2 pith:HGREYYVS submitted 2023-10-18 cs.RO cs.CLcs.CV

classification cs.ROcs.CLcs.CV
keywords long-horizonmanipulationtasksbenchmarkllmsmethodsmodelsreasoning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The convergence of embodied agents and large language models (LLMs) has brought significant advancements to embodied instruction following. Particularly, the strong reasoning capabilities of LLMs make it possible for robots to perform long-horizon tasks without expensive annotated demonstrations. However, public benchmarks for testing the long-horizon reasoning capabilities of language-conditioned robots in various scenarios are still missing. To fill this gap, this work focuses on the tabletop manipulation task and releases a simulation benchmark, \textit{LoHoRavens}, which covers various long-horizon reasoning aspects spanning color, size, space, arithmetics and reference. Furthermore, there is a key modality bridging problem for long-horizon manipulation tasks with LLMs: how to incorporate the observation feedback during robot execution for the LLM's closed-loop planning, which is however less studied by prior work. We investigate two methods of bridging the modality gap: caption generation and learnable interface for incorporating explicit and implicit observation feedback to the LLM, respectively. These methods serve as the two baselines for our proposed benchmark. Experiments show that both methods struggle to solve some tasks, indicating long-horizon manipulation tasks are still challenging for current popular models. We expect the proposed public benchmark and baselines can help the community develop better models for long-horizon tabletop manipulation tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation

    cs.RO 2025-05 conditional novelty 7.0 of 10

    PartInstruct is a new large-scale simulated benchmark with part-level language instructions and training demonstrations; current robot policies achieve at most 31.72% average success on it.

  2. TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A training-free pipeline generates instance-level, physically interactive 3D tabletop scenes from text or one image, with a differentiable rotation optimizer and top-view spatial alignment for collision-free layouts.

  3. RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A hierarchical pipeline decomposes long-horizon robot instructions into keyframes, interpolates between them, and regresses joint states from the generated video, reaching 67.4% success on simulated long-horizon tasks.

  4. LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A unified vision-language-action model that emits a sub-task description followed by a discrete action token outperforms modular and action-only baselines on simulated long-horizon tabletop tasks.

Pith tools