Pith. sign in

REVIEW 4 cited by

Wonderful Team: Zero-Shot Physical Task Planning with Visual LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.19094 v6 pith:CWNBWH3Q submitted 2024-07-26 cs.AI cs.RO

classification cs.AIcs.RO
keywords planninghigh-levelaverageimprovementmethodsrobotictasktasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Wonderful Team, a multi-agent Vision Large Language Model (VLLM) framework for executing high-level robotic planning in a zero-shot regime. In our context, zero-shot high-level planning means that for a novel environment, we provide a VLLM with an image of the robot's surroundings and a task description, and the VLLM outputs the sequence of actions necessary for the robot to complete the task. Unlike previous methods for high-level visual planning for robotic manipulation, our method uses VLLMs for the entire planning process, enabling a more tightly integrated loop between perception, control, and planning. As a result, Wonderful Team's performance on real-world semantic and physical planning tasks often exceeds methods that rely on separate vision systems. For example, we see an average 40% success rate improvement on VimaBench over prior methods such as NLaP, an average 30% improvement over Trajectory Generators on tasks from the Trajectory Generator paper, including drawing and wiping a plate, and an average 70% improvement over Trajectory Generators on a new set of semantic reasoning tasks including environment rearrangement with implicit linguistic constraints. We hope these results highlight the rapid improvements of VLLMs in the past year, and motivate the community to consider VLLMs as an option for some high-level robotic planning problems in the future.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Long-Term Memory for VLA-based Agents in Open-World Task Execution

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    ChemBot adds dual-layer memory and future-state asynchronous inference to VLA models, enabling better long-horizon success in chemical lab automation on collaborative robots.

  2. EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues

    cs.CV 2024-12 conditional novelty 6.0 of 10

    EarthDial is a 4B-parameter remote sensing chatbot trained on 11.11M instruction pairs to handle multi-resolution, multi-spectral, and multi-temporal satellite imagery, and it reports gains over prior VLMs on dozens o...

  3. RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World

    cs.RO 2024-11 reject novelty 4.0 of 10

    A skill-centric hierarchical framework with a unified vision-language-action model executes new tasks by recombining eight meta-skills, reporting up to 50 percentage points higher success than task-centric baselines.

  4. Visual Large Language Models for Generalized and Specialized Applications

    cs.CV 2025-01 conditional novelty 3.0 of 10

    This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.

Pith tools