Pith. sign in

REVIEW 1 cited by

Visually-Grounded Planning without Vision: Language Models Infer Detailed Plans from High-level Instructions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.14259 v2 pith:UWDEXLHT submitted 2020-09-29 cs.CL cs.CV

classification cs.CLcs.CV
keywords languagevirtualdirectivesenvironmentmodelsmulti-stepvisualbest-performing
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The recently proposed ALFRED challenge task aims for a virtual robotic agent to complete complex multi-step everyday tasks in a virtual home environment from high-level natural language directives, such as "put a hot piece of bread on a plate". Currently, the best-performing models are able to complete less than 5% of these tasks successfully. In this work we focus on modeling the translation problem of converting natural language directives into detailed multi-step sequences of actions that accomplish those goals in the virtual environment. We empirically demonstrate that it is possible to generate gold multi-step plans from language directives alone without any visual input in 26% of unseen cases. When a small amount of visual information is incorporated, namely the starting location in the virtual environment, our best-performing GPT-2 model successfully generates gold command sequences in 58% of cases. Our results suggest that contextualized language models may provide strong visual semantic planning modules for grounded virtual agents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MuTRAP: Multi-trigger Trojans Attacking Robot Task Planning Systems

    cs.RO 2025-04 conditional novelty 6.0 of 10

    A two-step soft-prompt backdoor attack called Robo-Troj (listed as MuTRAP on arXiv) makes LLM-based robot planners emit malicious plans when hidden trigger words are present, with near-perfect attack success.

Pith tools