REVIEW 2 cited by
Few-shot Subgoal Planning with Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Pre-trained large language models have shown successful progress in many language understanding benchmarks. This work explores the capability of these models to predict actionable plans in real-world environments. Given a text instruction, we show that language priors encoded in pre-trained language models allow us to infer fine-grained subgoal sequences. In contrast to recent methods which make strong assumptions about subgoal supervision, our experiments show that language models can infer detailed subgoal sequences from few training sequences without any fine-tuning. We further propose a simple strategy to re-rank language model predictions based on interaction and feedback from the environment. Combined with pre-trained navigation and visual reasoning components, our approach demonstrates competitive performance on subgoal prediction and task completion in the ALFRED benchmark compared to prior methods that assume more subgoal supervision.
Forward citations
Cited by 2 Pith papers
-
MuTRAP: Multi-trigger Trojans Attacking Robot Task Planning Systems
A two-step soft-prompt backdoor attack called Robo-Troj (listed as MuTRAP on arXiv) makes LLM-based robot planners emit malicious plans when hidden trigger words are present, with near-perfect attack success.
-
LLM-Based Offline Learning for Embodied Agents via Consistency-Guided Reward Ensemble
CoREN uses an LLM offline to estimate dense action rewards, filters them through three consistency checks, and aligns them to sparse success labels to train a small offline RL agent for household instruction-following tasks.
Discussion (0). Continue with ORCID to comment.