Pith. sign in

REVIEW 2 cited by

Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.05829 v1 pith:EEJUYERT submitted 2019-09-12 cs.LG cs.AIcs.CVcs.ROstat.ML

classification cs.LGcs.AIcs.CVcs.ROstat.ML
keywords planningsubgoaltasksvisualapproachesclutteredforesightgeneration
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Video prediction models combined with planning algorithms have shown promise in enabling robots to learn to perform many vision-based tasks through only self-supervision, reaching novel goals in cluttered scenes with unseen objects. However, due to the compounding uncertainty in long horizon video prediction and poor scalability of sampling-based planning optimizers, one significant limitation of these approaches is the ability to plan over long horizons to reach distant goals. To that end, we propose a framework for subgoal generation and planning, hierarchical visual foresight (HVF), which generates subgoal images conditioned on a goal image, and uses them for planning. The subgoal images are directly optimized to decompose the task into easy to plan segments, and as a result, we observe that the method naturally identifies semantically meaningful states as subgoals. Across three out of four simulated vision-based manipulation tasks, we find that our method achieves nearly a 200% performance improvement over planning without subgoals and model-free RL approaches. Further, our experiments illustrate that our approach extends to real, cluttered visual scenes. Project page: https://sites.google.com/stanford.edu/hvf

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Extracting Visual Plans from Unlabeled Videos via Symbolic Guidance

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Vis2Plan extracts object symbols from unlabeled play videos with vision models, plans symbolically with A* search, and retrieves reachable real images as subgoals for a goal-conditioned robot policy.

  2. Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning

    cs.LG 2025-02 reject novelty 4.0 of 10

    VSC-RL combines VLM-generated subgoals with a subgoal-conditioned AWR-style RL objective and claims improved sample efficiency over DigiRL and WebRL on AitW and WebArena-Lite.

Pith tools