Pith. sign in

REVIEW 1 cited by

How to Avoid Being Eaten by a Grue: Structured Exploration Strategies for Textual Worlds

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.07409 v1 pith:424ARYA7 submitted 2020-06-12 cs.AI cs.CLcs.LGstat.ML

classification cs.AIcs.CLcs.LGstat.ML
keywords agentagentsovercomebertbottleneckseatenexplorationgames
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-based games are long puzzles or quests, characterized by a sequence of sparse and potentially deceptive rewards. They provide an ideal platform to develop agents that perceive and act upon the world using a combinatorially sized natural language state-action space. Standard Reinforcement Learning agents are poorly equipped to effectively explore such spaces and often struggle to overcome bottlenecks---states that agents are unable to pass through simply because they do not see the right action sequence enough times to be sufficiently reinforced. We introduce Q*BERT, an agent that learns to build a knowledge graph of the world by answering questions, which leads to greater sample efficiency. To overcome bottlenecks, we further introduce MC!Q*BERT an agent that uses an knowledge-graph-based intrinsic motivation to detect bottlenecks and a novel exploration strategy to efficiently learn a chain of policy modules to overcome them. We present an ablation study and results demonstrating how our method outperforms the current state-of-the-art on nine text games, including the popular game, Zork, where, for the first time, a learning agent gets past the bottleneck where the player is eaten by a Grue.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Odyssey of the Fittest: Can Agents Survive and Still Be Good?

    cs.AI 2025-02 reject novelty 6.0 of 10

    In an LLM-generated text survival game, a GPT-4o agent was reported to survive better and score more ethically than NEAT and SVI Bayesian agents, but the evaluation is circular because GPT-4o labels its own behavior.

Pith tools