Pith. sign in

REVIEW 11 cited by

BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.08272 v4 pith:YP7YTWYO submitted 2018-10-18 cs.AI cs.CL

classification cs.AIcs.CL
keywords languagelearningbabyaiplatformagentlevelscurrentefficiency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Allowing humans to interactively train artificial agents to understand language instructions is desirable for both practical and scientific reasons, but given the poor data efficiency of the current learning methods, this goal may require substantial research efforts. Here, we introduce the BabyAI research platform to support investigations towards including humans in the loop for grounded language learning. The BabyAI platform comprises an extensible suite of 19 levels of increasing difficulty. The levels gradually lead the agent towards acquiring a combinatorially rich synthetic language which is a proper subset of English. The platform also provides a heuristic expert agent for the purpose of simulating a human teacher. We report baseline results and estimate the amount of human involvement that would be required to train a neural network-based agent on some of the BabyAI levels. We put forward strong evidence that current deep learning methods are not yet sufficiently sample efficient when it comes to learning a language with compositional properties.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Long-Horizon Embodied Decision-Making via Multimodal Memory Compression

    cs.CV 2026-08 conditional novelty 6.0 of 10

    DunphyBench tests long-horizon, preference-driven house selection in virtual homes; MeMento, a preference-conditioned memory compressor, raises VLM agent accuracy by 7.18% and cuts memory by 85.38%.

  2. Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Relabeling LLM-agent trajectories with all goals actually achieved, plus action masking and reweighting, yields sample-efficient gains over SFT and DPO on ALFWorld, PlanCraft, and WebShop.

  3. From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Replaying teacher trajectory prefixes and explicitly optimizing the historical prefix tokens improves small-model agent success rates over distillation and response-only GRPO baselines in TextCraft, BabyAI, and ALFWorld.

  4. Enhancing Decision-Making of Large Language Models via Actor-Critic

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LAC improves LLM decision-making by computing action scores from token logits and combining them with the model's prior policy through a gradient-free KL-constrained update.

  5. From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A language-instruction embedding can replace the gradient-based inner loop of MAML, yielding competitive BabyAI performance with lower per-iteration wall-clock time.

  6. Dual-Process Atomic Skill Learning: Decoupling Semantic Reasoning and Real-Time Control

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Asynchronous dual-frequency hierarchical imitation learning with VQ skills and training-only latent diffusion improves compositional language-conditioned robot control and reduces skill codebook collapse.

  7. MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

    cs.AI 2025-07 reject novelty 5.0 of 10

    A new maze-navigation benchmark claims LLM spatial reasoning is language-dependent, with O3 exceptional and other models failing by looping, but the looping result is an artifact of the termination rule.

  8. ChatPD: An LLM-driven Paper-Dataset Networking System

    cs.DB 2025-05 conditional novelty 5.0 of 10

    ChatPD automatically builds a paper-dataset network by using LLMs to extract dataset mentions from papers and a graph-based algorithm to match them to known datasets, outperforming PapersWithCode in coverage.

  9. An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models

    cs.LG 2025-06 conditional novelty 4.0 of 10

    MultiNet provides an open-source benchmark, data SDK, evaluation harness, and adapted VLA models for assessing generalization across vision, language, and action tasks.

  10. Large Language Models for Planning: A Comprehensive and Systematic Survey

    cs.AI 2025-05 conditional novelty 3.0 of 10

    A structured survey of LLM planning methods, benchmarks, and interpretability work, organized around a three-way taxonomy.

  11. Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact

    cs.AI 2025-07 conditional novelty 2.0 of 10

    A broad survey arguing that AGI requires modular, memory-augmented, embodied architectures rather than scaled-up token prediction, with a brief proposal to decompose intelligence into five components.

Pith tools