Pith. sign in

REVIEW 6 cited by

OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15568 v3 pith:K7C5GIGD submitted 2024-05-24 cs.AI

classification cs.AI
keywords omni-epiclearningenvironmentstaskscodecreateinterestingmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open-ended and AI-generating algorithms aim to continuously generate and solve increasingly complex tasks indefinitely, offering a promising path toward more general intelligence. To accomplish this grand vision, learning must occur within a vast array of potential tasks. Existing approaches to automatically generating environments are constrained within manually predefined, often narrow distributions of environment, limiting their ability to create any learning environment. To address this limitation, we introduce a novel framework, OMNI-EPIC, that augments previous work in Open-endedness via Models of human Notions of Interestingness (OMNI) with Environments Programmed in Code (EPIC). OMNI-EPIC leverages foundation models to autonomously generate code specifying the next learnable (i.e., not too easy or difficult for the agent's current skill set) and interesting (e.g., worthwhile and novel) tasks. OMNI-EPIC generates both environments (e.g., an obstacle course) and reward functions (e.g., progress through the obstacle course quickly without touching red objects), enabling it, in principle, to create any simulatable learning task. We showcase the explosive creativity of OMNI-EPIC, which continuously innovates to suggest new, interesting learning challenges. We also highlight how OMNI-EPIC can adapt to reinforcement learning agents' learning progress, generating tasks that are of suitable difficulty. Overall, OMNI-EPIC can endlessly create learnable and interesting environments, further propelling the development of self-improving AI systems and AI-Generating Algorithms. Project website with videos: https://dub.sh/omniepic

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Octax is a JAX-based CHIP-8 emulator that runs thousands of parallel arcade environments on GPUs (350k steps/s) and supports LLM-generated games for RL training.

  2. How Should We Meta-Learn Reinforcement Learning Algorithms?

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A systematic comparison of black-box evolution, neural and symbolic distillation, and LLM-based proposal for meta-learning RL algorithms yields practical recommendations: warm-started LLM proposal is sample-efficient,...

  3. To Trade or Not to Trade: An Agentic Approach to Estimating Market Risk Improves Trading Decisions

    q-fin.ST 2025-07 conditional novelty 6.0 of 10

    LLM-discovered stochastic models of price paths provide risk metrics that improve trader-agent decisions, raising average Sharpe ratios from 0.88 to 1.40 in the paper's backtests.

  4. LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

    cs.AI 2026-07 conditional novelty 5.0 of 10

    LEMUR jointly learns a separate reward model for each teacher's preferences and uses them to train a population of multi-objective policies, beating baselines that merge feedback into one reward.

  5. Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games

    cs.AI 2025-06 conditional novelty 4.0 of 10

    An LM-driven feedback loop that tunes reward weights from scalar performance statistics reaches 80.4% lap success in a racing task, close to a human expert's peak of 93.6%.

  6. VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems

    cs.AI 2025-06 conditional novelty 4.0 of 10

    VoyagerVision combines GPT-4o with point-of-view screenshots in the Voyager Minecraft agent, producing 18 verified unit-test building successes out of 50 attempts and 2.75 unique structures per 50-step open-ended run.

Pith tools