Pith. sign in

REVIEW 4 cited by

A Definition of Continual Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.11046 v2 pith:B6ACLNIW submitted 2023-07-20 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningcontinualreinforcementagentsdefinitionproblemagentbest
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In a standard view of the reinforcement learning problem, an agent's goal is to efficiently identify a policy that maximizes long-term reward. However, this perspective is based on a restricted view of learning as finding a solution, rather than treating learning as endless adaptation. In contrast, continual reinforcement learning refers to the setting in which the best agents never stop learning. Despite the importance of continual reinforcement learning, the community lacks a simple definition of the problem that highlights its commitments and makes its primary concepts precise and clear. To this end, this paper is dedicated to carefully defining the continual reinforcement learning problem. We formalize the notion of agents that "never stop learning" through a new mathematical language for analyzing and cataloging agents. Using this new language, we define a continual learning agent as one that can be understood as carrying out an implicit search process indefinitely, and continual reinforcement learning as the setting in which the best agents are all continual learning agents. We provide two motivating examples, illustrating that traditional views of multi-task reinforcement learning and continual supervised learning are special cases of our definition. Collectively, these definitions and perspectives formalize many intuitive concepts at the heart of learning, and open new research pathways surrounding continual learning agents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 9 citations worldwide. Full citation record

  1. LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

    cs.AI 2026-07 conditional novelty 5.0 of 10

    LEMUR jointly learns a separate reward model for each teacher's preferences and uses them to train a population of multi-objective policies, beating baselines that merge feedback into one reward.

  2. Continual-RL for Generalization in Autonomous Racing on the RoboRacer Platform

    cs.RO 2026-07 conditional novelty 5.0 of 10

    SAC plus Continual Backpropagation, trained only on real multi-track data, fine-tunes in ~15 minutes on an unseen lower-friction RoboRacer track and outperforms MAP and MPC.

  3. Adaptive Multi-Horizon Reinforcement Learning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A gated mixture of Q-functions with different discount factors, trained with undiscounted Bellman error, adapts its temporal horizon in small MiniGrid tasks, but its theoretical justification is circular and baseline ...

  4. An Arbitration Control for an Ensemble of Diversified DQN variants in Continual Reinforcement Learning

    cs.LG 2025-09 conditional novelty 5.0 of 10

    ACED-DQN combines heterogeneous DQN variants with loss-based reliability weighting and experience assignment, but the paper's own ablation indicates that arbitration control is not the key factor behind the performance gain.

Pith tools