Pith. sign in

REVIEW 5 cited by

An empirical investigation of the challenges of real-world reinforcement learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.11881 v2 pith:QH3BTW33 submitted 2020-03-24 cs.LG cs.AI

classification cs.LGcs.AI
keywords challengesreal-worldlearningserieschallengeproposedreinforcementsome
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning (RL) has proven its worth in a series of artificial domains, and is beginning to show some successes in real-world scenarios. However, much of the research advances in RL are hard to leverage in real-world systems due to a series of assumptions that are rarely satisfied in practice. In this work, we identify and formalize a series of independent challenges that embody the difficulties that must be addressed for RL to be commonly deployed in real-world systems. For each challenge, we define it formally in the context of a Markov Decision Process, analyze the effects of the challenge on state-of-the-art learning algorithms, and present some existing attempts at tackling it. We believe that an approach that addresses our set of proposed challenges would be readily deployable in a large number of real world problems. Our proposed challenges are implemented in a suite of continuous control environments called the realworldrl-suite which we propose an as an open-source benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distributionally Robust Control via Stein Variational Inference for Contact-Rich Manipulation

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    SV-DRO evolves parameter particles via task-optimality-gap Stein gradients inside DRO-MPC, yielding up to 3× higher success on contact-rich manipulation under parametric uncertainty.

  2. HypEMBER: Hypernetwork-based Ensemble for Robust Policy Learning of Parametrized Dynamical Systems

    cs.LG 2026-07 conditional novelty 5.0 of 10

    HypEMBER joins hypernetwork-generated policies with an ensemble critic to improve robustness of reinforcement-learning controllers for parametrized dynamical systems.

  3. Gym4ReaL: A Suite for Benchmarking Real-World Reinforcement Learning

    cs.LG 2025-06 conditional novelty 4.0 of 10

    The paper presents Gym4ReaL, a benchmark suite of six realistic RL environments, and shows that standard PPO, DQN, Q-Learning, and SARSA agents beat simple rule-based baselines on most tasks.

  4. Towards Bio-inspired Heuristically Accelerated Reinforcement Learning for Adaptive Underwater Multi-Agents Behaviour

    cs.RO 2025-02 reject novelty 4.0 of 10

    PSO-guided exploration is added to MASAC and claimed to reduce training time for a 3-agent coverage planning task, without quantitative validation.

  5. A Survey of Reinforcement Learning for Optimization in Automation

    cs.LG 2025-02 conditional novelty 2.0 of 10

    A structured survey of reinforcement learning methods applied to optimization across manufacturing, energy, and robotics, with challenges and future directions.

Pith tools