REVIEW 5 cited by
EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Sample efficiency remains a crucial challenge in applying Reinforcement Learning (RL) to real-world tasks. While recent algorithms have made significant strides in improving sample efficiency, none have achieved consistently superior performance across diverse domains. In this paper, we introduce EfficientZero V2, a general framework designed for sample-efficient RL algorithms. We have expanded the performance of EfficientZero to multiple domains, encompassing both continuous and discrete actions, as well as visual and low-dimensional inputs. With a series of improvements we propose, EfficientZero V2 outperforms the current state-of-the-art (SOTA) by a significant margin in diverse tasks under the limited data setting. EfficientZero V2 exhibits a notable advancement over the prevailing general algorithm, DreamerV3, achieving superior outcomes in 50 of 66 evaluated tasks across diverse benchmarks, such as Atari 100k, Proprio Control, and Vision Control.
Forward citations
Cited by 5 Pith papers
-
MuJoCo Playground
An open-source, MJX-based robot learning framework with integrated batch rendering that provides fast training and demonstrates sim-to-real transfer on six robot platforms.
-
Human-Like Goalkeeping in a Realistic Football Simulation: a Sample-Efficient Reinforcement Learning Approach
A sample-efficient SAC-based method with replay-ratio resets, offline data bootstrapping, and expert-driven fine-tuning produces a goalkeeper that outperforms the built-in AI in EA SPORTS FC 25.
-
AXIOM: Learning to Play Games in Minutes with Expanding Object-Centric Models
AXIOM, a gradient-free active inference agent with growing and pruning object-centric mixture models, achieves better or similar reward than BBF and DreamerV3 after 10,000 interactions on the custom Gameworld 10k suite.
-
Hadamax Encoding: Elevating Performance in Model-Free Atari
Hadamax, a Hadamard-product and max-pooling encoder, improves PQN's median human-normalized Atari-57 score by about 80% with no algorithmic changes.
-
Sample-Efficient Reinforcement Learning Controller for Deep Brain Stimulation in Parkinson's Disease
A DDPG-based adaptive DBS controller with a predictive reward model and Gumbel-Softmax exploration suppresses beta power faster than standard DDPG in simulation and survives FP16 quantization.
Discussion (0). Continue with ORCID to comment.