REVIEW 7 cited by
Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Most reinforcement learning (RL) algorithms assume online access to the environment, in which one may readily interleave updates to the policy with experience collection using that policy. However, in many real-world applications such as health, education, dialogue agents, and robotics, the cost or potential risk of deploying a new data-collection policy is high, to the point that it can become prohibitive to update the data-collection policy more than a few times during learning. With this view, we propose a novel concept of deployment efficiency, measuring the number of distinct data-collection policies that are used during policy learning. We observe that na\"{i}vely applying existing model-free offline RL algorithms recursively does not lead to a practical deployment-efficient and sample-efficient algorithm. We propose a novel model-based algorithm, Behavior-Regularized Model-ENsemble (BREMEN) that can effectively optimize a policy offline using 10-20 times fewer data than prior works. Furthermore, the recursive application of BREMEN is able to achieve impressive deployment efficiency while maintaining the same or better sample efficiency, learning successful policies from scratch on simulated robotic environments with only 5-10 deployments, compared to typical values of hundreds to millions in standard RL baselines. Codes and pre-trained models are available at https://github.com/matsuolab/BREMEN .
Forward citations
Cited by 7 Pith papers
-
Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL
CSDG modifies the offline Bellman backup by scaling a smoothed convex-hull-neighborhood correction against an in-sample expectile target and reports strong D4RL aggregate performance.
-
Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation
CQ2L uses Morse-network uncertainty to select in-distribution action queries and to scale CQL's regularization, reporting higher D4RL scores than the prior OAP method.
-
Improving Sequential Recommenders through Counterfactual Augmentation of System Exposure
A decision transformer trained on logged plus counterfactually augmented exposure sequences outperforms sequential recommendation baselines on ZhihuRec, Tenrec, and KuaiRand in the reported experiments.
-
Unsupervised Data Generation for Offline Reinforcement Learning: A Perspective from Model
Training a diverse set of policies, relabeling their experience with the target reward, and selecting the highest-return buffer improves model-based offline RL on unknown tasks, supported by a Wasserstein-distance analysis.
-
A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics
A three-phase taxonomy of physical risk control for foundation-model-enabled robots, with identified research gaps.
-
Iterative Batch Reinforcement Learning via Safe Diversified Model-based Policy Search
An iterative batch RL method combining model-based policy search with minimum pairwise trajectory diversity and behavior-based safety constraints speeds up cost reduction across batch iterations in Industrial Benchmar...
-
A Survey of Reinforcement Learning for Optimization in Automation
A structured survey of reinforcement learning methods applied to optimization across manufacturing, energy, and robotics, with challenges and future directions.
Discussion (0). Continue with ORCID to comment.