Pith. sign in

REVIEW 7 cited by

Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.03647 v2 pith:E3SNX2HM submitted 2020-06-05 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords policylearningbremendata-collectionefficiencyofflinealgorithmalgorithms
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Most reinforcement learning (RL) algorithms assume online access to the environment, in which one may readily interleave updates to the policy with experience collection using that policy. However, in many real-world applications such as health, education, dialogue agents, and robotics, the cost or potential risk of deploying a new data-collection policy is high, to the point that it can become prohibitive to update the data-collection policy more than a few times during learning. With this view, we propose a novel concept of deployment efficiency, measuring the number of distinct data-collection policies that are used during policy learning. We observe that na\"{i}vely applying existing model-free offline RL algorithms recursively does not lead to a practical deployment-efficient and sample-efficient algorithm. We propose a novel model-based algorithm, Behavior-Regularized Model-ENsemble (BREMEN) that can effectively optimize a policy offline using 10-20 times fewer data than prior works. Furthermore, the recursive application of BREMEN is able to achieve impressive deployment efficiency while maintaining the same or better sample efficiency, learning successful policies from scratch on simulated robotic environments with only 5-10 deployments, compared to typical values of hundreds to millions in standard RL baselines. Codes and pre-trained models are available at https://github.com/matsuolab/BREMEN .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL

    cs.LG 2026-08 conditional novelty 6.0 of 10

    CSDG modifies the offline Bellman backup by scaling a smoothed convex-hull-neighborhood correction against an in-sample expectile target and reports strong D4RL aggregate performance.

  2. Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation

    cs.LG 2026-07 reject novelty 6.0 of 10

    CQ2L uses Morse-network uncertainty to select in-distribution action queries and to scale CQL's regularization, reporting higher D4RL scores than the prior OAP method.

  3. Improving Sequential Recommenders through Counterfactual Augmentation of System Exposure

    cs.IR 2025-04 reject novelty 6.0 of 10

    A decision transformer trained on logged plus counterfactually augmented exposure sequences outperforms sequential recommendation baselines on ZhihuRec, Tenrec, and KuaiRand in the reported experiments.

  4. Unsupervised Data Generation for Offline Reinforcement Learning: A Perspective from Model

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Training a diverse set of policies, relabeling their experience with the target reward, and selecting the highest-return buffer improves model-based offline RL on unknown tasks, supported by a Wasserstein-distance analysis.

  5. A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics

    cs.RO 2025-05 conditional novelty 5.0 of 10

    A three-phase taxonomy of physical risk control for foundation-model-enabled robots, with identified research gaps.

  6. Iterative Batch Reinforcement Learning via Safe Diversified Model-based Policy Search

    cs.LG 2024-11 conditional novelty 4.0 of 10

    An iterative batch RL method combining model-based policy search with minimum pairwise trajectory diversity and behavior-based safety constraints speeds up cost reduction across batch iterations in Industrial Benchmar...

  7. A Survey of Reinforcement Learning for Optimization in Automation

    cs.LG 2025-02 conditional novelty 2.0 of 10

    A structured survey of reinforcement learning methods applied to optimization across manufacturing, energy, and robotics, with challenges and future directions.

Pith tools