Pith. sign in

REVIEW 1 cited by

ORL: Reinforcement Learning Benchmarks for Online Stochastic Optimization Problems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.10641 v2 pith:VZTCFMIW submitted 2019-11-24 cs.LG cs.AImath.OC

classification cs.LGcs.AImath.OC
keywords problemsalgorithmsapproachesbenchmarkslearningonlineoptimizationperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement Learning (RL) has achieved state-of-the-art results in domains such as robotics and games. We build on this previous work by applying RL algorithms to a selection of canonical online stochastic optimization problems with a range of practical applications: Bin Packing, Newsvendor, and Vehicle Routing. While there is a nascent literature that applies RL to these problems, there are no commonly accepted benchmarks which can be used to compare proposed approaches rigorously in terms of performance, scale, or generalizability. This paper aims to fill that gap. For each problem we apply both standard approaches as well as newer RL algorithms and analyze results. In each case, the performance of the trained RL policy is competitive with or superior to the corresponding baselines, while not requiring much in the way of domain knowledge. This highlights the potential of RL in real-world dynamic resource allocation problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Structure-Informed Deep Reinforcement Learning for Inventory Management

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A generic DirectBackprop deep RL policy, trained only on historical demand across many products, matches or beats classical inventory heuristics in five problem settings, and structural monotonicity penalties improve ...

Pith tools