Pith. sign in

REVIEW 1 cited by

Comparing Deep Reinforcement Learning Algorithms in Two-Echelon Supply Chains

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.09603 v3 pith:B6QNPTVA submitted 2022-04-20 cs.LG cs.AImath.OC

classification cs.LGcs.AImath.OC
keywords supplyalgorithmschaindeeplearningreinforcementinventorymanagement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this study, we analyze and compare the performance of state-of-the-art deep reinforcement learning algorithms for solving the supply chain inventory management problem. This complex sequential decision-making problem consists of determining the optimal quantity of products to be produced and shipped across different warehouses over a given time horizon. In particular, we present a mathematical formulation of a two-echelon supply chain environment with stochastic and seasonal demand, which allows managing an arbitrary number of warehouses and product types. Through a rich set of numerical experiments, we compare the performance of different deep reinforcement learning algorithms under various supply chain structures, topologies, demands, capacities, and costs. The results of the experimental plan indicate that deep reinforcement learning algorithms outperform traditional inventory management strategies, such as the static (s, Q)-policy. Furthermore, this study provides detailed insight into the design and development of an open-source software library that provides a customizable environment for solving the supply chain inventory management problem using a wide range of data-driven approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Collaborating in a competitive world: Heterogeneous Multi-Agent Decision Making in Symbiotic Supply Chain Environments

    cs.MA 2025-01 conditional novelty 5.0 of 10

    Separate per-node policies reduce the bullwhip effect in a simulated two-node supply chain, but a single shared policy earns more in low-demand settings; SAC beats PPO in high demand.

Pith tools