Pith. sign in

REVIEW 2 cited by

Differentiable Discrete Event Simulation for Queuing Network Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.03740 v1 pith:6SBDXBTW submitted 2024-09-05 cs.LG cs.SYeess.SYmath.OC

classification cs.LGcs.SYeess.SYmath.OC
keywords controlpolicydiscreteeventnetworkqueueingsystemschallenges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Queuing network control is essential for managing congestion in job-processing systems such as service systems, communication networks, and manufacturing processes. Despite growing interest in applying reinforcement learning (RL) techniques, queueing network control poses distinct challenges, including high stochasticity, large state and action spaces, and lack of stability. To tackle these challenges, we propose a scalable framework for policy optimization based on differentiable discrete event simulation. Our main insight is that by implementing a well-designed smoothing technique for discrete event dynamics, we can compute pathwise policy gradients for large-scale queueing networks using auto-differentiation software (e.g., Tensorflow, PyTorch) and GPU parallelization. Through extensive empirical experiments, we observe that our policy gradient estimators are several orders of magnitude more accurate than typical REINFORCE-based estimators. In addition, We propose a new policy architecture, which drastically improves stability while maintaining the flexibility of neural-network policies. In a wide variety of scheduling and admission control tasks, we demonstrate that training control policies with pathwise gradients leads to a 50-1000x improvement in sample efficiency over state-of-the-art RL methods. Unlike prior tailored approaches to queueing, our methods can flexibly handle realistic scenarios, including systems operating in non-stationary environments and those with non-exponential interarrival/service times.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hard Constraints, Smooth Gradients: Learning Feasible Inventory Policies via Differentiable Projection

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A differentiable QP-projection plus dual-aware rounding layer lets deep RL policies enforce interdependent hard constraints, achieving near-optimal cost on small instances and 2.5–3.2% savings on an ASML case study.

  2. A Planning Framework for Adaptive Labeling

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A planning framework for adaptive labeling where a smoothed auto-differential policy gradient (Smoothed-Autodiff) selects batches to minimize final posterior uncertainty, outperforming active-learning heuristics and R...

Pith tools