REVIEW 2 cited by
Iroko: A Framework to Prototype Reinforcement Learning for Data Center Traffic Control
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent networking research has identified that data-driven congestion control (CC) can be more efficient than traditional CC in TCP. Deep reinforcement learning (RL), in particular, has the potential to learn optimal network policies. However, RL suffers from instability and over-fitting, deficiencies which so far render it unacceptable for use in datacenter networks. In this paper, we analyze the requirements for RL to succeed in the datacenter context. We present a new emulator, Iroko, which we developed to support different network topologies, congestion control algorithms, and deployment scenarios. Iroko interfaces with the OpenAI gym toolkit, which allows for fast and fair evaluation of different RL and traditional CC algorithms under the same conditions. We present initial benchmarks on three deep RL algorithms compared to TCP New Vegas and DCTCP. Our results show that these algorithms are able to learn a CC policy which exceeds the performance of TCP New Vegas on a dumbbell and fat-tree topology. We make our emulator open-source and publicly available: https://github.com/dcgym/iroko
Forward citations
Cited by 2 Pith papers
-
ConfigTron: Tackling network diversity with heterogeneous configurations
ConfigTron uses a contextual multi-armed bandit to learn per-network-class TCP and HTTP configurations online, cutting median page load time by up to 19% in simulation and 8-10% in a live deployment.
-
A View on Deep Reinforcement Learning in System Optimization
A survey of deep RL in system optimization finds that many published solutions lack strong baselines, reproducibility, and comparisons to random search, and proposes evaluation questions to address this.
Discussion (0). Continue with ORCID to comment.