Pith. sign in

REVIEW 2 cited by

Diagnosing Reinforcement Learning for Traffic Signal Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.04716 v1 pith:SMWJTN7S submitted 2019-05-12 cs.LG cs.AI

classification cs.LGcs.AI
keywords trafficcontrolsignalrewardstatelearningapproachesdesign
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the increasing availability of traffic data and advance of deep reinforcement learning techniques, there is an emerging trend of employing reinforcement learning (RL) for traffic signal control. A key question for applying RL to traffic signal control is how to define the reward and state. The ultimate objective in traffic signal control is to minimize the travel time, which is difficult to reach directly. Hence, existing studies often define reward as an ad-hoc weighted linear combination of several traffic measures. However, there is no guarantee that the travel time will be optimized with the reward. In addition, recent RL approaches use more complicated state (e.g., image) in order to describe the full traffic situation. However, none of the existing studies has discussed whether such a complex state representation is necessary. This extra complexity may lead to significantly slower learning process but may not necessarily bring significant performance gain. In this paper, we propose to re-examine the RL approaches through the lens of classic transportation theory. We ask the following questions: (1) How should we design the reward so that one can guarantee to minimize the travel time? (2) How to design a state representation which is concise yet sufficient to obtain the optimal solution? Our proposed method LIT is theoretically supported by the classic traffic signal control methods in transportation field. LIT has a very simple state and reward design, thus can serve as a building block for future RL approaches to traffic signal control. Extensive experiments on both synthetic and real datasets show that our method significantly outperforms the state-of-the-art traffic signal control methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GPLight+: A Genetic Programming Method for Learning Symmetric Traffic Signal Control Policy

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Enforcing symmetry in GP-evolved phase urgency functions, by sharing a single subtree between the two turn movements and summing their urgencies, improves traffic signal control performance on 5 of 6 benchmark datasets.

  2. Joint-Local Grounded Action Transformation for Sim-to-Real Transfer in Multi-Agent Traffic Control

    cs.LG 2025-07 conditional novelty 6.0 of 10

    JL-GAT extends grounded action transformation to multi-agent traffic signal control by feeding each agent's grounding models with neighboring state and action information, reducing the sim-to-real gap in simulated rai...

Pith tools