Pith. sign in

REVIEW 1 cited by

Exploratory Combinatorial Optimization with Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.04063 v2 pith:RI4PRR3Y submitted 2019-09-09 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords combinatorialoptimizationlearningagentapproacheco-dqnexploratorygraph
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many real-world problems can be reduced to combinatorial optimization on a graph, where the subset or ordering of vertices that maximize some objective function must be found. With such tasks often NP-hard and analytically intractable, reinforcement learning (RL) has shown promise as a framework with which efficient heuristic methods to tackle these problems can be learned. Previous works construct the solution subset incrementally, adding one element at a time, however, the irreversible nature of this approach prevents the agent from revising its earlier decisions, which may be necessary given the complexity of the optimization task. We instead propose that the agent should seek to continuously improve the solution by learning to explore at test time. Our approach of exploratory combinatorial optimization (ECO-DQN) is, in principle, applicable to any combinatorial problem that can be defined on a graph. Experimentally, we show our method to produce state-of-the-art RL performance on the Maximum Cut problem. Moreover, because ECO-DQN can start from any arbitrary configuration, it can be combined with other search methods to further improve performance, which we demonstrate using a simple random search.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Nonlocal Monte Carlo via Reinforcement Learning

    cs.LG 2025-08 conditional novelty 6.0 of 10

    A reinforcement-learning-trained policy for selecting nonlocal cluster moves improves a Monte Carlo solver for hard 4-SAT benchmarks over simulated annealing.

Pith tools