Pith. sign in

REVIEW 2 cited by

Graph neural induction of value iteration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.12604 v1 pith:JNMCIG3C submitted 2020-09-26 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords iterationvaluegraphnetworkneuralplanningacrossbeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many reinforcement learning tasks can benefit from explicit planning based on an internal model of the environment. Previously, such planning components have been incorporated through a neural network that partially aligns with the computational graph of value iteration. Such network have so far been focused on restrictive environments (e.g. grid-worlds), and modelled the planning procedure only indirectly. We relax these constraints, proposing a graph neural network (GNN) that executes the value iteration (VI) algorithm, across arbitrary environment models, with direct supervision on the intermediate steps of VI. The results indicate that GNNs are able to model value iteration accurately, recovering favourable metrics and policies across a variety of out-of-distribution tests. This suggests that GNN executors with strong supervision are a viable component within deep reinforcement learning systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Primal-Dual Neural Algorithmic Reasoning

    cs.LG 2025-05 conditional novelty 7.0 of 10

    A GNN framework that simulates primal-dual approximation algorithms for NP-hard problems and, with small-instance optimal labels, can beat the algorithm it learns.

  2. Unrolling Dynamic Programming via Graph Filters

    cs.AI 2025-07 conditional novelty 5.0 of 10

    BellNet learns graph-filter coefficients for truncated policy iteration and approximates optimal policies in fewer steps than classical DP on grid-world tasks.

Pith tools