Pith. sign in

REVIEW 4 cited by

Pointer Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1506.03134 v2 pith:XL3ZF24V submitted 2015-06-09 stat.ML cs.CGcs.LGcs.NE

classification stat.MLcs.CGcs.LGcs.NE
keywords attentionoutputproblemsinputneuralvariablepointersequence
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce a new neural architecture to learn the conditional probability of an output sequence with elements that are discrete tokens corresponding to positions in an input sequence. Such problems cannot be trivially addressed by existent approaches such as sequence-to-sequence and Neural Turing Machines, because the number of target classes in each step of the output depends on the length of the input, which is variable. Problems such as sorting variable sized sequences, and various combinatorial optimization problems belong to this class. Our model solves the problem of variable size output dictionaries using a recently proposed mechanism of neural attention. It differs from the previous attention attempts in that, instead of using attention to blend hidden units of an encoder to a context vector at each decoder step, it uses attention as a pointer to select a member of the input sequence as the output. We call this architecture a Pointer Net (Ptr-Net). We show Ptr-Nets can be used to learn approximate solutions to three challenging geometric problems -- finding planar convex hulls, computing Delaunay triangulations, and the planar Travelling Salesman Problem -- using training examples alone. Ptr-Nets not only improve over sequence-to-sequence with input attention, but also allow us to generalize to variable size output dictionaries. We show that the learnt models generalize beyond the maximum lengths they were trained on. We hope our results on these tasks will encourage a broader exploration of neural learning for discrete problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 134 citations worldwide. Full citation record

  1. DEGR: Dual Exploration-Driven Generative Re-Ranking for Adaptive Cross-Request Context Bridging

    cs.IR 2026-08 conditional novelty 6.0 of 10

    DEGR trains a generative re-ranker with a learned reward that balances immediate clicks against exploratory browsing, and reports modest online gains on JD's homepage.

  2. Graph Optimization Foundation Model: Tokenizing Graph via A Language-Model Paradigm

    cs.LG 2025-09 reject novelty 5.0 of 10

    A per-graph BERT-style masked random-walk model is repurposed to generate shortest paths and tours, with mixed quality versus classical solvers and no cross-graph transfer evaluation.

  3. Efficient End-to-End Learning for Decision-Making: A Meta-Optimization Approach

    cs.LG 2025-05 conditional novelty 5.0 of 10

    ProjectNet learns a matrix-parameterized update rule that approximates optimization solutions in a few forward steps, and using this surrogate in end-to-end training cuts training time by 2 to 10 times while keeping d...

  4. Machine Reading Comprehension: a Literature Review

    cs.CL 2019-06 unverdicted novelty 1.0 of 10

    A 2019 survey of machine reading comprehension corpora and methods.

Pith tools