Pith. sign in

Paper Citation Record · LEDGER

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization

As of 13 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 0 inbound Pith citation observations for arXiv:2608.00296.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00296 v1

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:49:55.058885Z

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e756a73b-31a5-4491-b22d-00e0f686a065 · outbound

This paper cites 2020 , eprint =.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization 2020 , eprint =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:54.637867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:54.637867Z digest=sha256:26dd426ee540aaefa525c0f2c20f2a1331aeaf1957c24f15a7916d5f5c8533e7

Observation d4c827ae-409e-4e16-a647-6c8960841a67 · outbound

This paper cites Winner Takes It All: Training Performant.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization Winner Takes It All: Training Performant

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:54.685712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:54.685712Z digest=sha256:b37e562c543d9f62648063bf739e3b8b880f78dc11abf129a6714d9743f8c3a7

Observation 3cf052a6-a3e8-499e-bf48-cd60dc73ef02 · outbound

This paper cites Leader Reward for.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization Leader Reward for

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:54.722456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:54.722456Z digest=sha256:90384b67b98bb16dbf55a1b46c45298b5bc96404a439f8606b81c10bd60ff5c6

Observation 8abf02ee-ced8-4cd6-b11e-0bc74bdc121a · outbound

This paper cites 2025 , eprint =.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization 2025 , eprint =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:54.767551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:54.767551Z digest=sha256:0d20ae17c995391055aa56dee210c90520d7417b8ee82c20495ed0dd2ef93029

Observation 006768af-557e-43b9-92c8-62ce1a4519a5 · outbound

This paper cites On Advantage Estimates for Max@K Policy Gradients.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization On Advantage Estimates for Max@K Policy Gradients

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:54.921367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:54.921367Z digest=sha256:c423fd4ad4c743173fa1626784ae6b1c8420815f3c1cf29fa92d4c50a6534611

Observation f65710a9-e0cd-46d2-9e9c-bdffc8fa8711 · outbound

This paper cites OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:55.058885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:55.058885Z digest=sha256:d44859ee3a5bada8c6396cb8b05adf9bdaff0389c70bdd849dfcbfb414d1b42c

Pith citing papers

No inbound Pith citation observations are available.