Pith. sign in

Paper Citation Record · LEDGER

Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:1810.02525.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1810.02525 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.246831Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T14:27:03.775857Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e860f8dc-57df-468f-a7eb-c6b993051fe4 · inbound

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition cites this paper.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.246831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.246831Z digest=sha256:4a4e2e79d3ede591b547994676c50c142d9b62b672c3f7b97c118c1dabe23c32

Observation 462eb6ae-57ea-4a68-a097-ca767aef6757 · inbound

Segmenting Action-Value Functions Over Time-Scales in SARSA via TD($\Delta$) cites this paper.

Segmenting Action-Value Functions Over Time-Scales in SARSA via TD($\Delta$) Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:27.468115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:27.468115Z digest=sha256:ab34f3a07f8e9c87048685e09285f1d1d4b139e03cc22d6bc334821bee110842

Observation e6609016-7c4d-4f2b-a22f-87b250450dde · inbound

Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps cites this paper.

Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T05:51:10.312034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:51:10.312034Z digest=sha256:7d3548bf4ff48e88548e2f376c0d602a2ecbb0b50fb2c4042f588947e014a4c2

Observation 99c6a4c4-ac1e-404c-9774-4c117adfa27d · inbound

Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers cites this paper.

Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:20.556615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:20.556615Z digest=sha256:32e707a5be8f915840f522a6341f270460f0e086defc743a810813bdc6879808

Observation bdc9fabe-22f1-449a-a24e-77e7fc166b6c · inbound

Reweighting Adversarial Networks for Unbinned Unfolding cites this paper.

Reweighting Adversarial Networks for Unbinned Unfolding Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-07-02T14:27:03.777137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T00:28:42.520230Z digest=sha256:d821044a7a283b6313eaf1e0237264bba507929fbb2b11725266cd69c398bd9a