Pith. sign in

Paper Citation Record · LEDGER

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient

As of 17 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2509.02737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02737 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:41:31.820379Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d35bf1a-e57d-4d8f-8a1b-5dc505c3d3a9 · outbound

This paper cites Contrastive policy gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Contrastive policy gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:41:32.224570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:41:31.753623Z digest=sha256:f1fd8047852b148750361c44c0f06b1b7f75d8b49de437ff4c626366bea3d3b4

Observation f8e58389-6fe2-4ffb-9a72-4a59c7d60f94 · outbound

This paper cites Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central Path.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central Path

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.766286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.766286Z digest=sha256:bd1de122c10608976f0fe0a1483fe856064f1f3d021e4c7e5982cc167c581ba4

Observation 9d7b9e5a-49bb-47da-be6b-bf81ae7030e3 · outbound

This paper cites All experimental results are average over20 random seeds and we choose the best from three learning rates.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient All experimental results are average over20 random seeds and we choose the best from three learning rates

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:41:32.117027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:41:31.820379Z digest=sha256:f958ea094f629023989030518d67aaaaf0a42b46769dc0d34a8b577cea6a67c1

Observation bedbb2b9-5aee-49cc-b0bc-574766382c09 · outbound

This paper cites Inducing Neural Collapse to a Fixed Hierarchy-Aware Frame for Reducing Mistake Severity.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Inducing Neural Collapse to a Fixed Hierarchy-Aware Frame for Reducing Mistake Severity

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:41:31.973875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:41:31.773745Z digest=sha256:f637cacedaf04ab3d232f95046f6ce8f8949f5f83e2473f1faffa09dad62e30d

Observation 45e855d4-ff40-48cc-9e16-7eb3a5b7474d · outbound

This paper cites Continuous control with deep reinforcement learning.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Continuous control with deep reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.777647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.777647Z digest=sha256:0589253ae08ce13daf538e1c44cbb97f02b863fb7280604023adc91b986c77a0

Observation 9cb96999-d049-4458-a9c1-d8ad2bd83976 · outbound

This paper cites Neural Collapse with Cross-Entropy Loss.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Neural Collapse with Cross-Entropy Loss

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.781429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.781429Z digest=sha256:a1d68d55cba6379fe1a884962f338d755a00be150d549c7014e265fca648e4e1

Observation 7667d8f6-d6e8-4567-af9b-59a8740588ff · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Asynchronous methods for deep reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.789604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.789604Z digest=sha256:8176aeacb51ccf66fc02fdba4dde734d77d007c224650958266b6a639472f1a8

Observation a05b1657-feca-4682-b4ec-62170104d518 · outbound

This paper cites Instruction Tuning with GPT-4.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Instruction Tuning with GPT-4

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.792933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.792933Z digest=sha256:58febfe1e95fd43d53228498937a2e7a6f70d89d72e131c83e3de3b8c5bce5d4

Observation 53350eeb-dbf1-454d-a348-85532088681d · outbound

This paper cites Linguistic Collapse: Neural Collapse in (Large) Language Models.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Linguistic Collapse: Neural Collapse in (Large) Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.802573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.802573Z digest=sha256:650266ff22ccea075e87566e9a9d1c0e079f6fb69d6dadec4a315e80aa226a07

Observation 420bd0d2-310f-46e7-b139-d779ab608147 · outbound

This paper cites Huaqing Xiong, Tengyu Xu, Lin Zhao, Yingbin Liang, and Wei Zhang.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Huaqing Xiong, Tengyu Xu, Lin Zhao, Yingbin Liang, and Wei Zhang

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.807011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.807011Z digest=sha256:4fd5e7d5938850b44e437801b579a52dc53afd4edda9db21d3ff287531dae998

Observation aa3c2ad1-eed3-4b28-a8e7-0e04cf3f2c4d · outbound

This paper cites Convergence and iteration complexity of policy gradient method for infinite-horizon reinforcement learning.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Convergence and iteration complexity of policy gradient method for infinite-horizon reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:41:32.167884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:41:31.811116Z digest=sha256:17748d9d69fb19eddb9b34a2ece21f85b7a4fae48be4612bfe04898915b916c3

Observation 831869e0-15e0-4b83-a0d7-865c67f2d115 · outbound

This paper cites an unresolved cited work.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:41:32.143636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:41:31.815311Z digest=sha256:62211b36b522debc1e6b63b3fb237140fb126e1bf681e26ca11eea7aec23ad10

Observation 02e43c6a-ac4c-4afe-82f7-5a13f2499e98 · outbound

This paper cites On the Global Convergence of Natural Actor-Critic with Two-layer Neural Network Parametrization.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient On the Global Convergence of Natural Actor-Critic with Two-layer Neural Network Parametrization

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.761656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.761656Z digest=sha256:57cc37564149ea929f3ac68d768d9ac1d68e6d38c5c9c3f013bc79613cd5bf6d

Observation 9966f96a-fc25-44d9-8b33-5bf80b3e4f46 · outbound

This paper cites Learning classifiers on positive and unlabeled data with policy gradient.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Learning classifiers on positive and unlabeled data with policy gradient

Reference 1999

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:41:32.190806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:41:31.770274Z digest=sha256:15b3264e8f5ad3199edf5be3a16a10dcb92ff365de3874a46535b53e6e534d75

Observation 05fc2478-8a1a-49cc-8652-54b738a9b238 · outbound

This paper cites Global Optimality Guarantees For Policy Gradient Methods.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Global Optimality Guarantees For Policy Gradient Methods

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.745985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.745985Z digest=sha256:eba2b60fc17a5e29947a0d15ab0359a87fb1066b87cca4f78af7ae7150f67b36

Observation 1e0d5b91-927e-4028-9c8c-657d004fab70 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Proximal Policy Optimization Algorithms

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.796790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.796790Z digest=sha256:cf27c0f1a4eb3b0f8705d9ddd9034c95ee8af51de30fce7e732d489c3745572d

Observation cfc20b73-a249-4daf-a3be-f04b21089ac2 · outbound

This paper cites TD Convergence: An Optimization Perspective.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient TD Convergence: An Optimization Perspective

Reference 2017

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:41:32.102065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:41:31.738118Z digest=sha256:ca4bf7af06fcf5cc9049de27e4f3664c2f62f1dbc8693babd48c000883a00b4c

Observation 0c029449-dc7f-41c3-9c3b-709a55ba1910 · outbound

This paper cites Understanding of a convolutional neural network.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Understanding of a convolutional neural network

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:41:32.240361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:41:31.733739Z digest=sha256:fb21fa383574810a85a46d829a0a196a082cd46c165b05ef108568189316caf0

Observation 1f8e7448-b3dd-43dd-86ff-fb971f93b4d7 · outbound

This paper cites Neural collapse with unconstrained features.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Neural collapse with unconstrained features

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.785414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.785414Z digest=sha256:6271359bc20e064beb97e30db6d967f1aa1cf35db1cf38f0934ff37d9f5473b5

Observation b845a436-ed3a-4a49-ba40-bab28aabbdeb · outbound

This paper cites OpenAI Gym.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient OpenAI Gym

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.749577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.749577Z digest=sha256:16bd019dc0e4b69f90739ace23844d714566f03fe1e737818d96dd638d2f5313

Observation cdeee453-98d9-47a6-a565-64e2b28942a6 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient A General Language Assistant as a Laboratory for Alignment

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:31.742107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:31.742107Z digest=sha256:5e7b9bf1c0c4c422a86964ac5f07e061527af5a9855014c5c2381a195e56658a

Observation aed67a22-9d53-41f1-95b1-e161316c5c3a · outbound

This paper cites doi: 10.18653/v1/2024.emnlp-main.1190.

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient doi: 10.18653/v1/2024.emnlp-main.1190

Reference 2024

Resolution
verified exact
doi, observed 2026-08-15T16:41:31.876373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:41:31.757834Z digest=sha256:f7b2437e5d73e4b6a691dfeb7c6b9b7503e7cbb7c9a4758a262f6228cacbe583

Pith citing papers

No inbound Pith citation observations are available.