Pith. sign in

Paper Citation Record · LEDGER

The Geometry of Nonlinear Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2509.01432.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01432 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:37:43.138992Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T09:24:21.047027Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact4
  • verified fuzzy11
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7017f336-dd9e-45ff-8230-88b646aeb530 · outbound

This paper cites Maximum a Posteriori Policy Optimisation.

The Geometry of Nonlinear Reinforcement Learning Maximum a Posteriori Policy Optimisation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.058718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.058718Z digest=sha256:b0afcd39c2e0687628b12cb7703b27496fc14b87ebbb11c0504a92f4354896d4

Observation 4a833dc7-94fc-4146-b840-4d5c48f00440 · outbound

This paper cites an unresolved cited work.

The Geometry of Nonlinear Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:37:45.041818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:42.672881Z digest=sha256:e3969ea095da4ebfbdff146937167449f7ceaf4ea4833ffbb02e2967de36d726

Observation 00258e82-80e4-45ea-8e95-57830cc71bbb · outbound

This paper cites Embedding Safety into RL: A New Take on Trust Region Methods.

The Geometry of Nonlinear Reinforcement Learning Embedding Safety into RL: A New Take on Trust Region Methods

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:37:43.877379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:41.538793Z digest=sha256:45bd938a492f5271084b0fa927faa6e75aadbc32b8ca185973e44168a26db9d0

Observation 7a1ef6a6-4859-4cd6-84de-2a19d6439ba4 · outbound

This paper cites Central Path Proximal Policy Optimization.

The Geometry of Nonlinear Reinforcement Learning Central Path Proximal Policy Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.630698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.630698Z digest=sha256:6ed7aee768461b465917c2aa023331f277636f827a19f7a951fdac53a55211d6

Observation da82822d-3ba2-44a5-917e-2123bfae2ef0 · outbound

This paper cites Challenging Common Assumptions in Convex Reinforcement Learning.

The Geometry of Nonlinear Reinforcement Learning Challenging Common Assumptions in Convex Reinforcement Learning

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T12:37:43.566193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:41.781429Z digest=sha256:8c09e21ce60d910c7ec7a2b8f87366b373c96f93ba0b7d795a7fcf7f802cbb09

Observation 9b485471-8f66-456a-bb8a-ff4990223dc6 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

The Geometry of Nonlinear Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.925676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.925676Z digest=sha256:a76e634cfe9bdd65b80479e566bf8b42d71e1650d544118ce7a73cc291a4e3e3

Observation cd88364a-c138-46fb-9d54-6033d95ee3aa · outbound

This paper cites ∇θ log π(a′|s′) X s,a Mπ(s, a|s′, a′)rπ(s, a) # (30) = (1 − γ)Es′,a′∼ω.

The Geometry of Nonlinear Reinforcement Learning ∇θ log π(a′|s′) X s,a Mπ(s, a|s′, a′)rπ(s, a) # (30) = (1 − γ)Es′,a′∼ω

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.051249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:42.618804Z digest=sha256:ee01125f2c21f478d1e7bf0bc6e13684b0d6d9894f5a6a79f796377c2a30b0be

Observation aeb9da49-ed6c-420a-af72-5dde93a64e1d · outbound

This paper cites Mirror descent policy optimization.

The Geometry of Nonlinear Reinforcement Learning Mirror descent policy optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.091767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:42.255450Z digest=sha256:64b8d82be0047ac5e125240e6d266555b948550441a2c6a2574a8607f2d2fa4f

Observation 0b3b05ba-0956-46ea-af58-317195781842 · outbound

This paper cites Policy Mirror Descent for Regularized Reinforcement Learning: A Generalized Framework with Linear Convergence.

The Geometry of Nonlinear Reinforcement Learning Policy Mirror Descent for Regularized Reinforcement Learning: A Generalized Framework with Linear Convergence

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:42.367138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:42.367138Z digest=sha256:581d48b097f1db785e2043b9d5ae5d4565aea07c55c8753a1f25b20e3b84f1ad

Observation a9aeef1b-a7fe-40fa-8600-2d5e8c36dc54 · outbound

This paper cites Interestingly, it also appears in the differential of the map between policy and state-action spaces.

The Geometry of Nonlinear Reinforcement Learning Interestingly, it also appears in the differential of the map between policy and state-action spaces

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.061885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:42.537762Z digest=sha256:327a9a01da61bf50f30b5e00572ad6ec3053ea49c8b75346ed38ca7d17d0fb84

Observation c10c0bf0-b641-4ffe-83d7-6b67e25d83fe · outbound

This paper cites The reward isrπ(s, a) = P i zi[log pπ(i|s) − log zi].

The Geometry of Nonlinear Reinforcement Learning The reward isrπ(s, a) = P i zi[log pπ(i|s) − log zi]

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.031379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:42.757521Z digest=sha256:2ce4f12c2af9c9d51b61450ebf404dbf4b3fce9c1abbdc82b58f18d881fc5147

Observation bf70a841-3077-48e5-9eeb-61078b2c0de7 · outbound

This paper cites an unresolved cited work.

The Geometry of Nonlinear Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:37:45.020228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:42.843902Z digest=sha256:afe5c7e7d3883bd189c90b433e5bb1ef46f886e0e7f413b8ad54552afdb564e5

Observation e2bf02d1-78b0-4fe9-b301-dc9c0ef871f9 · outbound

This paper cites an unresolved cited work.

The Geometry of Nonlinear Reinforcement Learning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:37:45.010369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:42.902934Z digest=sha256:ab7e7fe334b499a249c0062eac753850d71978023a8a6cd72c1f6dfe1edcb8c6

Observation e82d5808-1bfc-4614-9f81-0debb65398b8 · outbound

This paper cites This results in anintractable policy divergence with Hessian 16 GTML 2025 HC(θ) = Es∼ωπ F (θ) + X i βiϕ′′(bi − Vci (θ))∇2 θVci (θ) θ=θk.

The Geometry of Nonlinear Reinforcement Learning This results in anintractable policy divergence with Hessian 16 GTML 2025 HC(θ) = Es∼ωπ F (θ) + X i βiϕ′′(bi − Vci (θ))∇2 θVci (θ) θ=θk

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:44.999360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:42.993772Z digest=sha256:c7715a47022c204b71333280da197e8f44a70360a8737c4306547a0cd8bb9cb4

Observation c0433396-15d7-4f67-b278-5d0b262b1b41 · outbound

This paper cites 17 GTML 2025 When this standard geometry is restricted to the manifoldΩ, i.e.

The Geometry of Nonlinear Reinforcement Learning 17 GTML 2025 When this standard geometry is restricted to the manifoldΩ, i.e

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:44.988359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:43.077827Z digest=sha256:90931ad18e9bc1dd562df5c6c0087f11a8204b5e490a62cb38bea1b623db866d

Observation 69823b71-98d0-4259-bac2-dcf6352681f2 · outbound

This paper cites Definition C.1(Successor Representation).

The Geometry of Nonlinear Reinforcement Learning Definition C.1(Successor Representation)

Reference 1993

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.071461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:42.484428Z digest=sha256:29b360e48a80c9595e4bfd70f4bfe6ba73696e8d0d2d515c2093493a8664710f

Observation d3426bb7-af16-429d-8bb2-cc5bfcfd3252 · outbound

This paper cites Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints.

The Geometry of Nonlinear Reinforcement Learning Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints

Reference 1994

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:37:44.142127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:41.477513Z digest=sha256:b2fbd5c40ccd80934908c35166f507d9f821d3a34649291abf082bf4a5951fc9

Observation b1fc9aea-a97a-4876-811b-c0960063ca8f · outbound

This paper cites Diversity is All You Need: Learning Skills without a Reward Function.

The Geometry of Nonlinear Reinforcement Learning Diversity is All You Need: Learning Skills without a Reward Function

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.228201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.228201Z digest=sha256:db518edf13b59eb54de531500d070dfcbf259dbac6a9f50fb2bd8bb5a4468b3b

Observation 14ad5aa8-9218-485c-9508-bfe56b4e9f54 · outbound

This paper cites Motivation for a General Hessian FrameworkThe existence of at least two distinct, natural geometries on the same occupancy manifold is a key insight.

The Geometry of Nonlinear Reinforcement Learning Motivation for a General Hessian FrameworkThe existence of at least two distinct, natural geometries on the same occupancy manifold is a key insight

Reference 2001

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:44.977412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:43.138992Z digest=sha256:aa8ba366d38804292931f978d8c40fa59d1160026f7969df7bb7438e7db8e9d3

Observation 785880e8-847f-4a4d-b36e-ed8247e1eb0f · outbound

This paper cites Discovering Diverse Nearly Optimal Policies with Successor Features.

The Geometry of Nonlinear Reinforcement Learning Discovering Diverse Nearly Optimal Policies with Successor Features

Reference 2006

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:37:43.327076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:42.318169Z digest=sha256:08d195cad79637fa366c59b5cc4770fa8715a44c4ad9dba85c944050d9c552c7

Observation 577f2040-8e0c-4fc8-84f9-032c966a58e6 · outbound

This paper cites Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning.

The Geometry of Nonlinear Reinforcement Learning Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning

Reference 2007

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.100499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:42.172013Z digest=sha256:92db2988e1ca2bb784729040ff0bad2fd9c80ed0f14622446e877f313ac6d8b2

Observation e2546c89-beba-4e33-9484-133f4faf63b9 · outbound

This paper cites ∞X t=0 γtf (st, at) # = Es,a∼dµ π [f (s, a)] (5) 8 GTML 2025 Proof. (1 − γ)Eτ ∼π,µ.

The Geometry of Nonlinear Reinforcement Learning ∞X t=0 γtf (st, at) # = Es,a∼dµ π [f (s, a)] (5) 8 GTML 2025 Proof. (1 − γ)Eτ ∼π,µ

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.081576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:42.447186Z digest=sha256:17ef6e0a9398595e5551bb5c01dae7a10a2a509725080f7c62e7ddd6d5877b15

Observation e23415ff-ead5-4ac1-be5c-778fe12aa0a4 · outbound

This paper cites an unresolved cited work.

The Geometry of Nonlinear Reinforcement Learning Unresolved cited work

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:42.043518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:42.043518Z digest=sha256:3e4753d7a5e698ec66bbe062f7cbf488004af851b772c3736543c348572d4808

Observation 40b258de-e9bb-4761-8430-f7f02e90b5ad · outbound

This paper cites Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features Critics.

The Geometry of Nonlinear Reinforcement Learning Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features Critics

Reference 2016

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:37:44.503773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:41.368134Z digest=sha256:747c38fa7e34a0be63de338d7bd7a6003d8b0c53d7c7d7a7c5de8abe62b41507

Observation b5e69283-05ea-44f4-a252-07ef7c669a48 · outbound

This paper cites Variational Intrinsic Control.

The Geometry of Nonlinear Reinforcement Learning Variational Intrinsic Control

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.314647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.314647Z digest=sha256:fd9359f8b6186c68460a7eb25a08e877274d3d48337b4a3efd38f6df9474b79b

Observation 6b51656a-4d66-476c-bf2c-b372f13c1876 · outbound

This paper cites cc/paper_files/paper/2019/file/873be0705c80679f2c71fbf4d872df59-Paper.pdf.

The Geometry of Nonlinear Reinforcement Learning cc/paper_files/paper/2019/file/873be0705c80679f2c71fbf4d872df59-Paper.pdf

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.109888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:41.701622Z digest=sha256:77ab0611962c4036771a47915a4ed4f4f0b08ce013e520bb79978da0a464135f

Observation fe7a6825-0ebf-4860-99c5-caa4cffe3680 · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

The Geometry of Nonlinear Reinforcement Learning A unified view of entropy-regularized Markov decision processes

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.845185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.845185Z digest=sha256:b7ffc7bba53afbb1d3809db196d3829e611e07c7eda63cbbac2109897c67b1ec

Observation 8ba514d3-52d7-4f4c-b52e-f0152f3eb758 · outbound

This paper cites Concave Utility Reinforcement Learning: the Mean-Field Game Viewpoint.

The Geometry of Nonlinear Reinforcement Learning Concave Utility Reinforcement Learning: the Mean-Field Game Viewpoint

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T12:37:44.810808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:37:41.266986Z digest=sha256:ab464d1ac0061f73583b5aacb1f62a6f4919dd8d1da5762d225f4a39d7e5d7c4

Observation 0003b65f-ac2a-4d90-bd5a-7a38a7dc51b4 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

The Geometry of Nonlinear Reinforcement Learning Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.141219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.141219Z digest=sha256:0e5e229a59c02690c28e4002c574734e2628446504f6501c3cc7f88a137d678a

Observation cbecf1e9-e25b-42d0-a701-e49ea4ca7953 · outbound

This paper cites Fast Task Inference with Variational Intrinsic Successor Features.

The Geometry of Nonlinear Reinforcement Learning Fast Task Inference with Variational Intrinsic Successor Features

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.422027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.422027Z digest=sha256:cf1a6820cd3a30721c94342dbe726307e62d3895cde8c31d7aeeb6c48c8ddee9

Observation b588e5e9-9603-43e3-81cf-0272262e8a24 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Geometry of Nonlinear Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:42.090426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:42.090426Z digest=sha256:bbb0b0d1d2603265a604562329c89cab6eb3910f0303f22f77f21039c8c1fd21

Pith citing papers

Observation ca580af9-5ac7-4c3c-93c9-5f55179017f4 · inbound

Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs cites this paper.

Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs The Geometry of Nonlinear Reinforcement Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:24:31.577549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T09:24:21.047027Z digest=sha256:62e24605d5449e0de0030f0dbcda8cc30aad45df82f8d06520d142ab92b83482