Pith. sign in

Paper Citation Record · LEDGER

Efficient Hypergradient Descent for Inverse Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2608.11052.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11052 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:26:02.942258Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact5
  • verified fuzzy5
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28a544c9-ce1c-46a3-9b4c-2e9a0bb31ec3 · outbound

This paper cites Therefore, αDKL(epπθ ∥epϕ) =E τ∼epπθ " ∞X t=1 (αlogπ θ(at |s t)−r ϕ(st, at)) # +αlogZ ϕ.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Therefore, αDKL(epπθ ∥epϕ) =E τ∼epπθ " ∞X t=1 (αlogπ θ(at |s t)−r ϕ(st, at)) # +αlogZ ϕ

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:26:03.543094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:26:02.934443Z digest=sha256:1c47f4b2f25a16b84c23b2f21a10faf02503f4b444f1e5b6b8a0cf7d0eb23eee

Observation 8d26906d-e9d7-4351-9676-5ae542388b98 · outbound

This paper cites ∞X t=1 γt−1 logπ θ(at |s t) # =−E τ∼p expert.

Efficient Hypergradient Descent for Inverse Reinforcement Learning ∞X t=1 γt−1 logπ θ(at |s t) # =−E τ∼p expert

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:26:03.530082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:26:02.938246Z digest=sha256:bbdc3d576c4da8f7844caf523d1cf2bad2bf0868828b994e9d9a987b800c4553

Observation 99918331-970a-4806-b099-f850cac23993 · outbound

This paper cites Natural hypergradient de- scent: Algorithm design, convergence analysis, and parallel implementation.arXiv preprint arXiv:2602.10905,.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Natural hypergradient de- scent: Algorithm design, convergence analysis, and parallel implementation.arXiv preprint arXiv:2602.10905,

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-12T11:26:03.319693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:26:02.906229Z digest=sha256:68a57242014d758eda180f2f230097a161b011843c5e1e1e3654d1d46e3d87e5

Observation a6300e82-9013-4f85-b8f3-27690ad68883 · outbound

This paper cites Natural Policy Gradients In Reinforcement Learning Explained.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Natural Policy Gradients In Reinforcement Learning Explained

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:26:03.105051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:26:02.918211Z digest=sha256:4b9131c16375e60d04c6ca3ffe87d34c6b9252ebb35851f8afbd6b09c44c77a7

Observation a2d1dcd2-a8ef-4ae1-8586-5ac449da702a · outbound

This paper cites ∂g ∂θ θ⋆(ϕ),ϕ #−1 ∂g ∂ϕ θ⋆(ϕ),ϕ . Since ∂g ∂θ = ∂2Linner ∂θ 2 , ∂g ∂ϕ = ∂2Linner ∂θ∂ϕ , we get dθ⋆ dϕ ϕ =−.

Efficient Hypergradient Descent for Inverse Reinforcement Learning ∂g ∂θ θ⋆(ϕ),ϕ #−1 ∂g ∂ϕ θ⋆(ϕ),ϕ . Since ∂g ∂θ = ∂2Linner ∂θ 2 , ∂g ∂ϕ = ∂2Linner ∂θ∂ϕ , we get dθ⋆ dϕ ϕ =−

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:26:03.554667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:26:02.930040Z digest=sha256:6101bf40a3c3de63824bcb06e0f47f0556522b06826ae9f838314f4bbb7bf2ab

Observation 835c4e51-022e-47ac-b652-c035985b099a · outbound

This paper cites ∞X t=1 gt(τ)g t(τ) ⊤ # . Finally, applying Lemma 4.1 componentwise gives Eτ∼epπθ.

Efficient Hypergradient Descent for Inverse Reinforcement Learning ∞X t=1 gt(τ)g t(τ) ⊤ # . Finally, applying Lemma 4.1 componentwise gives Eτ∼epπθ

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:26:03.516484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:26:02.942258Z digest=sha256:1016532e288d847d7f961f52de44ef53a49aaa0592b083db047d7701754d19d8

Observation 77b36f46-a84c-4f6c-8c3f-4d103b32b54e · outbound

This paper cites Souradip Chakraborty, Amrit Bedi, Alec Koppel, Huazheng Wang, Dinesh Manocha, Mengdi Wang, and Furong Huang.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Souradip Chakraborty, Amrit Bedi, Alec Koppel, Huazheng Wang, Dinesh Manocha, Mengdi Wang, and Furong Huang

Reference 2014

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T11:26:03.505054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:26:02.885600Z digest=sha256:a4db276a5e68a83e50f70438001836264ac196809979c1b8581c88c663ae04d0

Observation c70af340-f658-41db-9ef7-2cf63e31a3b2 · outbound

This paper cites Rank-1 ap- proximation of inverse fisher for natural policy gradients in deep reinforcement learning.arXiv preprint arXiv:2601.18626,.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Rank-1 ap- proximation of inverse fisher for natural policy gradients in deep reinforcement learning.arXiv preprint arXiv:2601.18626,

Reference 2015

Resolution
verified exact
raw_fallback, observed 2026-08-12T11:26:03.397415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:26:02.902323Z digest=sha256:f7cca8f98998551a9e00a27e9eaa4165d76f12ebd5fe3ad0ed78cac1d9911799

Observation d5f7b85f-c790-4f6b-be2c-172610441a12 · outbound

This paper cites Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T11:26:02.922372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:26:02.922372Z digest=sha256:29182e779c59645d1d464a1da6425eccbbe257f1b4f7eccc0204e6308eb16184

Observation 9f09e1a8-f956-4cb0-b6ec-e6f769f93771 · outbound

This paper cites Frequent Directions : Simple and Deterministic Matrix Sketching.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Frequent Directions : Simple and Deterministic Matrix Sketching

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T11:26:02.897810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:26:02.897810Z digest=sha256:04955fbfda774272c34737fa48a374aac288abe92a81c72ad5df0c219f89366e

Observation 53da9989-9480-4f50-bfb5-83c2cef374ac · outbound

This paper cites Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T11:26:02.909932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:26:02.909932Z digest=sha256:18badf6cdc8574e149bbb57847ec41d2ef720901408933ca9b0d35d60dcdd626

Observation 444421fd-0641-4f09-ab05-109e196003cc · outbound

This paper cites Explaining and Preventing Alignment Collapse in Iterative RLHF.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Explaining and Preventing Alignment Collapse in Iterative RLHF

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:26:03.437865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:26:02.889788Z digest=sha256:5761aec01cadcd954893ababfc1a66abb18e4a0eba1199908f1d4a4484620a90

Observation 239f0717-51cb-4e91-9462-21904e9e9e82 · outbound

This paper cites Then the gradient of the induced outer objective eLouter(ϕ) :=L outer(θ⋆(ϕ)) is given by ∇ϕ eLouter ϕ =− ∂2Linner ∂ϕ∂θ θ⋆(ϕ),ϕ " ∂2Linner ∂θ 2 θ⋆(ϕ),ϕ #−1 ∇θLouter|θ⋆(ϕ).

Efficient Hypergradient Descent for Inverse Reinforcement Learning Then the gradient of the induced outer objective eLouter(ϕ) :=L outer(θ⋆(ϕ)) is given by ∇ϕ eLouter ϕ =− ∂2Linner ∂ϕ∂θ θ⋆(ϕ),ϕ " ∂2Linner ∂θ 2 θ⋆(ϕ),ϕ #−1 ∇θLouter|θ⋆(ϕ)

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:26:03.568289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:26:02.926453Z digest=sha256:05426f8d5e65c6f2c79ea369e8725cbdbe9e2c318f60350f2d856677912fba74

Observation e58c82b5-d24e-49e6-9a5d-35eefb8e4fc7 · outbound

This paper cites Scalable linucb: Low-rank design matrix updates for recommenders with large action spaces.arXiv preprint arXiv:2510.19349,.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Scalable linucb: Low-rank design matrix updates for recommenders with large action spaces.arXiv preprint arXiv:2510.19349,

Reference 2025

Resolution
verified exact
raw_fallback, observed 2026-08-12T11:26:03.214880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:26:02.914384Z digest=sha256:e7c4e877b7752cbf358b6f7dd2c1796a1360e62f97d241b7deff9489a51135f1

Observation b77cffe1-f9a4-4743-bc39-6893c3c913f5 · outbound

This paper cites Approximation Methods for Bilevel Programming.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Approximation Methods for Bilevel Programming

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-12T11:26:02.893846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:26:02.893846Z digest=sha256:74b6bf6ced1d13de85bc1001c2fc743b7c9ea98bc328b6fd1a12fe611f79262a

Pith citing papers

No inbound Pith citation observations are available.