Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:24:51.822886Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2504.20887.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:24:51.822886Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T05:00:42.106290Z
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7fbd467b-7cc3-4d36-b8c7-171c03715c7e · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c93b054f-e52b-49f7-b71d-be1852cc8eb5 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Entropic value-at-risk: A new coherent risk measure
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1b13c10c-5166-4dd1-a500-14b8c8a98d5c · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Coherent measures of risk
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 534dfed5-9f78-4013-8612-7498e2625df5 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation and Ott, J
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3d411935-5ef4-47bb-abd3-b21ee96ae61b · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation G., Dabney, W., and Munos, R
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 36cd2613-7316-4256-8d87-71d7f8620e74 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation OpenAI Gym
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001b0b41-e592-42ee-9920-a4f4f0a5829c · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Risk-constrained reinforcement learning with percentile risk criteria
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eb44e88-ba54-4cf0-9341-900c0b4bd962 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Implicit quantile networks for distributional reinforcement learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ef75e414-19e6-4d21-b557-dbb2109a11e8 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation A., Krass, D., and Ross, K
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 096752ef-d711-45f0-89e0-7949217b2328 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Efficient risk-averse reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5f691eff-8ab2-4dd1-a104-588dbecf0747 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Contextual Markov Decision Processes
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1f27cc6-d9e4-4585-949e-87187563aecf · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Being optimistic to be conservative: Quickly learning a cvar policy
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 225754bb-9d06-4566-b10f-98070d457f42 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation and Min, S
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d9029d2f-6c68-4c0e-890d-a9b4825b0716 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation and Ghavamzadeh, M
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4909ab9b-f809-4198-8c88-7533f4cd14eb · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dd044ddb-b837-4444-ba6f-46d27b4f765d · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation An alternative to variance: Gini deviation for risk-averse policy gradient
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a775af44-f3f0-47f6-9feb-c374b7beb1a9 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation A simple mixture policy parameterization for improving sample efficiency of cvar optimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 599c907c-27fd-44bd-8c44-e964fe44ad02 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9712860f-e0bb-4165-aa1c-3e16c638f894 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Risk averse robust adversarial reinforcement learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3303bbc5-905c-44ad-8210-25fcb97369cb · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Robust adversarial reinforcement learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c5fc1ffc-929b-4b45-b5bd-8959c6f0c675 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation and Ghavamzadeh, M
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7d701ba4-646c-4c66-a2ad-ce14249f14f1 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Risk-averse bayes-adaptive reinforcement learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 04b1bb58-f80b-4f32-9a7e-4b1188b7decf · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation T., Uryasev, S., et al
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3a30c3d-c60c-4ecd-9b37-0266d053624b · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Risk-averse dynamic programming for markov decision processes
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9176156-48a1-4f6d-94e9-75c5638d2820 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Td algorithm for the variance of return and mean-variance reinforcement learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f100aa98-1d27-47bf-83e6-ece8e05748a0 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Learning risk-aware quadrupedal locomotion using distributional reinforcement learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a33fdc01-e9cc-4a60-bb13-25689409e5ee · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Proximal policy optimization algorithms, 2017
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1eeef5a-2ce1-4030-8aaf-708b437d5703 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f2eb2229-f88a-4a4d-b4b7-db0e9a63932a · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Responsive safety in reinforcement learning by pid lagrangian methods
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fea598fa-a8d1-4774-be37-2c3cdb39491e · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d484be4-8971-4cfa-b5f6-ef2963eca8ea · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Optimizing the cvar via sampling
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5d4c26ee-aaef-499a-9115-6b59d715d2b5 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Worst Cases Policy Gradients
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b25a8f75-294b-41c0-9312-e85ac17e675b · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Gymnasium: A Standard Interface for Reinforcement Learning Environments
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf327413-a44a-4ee4-9938-e4befbed8c43 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9c4a3620-2e61-44a5-ab2c-4a4332aa8ba2 · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Risk-aware reward shaping of reinforcement learning agents for autonomous driving
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bbe21a40-01b8-4e58-bba6-008f881fadaa · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Off-Policy Primal-Dual Safe Reinforcement Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 250cdd10-2b4a-4d87-87fd-d34c3fdfe57e · outbound
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation D., Tindemans, S
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 729129bc-3353-48b3-833d-1f61bc861859 · inbound
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.