Pith. sign in

Paper Citation Record · LEDGER

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation

As of 21 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2504.20887.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.20887 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:24:51.822886Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T05:00:42.106290Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7fbd467b-7cc3-4d36-b8c7-171c03715c7e · outbound

This paper cites write newline.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.684717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.684717Z digest=sha256:57b931b1d1577afbef56e916e48eb83b2189f8593ac6d1453d48108eee9817ad

Observation c93b054f-e52b-49f7-b71d-be1852cc8eb5 · outbound

This paper cites Entropic value-at-risk: A new coherent risk measure.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Entropic value-at-risk: A new coherent risk measure

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.256237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.690040Z digest=sha256:0925b70999b01083aa84ac89016656b71b81eaa924ed03e31b189fd09b9f5493

Observation 1b13c10c-5166-4dd1-a500-14b8c8a98d5c · outbound

This paper cites Coherent measures of risk.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Coherent measures of risk

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.245978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.693799Z digest=sha256:8e9645344c6c59c042bb7dbc6cb8344b42034c603b396d23286197ad6a85dfe1

Observation 534dfed5-9f78-4013-8612-7498e2625df5 · outbound

This paper cites and Ott, J.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation and Ott, J

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.235758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.697568Z digest=sha256:719c9f340f0b07f4de86a70e97fd02a68af163823ce5b9f774408d32a39860ab

Observation 3d411935-5ef4-47bb-abd3-b21ee96ae61b · outbound

This paper cites G., Dabney, W., and Munos, R.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation G., Dabney, W., and Munos, R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.224895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.701811Z digest=sha256:ed481eee7392694d04169bdd7578c49dbfe79f2bfb3966fc05476f82a2d04f95

Observation 36cd2613-7316-4256-8d87-71d7f8620e74 · outbound

This paper cites OpenAI Gym.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation OpenAI Gym

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.705527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.705527Z digest=sha256:a0aa8102eb64c9f8dda1e29a5709d87a25c2b2b45d982be9f958742aafab47a7

Observation 001b0b41-e592-42ee-9920-a4f4f0a5829c · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Risk-constrained reinforcement learning with percentile risk criteria

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.709484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.709484Z digest=sha256:8cc61927078a3a5ddc0005cde71f8de10b5e874fc9becd3b295288f2ca28b9f3

Observation 3eb44e88-ba54-4cf0-9341-900c0b4bd962 · outbound

This paper cites Implicit quantile networks for distributional reinforcement learning.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Implicit quantile networks for distributional reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.208327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.713399Z digest=sha256:c882b4ce06bac2d08167d521bd2de56742eebf7fbea8d7a966f90ee9ed3ded43

Observation ef75e414-19e6-4d21-b557-dbb2109a11e8 · outbound

This paper cites A., Krass, D., and Ross, K.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation A., Krass, D., and Ross, K

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.198590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.716837Z digest=sha256:de4c557b5c64834a5c11f6df1490fe104358f1676b51bde731e1e3f19c430684

Observation 096752ef-d711-45f0-89e0-7949217b2328 · outbound

This paper cites Efficient risk-averse reinforcement learning.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Efficient risk-averse reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.187298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.720123Z digest=sha256:d7c4af22f5f24ef69d45450394357e048a5cdb74dab75cb306068791a6381117

Observation 5f691eff-8ab2-4dd1-a104-588dbecf0747 · outbound

This paper cites Contextual Markov Decision Processes.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Contextual Markov Decision Processes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.723782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.723782Z digest=sha256:6fbbc152775d5d8c40a5e7ee349aef1f46b8e93dd4c45b0b2cc4bea63a49c9ec

Observation c1f27cc6-d9e4-4585-949e-87187563aecf · outbound

This paper cites Being optimistic to be conservative: Quickly learning a cvar policy.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Being optimistic to be conservative: Quickly learning a cvar policy

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.728072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.728072Z digest=sha256:a9b0a177346570247e2b32874aac92daeb56f8f1c3934a5050a85aa797a1732f

Observation 225754bb-9d06-4566-b10f-98070d457f42 · outbound

This paper cites and Min, S.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation and Min, S

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.175018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.731740Z digest=sha256:e1e3bbe61c7a972b68d0bbcb591c5d716a24225e67f2b1b4333a8fecedab8c3c

Observation d9029d2f-6c68-4c0e-890d-a9b4825b0716 · outbound

This paper cites and Ghavamzadeh, M.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation and Ghavamzadeh, M

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.164267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.735247Z digest=sha256:7d10b044d0029d24812a863e8f166a68037207fd986caf1bf714dcace691b528

Observation 4909ab9b-f809-4198-8c88-7533f4cd14eb · outbound

This paper cites an unresolved cited work.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:24:52.153287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.738823Z digest=sha256:e256b14855e43d7b7ed2a91ab08cc7275ce8a4b05f83ec6e02413a310bc5cd40

Observation dd044ddb-b837-4444-ba6f-46d27b4f765d · outbound

This paper cites An alternative to variance: Gini deviation for risk-averse policy gradient.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation An alternative to variance: Gini deviation for risk-averse policy gradient

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.142878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.742553Z digest=sha256:77d6a8edfe174c77c34c42d28e4809ae6ea09e2df778f10d4376beca0a412967

Observation a775af44-f3f0-47f6-9feb-c374b7beb1a9 · outbound

This paper cites A simple mixture policy parameterization for improving sample efficiency of cvar optimization.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation A simple mixture policy parameterization for improving sample efficiency of cvar optimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.131793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.746232Z digest=sha256:a82f2541b094cc47ff7a4b629c2c0f754561c496e93317ec20f2983edbe05525

Observation 599c907c-27fd-44bd-8c44-e964fe44ad02 · outbound

This paper cites DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.750051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.750051Z digest=sha256:8f2fd8d335c631a75067774ab59c2c26a2f9c920ebb754a3c259d8085de3a5ac

Observation 9712860f-e0bb-4165-aa1c-3e16c638f894 · outbound

This paper cites Risk averse robust adversarial reinforcement learning.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Risk averse robust adversarial reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.120848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.753961Z digest=sha256:9ff0f51494d094087be1102710f41e5239a0a915b610d7b245b6266f62a20d61

Observation 3303bbc5-905c-44ad-8210-25fcb97369cb · outbound

This paper cites Robust adversarial reinforcement learning.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Robust adversarial reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.110374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.757713Z digest=sha256:41e4c619134eaad21bf86b61797f469918abea362da0c119c9d9d53f408be63a

Observation c5fc1ffc-929b-4b45-b5bd-8959c6f0c675 · outbound

This paper cites and Ghavamzadeh, M.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation and Ghavamzadeh, M

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.098851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.761373Z digest=sha256:6c4ff9dd22d84404e9eed545e01f72947961e3c7596239d8b18a855edd2ba0f2

Observation 7d701ba4-646c-4c66-a2ad-ce14249f14f1 · outbound

This paper cites Risk-averse bayes-adaptive reinforcement learning.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Risk-averse bayes-adaptive reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.087717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.764894Z digest=sha256:eb04872ff0ad66e6ffe9ce44634432ca993970cd44ec4c25a6d6101116bab1ed

Observation 04b1bb58-f80b-4f32-9a7e-4b1188b7decf · outbound

This paper cites T., Uryasev, S., et al.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation T., Uryasev, S., et al

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.768761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.768761Z digest=sha256:a5b45010936bde5ea1a634efbd4928c5d3f32317bf006ec9ecf09402a5294596

Observation a3a30c3d-c60c-4ecd-9b37-0266d053624b · outbound

This paper cites Risk-averse dynamic programming for markov decision processes.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Risk-averse dynamic programming for markov decision processes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.772344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.772344Z digest=sha256:9c8bf17b0f2e9c4527ef23bf2ab87fe8e61dcdab5096a11ef94a01c91c6a80ad

Observation d9176156-48a1-4f6d-94e9-75c5638d2820 · outbound

This paper cites Td algorithm for the variance of return and mean-variance reinforcement learning.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Td algorithm for the variance of return and mean-variance reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.064073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.776702Z digest=sha256:f1b98dd46804c6003f330514a59a3330c8285c6c9fea1ff580ab3f24d4a5fd29

Observation f100aa98-1d27-47bf-83e6-ece8e05748a0 · outbound

This paper cites Learning risk-aware quadrupedal locomotion using distributional reinforcement learning.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Learning risk-aware quadrupedal locomotion using distributional reinforcement learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.780739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.780739Z digest=sha256:423b7ec993a02794422ff493e6b9a9b24940866019db1e7b641778557929485a

Observation a33fdc01-e9cc-4a60-bb13-25689409e5ee · outbound

This paper cites Proximal policy optimization algorithms, 2017.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Proximal policy optimization algorithms, 2017

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.784208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.784208Z digest=sha256:2d002f7317bf8c58e8265be7cabd4b73159013cb778aa4447a4d00b63067ea46

Observation a1eeef5a-2ce1-4030-8aaf-708b437d5703 · outbound

This paper cites an unresolved cited work.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:24:52.046873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.787837Z digest=sha256:257fd0d6671c556bda8d3d621ba3bf64ba93f6d37dfd797022665fce27d79c8a

Observation f2eb2229-f88a-4a4d-b4b7-db0e9a63932a · outbound

This paper cites Responsive safety in reinforcement learning by pid lagrangian methods.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Responsive safety in reinforcement learning by pid lagrangian methods

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.791196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.791196Z digest=sha256:664e72027dffbe9df4cfe35b603456c354fbf41bbeda22d3c57cfdbe005ead4c

Observation fea598fa-a8d1-4774-be37-2c3cdb39491e · outbound

This paper cites an unresolved cited work.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.794615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.794615Z digest=sha256:b34273ca149dfe4bfab5e71482a73c5585421cd52d1ce4853fecf4deaddb6b72

Observation 7d484be4-8971-4cfa-b5f6-ef2963eca8ea · outbound

This paper cites Optimizing the cvar via sampling.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Optimizing the cvar via sampling

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.023700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.799693Z digest=sha256:473d82e709b315a903e9e1d5ba5d374abb6156382e1bb8086c0312d1a5b3a3a8

Observation 5d4c26ee-aaef-499a-9115-6b59d715d2b5 · outbound

This paper cites Worst Cases Policy Gradients.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Worst Cases Policy Gradients

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.803063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.803063Z digest=sha256:bd8848595dd23abebf61bf66535f94d3c0a1663736e68ceb5e6f3e631b62489e

Observation b25a8f75-294b-41c0-9312-e85ac17e675b · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.807162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.807162Z digest=sha256:d9621b1325f95bd0d716c7e87b0a098917da95719d899d4874bd040ba2e0417d

Observation cf327413-a44a-4ee4-9938-e4befbed8c43 · outbound

This paper cites an unresolved cited work.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:24:52.012377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.810844Z digest=sha256:bb761c7a81f57691529a44a3530af24c3c8c988275d975d7f04b08bc3634d64b

Observation 9c4a3620-2e61-44a5-ab2c-4a4332aa8ba2 · outbound

This paper cites Risk-aware reward shaping of reinforcement learning agents for autonomous driving.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Risk-aware reward shaping of reinforcement learning agents for autonomous driving

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:52.001336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.814178Z digest=sha256:e91f2dd91a6bc0d3f2b21b03614dee4d066e9afb1f90bc34b771f58e97598752

Observation bbe21a40-01b8-4e58-bba6-008f881fadaa · outbound

This paper cites Off-Policy Primal-Dual Safe Reinforcement Learning.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation Off-Policy Primal-Dual Safe Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:51.818977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:51.818977Z digest=sha256:8a063ae8eadb1bbd2d7fbfda1bba2cf881a57335afb210e09045c9a1a635789a

Observation 250cdd10-2b4a-4d87-87fd-d34c3fdfe57e · outbound

This paper cites D., Tindemans, S.

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation D., Tindemans, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:24:51.988906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T05:24:51.822886Z digest=sha256:10d0b7119cb106080276517f3c988b44268052704cb8ab97295634bbf66ddee4

Pith citing papers

Observation 729129bc-3353-48b3-833d-1f61bc861859 · inbound

Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity cites this paper.

Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T05:00:42.106290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:00:42.106290Z digest=sha256:c547df0492638e1174d696cd11cbcffe588a67a51a713b38b9e50bd86b080133