Pith. sign in

Paper Citation Record · LEDGER

Risk-aware Direct Preference Optimization under Nested Risk Measure

As of 8 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2505.20359.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20359 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:24.417982Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 266962d9-0de5-4335-907b-edde5bb7c2c1 · outbound

This paper cites Deep reinforcement learning from human preferences.

Risk-aware Direct Preference Optimization under Nested Risk Measure Deep reinforcement learning from human preferences

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:32.477197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.555869Z digest=sha256:111ca9554bd79b763639352ac259eaa9e22e19e2c264618fb5aaba0d11612e2f

Observation 285f0b01-dcc3-4a49-b6fc-1ed7e24fb668 · outbound

This paper cites Training language models to follow instructions with human feedback.

Risk-aware Direct Preference Optimization under Nested Risk Measure Training language models to follow instructions with human feedback

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:32.298990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.609106Z digest=sha256:14c2b944562d9160e41f079ba63995a900ccc086ed7cd78ecf3e350d0f41d025

Observation 38a75ca1-948a-406c-81f7-5adadf20053f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Risk-aware Direct Preference Optimization under Nested Risk Measure Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:18.705132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:18.705132Z digest=sha256:c32000e74535cfa47844145782ddec1d3b79abc6f024b92415693e2b63bbedbf

Observation 3b4f0662-34d9-4000-b1e0-c93c88229940 · outbound

This paper cites Preference ranking optimization for human alignment.

Risk-aware Direct Preference Optimization under Nested Risk Measure Preference ranking optimization for human alignment

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:32.168523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.798998Z digest=sha256:1fff5d5d1bea9199ceb3786e9d5fca8a44a3edcd851aec19ace299b62d95a3fd

Observation 56a70f41-33c8-4d05-b0b9-dac4aa0d49bd · outbound

This paper cites Direct preference optimization: your language model is secretly a reward model.

Risk-aware Direct Preference Optimization under Nested Risk Measure Direct preference optimization: your language model is secretly a reward model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.995327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:18.934357Z digest=sha256:21447ae6f134b8deadc626a6e8042dc7e9adb831914a2e06109e79271e8103f5

Observation 28c46a98-b358-433f-b7ed-3b179786fa5f · outbound

This paper cites Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints.

Risk-aware Direct Preference Optimization under Nested Risk Measure Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.866830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:19.071460Z digest=sha256:8bbeec526d919d0f08221633ee7b9d5522d820a3eaff9acfff62a85ac95c6e6a

Observation 4694cf76-44d0-485d-b000-c4f32623e342 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Risk-aware Direct Preference Optimization under Nested Risk Measure A general theoretical paradigm to understand learning from human preferences

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.742338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:19.203169Z digest=sha256:00f6cbbc8d2229ab7cf40612d434ea1e8318d25eee01711b144b7a2a6acc3b5a

Observation 1acd4098-80ed-4623-bbf1-898a97b2c34f · outbound

This paper cites Robust Preference Optimization through Reward Model Distillation.

Risk-aware Direct Preference Optimization under Nested Risk Measure Robust Preference Optimization through Reward Model Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:19.345978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:19.345978Z digest=sha256:ce4328eb7a2be3e0811de12aab09e6f9278d0414ee0bba49fd6cad3a251f6930

Observation 6748e5cc-8e5e-46e6-8f2f-fd9d81909661 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Risk-aware Direct Preference Optimization under Nested Risk Measure SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:19.460292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:19.460292Z digest=sha256:f4c02463a9348a36d1b38bd169414b76cc32a233a74503e340fe92824a1c0302

Observation 2e749cdd-98e1-442a-bcbc-3b0a56471f0e · outbound

This paper cites Token- level direct preference optimization.

Risk-aware Direct Preference Optimization under Nested Risk Measure Token- level direct preference optimization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.639826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:19.595079Z digest=sha256:8e82254c0af89a2324b6f2b05780467f555c9948feb76907bc152f61ce8097c7

Observation 628382ae-80bc-4de9-9d04-301eec9e6575 · outbound

This paper cites Trust region policy optimization.

Risk-aware Direct Preference Optimization under Nested Risk Measure Trust region policy optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:19.714328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:19.714328Z digest=sha256:81cffefaaa2ae70d94189e5e9731d84b866b417e1513da66863f5c63c0d9dbc0

Observation 41a12a3d-99c1-43a9-a0e0-857624ca58b5 · outbound

This paper cites Deep reinforcement learning: A brief survey.

Risk-aware Direct Preference Optimization under Nested Risk Measure Deep reinforcement learning: A brief survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:19.881577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:19.881577Z digest=sha256:cbb0bc2ca75b772e3bb2a13b52ac86d375899e9da0a7c8c85b0f80df07ff5e53

Observation f60c15d1-a973-4088-a52a-b246dcbf9a86 · outbound

This paper cites An introduction to deep reinforcement learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure An introduction to deep reinforcement learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:20.030535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:20.030535Z digest=sha256:5cfec114f3311a9d1e867f08a4d9000e6aaf208506adb78807a33b661e3cdc84

Observation 019c4f99-7a71-4f5d-a89f-f18fe1d21e10 · outbound

This paper cites Robust risk-aware reinforcement learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure Robust risk-aware reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.441474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:20.197325Z digest=sha256:f7dd7c6043b52bd54c4ea45cc19aa073313d46b012680f0e443725534f869443

Observation e1b1f053-a824-4fc5-94cc-ebf63d3cfc0a · outbound

This paper cites Risk-averse policy optimization via risk-neutral policy optimization.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-averse policy optimization via risk-neutral policy optimization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.321843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:20.309724Z digest=sha256:f4f1ef4bbaf46d6de6161f95e54f6699b108f6d049909ced04cc48086a575e2d

Observation 80b1104b-33e1-4da0-baa3-3799545f81f0 · outbound

This paper cites Risk-aware controller for autonomous vehicles using model-based collision prediction and reinforcement learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-aware controller for autonomous vehicles using model-based collision prediction and reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.198241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:20.466741Z digest=sha256:675cd7e47e37c3d50ccd7b035ad06c5eff800918b2f882829f07c663b42b6e14

Observation 61e19d94-d1c2-4b6f-996d-c09147e1a57b · outbound

This paper cites Risk-averse fine-tuning of large language models.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-averse fine-tuning of large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.091831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:20.594897Z digest=sha256:eeef83c1c02430f08dbd2c2a471cca5d355321536bd3ac297e4ed66931060c38

Observation 9c2fec67-6fc3-4b51-8f77-c14aac02b2ad · outbound

This paper cites Model alignment as prospect theoretic optimization.

Risk-aware Direct Preference Optimization under Nested Risk Measure Model alignment as prospect theoretic optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.986076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:20.752174Z digest=sha256:7b219b3782e09dd80268dfefd63d387a9134fb598a7e7c6c1dd7a57dd1cab788

Observation 267f7619-4325-47a3-afdf-2820c0c843b5 · outbound

This paper cites Advances in prospect theory: Cumulative representation of uncertainty.

Risk-aware Direct Preference Optimization under Nested Risk Measure Advances in prospect theory: Cumulative representation of uncertainty

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.894237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:20.878414Z digest=sha256:81737e3b10338c286122e287940d1e6732b502e96d30d6df25c9802f3c028e84

Observation 8eb6289a-eb4e-4873-8a7c-3bcf57af824f · outbound

This paper cites Rank analysis of incomplete block designs: I.

Risk-aware Direct Preference Optimization under Nested Risk Measure Rank analysis of incomplete block designs: I

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:21.101329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:21.101329Z digest=sha256:0b03b2036eb9e34e3055a391ee9fd1a71ba3299b82a43c978c8a99bac112746a

Observation 046ef407-7718-4744-9e95-3dce0d99a81c · outbound

This paper cites More risk-sensitive markov decision processes.

Risk-aware Direct Preference Optimization under Nested Risk Measure More risk-sensitive markov decision processes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.669560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:21.205496Z digest=sha256:395d6d259e41817f429d422e4f04ce739acf7071ec6dd65cdace86d202d839e3

Observation 3408fba7-558e-4add-ac04-21632c2eb58d · outbound

This paper cites Risk-averse autonomous systems: A brief history and recent developments from the perspective of optimal control.Artificial Intelligence, 311:103743, 2022.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-averse autonomous systems: A brief history and recent developments from the perspective of optimal control.Artificial Intelligence, 311:103743, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.518194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:21.317906Z digest=sha256:a713801427c3662fe60b5bb05126afd51b44d529cdef3e8b08f901f504d127f1

Observation 969f46a1-2a6a-4576-a4bc-d0fe256675c2 · outbound

This paper cites Thinking coherently.

Risk-aware Direct Preference Optimization under Nested Risk Measure Thinking coherently

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.345849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:21.373720Z digest=sha256:df79d824f121e907041651d04b7097b2b71431f446ce53431148668e24d298fa

Observation aad407c8-9b36-4909-b2f0-22888e191e79 · outbound

This paper cites Optimization of conditional value-at-risk.

Risk-aware Direct Preference Optimization under Nested Risk Measure Optimization of conditional value-at-risk

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.213248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:21.432352Z digest=sha256:22c560bde9401b79479898e638110a2fbbbc3e9de5e882be6aa68d78eca2d847

Observation 5677c95c-c45a-4d81-be50-685ee6e45259 · outbound

This paper cites Risk-sensitive and robust decision- making: a cvar optimization approach.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-sensitive and robust decision- making: a cvar optimization approach

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.125316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:21.477336Z digest=sha256:489d89fcab2793d45fb554d36be340cffa57005f28783b427ad13afab9de8020

Observation 27664587-af4a-4cd0-a591-f7eb4d4d4d58 · outbound

This paper cites Convex measures of risk and trading constraints.

Risk-aware Direct Preference Optimization under Nested Risk Measure Convex measures of risk and trading constraints

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:29.998074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:21.532142Z digest=sha256:774a84687b10d4a30ccbc374e3de6e82c1d952c435689fde1f7696f208b9ba5c

Observation e883111d-e3b5-471a-8f7c-9a8adbfd4799 · outbound

This paper cites Entropic risk optimization in discounted mdps.

Risk-aware Direct Preference Optimization under Nested Risk Measure Entropic risk optimization in discounted mdps

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:29.757999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:21.594078Z digest=sha256:a3d1442f3c1f46fc3c9bfcea86ad7a9dcefd3ed9136782dd445340287bd00bca

Observation 5116fd70-e7fd-46bc-867a-ba14c3ef4064 · outbound

This paper cites Risk-sensitive reinforcement learning: Near-optimal risk-sample tradeoff in regret.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-sensitive reinforcement learning: Near-optimal risk-sample tradeoff in regret

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:29.466612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:21.642270Z digest=sha256:5c2a62324eb6441eccb6e28ce718da88dc244af2d5049e8fecc9a5f362c94c7c

Observation f8f48a59-0ae5-4b6f-ab37-d8db67b42722 · outbound

This paper cites Provably efficient iterated cvar reinforcement learning with function approximation and human feedback.

Risk-aware Direct Preference Optimization under Nested Risk Measure Provably efficient iterated cvar reinforcement learning with function approximation and human feedback

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:29.226939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:21.717663Z digest=sha256:2f4b127ce02af91761555c5d773429534f73a515359c88b51530cc3d0fd6f4c8

Observation cd7353c2-64ae-46b7-a2db-1cba3c3a5c81 · outbound

This paper cites Ra-pbrl: Provably efficient risk-aware preference-based reinforcement learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure Ra-pbrl: Provably efficient risk-aware preference-based reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:28.987173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:21.798858Z digest=sha256:297897d14605086098c5d065c6d31a691ad23a7a4f2792529798abbaf269728b

Observation fbc31907-6796-4714-a501-2f634b0f1df0 · outbound

This paper cites Entropic risk optimization in discounted mdps.

Risk-aware Direct Preference Optimization under Nested Risk Measure Entropic risk optimization in discounted mdps

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:28.709483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:21.886758Z digest=sha256:f5870bfce82a79aea5143159b0a30522b478cd4fd42c28b0bb33b0ea10b816ec

Observation 04727585-b3ef-4b86-8cdd-c5b37b83ca81 · outbound

This paper cites Learning word vectors for sentiment analysis.

Risk-aware Direct Preference Optimization under Nested Risk Measure Learning word vectors for sentiment analysis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:28.435952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:21.973558Z digest=sha256:9e64d1d3a091d66e244906491d2bdd9a51c04008328cc862e9b1241a1b2a33a7

Observation 4818b965-e2b0-4826-922e-1f8e6ca1a900 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Risk-aware Direct Preference Optimization under Nested Risk Measure Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.071225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.071225Z digest=sha256:7265f57b7cade73cfd01a53d560a476eb5c1b89bce58c6b5384b70f08857121d

Observation 990f0832-084c-4bab-a0aa-4bfc1ec7618d · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Risk-aware Direct Preference Optimization under Nested Risk Measure Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.175187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.175187Z digest=sha256:ea88b61cc6e0f6ea1025f7f189fcc417ce757e3d5bad5798dcee8014ece21f14

Observation afc58c1c-08da-478c-b814-543f0ef3790d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Risk-aware Direct Preference Optimization under Nested Risk Measure Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.269080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.269080Z digest=sha256:b8ebade8972bf1e10e892901518061e46ea5fd8f78922bded8b3ba79fc501104

Observation edd5acce-2221-4fbe-98e5-61ef5cb0f7a9 · outbound

This paper cites Language models are unsupervised multitask learners.

Risk-aware Direct Preference Optimization under Nested Risk Measure Language models are unsupervised multitask learners

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.384512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.384512Z digest=sha256:d38a837cae94bc1f479140bfbb3be0a59f0c58025e6bf895ac54d1f4d3044d3a

Observation 55b57b8d-0628-4271-910d-60cdb8250eb7 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Risk-aware Direct Preference Optimization under Nested Risk Measure Pythia: A suite for analyzing large language models across training and scaling

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:28.157463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:22.479912Z digest=sha256:39211818ba4ebe1b83333f98ce79857fa4e9bfbcd0a34e3ceee70e1e44b867ae

Observation be60ffa7-ae5f-4129-a599-cc034230db8d · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

Risk-aware Direct Preference Optimization under Nested Risk Measure Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.576261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.576261Z digest=sha256:82829f82df89c3dcd83292d48a879d04630fdc4e25b34baf4ab90261389c2c16

Observation 2c902023-c0ca-42ef-9ffb-bd7900bc5260 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Risk-aware Direct Preference Optimization under Nested Risk Measure Safe rlhf: Safe reinforcement learning from human feedback

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.953097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:22.651044Z digest=sha256:07c2997b143b28414fb7c74dec91f98bb00a9ae70468376ac7bfa8ff9815b991

Observation 5c888fa5-67c1-4b57-a5eb-a607d4ba7166 · outbound

This paper cites Challenges and Future Directions of Data-Centric AI Alignment.

Risk-aware Direct Preference Optimization under Nested Risk Measure Challenges and Future Directions of Data-Centric AI Alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.733585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.733585Z digest=sha256:a69d396c115f95e370866a2252fb5acf8b692ba124002c0432598b0bfc6f6f44

Observation 9f732037-9151-42d8-9766-eab794e6bd58 · outbound

This paper cites Fine-grained human feedback gives better rewards for language model training.

Risk-aware Direct Preference Optimization under Nested Risk Measure Fine-grained human feedback gives better rewards for language model training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.782998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:22.802805Z digest=sha256:269e3b7f820500e6a9b73cfd9d24ebc2a616f658f38dac1b6d3dfed21ab97e60

Observation 453b2863-17e6-4044-a57b-42e25ce55c33 · outbound

This paper cites Token-level Proximal Policy Optimization for Query Generation.

Risk-aware Direct Preference Optimization under Nested Risk Measure Token-level Proximal Policy Optimization for Query Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:15:24.647583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:22.850475Z digest=sha256:1b798ed5e5ae64271ef5b3cb68f6a13ba0d3d4a3361e2d4ca65d09cfe4c41a9b

Observation 210d9d09-36b8-4659-8815-dbd3d517a8d3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Risk-aware Direct Preference Optimization under Nested Risk Measure DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.914266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.914266Z digest=sha256:627145a54f0ea8048301415567f5d51de7025779ed00307d0cf7a095033d7dac

Observation f9ed9aca-f69c-4703-b401-a43c91b3fe3b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.964183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.964183Z digest=sha256:8237e2d6499d6885196b0308a2f86ab7c624eced0a0ca6add6506af8154095a3

Observation 27b7eca0-57dd-41bb-95bf-66977aa009c7 · outbound

This paper cites Rusu, Joel Veness, Marc G.

Risk-aware Direct Preference Optimization under Nested Risk Measure Rusu, Joel Veness, Marc G

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:23.064520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:23.064520Z digest=sha256:5c4ab62bb76005805de3a922f2c0b31137e06c02d23b72846b3ee05c0510b84d

Observation 3d8eee37-33f2-4a77-85b6-1be95e1961b7 · outbound

This paper cites Reinforcement learning in robotic applications: a comprehensive survey.

Risk-aware Direct Preference Optimization under Nested Risk Measure Reinforcement learning in robotic applications: a comprehensive survey

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:23.142797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:23.142797Z digest=sha256:a687c370a78f721d357a4ca86311c293e0ac26681768134a20c6e6caf8ab3bc9

Observation f0221ff5-480a-474c-bdda-4bbdfb0f7627 · outbound

This paper cites A review of safe reinforcement learning: Methods, theory and applications.

Risk-aware Direct Preference Optimization under Nested Risk Measure A review of safe reinforcement learning: Methods, theory and applications

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.602650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:23.279032Z digest=sha256:c9a66e4f109fafe9e84e76f8129832176435e5fcc210957acfd8832a27e038bc

Observation f82a1cbf-d09e-4721-af42-baa7dbf92d6b · outbound

This paper cites Risk-sensitive reinforcement learning with function approximation: A debiasing approach.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-sensitive reinforcement learning with function approximation: A debiasing approach

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.480888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:23.383108Z digest=sha256:a073ff20112723f70c0dc320a8f2115dfd1def441fe5510036232a4f26435f81

Observation 24df7347-0acc-4d65-a3b6-311d8275634b · outbound

This paper cites Regret bounds for risk- sensitive reinforcement learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure Regret bounds for risk- sensitive reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.245057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:23.481362Z digest=sha256:783c359b4af3021de90122bb7d28c42eb2343043e5a4c59fd7ca5225d540792d

Observation f58e3802-5610-4ee6-9723-da7b420fa2a6 · outbound

This paper cites Near-minimax-optimal risk-sensitive reinforce- ment learning with cvar.

Risk-aware Direct Preference Optimization under Nested Risk Measure Near-minimax-optimal risk-sensitive reinforce- ment learning with cvar

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:26.980660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:23.599532Z digest=sha256:9fcdb23792845a9f391b71c29eab8d762583d595de27a09f0e9f935a4ce07c52

Observation af5d9855-4ba9-46ce-b257-88718fb48075 · outbound

This paper cites Provably efficient risk-sensitive reinforcement learning: Iterated cvar and worst path.

Risk-aware Direct Preference Optimization under Nested Risk Measure Provably efficient risk-sensitive reinforcement learning: Iterated cvar and worst path

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:26.734868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:23.675659Z digest=sha256:7a455989ff2fdf22a5b270b5090f63b802110e7218b84c00e72b6272795c501a

Observation 70dc62a2-eda3-4005-b2a4-3df23f485650 · outbound

This paper cites Group robust preference optimization in reward-free rlhf.

Risk-aware Direct Preference Optimization under Nested Risk Measure Group robust preference optimization in reward-free rlhf

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:26.458476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:23.778098Z digest=sha256:6effb65d48c91f9171216791bf0a436073e25e8f35d6f9f8f80e62e0c286ed78

Observation 3073a4cc-e426-48f6-bd5f-f2bedbe07e9e · outbound

This paper cites Preference learning of latent decision utilities with a human-like model of preferential choice.

Risk-aware Direct Preference Optimization under Nested Risk Measure Preference learning of latent decision utilities with a human-like model of preferential choice

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:26.197413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:23.884353Z digest=sha256:6477855d0b11d381512ad66e9f2f72c906d2502878648075c7a96ca34569c990

Observation f8f11223-2e07-44ab-b68a-84d7639b6110 · outbound

This paper cites Robust reinforcement learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure Robust reinforcement learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:23.965754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:23.965754Z digest=sha256:21de0d1621f5d21bbea7aa33aadebb52ee849631f8b8077c436b8546c65db73d

Observation c9c72cd9-3543-43cf-8eb5-39307002d122 · outbound

This paper cites Hamilton–jacobi reachability: Some recent theoretical advances and applications in unmanned airspace management.

Risk-aware Direct Preference Optimization under Nested Risk Measure Hamilton–jacobi reachability: Some recent theoretical advances and applications in unmanned airspace management

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:25.942242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:24.031706Z digest=sha256:cbfd6c4056bfd157aa4b619098272eb92b195f5c414f923757cfc7c86ecd0997

Observation c25ac1db-d5b9-4181-a928-fc07e0521b8a · outbound

This paper cites Coherent measures of risk.

Risk-aware Direct Preference Optimization under Nested Risk Measure Coherent measures of risk

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:25.638808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:24.147492Z digest=sha256:1b09f59ab04100f8bda43dbd5a89b7abf2a7b476922f3ab0c0e4d8d67c5e29b3

Observation 4b01fe23-d627-4dd6-9598-e16e97fe3144 · outbound

This paper cites Equivalence notions and model minimization in markov decision processes.

Risk-aware Direct Preference Optimization under Nested Risk Measure Equivalence notions and model minimization in markov decision processes

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:25.411257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:24.228406Z digest=sha256:db9461175aa59939945340089fa87611f40bf81590c88952a074a455ec8b1a93

Observation db018505-bb46-4dd5-8dd1-cfeeb1fe5048 · outbound

This paper cites Learning markov network structure with decision trees.

Risk-aware Direct Preference Optimization under Nested Risk Measure Learning markov network structure with decision trees

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:25.250218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:24.321204Z digest=sha256:4aaf7ff32bc309bc3f8d47ff10879098af11f4e4837e35ed21dce88aad12bea8

Observation 459a203a-5ae4-4887-83ed-acbff5a8a225 · outbound

This paper cites ∞X t=1 γt−1 R x, y<t , yt + γ Φµ ˜Vπ x, y<t+1 − ˜Vπ([x]) # =Eτ |π′.

Risk-aware Direct Preference Optimization under Nested Risk Measure ∞X t=1 γt−1 R x, y<t , yt + γ Φµ ˜Vπ x, y<t+1 − ˜Vπ([x]) # =Eτ |π′

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:25.040781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:24.417982Z digest=sha256:c203d7107bddf3f4aa1e5d233de27b128a4eb60d4d5b0eeaa802f6f2be612cdc

Pith citing papers

No inbound Pith citation observations are available.