Pith. sign in

Paper Citation Record · LEDGER

Risk-aware Direct Preference Optimization under Nested Risk Measure

As of 14 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2505.20359.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20359 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:24.417982Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 266962d9-0de5-4335-907b-edde5bb7c2c1 · outbound

This paper cites Deep reinforcement learning from human preferences.

Risk-aware Direct Preference Optimization under Nested Risk Measure Deep reinforcement learning from human preferences

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:32.477197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:18.555869Z digest=sha256:88735b20006c4048183d589bbe0be5a6d600f5fdf4a16e8183e10b0ce3564922

Observation 285f0b01-dcc3-4a49-b6fc-1ed7e24fb668 · outbound

This paper cites Training language models to follow instructions with human feedback.

Risk-aware Direct Preference Optimization under Nested Risk Measure Training language models to follow instructions with human feedback

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:32.298990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:18.609106Z digest=sha256:088ad4507fb2d1bc30c99a2fb3de253d497d27130e263409bb34f005a34130ed

Observation 38a75ca1-948a-406c-81f7-5adadf20053f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Risk-aware Direct Preference Optimization under Nested Risk Measure Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:18.705132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:18.705132Z digest=sha256:7c0a0e9ba4fec5e27538a7c1089442ecf7c47407ff8c8065214506c449d1501b

Observation 3b4f0662-34d9-4000-b1e0-c93c88229940 · outbound

This paper cites Preference ranking optimization for human alignment.

Risk-aware Direct Preference Optimization under Nested Risk Measure Preference ranking optimization for human alignment

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:32.168523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:18.798998Z digest=sha256:6cf9faa589211275b6dcef7f3d792becde34b17f7a01c6d674853974a4000ac7

Observation 56a70f41-33c8-4d05-b0b9-dac4aa0d49bd · outbound

This paper cites Direct preference optimization: your language model is secretly a reward model.

Risk-aware Direct Preference Optimization under Nested Risk Measure Direct preference optimization: your language model is secretly a reward model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.995327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:18.934357Z digest=sha256:3d905d3e091cf869350aa30b95daf7fd1b6a88aa821becd85d1d2d88f2262bcc

Observation 28c46a98-b358-433f-b7ed-3b179786fa5f · outbound

This paper cites Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints.

Risk-aware Direct Preference Optimization under Nested Risk Measure Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.866830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:19.071460Z digest=sha256:4b9e128b2d6522dcea38e8b15bd9f103076c8d7cbe06ebc14008aa53fb144cb1

Observation 4694cf76-44d0-485d-b000-c4f32623e342 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Risk-aware Direct Preference Optimization under Nested Risk Measure A general theoretical paradigm to understand learning from human preferences

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.742338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:19.203169Z digest=sha256:48d6e239d35001cbe831c2df55d516f65581f6bbf0533afee2ed21da735e761d

Observation 1acd4098-80ed-4623-bbf1-898a97b2c34f · outbound

This paper cites Robust Preference Optimization through Reward Model Distillation.

Risk-aware Direct Preference Optimization under Nested Risk Measure Robust Preference Optimization through Reward Model Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:19.345978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:19.345978Z digest=sha256:df55ae44df6a02c2cbeded14171f4a9565c167a8e18f42dac02d8f14c0629e16

Observation 6748e5cc-8e5e-46e6-8f2f-fd9d81909661 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Risk-aware Direct Preference Optimization under Nested Risk Measure SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:19.460292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:19.460292Z digest=sha256:4861e314a8b7ebe1ef5f6e63884412c26aee8d513127776a820625aaea13b361

Observation 2e749cdd-98e1-442a-bcbc-3b0a56471f0e · outbound

This paper cites Token- level direct preference optimization.

Risk-aware Direct Preference Optimization under Nested Risk Measure Token- level direct preference optimization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.639826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:19.595079Z digest=sha256:28402da6097a3c37628e11b392ec9203979e289503a602a3fe433da709c66d77

Observation 628382ae-80bc-4de9-9d04-301eec9e6575 · outbound

This paper cites Trust region policy optimization.

Risk-aware Direct Preference Optimization under Nested Risk Measure Trust region policy optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:19.714328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:19.714328Z digest=sha256:5a45154785ac5b7e01086f90d868486eea055d9123652f16497902c65f640581

Observation 41a12a3d-99c1-43a9-a0e0-857624ca58b5 · outbound

This paper cites Deep reinforcement learning: A brief survey.

Risk-aware Direct Preference Optimization under Nested Risk Measure Deep reinforcement learning: A brief survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:19.881577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:19.881577Z digest=sha256:89de841bbc8d6f8548e41f56c2f398cdecd924331bd582d780cbe225bbd599ec

Observation f60c15d1-a973-4088-a52a-b246dcbf9a86 · outbound

This paper cites An introduction to deep reinforcement learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure An introduction to deep reinforcement learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:20.030535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:20.030535Z digest=sha256:bea52d2813612e3fe8dc63812e1506c3e6fed158e82638598d42c79500a8339f

Observation 019c4f99-7a71-4f5d-a89f-f18fe1d21e10 · outbound

This paper cites Robust risk-aware reinforcement learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure Robust risk-aware reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.441474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:20.197325Z digest=sha256:c3519a8c833fefbfc4e9b7e97fb344f8df7384f04d366ee29e56c43b0a137d66

Observation e1b1f053-a824-4fc5-94cc-ebf63d3cfc0a · outbound

This paper cites Risk-averse policy optimization via risk-neutral policy optimization.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-averse policy optimization via risk-neutral policy optimization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.321843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:20.309724Z digest=sha256:28b0facaae7eccf0ec997ad5e4a13ce56a14d766a09badef9961e9823d865678

Observation 80b1104b-33e1-4da0-baa3-3799545f81f0 · outbound

This paper cites Risk-aware controller for autonomous vehicles using model-based collision prediction and reinforcement learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-aware controller for autonomous vehicles using model-based collision prediction and reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.198241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:20.466741Z digest=sha256:4006a68a4d936716226d65d8038f9d0edd27c0a5caf803ee7fd1fa0430efde6f

Observation 61e19d94-d1c2-4b6f-996d-c09147e1a57b · outbound

This paper cites Risk-averse fine-tuning of large language models.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-averse fine-tuning of large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:31.091831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:20.594897Z digest=sha256:84e49171c3c74329bf9a6568135a9cf0e369aea7242b7d6dd8c20f3b4cffcab6

Observation 9c2fec67-6fc3-4b51-8f77-c14aac02b2ad · outbound

This paper cites Model alignment as prospect theoretic optimization.

Risk-aware Direct Preference Optimization under Nested Risk Measure Model alignment as prospect theoretic optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.986076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:20.752174Z digest=sha256:41958f9fda9d0b705c15f5ceb51ba16f7810148b0f1c2d1dcca036e00f172c96

Observation 267f7619-4325-47a3-afdf-2820c0c843b5 · outbound

This paper cites Advances in prospect theory: Cumulative representation of uncertainty.

Risk-aware Direct Preference Optimization under Nested Risk Measure Advances in prospect theory: Cumulative representation of uncertainty

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.894237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:20.878414Z digest=sha256:0839dcc5ce0a5fd787711a5a9c94cac330c1918423d05c4a89041e71eef22e05

Observation 8eb6289a-eb4e-4873-8a7c-3bcf57af824f · outbound

This paper cites Rank analysis of incomplete block designs: I.

Risk-aware Direct Preference Optimization under Nested Risk Measure Rank analysis of incomplete block designs: I

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:21.101329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:21.101329Z digest=sha256:61d877f2a99caa941eacaaa8fc013d9c4699356d07acaad07da72e38b6b453e6

Observation 046ef407-7718-4744-9e95-3dce0d99a81c · outbound

This paper cites More risk-sensitive markov decision processes.

Risk-aware Direct Preference Optimization under Nested Risk Measure More risk-sensitive markov decision processes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.669560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:21.205496Z digest=sha256:8b41bb48230b7bb86b5547f038fc5a0873dd5d5f18a140e93587f64bff09ae0e

Observation 3408fba7-558e-4add-ac04-21632c2eb58d · outbound

This paper cites Risk-averse autonomous systems: A brief history and recent developments from the perspective of optimal control.Artificial Intelligence, 311:103743, 2022.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-averse autonomous systems: A brief history and recent developments from the perspective of optimal control.Artificial Intelligence, 311:103743, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.518194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:21.317906Z digest=sha256:4bc046ae79732cb0d749aa38f8579461d530ee9968a18ff87ee7da431de49d40

Observation 969f46a1-2a6a-4576-a4bc-d0fe256675c2 · outbound

This paper cites Thinking coherently.

Risk-aware Direct Preference Optimization under Nested Risk Measure Thinking coherently

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.345849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:21.373720Z digest=sha256:57dce3d1f712d15e3adb1962405668658c4360487d2c543242868da183af9a3f

Observation aad407c8-9b36-4909-b2f0-22888e191e79 · outbound

This paper cites Optimization of conditional value-at-risk.

Risk-aware Direct Preference Optimization under Nested Risk Measure Optimization of conditional value-at-risk

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.213248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:21.432352Z digest=sha256:45cd15c93ce33c48bafae7b7f95fa78e450e251d9cd6a189a1bc3995928212f7

Observation 5677c95c-c45a-4d81-be50-685ee6e45259 · outbound

This paper cites Risk-sensitive and robust decision- making: a cvar optimization approach.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-sensitive and robust decision- making: a cvar optimization approach

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:30.125316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:21.477336Z digest=sha256:b4b6d90f901d47474e029560ac78efb02ae6f97e2b7219c9a8bb3db8cd2209e0

Observation 27664587-af4a-4cd0-a591-f7eb4d4d4d58 · outbound

This paper cites Convex measures of risk and trading constraints.

Risk-aware Direct Preference Optimization under Nested Risk Measure Convex measures of risk and trading constraints

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:29.998074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:21.532142Z digest=sha256:23ab1de25a091979f96a0117c304822ee1e607f5301895849d4e27b7b98edc71

Observation e883111d-e3b5-471a-8f7c-9a8adbfd4799 · outbound

This paper cites Entropic risk optimization in discounted mdps.

Risk-aware Direct Preference Optimization under Nested Risk Measure Entropic risk optimization in discounted mdps

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:29.757999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:21.594078Z digest=sha256:183ea85339499133da1b49f2e286c11498e049e0547e9c27f93920769a96ae58

Observation 5116fd70-e7fd-46bc-867a-ba14c3ef4064 · outbound

This paper cites Risk-sensitive reinforcement learning: Near-optimal risk-sample tradeoff in regret.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-sensitive reinforcement learning: Near-optimal risk-sample tradeoff in regret

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:29.466612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:21.642270Z digest=sha256:3371ae577c4f42abcc90c624b688f983fa958b6814eb11545c6daff99ea69254

Observation f8f48a59-0ae5-4b6f-ab37-d8db67b42722 · outbound

This paper cites Provably efficient iterated cvar reinforcement learning with function approximation and human feedback.

Risk-aware Direct Preference Optimization under Nested Risk Measure Provably efficient iterated cvar reinforcement learning with function approximation and human feedback

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:29.226939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:21.717663Z digest=sha256:54b8e911afe2b2927697ebf6ce2a37502b7b43bbac74737a34be606598860439

Observation cd7353c2-64ae-46b7-a2db-1cba3c3a5c81 · outbound

This paper cites Ra-pbrl: Provably efficient risk-aware preference-based reinforcement learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure Ra-pbrl: Provably efficient risk-aware preference-based reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:28.987173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:21.798858Z digest=sha256:6e2b8a96e812c9f3f773264e88cbaf32b0064b16881d71c0aa6df863cdea756b

Observation fbc31907-6796-4714-a501-2f634b0f1df0 · outbound

This paper cites Entropic risk optimization in discounted mdps.

Risk-aware Direct Preference Optimization under Nested Risk Measure Entropic risk optimization in discounted mdps

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:28.709483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:21.886758Z digest=sha256:efcc159f1d8e61266fd50ac400a7e8d55fd1703e1905340a68dfecc2d4d4fd88

Observation 04727585-b3ef-4b86-8cdd-c5b37b83ca81 · outbound

This paper cites Learning word vectors for sentiment analysis.

Risk-aware Direct Preference Optimization under Nested Risk Measure Learning word vectors for sentiment analysis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:28.435952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:21.973558Z digest=sha256:16bc70c681d84ef6e9740910c93fdf99ffc2f85c68f5ec3837b4e64e49a11ead

Observation 4818b965-e2b0-4826-922e-1f8e6ca1a900 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Risk-aware Direct Preference Optimization under Nested Risk Measure Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.071225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.071225Z digest=sha256:d153a9ffc9846600535540520ad2416a555d0dcc0d75fffeb0048c78954a83ce

Observation 990f0832-084c-4bab-a0aa-4bfc1ec7618d · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Risk-aware Direct Preference Optimization under Nested Risk Measure Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.175187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.175187Z digest=sha256:5affbccc44d5e6a2bb33741420f0837407251e16d0de0b80e75f00f9080a6671

Observation afc58c1c-08da-478c-b814-543f0ef3790d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Risk-aware Direct Preference Optimization under Nested Risk Measure Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.269080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.269080Z digest=sha256:3bf669d6c394176be2423eb227c26fda70b1d32fb8b589da05f917c96868f2e1

Observation edd5acce-2221-4fbe-98e5-61ef5cb0f7a9 · outbound

This paper cites Language models are unsupervised multitask learners.

Risk-aware Direct Preference Optimization under Nested Risk Measure Language models are unsupervised multitask learners

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.384512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.384512Z digest=sha256:4e58786154916a60fe5787bfa6566f0c568d070873466f9e7fe2d2e0fd9300d7

Observation 55b57b8d-0628-4271-910d-60cdb8250eb7 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Risk-aware Direct Preference Optimization under Nested Risk Measure Pythia: A suite for analyzing large language models across training and scaling

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:28.157463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:22.479912Z digest=sha256:17deb0ea03bf9a838d4a4f1e874df9adf769b7001af1591ad7e681b8f4dd9553

Observation be60ffa7-ae5f-4129-a599-cc034230db8d · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

Risk-aware Direct Preference Optimization under Nested Risk Measure Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.576261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.576261Z digest=sha256:8ba67ff322d2c3876df3b5b51141aad7093a627dc9fff3e2ce41717c1184e03b

Observation 2c902023-c0ca-42ef-9ffb-bd7900bc5260 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Risk-aware Direct Preference Optimization under Nested Risk Measure Safe rlhf: Safe reinforcement learning from human feedback

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.953097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:22.651044Z digest=sha256:79d3c2b41531cadcd8efe2fc27f296f7a1178813dd43ad8430f2aec86ad79356

Observation 5c888fa5-67c1-4b57-a5eb-a607d4ba7166 · outbound

This paper cites Challenges and Future Directions of Data-Centric AI Alignment.

Risk-aware Direct Preference Optimization under Nested Risk Measure Challenges and Future Directions of Data-Centric AI Alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.733585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.733585Z digest=sha256:d198f4129bdc6a357e359b55c28ef316e50627f904f05cc4475c87fcddddafcf

Observation 9f732037-9151-42d8-9766-eab794e6bd58 · outbound

This paper cites Fine-grained human feedback gives better rewards for language model training.

Risk-aware Direct Preference Optimization under Nested Risk Measure Fine-grained human feedback gives better rewards for language model training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.782998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:22.802805Z digest=sha256:53445d3b693be5d985cb57859ad129cb332b192f6583bcd373d6b1cae4540718

Observation 453b2863-17e6-4044-a57b-42e25ce55c33 · outbound

This paper cites Token-level Proximal Policy Optimization for Query Generation.

Risk-aware Direct Preference Optimization under Nested Risk Measure Token-level Proximal Policy Optimization for Query Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:15:24.647583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:22.850475Z digest=sha256:42d87de04929bcfeb5d8348c6f9bd48e1b66f0aead72959bd0e8b489b11b0dda

Observation 210d9d09-36b8-4659-8815-dbd3d517a8d3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Risk-aware Direct Preference Optimization under Nested Risk Measure DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.914266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.914266Z digest=sha256:44ee8e62d663795cd8baa3b3138d2774eb326e35c295cab6ea29d058a666d4c5

Observation f9ed9aca-f69c-4703-b401-a43c91b3fe3b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.964183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.964183Z digest=sha256:4fa165626cdfd855b6f1931ad14b656958fc461f98a0dd20228f1245101f6852

Observation 27b7eca0-57dd-41bb-95bf-66977aa009c7 · outbound

This paper cites Rusu, Joel Veness, Marc G.

Risk-aware Direct Preference Optimization under Nested Risk Measure Rusu, Joel Veness, Marc G

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:23.064520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:23.064520Z digest=sha256:a2a9b1aaeb817457b5cd7b5b81f5f7441cbfcb19669606c40aabc92d85f532ca

Observation 3d8eee37-33f2-4a77-85b6-1be95e1961b7 · outbound

This paper cites Reinforcement learning in robotic applications: a comprehensive survey.

Risk-aware Direct Preference Optimization under Nested Risk Measure Reinforcement learning in robotic applications: a comprehensive survey

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:23.142797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:23.142797Z digest=sha256:6adc0ac99d06e3d002e116fc89f5aef2cd70ef6e60200e969404e3db3f036327

Observation f0221ff5-480a-474c-bdda-4bbdfb0f7627 · outbound

This paper cites A review of safe reinforcement learning: Methods, theory and applications.

Risk-aware Direct Preference Optimization under Nested Risk Measure A review of safe reinforcement learning: Methods, theory and applications

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.602650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:23.279032Z digest=sha256:cb2386deca4177fe8917a0af2627b743c56326f0f4b3f67832a0f1aec2a2ead8

Observation f82a1cbf-d09e-4721-af42-baa7dbf92d6b · outbound

This paper cites Risk-sensitive reinforcement learning with function approximation: A debiasing approach.

Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-sensitive reinforcement learning with function approximation: A debiasing approach

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.480888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:23.383108Z digest=sha256:1ce23072d7e51b6c550ae8890a94e93c23392d6c48bbf4eab1ac611d32c7dee9

Observation 24df7347-0acc-4d65-a3b6-311d8275634b · outbound

This paper cites Regret bounds for risk- sensitive reinforcement learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure Regret bounds for risk- sensitive reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:27.245057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:23.481362Z digest=sha256:276e531bab7733b3fc2ad1e001c4fc99cd69e2f3e7b5f70c27af6b81d4d53e20

Observation f58e3802-5610-4ee6-9723-da7b420fa2a6 · outbound

This paper cites Near-minimax-optimal risk-sensitive reinforce- ment learning with cvar.

Risk-aware Direct Preference Optimization under Nested Risk Measure Near-minimax-optimal risk-sensitive reinforce- ment learning with cvar

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:26.980660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:23.599532Z digest=sha256:d534ef75f9f3c469e2e485c4c999e453c97d27176074a296be3fa03a34b8ac3c

Observation af5d9855-4ba9-46ce-b257-88718fb48075 · outbound

This paper cites Provably efficient risk-sensitive reinforcement learning: Iterated cvar and worst path.

Risk-aware Direct Preference Optimization under Nested Risk Measure Provably efficient risk-sensitive reinforcement learning: Iterated cvar and worst path

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:26.734868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:23.675659Z digest=sha256:264c4863404cfd4da91685a4505f06a916402173d56309d3225f20dc6ca00c2a

Observation 70dc62a2-eda3-4005-b2a4-3df23f485650 · outbound

This paper cites Group robust preference optimization in reward-free rlhf.

Risk-aware Direct Preference Optimization under Nested Risk Measure Group robust preference optimization in reward-free rlhf

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:26.458476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:23.778098Z digest=sha256:2a4e10d8db1b0e003cd72412d9e0aea34c59b0e2977618b6502aae6c7454c6c8

Observation 3073a4cc-e426-48f6-bd5f-f2bedbe07e9e · outbound

This paper cites Preference learning of latent decision utilities with a human-like model of preferential choice.

Risk-aware Direct Preference Optimization under Nested Risk Measure Preference learning of latent decision utilities with a human-like model of preferential choice

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:26.197413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:23.884353Z digest=sha256:a677aabf5dcc6a65d32bf2b9f0ec68afe491a414587d6c9859f215ba5ff14703

Observation f8f11223-2e07-44ab-b68a-84d7639b6110 · outbound

This paper cites Robust reinforcement learning.

Risk-aware Direct Preference Optimization under Nested Risk Measure Robust reinforcement learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:23.965754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:23.965754Z digest=sha256:40da639031a8e11ce0b9a2983240a6e98b1260be8c5d7f552f41d60ea24394a6

Observation c9c72cd9-3543-43cf-8eb5-39307002d122 · outbound

This paper cites Hamilton–jacobi reachability: Some recent theoretical advances and applications in unmanned airspace management.

Risk-aware Direct Preference Optimization under Nested Risk Measure Hamilton–jacobi reachability: Some recent theoretical advances and applications in unmanned airspace management

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:25.942242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:24.031706Z digest=sha256:c92eb886a7ade93a6eec882a5e48cf23b1755e3f080fd822aa159f5b9c797cb2

Observation c25ac1db-d5b9-4181-a928-fc07e0521b8a · outbound

This paper cites Coherent measures of risk.

Risk-aware Direct Preference Optimization under Nested Risk Measure Coherent measures of risk

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:25.638808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:24.147492Z digest=sha256:5db73263fcf4b27e56b4c8fe7898be2cf0a2ff6b5f09f10271a62cccae3fa5f1

Observation 4b01fe23-d627-4dd6-9598-e16e97fe3144 · outbound

This paper cites Equivalence notions and model minimization in markov decision processes.

Risk-aware Direct Preference Optimization under Nested Risk Measure Equivalence notions and model minimization in markov decision processes

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:25.411257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:24.228406Z digest=sha256:8ee8f01e548d23d369b56dfbe0498b0a73df41935f6ee8cbc0a2608c0a930e31

Observation db018505-bb46-4dd5-8dd1-cfeeb1fe5048 · outbound

This paper cites Learning markov network structure with decision trees.

Risk-aware Direct Preference Optimization under Nested Risk Measure Learning markov network structure with decision trees

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:25.250218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:24.321204Z digest=sha256:ae032fd5f5b01670ad3536b6408906da296c3ed1f70fc51fa7ff05181fc96598

Observation 459a203a-5ae4-4887-83ed-acbff5a8a225 · outbound

This paper cites ∞X t=1 γt−1 R x, y<t , yt + γ Φµ ˜Vπ x, y<t+1 − ˜Vπ([x]) # =Eτ |π′.

Risk-aware Direct Preference Optimization under Nested Risk Measure ∞X t=1 γt−1 R x, y<t , yt + γ Φµ ˜Vπ x, y<t+1 − ˜Vπ([x]) # =Eτ |π′

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:25.040781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:24.417982Z digest=sha256:99aa8265e9138fe005ec0a114c6592f21ea7171cab99f354faa0173b47f701a1

Pith citing papers

No inbound Pith citation observations are available.