Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:24.417982Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2505.20359.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:24.417982Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 266962d9-0de5-4335-907b-edde5bb7c2c1 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Deep reinforcement learning from human preferences
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 285f0b01-dcc3-4a49-b6fc-1ed7e24fb668 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Training language models to follow instructions with human feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 38a75ca1-948a-406c-81f7-5adadf20053f · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b4f0662-34d9-4000-b1e0-c93c88229940 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Preference ranking optimization for human alignment
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56a70f41-33c8-4d05-b0b9-dac4aa0d49bd · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Direct preference optimization: your language model is secretly a reward model
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 28c46a98-b358-433f-b7ed-3b179786fa5f · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4694cf76-44d0-485d-b000-c4f32623e342 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure A general theoretical paradigm to understand learning from human preferences
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1acd4098-80ed-4623-bbf1-898a97b2c34f · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Robust Preference Optimization through Reward Model Distillation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6748e5cc-8e5e-46e6-8f2f-fd9d81909661 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e749cdd-98e1-442a-bcbc-3b0a56471f0e · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Token- level direct preference optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 628382ae-80bc-4de9-9d04-301eec9e6575 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Trust region policy optimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a12a3d-99c1-43a9-a0e0-857624ca58b5 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Deep reinforcement learning: A brief survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f60c15d1-a973-4088-a52a-b246dcbf9a86 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure An introduction to deep reinforcement learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 019c4f99-7a71-4f5d-a89f-f18fe1d21e10 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Robust risk-aware reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e1b1f053-a824-4fc5-94cc-ebf63d3cfc0a · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-averse policy optimization via risk-neutral policy optimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 80b1104b-33e1-4da0-baa3-3799545f81f0 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-aware controller for autonomous vehicles using model-based collision prediction and reinforcement learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 61e19d94-d1c2-4b6f-996d-c09147e1a57b · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-averse fine-tuning of large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9c2fec67-6fc3-4b51-8f77-c14aac02b2ad · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Model alignment as prospect theoretic optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 267f7619-4325-47a3-afdf-2820c0c843b5 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Advances in prospect theory: Cumulative representation of uncertainty
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8eb6289a-eb4e-4873-8a7c-3bcf57af824f · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Rank analysis of incomplete block designs: I
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 046ef407-7718-4744-9e95-3dce0d99a81c · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure More risk-sensitive markov decision processes
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3408fba7-558e-4add-ac04-21632c2eb58d · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-averse autonomous systems: A brief history and recent developments from the perspective of optimal control.Artificial Intelligence, 311:103743, 2022
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 969f46a1-2a6a-4576-a4bc-d0fe256675c2 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Thinking coherently
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aad407c8-9b36-4909-b2f0-22888e191e79 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Optimization of conditional value-at-risk
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5677c95c-c45a-4d81-be50-685ee6e45259 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-sensitive and robust decision- making: a cvar optimization approach
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 27664587-af4a-4cd0-a591-f7eb4d4d4d58 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Convex measures of risk and trading constraints
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e883111d-e3b5-471a-8f7c-9a8adbfd4799 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Entropic risk optimization in discounted mdps
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5116fd70-e7fd-46bc-867a-ba14c3ef4064 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-sensitive reinforcement learning: Near-optimal risk-sample tradeoff in regret
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8f48a59-0ae5-4b6f-ab37-d8db67b42722 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Provably efficient iterated cvar reinforcement learning with function approximation and human feedback
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cd7353c2-64ae-46b7-a2db-1cba3c3a5c81 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Ra-pbrl: Provably efficient risk-aware preference-based reinforcement learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fbc31907-6796-4714-a501-2f634b0f1df0 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Entropic risk optimization in discounted mdps
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 04727585-b3ef-4b86-8cdd-c5b37b83ca81 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Learning word vectors for sentiment analysis
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4818b965-e2b0-4826-922e-1f8e6ca1a900 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 990f0832-084c-4bab-a0aa-4bfc1ec7618d · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afc58c1c-08da-478c-b814-543f0ef3790d · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Proximal Policy Optimization Algorithms
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edd5acce-2221-4fbe-98e5-61ef5cb0f7a9 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Language models are unsupervised multitask learners
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55b57b8d-0628-4271-910d-60cdb8250eb7 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Pythia: A suite for analyzing large language models across training and scaling
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation be60ffa7-ae5f-4129-a599-cc034230db8d · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c902023-c0ca-42ef-9ffb-bd7900bc5260 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Safe rlhf: Safe reinforcement learning from human feedback
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5c888fa5-67c1-4b57-a5eb-a607d4ba7166 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Challenges and Future Directions of Data-Centric AI Alignment
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f732037-9151-42d8-9766-eab794e6bd58 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Fine-grained human feedback gives better rewards for language model training
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 453b2863-17e6-4044-a57b-42e25ce55c33 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Token-level Proximal Policy Optimization for Query Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 210d9d09-36b8-4659-8815-dbd3d517a8d3 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9ed9aca-f69c-4703-b401-a43c91b3fe3b · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27b7eca0-57dd-41bb-95bf-66977aa009c7 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Rusu, Joel Veness, Marc G
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d8eee37-33f2-4a77-85b6-1be95e1961b7 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Reinforcement learning in robotic applications: a comprehensive survey
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0221ff5-480a-474c-bdda-4bbdfb0f7627 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure A review of safe reinforcement learning: Methods, theory and applications
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f82a1cbf-d09e-4721-af42-baa7dbf92d6b · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Risk-sensitive reinforcement learning with function approximation: A debiasing approach
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 24df7347-0acc-4d65-a3b6-311d8275634b · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Regret bounds for risk- sensitive reinforcement learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f58e3802-5610-4ee6-9723-da7b420fa2a6 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Near-minimax-optimal risk-sensitive reinforce- ment learning with cvar
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af5d9855-4ba9-46ce-b257-88718fb48075 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Provably efficient risk-sensitive reinforcement learning: Iterated cvar and worst path
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 70dc62a2-eda3-4005-b2a4-3df23f485650 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Group robust preference optimization in reward-free rlhf
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3073a4cc-e426-48f6-bd5f-f2bedbe07e9e · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Preference learning of latent decision utilities with a human-like model of preferential choice
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8f11223-2e07-44ab-b68a-84d7639b6110 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Robust reinforcement learning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9c72cd9-3543-43cf-8eb5-39307002d122 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Hamilton–jacobi reachability: Some recent theoretical advances and applications in unmanned airspace management
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c25ac1db-d5b9-4181-a928-fc07e0521b8a · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Coherent measures of risk
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4b01fe23-d627-4dd6-9598-e16e97fe3144 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Equivalence notions and model minimization in markov decision processes
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation db018505-bb46-4dd5-8dd1-cfeeb1fe5048 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure Learning markov network structure with decision trees
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 459a203a-5ae4-4887-83ed-acbff5a8a225 · outbound
Risk-aware Direct Preference Optimization under Nested Risk Measure ∞X t=1 γt−1 R x, y<t , yt + γ Φµ ˜Vπ x, y<t+1 − ˜Vπ([x]) # =Eτ |π′
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.