Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T18:32:21.331060Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2502.10505.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T18:32:21.331060Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d6180138-a454-4301-9e33-643e150b0bb8 · outbound
Preference learning made easy: Everything should be understood through win rate Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1b074722-1adc-408b-b4cb-9b3f3040dc52 · outbound
Preference learning made easy: Everything should be understood through win rate Calibration and consistency of adversarial surrogate losses
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c251466-16c3-4b63-b3ac-5d2ad985737c · outbound
Preference learning made easy: Everything should be understood through win rate Multi-class h -consistency bounds
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c2cc416-e36f-4b17-94ca-93af7a38517c · outbound
Preference learning made easy: Everything should be understood through win rate G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d8ab9d49-91d5-49cd-a64c-2dc628851e04 · outbound
Preference learning made easy: Everything should be understood through win rate Training a helpful and harmless assistant with reinforcement learning from human feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bad98efd-8a4b-48f9-9230-63a8f1eee886 · outbound
Preference learning made easy: Everything should be understood through win rate Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e77688e-f12f-4247-b719-10fae9e6a0b8 · outbound
Preference learning made easy: Everything should be understood through win rate Quantile Filtered Imitation Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 43afe68a-4ec2-412e-85dc-3230471c94a9 · outbound
Preference learning made easy: Everything should be understood through win rate Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f07e96b-2159-4740-a1d1-875239dda663 · outbound
Preference learning made easy: Everything should be understood through win rate Human Alignment of Large Language Models through Online Preference Optimisation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8107c9da-f2c5-40f5-8b2f-330139a41e75 · outbound
Preference learning made easy: Everything should be understood through win rate H., Chen, X., Zhang, Q., Ranganath, R., and Cho, K
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2902befc-3e72-43b6-acfe-88eb4cef412d · outbound
Preference learning made easy: Everything should be understood through win rate Deep reinforcement learning from human preferences
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b619d067-3659-4495-ad93-2921e730fb13 · outbound
Preference learning made easy: Everything should be understood through win rate Raft: Reward ranked finetuning for generative foundation model alignment
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 51dd05ae-79b1-41fc-ba02-c0308782ba0f · outbound
Preference learning made easy: Everything should be understood through win rate Understanding dataset difficulty with v-usable information
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5d841ae-b317-4623-9930-c2b9d97c89b4 · outbound
Preference learning made easy: Everything should be understood through win rate Bonbon alignment for large language models and the sweetness of best-of-n sampling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b64d9244-8006-498c-96d9-4831e68de4b7 · outbound
Preference learning made easy: Everything should be understood through win rate L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5724ed26-706e-4021-a240-b524f4676b17 · outbound
Preference learning made easy: Everything should be understood through win rate Direct Language Model Alignment from Online AI Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ae999b-8da0-42a5-b001-6d340fead544 · outbound
Preference learning made easy: Everything should be understood through win rate $f$-PO: Generalizing Preference Optimization with $f$-divergence Minimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3f076e6-6916-4baf-95d1-94ee9fcb0d54 · outbound
Preference learning made easy: Everything should be understood through win rate Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52740ac5-432c-4efb-870d-81ff9b8c7ff2 · outbound
Preference learning made easy: Everything should be understood through win rate A., Choi, Y., and Hajishirzi, H
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 42270018-9ef5-4abf-b241-7186c0528296 · outbound
Preference learning made easy: Everything should be understood through win rate Towards Efficient Exact Optimization of Language Model Alignment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb3b3895-c27c-4a72-9c09-74f8b2d8a384 · outbound
Preference learning made easy: Everything should be understood through win rate A Survey on Human Preference Learning for Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ba482e-7eed-4cf0-ba69-08f0200da95b · outbound
Preference learning made easy: Everything should be understood through win rate A survey of reinforcement learning from human feedback
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15754356-b360-4f9e-948e-14bdfb38a397 · outbound
Preference learning made easy: Everything should be understood through win rate Understanding the effects of rlhf on llm generalisation and diversity
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d73ad913-60ec-4daa-947a-c504d2031a5d · outbound
Preference learning made easy: Everything should be understood through win rate OpenAssistant Conversations -- Democratizing Large Language Model Alignment
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 103738c8-633e-451f-be24-7474252aa77b · outbound
Preference learning made easy: Everything should be understood through win rate Huggingface h4 stack exchange preference dataset
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 367c276c-8757-48fc-ba82-3f127d3ae281 · outbound
Preference learning made easy: Everything should be understood through win rate Self-Alignment with Instruction Backtranslation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 959bff62-81b5-428d-a8df-db9dabe0082d · outbound
Preference learning made easy: Everything should be understood through win rate Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 14b5bd41-343a-4f2a-b5ca-7ae15d2060c1 · outbound
Preference learning made easy: Everything should be understood through win rate Large Language Models: A Survey
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74293d31-b9a8-4b0c-98c0-135fa23159a2 · outbound
Preference learning made easy: Everything should be understood through win rate Monte carlo gradient estimation in machine learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39550b1e-bb78-4ec1-85ed-2813f2bada54 · outbound
Preference learning made easy: Everything should be understood through win rate Monte carlo gradient estimation in machine learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 465c40fd-57e9-44af-9f9e-309cf5c2cfec · outbound
Preference learning made easy: Everything should be understood through win rate G., Rowland, M., Guo, Z
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20c613a0-c772-4fd1-86f1-899de5cc54a6 · outbound
Preference learning made easy: Everything should be understood through win rate A., Lindsten, F., and Blei, D
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5336bf60-2085-4690-9267-135c1e51011f · outbound
Preference learning made easy: Everything should be understood through win rate Gpt-4 technical report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 816e2397-35a3-4c92-9c0f-9672f797b1e2 · outbound
Preference learning made easy: Everything should be understood through win rate Training language models to follow instructions with human feedback
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd5fa80e-fc39-44b0-87a1-6a881f593f98 · outbound
Preference learning made easy: Everything should be understood through win rate Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85b9a4a5-97aa-499d-8585-18d3b2556cd4 · outbound
Preference learning made easy: Everything should be understood through win rate Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5b4a645-6411-48d5-a1b2-a2dc0a364c3c · outbound
Preference learning made easy: Everything should be understood through win rate From r to q^* : Your language model is secretly a q-function
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1b439a24-3a82-4014-9314-a15db84d1e18 · outbound
Preference learning made easy: Everything should be understood through win rate D., Ermon, S., and Finn, C
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 05e21823-4197-4c1b-8316-77d5cd106f45 · outbound
Preference learning made easy: Everything should be understood through win rate Black box variational inference
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2b6830b4-152f-418b-afd5-34dffd1e055f · outbound
Preference learning made easy: Everything should be understood through win rate Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a389f205-5390-4018-9842-8d8291d01b0c · outbound
Preference learning made easy: Everything should be understood through win rate Vanishing gradients in reinforcement finetuning of language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cacef77f-dca8-49dd-a1ec-72cd0c5e9daf · outbound
Preference learning made easy: Everything should be understood through win rate Direct nash optimization: Teaching language models to self-improve with general preferences, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 64519d34-4c42-4d2f-a260-b89f02cfed61 · outbound
Preference learning made easy: Everything should be understood through win rate Direct Alignment with Heterogeneous Preferences
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 823f03aa-ab4a-4e29-9d59-69a72139f1aa · outbound
Preference learning made easy: Everything should be understood through win rate A Long Way to Go: Investigating Length Correlations in RLHF
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fea3bb9f-372c-4b55-89ef-4ac8e716d64b · outbound
Preference learning made easy: Everything should be understood through win rate Distributional preference learning: Understanding and accounting for hidden context in rlhf
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d577a9f7-535e-42c8-8c1a-7fde56c8e413 · outbound
Preference learning made easy: Everything should be understood through win rate How to compare different loss functions and their risks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3b65fc6e-bc9e-44f5-a051-1c4c21f62af3 · outbound
Preference learning made easy: Everything should be understood through win rate Learning to summarize from human feedback
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9667a6d-961a-4104-b483-ad8b2d1fbb5f · outbound
Preference learning made easy: Everything should be understood through win rate S., and Agarwal, A
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 59644008-f9e1-4566-9f7d-bb638ceb40cd · outbound
Preference learning made easy: Everything should be understood through win rate Preference fine-tuning of llms should leverage suboptimal, on-policy data
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a8b1c143-4b83-49ea-bf47-4462a0fc4edb · outbound
Preference learning made easy: Everything should be understood through win rate D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8a4e59a3-d701-4d99-98be-dd44878199ec · outbound
Preference learning made easy: Everything should be understood through win rate Trl: Transformer reinforcement learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 14b72f64-ca56-407f-aee4-3b0792b5404b · outbound
Preference learning made easy: Everything should be understood through win rate Enabling language models to implicitly learn self-improvement
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 622b1c22-0d0e-4116-9b0c-ca6b54aefe74 · outbound
Preference learning made easy: Everything should be understood through win rate Transforming and combining rewards for aligning large language models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 14cf1242-2bb1-4135-80d9-aa4b0aefe821 · outbound
Preference learning made easy: Everything should be understood through win rate Policy gradient algorithms
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 74368c74-59ad-45f1-a388-e08c26b885b0 · outbound
Preference learning made easy: Everything should be understood through win rate V., Murray, K., and Kim, Y
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97ad70a4-9085-4d57-ad7e-f3d6738836c3 · outbound
Preference learning made easy: Everything should be understood through win rate Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00f98043-ba0a-4e26-a390-a345ade3e382 · outbound
Preference learning made easy: Everything should be understood through win rate Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e8863dd-e6f4-4a5c-8ebe-e6fde31070dd · outbound
Preference learning made easy: Everything should be understood through win rate Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 862eb67f-a726-4035-9669-b97892d90489 · outbound
Preference learning made easy: Everything should be understood through win rate Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e9610168-ebbd-4b38-b0c3-8b7fab8f24bc · outbound
Preference learning made easy: Everything should be understood through win rate I., and Jiao, J
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4419939a-a0f5-49a7-b640-ecc4d607b2e2 · outbound
Preference learning made easy: Everything should be understood through win rate write newline
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.