Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T23:03:53.312441Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 3 inbound Pith citation observations for arXiv:2502.04270.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T23:03:53.312441Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:34:25.000468Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T01:09:19.294881Z
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd111811-bc5b-49eb-bf8e-46c0ebfbf20b · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b78b1f42-43ce-4d99-8b4a-a286343de7de · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6bedace-62c6-40fa-b99b-9e3ddcc879ca · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling G., Guo, Z
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d366c147-6187-445b-a00b-cad225c9df7c · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 838970c1-5395-4dac-8aa1-c3ae94305c32 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 947a26fa-6a17-409c-b893-4e575b955a2a · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63871527-a189-4f20-b578-52c2574926f1 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Off-Policy Actor-Critic
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 761777f9-ca75-45c0-b01f-9458df276dd2 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Raft: Reward ranked finetuning for generative foundation model alignment
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2c0eab95-ced4-4f35-ac75-f71210376015 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Rlhf workflow: From reward modeling to online rlhf, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8ad2621a-5c92-4661-9397-cf82252950d7 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1150a4df-239e-4938-916c-1c35dff9eef5 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling KTO: Model Alignment as Prospect Theoretic Optimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e40fb42b-1869-4cf7-ae12-60c7794bb127 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Scaling laws for reward model overoptimization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6cc714cd-169f-4aea-b35e-ed6641c58a34 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling L., and Baxter, J
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 91269a7e-356b-475a-a7ee-5548ece0c4cd · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling S., Lillicrap, T., Turner, R
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 45ccfa6a-012b-4955-a9d8-ff9a8851a250 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Direct Language Model Alignment from Online AI Feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9915083f-08c0-4946-94d3-e01d97f12f3e · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b963b863-fbed-47c6-a40a-e41b21635caf · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Reinforcement Learning from Human Feedback with Active Queries
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bdf1632-d609-4672-8d08-64d1603b2c39 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling and Uehara, M
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fd3d0c60-2e79-4346-8888-ecfcda468890 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 84c754f6-4510-434f-9f67-a9ebe0961861 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f6dfccbc-8622-4394-a503-5663da3b61b8 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Let's Verify Step by Step
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0405fa58-2304-4dc4-b7ec-12ea28d1299a · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ed352f6a-cebd-4543-a119-bf15d9a7874b · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ae139a1-50fd-426e-833d-3e0c37fbb40f · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Decoding-time Realignment of Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43773dbb-d1c3-4087-a2c6-1ad8d87638cc · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling J., and Liu, J
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d5bd172c-6ebc-44d1-a806-25d0f200b1b2 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Sample-Efficient Alignment for LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21a221cf-c917-4b45-a6cd-1a665e0ac4c6 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Sample Efficient Preference Alignment in LLMs via Active Exploration
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 391fe1c4-7614-4114-af2f-2c2254b3ff68 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Active preference learning for large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4c2ea069-f21c-4644-a552-e539434b2813 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Training language models to follow instructions with human feedback
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4e0b24-88bd-4097-b46c-05ca1cf15a58 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling D., Ermon, S., and Finn, C
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31a84014-6480-4388-aaaa-8d299a07ea28 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Optimal Design for Reward Modeling in RLHF
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation addf575c-03f8-4356-b91d-4568330cad1a · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Proximal Policy Optimization Algorithms
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00c51ad8-6c46-4b97-819d-695d06df7778 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling The Crucial Role of Samplers in Online Direct Preference Optimization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94df5939-c90a-40eb-8662-ed50ecc27064 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Deterministic policy gradient algorithms
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c0876af-7dad-40b8-a74d-14472e586c96 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling S., McAllester, D., Singh, S., and Mansour, Y
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f14a3650-1869-402e-8243-0d3f6210273e · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 271913c1-26bf-46d4-b332-b737d0649152 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c180d6a-0f93-4327-9a29-0f37aeb3f691 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50d3db4a-b21c-490e-9b23-72a1243ea515 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2179c1ad-d998-4e90-9b57-cec8f33112ff · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 64b1dcd3-0cac-4cc7-a338-22b0006d5f07 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 564d187d-2518-475f-9d04-2b782da60d06 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Inference Scaling for Long-Context Retrieval Augmented Generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c09f1447-2d90-4e96-bf86-e8d9d5c36f80 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a929772c-ed3a-491f-bbb6-b01a47469ff6 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling @esa (Ref
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c99ebb6-acea-4927-83e3-1c108421f5c1 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9d65432-a03b-4fc4-937b-532a80d351f6 · outbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afdb69b1-e454-4c5a-b4ef-a50ba8294804 · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities PILAF: Optimal Human Preference Sampling for Reward Modeling
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17a7ff4f-ca29-4304-8976-aa512fab2665 · inbound
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses PILAF: Optimal Human Preference Sampling for Reward Modeling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation edfe710a-2dec-40be-a4fb-e9c1886a31d9 · inbound
Which Pairs to Compare for LLM Post-Training? PILAF: Optimal Human Preference Sampling for Reward Modeling
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.