Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:32:47.980568Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2502.04357.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:32:47.980568Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4ec1d8ef-30e2-4d74-a1e8-a9787bb5b65f · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Scalable Ensembling For Mitigating Reward Overoptimisation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a45fd102-318c-4a34-ac81-ab1bff7d3bb1 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740ba855-2b7f-4785-8f0a-a938e5a28e0b · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs RLHF Workflow: From Reward Modeling to Online RLHF
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63b25d2b-3b20-4b58-be64-a4dc1fdbd324 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Universal Language Model Fine-tuning for Text Classification
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dc45675-9d25-4c49-be35-f7f3e5710963 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8610c8a-67c7-4f11-894a-7304a70c704c · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c01cb5-4458-43ae-9a23-55364f6abbff · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5469c96b-8161-467c-ab2f-18e5acfa753a · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Generative Reward Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d6b8332-5e9b-4309-bc3c-b8e45aa4390d · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Efficient Estimation of Word Representations in Vector Space
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 535875a3-e053-4e8f-8882-9686a245d96a · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Active Preference Learning for Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2610df-38d2-44ef-95a4-26cf73d61a6f · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0376cf34-9ce7-4b62-a313-8f94310bd832 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d499de53-74e5-4414-9332-c343e3d29346 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff9dfb75-04a8-4166-a8ae-d6d1ef3c2093 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs WaveNet: A Generative Model for Raw Audio
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5842996-e851-443c-bc8d-630bf3aa9e5c · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d3c12fb-96dc-4564-94f4-9bdabdee110e · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 092df6bd-e88b-4aae-885b-4a2b3ff2ef0a · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0de104c-e2b0-423f-8b48-3982a8e8f688 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45a48ba9-d9d9-43f3-b378-08ce8b1b1ee4 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e238bb2b-38b7-4b13-ba6f-0de4e32cb726 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da49c087-30ab-44db-a2ae-fcca4e1ecc80 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Red Teaming Language Models with Language Models
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d98f3d92-adc3-4ff9-ade5-ae8b43c7dd97 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Mitigating the alignment tax of rlhf
Reference 2015
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93890cd6-f7ca-4de8-91e6-db57fbf7eceb · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5258d64d-76db-48d4-accb-5a59a590a8eb · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Reward Model Ensembles Help Mitigate Overoptimization
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84a95113-7295-4417-9801-6b5db597809e · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs LoRA: Low-Rank Adaptation of Large Language Models
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8910208e-b3f3-44f8-a999-9658650905ba · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs OpenAI o1 System Card
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7daa53e-fbc0-42f7-9b19-acbe282a8b69 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Universal Sentence Encoder
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5702f126-caa3-4064-9d95-00667a1ce62f · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54600f9e-9665-49c2-8dac-a1a8bb733340 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs OffsetBias: Leveraging Debiased Data for Tuning Evaluators
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 911ebb87-91d6-44e2-b8eb-35204a5ed526 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e101d40-dd11-4c11-b33c-761b28fc3665 · outbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs The Llama 3 Herd of Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.