Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:01.415106Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 4 inbound Pith citation observations for arXiv:2505.12843.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:01.415106Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T07:28:18.050097Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-29T12:23:24.212503Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 687059e7-1487-40a1-974e-91027e8af4b3 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF A General Language Assistant as a Laboratory for Alignment
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca65930-4c38-49aa-9e68-b3412f374ba1 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f12f60ed-7165-4c6e-a46e-98f9aba21d7e · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Rank analysis of incomplete block designs: I
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8868765d-b6c9-4a2f-b86d-41ada71ae417 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Noise contrastive alignment of language models with explicit rewards, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 68bc6e29-fd45-4f2f-b856-df39cac60120 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF ODIN: Disentangled reward mitigates hacking in RLHF
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e1febd69-ddd5-41a2-9081-693e79f9f182 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbd7e727-5072-40ec-88c6-d92c1175db1b · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Deepseek-v3 technical report, 2025
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1d399348-7ffb-40c2-bc44-ac4d26134666 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96ae6cc-6e97-4c26-b659-3484f7806aa1 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Rlhf workflow: From reward modeling to online rlhf, 2024
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92b36bd0-66db-4942-a8bb-206cd3b71f28 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 895ee927-9f17-4fa9-96fb-e185860aa0ce · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Helping or herding? re- ward model ensembles mitigate but do not eliminate reward hacking
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5175fce3-dbec-4130-828f-6b4db048fca5 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF KTO: Model Alignment as Prospect Theoretic Optimization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00a8a2f2-ff85-4c42-92a9-3563c8715d8c · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Scaling laws for reward model overoptimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9f39062b-9580-4429-8b48-74aa5924b5a6 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF A general theoretical paradigm to understand learning from human preferences
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 091fd5b7-8e92-4857-99b4-3fe00829aab4 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF The llama 3 herd of models, 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd0117c1-6cf9-4226-8cb2-79708f860e0e · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe66a8c7-218f-4903-9363-d582253f066f · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Deep residual learning for im- age recognition
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 74341c34-490e-4297-a989-a09a27556a28 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF ORPO: Monolithic preference optimization without reference model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 461bde1b-68cd-415b-9a4b-068b4cb0c487 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Post-hoc reward calibra- tion: A case study on length bias
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 78169754-008c-4a47-93ab-923e4ca0c041 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Gonzalez, Hao Zhang, and Ion Stoica
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 021ab3bf-76ee-4e76-89d0-184714636678 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Openassistant conversations -democratizing large language model alignment
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2d3a95c6-e250-4a0b-8ff2-59dfa125a0ef · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73322ea2-4061-4ac9-aab0-9f0e9fb8f00b · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF RRM: Robust reward model training mitigates reward hacking
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ab8961fa-4a15-4bd6-b065-bbda41a69a15 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Simpo: Simple preference optimization with a reference-free reward
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a758781c-ef87-4653-ab11-23cc8fff87c8 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Expanding on what we missed with sycophancy, 2025
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 77de0340-a9e3-4520-94d6-15af306c4b2d · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Gpt-4 technical report, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5cd313b5-52c4-4f27-a8c4-4934d7f0dab3 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a0d94a4-fffc-452d-9a16-c2a02c146196 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Reward gaming in conditional text generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2627ea87-679b-48fb-893d-a7fc06ed9222 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF PyTorch: An Imperative Style, High-Performance Deep Learning Library
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d92e127b-ad5e-4f4c-a8c4-19e72c3bb138 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Notes on regression and inheritance in the case of two parents.Proceedings of the Royal Society of London, 58:240–242, 1895
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 37e87cc3-4853-45bc-90ff-18db27b3ab96 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Qwen2.5 technical report, 2025
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9a55eee8-8a33-4a5b-aa8e-d25da946b6d6 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Direct preference optimization: Your language model is secretly a reward model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b0b67e2-78f0-45f2-9c4e-f6777a4f5f11 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Warp: On the benefits of weight averaged rewarded policies, 2024
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c25c3d-c1a6-4e36-a4ea-024a25c267d7 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Warm: On the benefits of weight averaged reward models, 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d6da3a41-b5bf-438e-a7b8-50564ac0918e · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0b795c10-9d72-4afb-9121-51894a263dde · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Offline regularised reinforcement learning for large language models alignment, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 18fb0222-e90a-43ce-9ef6-c3ceb3a0cfba · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Proximal Policy Optimization Algorithms
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b50415-f300-4540-aae8-db9f96f76aee · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF BOND: Aligning LLMs with Best-of-N Distillation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a72128c-6975-459e-8d5c-fba75b26620a · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbd4fdcf-3d6d-4e65-96e9-9abeffdcf289 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f0ccf053-ec92-44ca-8847-175f6aa45da3 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF A Long Way to Go: Investigating Length Correlations in RLHF
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d44421d-76c1-4dc6-94a5-d28e21e8aac4 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6973f78-33e7-4b84-83db-5d1a703a0c78 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Learning to summarize with human feedback
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4a46a08-4386-4d1e-86f5-7e78004053c8 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 75dc0f5f-8422-4368-853e-d3a0f1184e20 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Llama 2: Open foundation and fine-tuned chat models, 2023
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff50780c-1bc8-4d80-824d-ddc9add43235 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Attention is all you need
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a4b8528-1e2e-42d9-8532-c0aef22b84d4 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Reward hacking in reinforcement learning, Nov 2024
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 08881065-5386-42bb-9196-9ca184a1626c · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Qwen2 technical report, 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe08b45c-ef18-4e96-8f0e-097c11063e79 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Dapo: An open-source llm reinforcement learning system at scale, 2025
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 947656c9-9a9e-41d4-9a11-c51b5d8c3a15 · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF From lists to emojis: How format bias affects model alignment, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b184816d-d7cd-4090-838b-07fa64a1387e · outbound
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Fine-tuning language models from human preferences, 2020.URL https://arxiv
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 83119876-bf91-4161-9c30-bee51499f1ec · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 01dc5012-628e-4f96-8630-5ab77989de2c · inbound
Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a1c70b58-51db-4535-98ad-f9d2cd75863d · inbound
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 12f1842d-a9bb-4ac9-a585-9b1f71f32c38 · inbound
Test-Time Scaling via Error Localization Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.