Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T17:46:29.322067Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 3 inbound Pith citation observations for arXiv:2502.00814.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T17:46:29.322067Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-19T09:40:58.067398Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T09:42:14.018543Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2e56f5ee-7e9f-43eb-a339-7cec12a716b0 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e687054-2d8a-4129-afdf-829a66dce354 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d39e875-6eee-42c7-b8b1-e0d907b9e23a · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1851354-1fa4-40e3-b08d-319ce8a0f1b7 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 639fa08c-a265-4ea5-8b7e-60be1ec09e59 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 895328f6-ddd7-41ce-bd66-3f7585ec7d44 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, and 13 others
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5e2afa0f-02b2-4f15-b42a-8db7b9025f3f · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Noise Contrastive Alignment of Language Models with Explicit Rewards
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b53fb7d2-d286-4eab-9a26-13e94c5e1500 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 90e4d899-9355-48da-9094-f0bca6d4f7ff · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01613714-7209-4ffd-9a3b-3301336890b3 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43b1f528-7d4a-43d9-a2b4-282d6622d2cb · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b68a731-4f52-4543-82ab-7af54c4d0614 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e9650ee8-d8bd-418b-aa70-d80e0581fa81 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling KTO: Model Alignment as Prospect Theoretic Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32ba3a61-ea14-4c2d-a423-e7d232e4e000 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48bd23ca-8018-44b6-bd0e-0f42d18fcf4d · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6f09260-d306-4ede-8fe5-6f7ccad15201 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a113081c-5738-4ae5-b925-73739651f750 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9fa1fa6-77fe-4932-b1a5-42ca979a86ca · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling GPT-4o System Card
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93666958-02ac-4f04-aad1-5e24a011bbba · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8189cc2d-8c2e-48a8-a87d-38025becc9b4 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Smith, Yejin Choi, and Hannaneh Hajishirzi
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6f5e6ebf-4922-4756-ba64-7020c4f606e1 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8781ca82-6335-4fac-8dfc-b7b335b52a1c · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfd90ed0-67ab-4ecf-a737-3d6991d3750e · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling o pf, Yannic Kilcher, Dimitri von R \
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c6244b72-6611-425a-9530-51d8f06b9be6 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d995181b-647d-4fa4-aa1f-7ba9fa2467c6 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 18fe0b8c-1553-4484-875b-0d72e477049e · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c49e95d-6763-498a-b9ae-8338a9fdac8a · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling RewardBench: Evaluating Reward Models for Language Modeling
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deed4d9b-04f2-4c85-aaf5-3b5df0d3df2c · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93579785-6e5e-480c-98ef-d69d6e2fc840 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f75e4305-dbf3-4c81-9146-00684ed25f74 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86bd11ef-ef1d-460a-9c8a-e014d7947641 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16dc55fe-fe52-4719-84a2-43320d48ddf1 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132237d9-77a9-4f51-8cea-c6efd0fd23e7 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5bb56b2a-80fe-4d09-bf07-d307e47eea2f · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Disentangling Length from Quality in Direct Preference Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1da4260-4dd8-41f0-beab-a4499c2decb1 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74ed2454-363a-430c-b1c0-d2f248191382 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988612c6-fb9c-4508-8a18-3d04be348795 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ff9b348c-7716-48ce-b735-f284b88385fd · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd98b722-028c-4980-8079-1ceae83db87f · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling WARM: On the Benefits of Weight Averaged Reward Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 372f1231-5464-4e35-a624-9d8cdafa3132 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Proximal Policy Optimization Algorithms
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f7d9d3-3b58-4538-a665-abd9caa2bc19 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20d7361c-5354-4e99-83d3-a0503e86928a · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5e01346b-9df2-4ca3-b2cd-0c293fc45593 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling A Long Way to Go: Investigating Length Correlations in RLHF
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1566380e-2250-4489-8b79-5214692ccddc · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 51a0f741-798b-4b3a-aec1-31155760caea · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c63ca5-29e6-45f7-be52-9c311174f437 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Gemini: A Family of Highly Capable Multimodal Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f4e3dc2-949e-4e33-a572-987686e0971a · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ef00cd-320d-4638-9330-b947142d06bf · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98cbafc3-e6d1-4391-ad86-82a413e2baba · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1a0f8728-8f31-4002-8b7b-e39d4c193fba · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Qwen2 Technical Report
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae705c90-d7b1-4065-86eb-f4a8027ba4af · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 671d4dca-2f6d-4b70-9a05-76c3579da0d1 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Following Length Constraints in Instructions
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83b96025-ddea-4fdc-8cb8-39c29ce3ccd9 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c349acb-28db-4358-b755-102c3c8626a4 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37292ebc-18f4-4ea9-91dd-c8adcabab1c6 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Fine-Tuning Language Models from Human Preferences
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34222efa-989c-4f61-ab43-5acd0b07059e · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling online" 'onlinestring :=
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f722cb3-3e73-4e09-abb8-92656311dea5 · outbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling write newline
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c42889-d3d5-4bcb-b8e2-d9ce36aaa618 · inbound
Exploring the Secondary Risks of Large Language Models Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 35d456f8-5298-4500-804d-5153ae5aaa8e · inbound
Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7f107fd6-152e-4728-859a-cb51dd97adeb · inbound
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.