Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:12:52.810257Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 10 inbound Pith citation observations for arXiv:2505.10527.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:12:52.810257Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:25:22.019686Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T00:07:27.940098Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dd4945dd-82cb-492c-9114-082e1378e27d · outbound
WorldPM: Scaling Human Preference Modeling A General Language Assistant as a Laboratory for Alignment
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fe14baf-8470-46ee-82a7-111fb8396b34 · outbound
WorldPM: Scaling Human Preference Modeling Program Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 038d8ad6-2151-455d-86af-2d48750a302f · outbound
WorldPM: Scaling Human Preference Modeling Qwen Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78d59d92-fe07-4506-a9a3-316ac83f7753 · outbound
WorldPM: Scaling Human Preference Modeling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4ab421e-f427-409a-b4c2-39ffd8ef1452 · outbound
WorldPM: Scaling Human Preference Modeling Constitutional AI: Harmlessness from AI Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d18cef2f-043a-4af2-aeb4-2142cd93b39f · outbound
WorldPM: Scaling Human Preference Modeling Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c15523b1-4a19-4dcf-a6f0-b574cf8f1a8c · outbound
WorldPM: Scaling Human Preference Modeling Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b407869-0a74-468d-ad8a-03bf4a5f4c32 · outbound
WorldPM: Scaling Human Preference Modeling Deep reinforcement learning from human preferences
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2cca691-02e8-4761-a1bc-2e5f04746e33 · outbound
WorldPM: Scaling Human Preference Modeling Reward Model Ensembles Help Mitigate Overoptimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c8dfd03-b2c8-42ba-9770-a3af028f6b7f · outbound
WorldPM: Scaling Human Preference Modeling Cram \'e r
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9e2dc8e8-9bd6-42ee-b1f0-a45a48687731 · outbound
WorldPM: Scaling Human Preference Modeling UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25581d1d-a1c4-44d6-bdf1-a56e93872eb4 · outbound
WorldPM: Scaling Human Preference Modeling The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5c2e0d6-c2b0-4cdd-81e0-35e196432e1a · outbound
WorldPM: Scaling Human Preference Modeling Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9edae22-3c0b-46a6-8a2f-6bcc74c1bbb9 · outbound
WorldPM: Scaling Human Preference Modeling Networks, crowds, and markets: Reasoning about a highly connected world, volume 1
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a1e8dd68-4f2c-469a-9845-439d8b3b5b15 · outbound
WorldPM: Scaling Human Preference Modeling Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56419043-d0e3-4983-ad43-cebbd94833a0 · outbound
WorldPM: Scaling Human Preference Modeling Athene-70b: Redefining the boundaries of post-training for open models, July 2024 a
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246eac89-1703-458f-8d1c-c5afa4793d67 · outbound
WorldPM: Scaling Human Preference Modeling How to Evaluate Reward Models for RLHF
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64056f1d-7309-43b0-bd39-97293d856249 · outbound
WorldPM: Scaling Human Preference Modeling Scaling laws for reward model overoptimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1a25115-ff8c-4727-84fa-310136e56e9e · outbound
WorldPM: Scaling Human Preference Modeling Zemel, Wieland Brendel, Matthias Bethge, and Felix Wichmann
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 60c429d7-0e1b-476c-9833-a885d173e835 · outbound
WorldPM: Scaling Human Preference Modeling Measuring Mathematical Problem Solving With the MATH Dataset
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3d2ec17-2845-421d-8bef-be598663dc91 · outbound
WorldPM: Scaling Human Preference Modeling Surface Form Competition: Why the Highest Probability Answer Isn't Always Right
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9b9f065-fc85-4afb-a579-b93c7823b9f4 · outbound
WorldPM: Scaling Human Preference Modeling Collaborative filtering for implicit feedback datasets
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6caf8d8f-3628-4ef2-8257-9cd018587043 · outbound
WorldPM: Scaling Human Preference Modeling Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2e0d544-b10c-42bb-8697-0bb6cdb470f1 · outbound
WorldPM: Scaling Human Preference Modeling Scaling Laws for Neural Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37859d0e-2136-4685-83c0-50360ea94db8 · outbound
WorldPM: Scaling Human Preference Modeling Smith, and Hannaneh Hajishirzi
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dc354db-4daf-4c07-b070-3bed702d1ab3 · outbound
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 00b5977d-921c-4773-ab1c-d00d38674d54 · outbound
WorldPM: Scaling Human Preference Modeling Gonzalez, and Ion Stoica
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 158002a8-73d6-4172-9fd1-c2b51571f46c · outbound
WorldPM: Scaling Human Preference Modeling ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01faa9ca-2b6a-43f3-b60a-92979cf005ca · outbound
WorldPM: Scaling Human Preference Modeling RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fd180ff-098e-42b7-a39c-3b787a326217 · outbound
WorldPM: Scaling Human Preference Modeling OctoPack: Instruction Tuning Code Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d186b9f-d975-4f0a-8f43-6bebc313e223 · outbound
WorldPM: Scaling Human Preference Modeling Training language models to follow instructions with human feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d62c5174-d338-49de-b53a-c6df2d7b9478 · outbound
WorldPM: Scaling Human Preference Modeling Maxwell Harper, and Joseph A
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c0978775-496a-4fb0-932b-a78846273445 · outbound
WorldPM: Scaling Human Preference Modeling OffsetBias: Leveraging Debiased Data for Tuning Evaluators
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb293d53-4df8-4b20-9de1-5c0af81d3f49 · outbound
WorldPM: Scaling Human Preference Modeling GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad4e9b09-21d8-48c5-9413-65c7648b1373 · outbound
WorldPM: Scaling Human Preference Modeling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e3d2f2f-2adb-418a-816b-b262f2d578a9 · outbound
WorldPM: Scaling Human Preference Modeling Learning to summarize with human feedback
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46458c3d-2c32-4ca9-b8fb-52607b01e985 · outbound
WorldPM: Scaling Human Preference Modeling Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a5b55ef-5e5e-4d10-8554-cb69fc0a421b · outbound
WorldPM: Scaling Human Preference Modeling Interpretable preferences via multi-objective reward modeling and mixture-of-experts
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41abadb6-c16f-48de-b069-8cef397bfd3d · outbound
WorldPM: Scaling Human Preference Modeling MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caed08d2-caed-4eec-b14c-a38746b092d5 · outbound
WorldPM: Scaling Human Preference Modeling HelpSteer2: Open-source dataset for training top-performing reward models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d4ad2b1-874d-4be8-a21c-f2565d598da1 · outbound
WorldPM: Scaling Human Preference Modeling Emergent Abilities of Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01a55ba0-bc37-4935-808b-11092538745d · outbound
WorldPM: Scaling Human Preference Modeling Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f256540d-e209-442c-9c08-be02850d0fe1 · outbound
WorldPM: Scaling Human Preference Modeling Qwen2 Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad59b7f5-6717-460b-8795-56c05bbad3ca · outbound
WorldPM: Scaling Human Preference Modeling Qwen2.5 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e9dce5-c17c-4dce-b01d-3cb7d9ee180b · outbound
WorldPM: Scaling Human Preference Modeling Evaluating large language models at evaluating instruction following
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3094a929-2a9b-4494-ae38-dc2201a933f1 · outbound
WorldPM: Scaling Human Preference Modeling Understanding deep learning requires rethinking generalization
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae01066-6b79-480c-b369-c48cf36c3d1e · outbound
WorldPM: Scaling Human Preference Modeling Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27dd7f08-22fb-434a-a112-2a496e1054c5 · outbound
WorldPM: Scaling Human Preference Modeling Secrets of RLHF in Large Language Models Part I: PPO
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 475d6f77-da7c-453e-975c-47bc36c51078 · outbound
WorldPM: Scaling Human Preference Modeling RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 234cd9c5-2b91-4154-b177-b17023842e3e · outbound
WorldPM: Scaling Human Preference Modeling Instruction-Following Evaluation for Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8678d4c-cfa5-494a-bfed-fd0f8a36e948 · outbound
WorldPM: Scaling Human Preference Modeling write newline
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b40fdb6-cf0b-46c9-bb56-43258e14e66d · outbound
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0420f27-bb26-4d8c-8baa-85f4abc1d253 · outbound
WorldPM: Scaling Human Preference Modeling Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187be880-bd1c-40ca-90cd-1d02e2f9bfb0 · outbound
WorldPM: Scaling Human Preference Modeling D sRGB DeXIfMM* i ͠ u @IDATx ]7 1ϳP<桔c iu0jP3
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ffe7d007-f520-4ac2-ab6f-069bcd1cb401 · inbound
Arch-Router: Aligning LLM Routing with Human Preferences WorldPM: Scaling Human Preference Modeling
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6887f61-aa82-4a10-9f0e-16a07b78309b · inbound
RewardDance: Reward Scaling in Visual Generation WorldPM: Scaling Human Preference Modeling
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f423c084-1aae-4047-82a9-ee19963d9dc0 · inbound
AI Can Learn Scientific Taste WorldPM: Scaling Human Preference Modeling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7f984f2-8052-45cb-abbe-747d8cfe1a63 · inbound
Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization WorldPM: Scaling Human Preference Modeling
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d8ebd339-9560-40c2-9808-0fb0e0b52c20 · inbound
Leveraging Verifier-Based Reinforcement Learning in Image Editing WorldPM: Scaling Human Preference Modeling
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b54261bb-8481-4f67-9bc9-04a6944e3572 · inbound
Leveraging Verifier-Based Reinforcement Learning in Image Editing WorldPM: Scaling Human Preference Modeling
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2619ad0c-fc3c-45c9-8374-4ed6ab95b580 · inbound
RewardHarness: Self-Evolving Agentic Post-Training WorldPM: Scaling Human Preference Modeling
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7808e38c-29ba-4e74-a93b-4aad4a8ebd20 · inbound
Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions WorldPM: Scaling Human Preference Modeling
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2d25d118-c7e0-45f1-8ef5-4658c4a9af16 · inbound
Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions WorldPM: Scaling Human Preference Modeling
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e834e349-dfe1-4c35-b6a7-699bc4d1b514 · inbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction WorldPM: Scaling Human Preference Modeling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.