Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:48:23.624843Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 2 inbound Pith citation observations for arXiv:2412.03187.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:48:23.624843Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:51:52.719416Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T20:11:10.907552Z
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fa2df7e7-2a38-4ae8-81f6-a448f0adc125 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c76c5867-7d6f-4880-b9a9-a7566584a428 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion A general theoretical paradigm to understand learning from human preferences
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bdb6a074-362c-4ba7-9477-72c4a49ccc0b · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Open LLM leaderboard, 2023
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e9b26107-200c-4ea6-8594-0de858eda617 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 987c2692-9544-4c07-b914-d8402c1e2058 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion InternLM2 Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 677a8ea8-e4b1-471b-9057-b5df9428d530 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Chatbot arena: An open platform for evaluating llms by human preference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 75da1c39-57b2-44ae-9b86-3e9eeef209c1 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Deep reinforcement learning from human preferences
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8efe2630-5fd6-480c-a619-03799a56b386 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Bam! born-again multi-task networks for natural language understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4263d49a-429a-4b04-bd54-bb3e996081d2 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dec28621-fc3c-495a-9e78-847e6f835a64 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Training Verifiers to Solve Math Word Problems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99746a63-d289-449a-850a-58eafb68ee0e · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion UltraFeedback : Boosting language models with high-quality feedback
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4775cba0-af72-4f46-8be1-4022df2dfb1e · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4205d85a-e57f-482f-a4ee-369bee78b7b4 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a5d6cd-c4ec-4816-86ff-08c7e97c2fbb · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Length-controlled alpacaeval: A simple debiasing of automatic evaluators
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8da83c02-4804-4967-9bb1-7eb79634432d · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion KTO : Model alignment as prospect theoretic optimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2baf45e3-cc3b-4972-a0f4-44b46327981e · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Mixture-of-LoRAs : An efficient multitask tuning method for large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b06abde6-cfcb-432d-8c9a-76e6b9a41b08 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Knowledge distillation: A survey
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f363a1-e440-4dc5-928a-8ccb7c20441c · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Measuring massive multitask language understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa7b7eef-46de-4541-a4d0-aab10c131917 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion ORPO : Monolithic preference optimization without reference model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 54fed2fb-f0d6-479c-9507-7c9b79833f0c · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Mistral 7B
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bb2d512-e04d-45e5-ba3b-8329bef43f20 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion LLM-Blender : Ensembling large language models with pairwise ranking and generative fusion
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 15c4804b-f6ab-4a43-9567-4db7889c7fe1 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Knowledge-augmented reasoning distillation for small language models in knowledge-intensive tasks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fd62c6fa-1447-42cd-9bfc-d72675134404 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Sparse upcycling: Training mixture-of-experts from dense checkpoints
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77196ec9-fd79-4e11-b815-dd2375a1c2a3 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion The Winograd schema challenge
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6fe8371d-ad98-4b69-a366-edf870852903 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff6ab5db-0b08-48ac-9aab-9da365e0c9c6 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Hashimoto
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bbb2b5b4-11a8-4b72-b300-044db28310ad · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion TruthfulQA : Measuring how models mimic human falsehoods
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f0d68b7d-1d6c-48be-8a21-650a12966267 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Merging models with fisher-weighted averaging
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f8078d2c-c058-4bb7-9de5-059fc087da29 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Pack of LLM s: Model fusion at test-time via perplexity optimization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bd9bc18c-2682-4fd9-836a-1b4bf1f82c66 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Sim PO : Simple preference optimization with a reference-free reward
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3e6431bc-985d-43d4-955b-60e0b6ec7616 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Disentangling Length from Quality in Direct Preference Optimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b82278-b6b4-4165-94e5-789218940319 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Manning, and Chelsea Finn
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 79fa6a87-56a2-4519-a243-40057d8fa9ce · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Aligning large and small language models via chain-of-thought reasoning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2d02388b-eada-4469-abc7-d2ef482a2652 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Gemma 2: Improving Open Language Models at a Practical Size
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95793e4b-f7db-4d41-b673-26c64e9f09cd · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Proximal Policy Optimization Algorithms
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e92aa4d9-474d-4f0c-9bc9-6f4ec3430183 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c75fea1-e282-4743-8301-7dbb6bd5d184 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion ProFuser : Progressive fusion of large language models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48ce57d4-346e-44b1-a248-00e74e220fa3 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Branch-Train-MiX : Mixing expert LLM s into a mixture-of-experts LLM
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5cb99833-72f3-4def-964f-2e007264acdd · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Preference fine-tuning of LLM s should leverage suboptimal, on-policy data
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 53601485-ab76-4eb5-b01a-d99134fecfb9 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbeb0572-5523-4118-999c-ed60c127d4cc · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Zephyr: Direct Distillation of LM Alignment
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db7ad270-3c0a-4181-bbf1-f9b9a35acae2 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Knowledge fusion of large language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7a3dab9d-288c-4681-b7b4-7da070f5dcbd · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Knowledge Fusion of Chat LLMs: A Preliminary Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8f05aee-af0b-445f-8e51-f5f5f1d691e1 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Interpretable preferences via multi-objective reward modeling and mixture-of-experts
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2a071f56-a858-42a1-a01a-5fb58ae6dcfe · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Mixture-of-Agents Enhances Large Language Model Capabilities
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8c57ab6-079c-48fc-8754-5faf0619f4f0 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion HelpSteer2: Open-source dataset for training top-performing reward models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c16990f-2056-473d-8554-5e27ca8829b0 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 508ae4d5-6415-4a6a-b645-a15ab15aebbd · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 19a098c0-1432-4deb-aeb6-0f4256f4e83e · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Is DPO superior to PPO for LLM alignment? A comprehensive study
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 003f0bc0-fab0-4a9b-99c9-63d6f8904f28 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Bridging the gap between different vocabularies for llm ensemble
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a8f40649-ac73-4eff-bd30-c88c3b870f7f · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Ties-merging: Resolving interference when merging models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3a4784d0-ecf0-4726-ac29-9a011c6c8908 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Qwen2 Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05fdfed6-020a-4d50-a75c-2237f08eba29 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Yi: Open Foundation Models by 01.AI
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8510248-c7dd-4e51-9425-be56280da948 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion RRHF : Rank responses to align language models with human feedback
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7bc69922-b776-439d-84b2-10e5fb0a2e97 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 4791--4800, 2019
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0f707a22-fc9c-4ed6-90c0-2268f4aafa11 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73386b1e-aaac-4504-9d68-1da14a69894e · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Judging LLM -as-a-judge with MT-Bench and Chatbot Arena
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a9eae4d2-50ac-4d9d-aaf1-130ba050a408 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion WPO : Enhancing RLHF with weighted preference optimization
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f3890c95-d620-4ec9-8fb3-9e6d9cd8036f · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Starling-7B : Improving helpfulness and harmlessness with RLAIF
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8fa8b405-7cdd-406c-857f-c41f0bfcb79f · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Fine-Tuning Language Models from Human Preferences
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ffb35a9-2081-4577-9e53-59e0adfdc756 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion write newline
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16127d6e-d884-4d43-a4b4-7dfe77098d2b · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion @esa (Ref
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834b00d9-0ed4-4f10-a813-3308cfe5dc31 · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Unresolved cited work
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d18745-b8f2-4f41-94bf-afcf3eaf9f6e · outbound
Weighted-Reward Preference Optimization for Implicit Model Fusion Unresolved cited work
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0e3fefa-5e6f-441f-8143-9a1d15942330 · inbound
An Overview and Discussion on Using Large Language Models for Implementation Generation of Solutions to Open-Ended Problems Weighted-Reward Preference Optimization for Implicit Model Fusion
Reference 143
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 048901a7-3233-4835-a467-f5e9c1f57323 · inbound
CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate Weighted-Reward Preference Optimization for Implicit Model Fusion
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.