Pith. sign in

Paper Citation Record · LEDGER

Weighted-Reward Preference Optimization for Implicit Model Fusion

As of 18 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 2 inbound Pith citation observations for arXiv:2412.03187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03187 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:48:23.624843Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:51:52.719416Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T20:11:10.907552Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa2df7e7-2a38-4ae8-81f6-a448f0adc125 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Weighted-Reward Preference Optimization for Implicit Model Fusion Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:20.804753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:20.804753Z digest=sha256:02d28e9881da8127d488c140d83f1c415d324a79f0d02d8127eb9b38d67fb194

Observation c76c5867-7d6f-4880-b9a9-a7566584a428 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Weighted-Reward Preference Optimization for Implicit Model Fusion A general theoretical paradigm to understand learning from human preferences

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:33.334877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:20.835140Z digest=sha256:dd44c3f3896e8dee4aa0bf64d55f9cf26caec95e4200854a847ea6b721a519fb

Observation bdb6a074-362c-4ba7-9477-72c4a49ccc0b · outbound

This paper cites Open LLM leaderboard, 2023.

Weighted-Reward Preference Optimization for Implicit Model Fusion Open LLM leaderboard, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:33.174881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:20.884774Z digest=sha256:1a06f6fe554ad87961e3c980f453bc93363443b183667e4a7621abd32c76c609

Observation e9b26107-200c-4ea6-8594-0de858eda617 · outbound

This paper cites an unresolved cited work.

Weighted-Reward Preference Optimization for Implicit Model Fusion Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:48:32.949019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:20.934756Z digest=sha256:64d55882590bc10b41c212e3b22e3b5ab5033138f9a19e915fa7f81762a88299

Observation 987c2692-9544-4c07-b914-d8402c1e2058 · outbound

This paper cites InternLM2 Technical Report.

Weighted-Reward Preference Optimization for Implicit Model Fusion InternLM2 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:20.994770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:20.994770Z digest=sha256:1aa4a663295ee364822a68f7548dd5fa4324471f93b870c135511fb25de7d820

Observation 677a8ea8-e4b1-471b-9057-b5df9428d530 · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

Weighted-Reward Preference Optimization for Implicit Model Fusion Chatbot arena: An open platform for evaluating llms by human preference

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:32.842939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.044765Z digest=sha256:1d66f7d001dc660ce7e612b997e9ed170461c7ab0e31280ed6fd130c574ada98

Observation 75da1c39-57b2-44ae-9b86-3e9eeef209c1 · outbound

This paper cites Deep reinforcement learning from human preferences.

Weighted-Reward Preference Optimization for Implicit Model Fusion Deep reinforcement learning from human preferences

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:32.691427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.104765Z digest=sha256:84009bdcbcaccfb80ed269ccee3793e69ce3844858d343d86f0ce877fc97b918

Observation 8efe2630-5fd6-480c-a619-03799a56b386 · outbound

This paper cites Bam! born-again multi-task networks for natural language understanding.

Weighted-Reward Preference Optimization for Implicit Model Fusion Bam! born-again multi-task networks for natural language understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:32.555244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.144759Z digest=sha256:4a1ea3e7e95612abaf834d5090db20534be36ff5162de5528fa5a41286988e94

Observation 4263d49a-429a-4b04-bd54-bb3e996081d2 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Weighted-Reward Preference Optimization for Implicit Model Fusion Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:21.194759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:21.194759Z digest=sha256:23668b936bcdc1bd20f9e3faf4ecc6c43ae25ddb3f0496cf0862e3a5a3aa5f96

Observation dec28621-fc3c-495a-9e78-847e6f835a64 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Weighted-Reward Preference Optimization for Implicit Model Fusion Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:21.244866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:21.244866Z digest=sha256:0209df321dd7e4b6c8f3cc4d92afb46220252b855284c72add31b28d17ac8044

Observation 99746a63-d289-449a-850a-58eafb68ee0e · outbound

This paper cites UltraFeedback : Boosting language models with high-quality feedback.

Weighted-Reward Preference Optimization for Implicit Model Fusion UltraFeedback : Boosting language models with high-quality feedback

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:32.434853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.284752Z digest=sha256:49db1493ac4a4a747ca44781acae4d6db1e890036ff503ce3fd0ec23da28a27f

Observation 4775cba0-af72-4f46-8be1-4022df2dfb1e · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

Weighted-Reward Preference Optimization for Implicit Model Fusion DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:21.328808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:21.328808Z digest=sha256:3eb813c1f00af4f2d7c96a65a48699aea575397c6d16a63fed960dcaca408227

Observation 4205d85a-e57f-482f-a4ee-369bee78b7b4 · outbound

This paper cites The Llama 3 Herd of Models.

Weighted-Reward Preference Optimization for Implicit Model Fusion The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:21.377609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:21.377609Z digest=sha256:fd2957bf1e441a6aa8d79296e35eadd6d7b4e1cd822761b0554f7dff1daa1bff

Observation 20a5d6cd-c4ec-4816-86ff-08c7e97c2fbb · outbound

This paper cites Length-controlled alpacaeval: A simple debiasing of automatic evaluators.

Weighted-Reward Preference Optimization for Implicit Model Fusion Length-controlled alpacaeval: A simple debiasing of automatic evaluators

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:32.266933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.434760Z digest=sha256:cd4bc33f07e1c0a2952b7cdc5800d3bc7118a9eff40fd8d26401f932d2dd428f

Observation 8da83c02-4804-4967-9bb1-7eb79634432d · outbound

This paper cites KTO : Model alignment as prospect theoretic optimization.

Weighted-Reward Preference Optimization for Implicit Model Fusion KTO : Model alignment as prospect theoretic optimization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:32.035409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.482693Z digest=sha256:7f0ac22cb423f852efc6bbb5485ff098f92cda3f39cefbb3caf39a549460192d

Observation 2baf45e3-cc3b-4972-a0f4-44b46327981e · outbound

This paper cites Mixture-of-LoRAs : An efficient multitask tuning method for large language models.

Weighted-Reward Preference Optimization for Implicit Model Fusion Mixture-of-LoRAs : An efficient multitask tuning method for large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:31.862767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.524750Z digest=sha256:11a3db9647a8e06f7f03335e2f677c45c39bc496796d59479a3ddb675f1f8574

Observation b06abde6-cfcb-432d-8c9a-76e6b9a41b08 · outbound

This paper cites Knowledge distillation: A survey.

Weighted-Reward Preference Optimization for Implicit Model Fusion Knowledge distillation: A survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:21.564921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:21.564921Z digest=sha256:897cb7a0380f008acb47e2e44946c4ee1ad4c6f976eab574c85cb0ec42c932f7

Observation a1f363a1-e440-4dc5-928a-8ccb7c20441c · outbound

This paper cites Measuring massive multitask language understanding.

Weighted-Reward Preference Optimization for Implicit Model Fusion Measuring massive multitask language understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:21.603599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:21.603599Z digest=sha256:e090471cef050a26e4ff2ef98e1e143d8ba703c15636101a29d77073109931fc

Observation aa7b7eef-46de-4541-a4d0-aab10c131917 · outbound

This paper cites ORPO : Monolithic preference optimization without reference model.

Weighted-Reward Preference Optimization for Implicit Model Fusion ORPO : Monolithic preference optimization without reference model

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:31.584902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.652923Z digest=sha256:5494ce9f9dbb30b64520a37cef377a86f7d2e78a9f543a28c87bdf29bb265c16

Observation 54fed2fb-f0d6-479c-9507-7c9b79833f0c · outbound

This paper cites Mistral 7B.

Weighted-Reward Preference Optimization for Implicit Model Fusion Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:21.684381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:21.684381Z digest=sha256:3f4626a474a9492aca769fef13a6ce7147ea4f0880ac78f5dc365c2d586201c8

Observation 4bb2d512-e04d-45e5-ba3b-8329bef43f20 · outbound

This paper cites LLM-Blender : Ensembling large language models with pairwise ranking and generative fusion.

Weighted-Reward Preference Optimization for Implicit Model Fusion LLM-Blender : Ensembling large language models with pairwise ranking and generative fusion

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:31.424752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.743745Z digest=sha256:f675c7c0566df0d180e15c7b39e4aebb508908b17c97b199cae0c57e6ec8b99f

Observation 15c4804b-f6ab-4a43-9567-4db7889c7fe1 · outbound

This paper cites Knowledge-augmented reasoning distillation for small language models in knowledge-intensive tasks.

Weighted-Reward Preference Optimization for Implicit Model Fusion Knowledge-augmented reasoning distillation for small language models in knowledge-intensive tasks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:31.284748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.794161Z digest=sha256:2f1baa58d5cff128ae2ece76a281237bc7f821e043e9a856cc16f02d364fecc0

Observation fd62c6fa-1447-42cd-9bfc-d72675134404 · outbound

This paper cites Sparse upcycling: Training mixture-of-experts from dense checkpoints.

Weighted-Reward Preference Optimization for Implicit Model Fusion Sparse upcycling: Training mixture-of-experts from dense checkpoints

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:21.832430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:21.832430Z digest=sha256:1f78b5f29ae1259a680d87a6d0aa09c170b4f174602a853ee96e2dede6ea3f17

Observation 77196ec9-fd79-4e11-b815-dd2375a1c2a3 · outbound

This paper cites The Winograd schema challenge.

Weighted-Reward Preference Optimization for Implicit Model Fusion The Winograd schema challenge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:31.095024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.845103Z digest=sha256:b349302ed54c53c043d7a43238398bf83248ce32bf935ee2a7574474f430cb86

Observation 6fe8371d-ad98-4b69-a366-edf870852903 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Weighted-Reward Preference Optimization for Implicit Model Fusion From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:21.873622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:21.873622Z digest=sha256:e55df68ad9e875ae8f00e3c9959fe95a8b1f248260302103846182de2cb6684c

Observation ff6ab5db-0b08-48ac-9aab-9da365e0c9c6 · outbound

This paper cites Hashimoto.

Weighted-Reward Preference Optimization for Implicit Model Fusion Hashimoto

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:30.865591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.920898Z digest=sha256:581e83adb0d40ed6cef20a6aaa8aea240a92498d9c796e47a701ec5e99c846ce

Observation bbb2b5b4-11a8-4b72-b300-044db28310ad · outbound

This paper cites TruthfulQA : Measuring how models mimic human falsehoods.

Weighted-Reward Preference Optimization for Implicit Model Fusion TruthfulQA : Measuring how models mimic human falsehoods

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:30.724753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.960673Z digest=sha256:3b48a4bcdd0f8abbff2d5c16b2b31d2616817ec3ebdfcc0aa77455d73bdbf685

Observation f0d68b7d-1d6c-48be-8a21-650a12966267 · outbound

This paper cites Merging models with fisher-weighted averaging.

Weighted-Reward Preference Optimization for Implicit Model Fusion Merging models with fisher-weighted averaging

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:30.623625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:21.994757Z digest=sha256:903a204f97964ee0e87831bb45c81dcc38e4c879513957f67cfdeee9136045ea

Observation f8078d2c-c058-4bb7-9de5-059fc087da29 · outbound

This paper cites Pack of LLM s: Model fusion at test-time via perplexity optimization.

Weighted-Reward Preference Optimization for Implicit Model Fusion Pack of LLM s: Model fusion at test-time via perplexity optimization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:30.463719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:22.034789Z digest=sha256:486b80458decbd54bb0783637d342a7fffc09bf67bb75eb059fef46b80529bef

Observation bd9bc18c-2682-4fd9-836a-1b4bf1f82c66 · outbound

This paper cites Sim PO : Simple preference optimization with a reference-free reward.

Weighted-Reward Preference Optimization for Implicit Model Fusion Sim PO : Simple preference optimization with a reference-free reward

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:30.351095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:22.084755Z digest=sha256:739bfa832be27eeaccd94bf0533ebf374803e5ed739f801b2b76b5cbba501ed7

Observation 3e6431bc-985d-43d4-955b-60e0b6ec7616 · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Weighted-Reward Preference Optimization for Implicit Model Fusion Disentangling Length from Quality in Direct Preference Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:22.144929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:22.144929Z digest=sha256:60cf8145ad26a41d9b35caa13325e2b24a9c6ac730aad76b50002bb55492a863

Observation 98b82278-b6b4-4165-94e5-789218940319 · outbound

This paper cites Manning, and Chelsea Finn.

Weighted-Reward Preference Optimization for Implicit Model Fusion Manning, and Chelsea Finn

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:30.094762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:22.194770Z digest=sha256:355fecd8328ab504b1394b7917f503bba24720f3720feaa805d6f0bcfdad0bcd

Observation 79fa6a87-56a2-4519-a243-40057d8fa9ce · outbound

This paper cites Aligning large and small language models via chain-of-thought reasoning.

Weighted-Reward Preference Optimization for Implicit Model Fusion Aligning large and small language models via chain-of-thought reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:29.914920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:22.244758Z digest=sha256:c4fcabd2aa1141de7f2341c9c48577783e0f163a8d8cf10bd5cf5c2ef92e912b

Observation 2d02388b-eada-4469-abc7-d2ef482a2652 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Weighted-Reward Preference Optimization for Implicit Model Fusion Gemma 2: Improving Open Language Models at a Practical Size

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:22.295318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:22.295318Z digest=sha256:1ba818593e704d03d1cae89814eef98456d5f34414011ebe7b73edbbb584dbca

Observation 95793e4b-f7db-4d41-b673-26c64e9f09cd · outbound

This paper cites Proximal Policy Optimization Algorithms.

Weighted-Reward Preference Optimization for Implicit Model Fusion Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:22.334184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:22.334184Z digest=sha256:0610a06788b3a1d4ad6130d4fd2ca423a65b038917a8c2f2a39fdfb810c665c2

Observation e92aa4d9-474d-4f0c-9bc9-6f4ec3430183 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Weighted-Reward Preference Optimization for Implicit Model Fusion DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:22.404833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:22.404833Z digest=sha256:8d80bc699f91342fa6536ac9c979a2ffe5aa7109b1962d3e0f4058590066bf21

Observation 6c75fea1-e282-4743-8301-7dbb6bd5d184 · outbound

This paper cites ProFuser : Progressive fusion of large language models.

Weighted-Reward Preference Optimization for Implicit Model Fusion ProFuser : Progressive fusion of large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:22.454900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:22.454900Z digest=sha256:fe3d63001f85e46d6a7a6d317ac7227a7b746f474f196f26cb862afd0d497c2d

Observation 48ce57d4-346e-44b1-a248-00e74e220fa3 · outbound

This paper cites Branch-Train-MiX : Mixing expert LLM s into a mixture-of-experts LLM.

Weighted-Reward Preference Optimization for Implicit Model Fusion Branch-Train-MiX : Mixing expert LLM s into a mixture-of-experts LLM

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:29.725257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:22.494850Z digest=sha256:ab0de94657e348fc70f5d9dab805f7d283b65049628e6cd59dc053731e804cc2

Observation 5cb99833-72f3-4def-964f-2e007264acdd · outbound

This paper cites Preference fine-tuning of LLM s should leverage suboptimal, on-policy data.

Weighted-Reward Preference Optimization for Implicit Model Fusion Preference fine-tuning of LLM s should leverage suboptimal, on-policy data

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:29.594758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:22.523428Z digest=sha256:9663669a3facf8e98e166862093e6d5a066dcbd0a6e4a9eeb7bf05a776f1cf3d

Observation 53601485-ab76-4eb5-b01a-d99134fecfb9 · outbound

This paper cites Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation.

Weighted-Reward Preference Optimization for Implicit Model Fusion Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:22.572244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:22.572244Z digest=sha256:fee5489ae8046417588121ef78160d3cfb2ea998bca42cba9d26559c7822eb26

Observation cbeb0572-5523-4118-999c-ed60c127d4cc · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Weighted-Reward Preference Optimization for Implicit Model Fusion Zephyr: Direct Distillation of LM Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:22.611041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:22.611041Z digest=sha256:e3086fc9565d424c1e25b7c72cc851774d049b9f5c993cfbefe8be3021eee8e5

Observation db7ad270-3c0a-4181-bbf1-f9b9a35acae2 · outbound

This paper cites Knowledge fusion of large language models.

Weighted-Reward Preference Optimization for Implicit Model Fusion Knowledge fusion of large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:29.458530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:22.639429Z digest=sha256:afcd674532f8ce6de301b4f5add0103734131dad3bb519f499d5174f95d9dc9c

Observation 7a3dab9d-288c-4681-b7b4-7da070f5dcbd · outbound

This paper cites Knowledge Fusion of Chat LLMs: A Preliminary Technical Report.

Weighted-Reward Preference Optimization for Implicit Model Fusion Knowledge Fusion of Chat LLMs: A Preliminary Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:22.689153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:22.689153Z digest=sha256:84fa1149ee0aa7e466d38b2ed075462b2f6a7b408226df648a1f912424786f82

Observation b8f05aee-af0b-445f-8e51-f5f5f1d691e1 · outbound

This paper cites Interpretable preferences via multi-objective reward modeling and mixture-of-experts.

Weighted-Reward Preference Optimization for Implicit Model Fusion Interpretable preferences via multi-objective reward modeling and mixture-of-experts

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:29.307077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:22.744765Z digest=sha256:9b036fee986be8551f3fb00aae0e013215baed3e024a8aa847b0828354c0b89d

Observation 2a071f56-a858-42a1-a01a-5fb58ae6dcfe · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

Weighted-Reward Preference Optimization for Implicit Model Fusion Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:22.787258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:22.787258Z digest=sha256:4d86b63f06b56192e1cc4940e80942fefd5cfb46c8617b95e265ee646a241283

Observation b8c57ab6-079c-48fc-8754-5faf0619f4f0 · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Weighted-Reward Preference Optimization for Implicit Model Fusion HelpSteer2: Open-source dataset for training top-performing reward models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:22.834763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:22.834763Z digest=sha256:9a784c13793d3ba7680316d0d66a7ca32f56a4f70c7d731b462f42a23d9e8d06

Observation 4c16990f-2056-473d-8554-5e27ca8829b0 · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Weighted-Reward Preference Optimization for Implicit Model Fusion Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:29.104760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:22.884755Z digest=sha256:f41b8a0612b94db4605cf84ef55069eae394bb7a66557c84104952fa2beb1155

Observation 508ae4d5-6415-4a6a-b645-a15ab15aebbd · outbound

This paper cites Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation.

Weighted-Reward Preference Optimization for Implicit Model Fusion Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:28.964747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:22.920817Z digest=sha256:67c5839a95fe37eab225684c36b40d3fbc01e17d995a66a269f37525d006f36b

Observation 19a098c0-1432-4deb-aeb6-0f4256f4e83e · outbound

This paper cites Is DPO superior to PPO for LLM alignment? A comprehensive study.

Weighted-Reward Preference Optimization for Implicit Model Fusion Is DPO superior to PPO for LLM alignment? A comprehensive study

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:28.804771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:22.954881Z digest=sha256:a631f3e2b77fac6189a3a37df144160bbe870e8ef06c49b1e20f5f2d18d3cc08

Observation 003f0bc0-fab0-4a9b-99c9-63d6f8904f28 · outbound

This paper cites Bridging the gap between different vocabularies for llm ensemble.

Weighted-Reward Preference Optimization for Implicit Model Fusion Bridging the gap between different vocabularies for llm ensemble

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:28.614759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:22.994861Z digest=sha256:89b48435a9d177c6e474b2553c31d62cae13a4dcf576f09c5d67391f92052fd4

Observation a8f40649-ac73-4eff-bd30-c88c3b870f7f · outbound

This paper cites Ties-merging: Resolving interference when merging models.

Weighted-Reward Preference Optimization for Implicit Model Fusion Ties-merging: Resolving interference when merging models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:28.334756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:23.025230Z digest=sha256:9e6feb42d3f89c3e4eaf9a5342ead10eb8a1eba8b0107a8eaec95e9305f10968

Observation 3a4784d0-ecf0-4726-ac29-9a011c6c8908 · outbound

This paper cites Qwen2 Technical Report.

Weighted-Reward Preference Optimization for Implicit Model Fusion Qwen2 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:23.061399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:23.061399Z digest=sha256:4873c5ae2fdc067fa087904c30b237db6b79cf6e66bd2d8920101ff5a85ab3e7

Observation 05fdfed6-020a-4d50-a75c-2237f08eba29 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Weighted-Reward Preference Optimization for Implicit Model Fusion Yi: Open Foundation Models by 01.AI

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:23.104877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:23.104877Z digest=sha256:2c7992decb5f35320d9367eabe3b687f0c1fdfb1cd64a892262a39be96bc7bf8

Observation a8510248-c7dd-4e51-9425-be56280da948 · outbound

This paper cites RRHF : Rank responses to align language models with human feedback.

Weighted-Reward Preference Optimization for Implicit Model Fusion RRHF : Rank responses to align language models with human feedback

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:28.062758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:23.152469Z digest=sha256:8295e412faab308c98c9c8d8f78d305e053b8d50950f9bade9a019a138c91f05

Observation 7bc69922-b776-439d-84b2-10e5fb0a2e97 · outbound

This paper cites H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 4791--4800, 2019.

Weighted-Reward Preference Optimization for Implicit Model Fusion H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 4791--4800, 2019

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:27.924755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:23.214763Z digest=sha256:cafe4c8761f8072809b76c912b623a6be7920c0f955533f2bce15cd1dc13c437

Observation 0f707a22-fc9c-4ed6-90c0-2268f4aafa11 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Weighted-Reward Preference Optimization for Implicit Model Fusion SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:23.264757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:23.264757Z digest=sha256:ffe88a92c841e22fd92428aa7917ea6c8d4e68ffb7a099c2b270555d55fe44d7

Observation 73386b1e-aaac-4504-9d68-1da14a69894e · outbound

This paper cites Judging LLM -as-a-judge with MT-Bench and Chatbot Arena.

Weighted-Reward Preference Optimization for Implicit Model Fusion Judging LLM -as-a-judge with MT-Bench and Chatbot Arena

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:27.787934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:23.315428Z digest=sha256:7a806b4781e828f6882de018c2305ad80d531c657b146132e29494d2631f8713

Observation a9eae4d2-50ac-4d9d-aaf1-130ba050a408 · outbound

This paper cites WPO : Enhancing RLHF with weighted preference optimization.

Weighted-Reward Preference Optimization for Implicit Model Fusion WPO : Enhancing RLHF with weighted preference optimization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:27.584759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:23.365264Z digest=sha256:ef95db8b7baec3ecff007499f1ff1fdcdc743b158d6d0e2f9fc92b4992982a4b

Observation f3890c95-d620-4ec9-8fb3-9e6d9cd8036f · outbound

This paper cites Starling-7B : Improving helpfulness and harmlessness with RLAIF.

Weighted-Reward Preference Optimization for Implicit Model Fusion Starling-7B : Improving helpfulness and harmlessness with RLAIF

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:48:27.433540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T22:48:23.404759Z digest=sha256:bf709baa61d0cb659356fcf38da0f75aa4786112279178e6d16ad2d4029be615

Observation 8fa8b405-7cdd-406c-857f-c41f0bfcb79f · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Weighted-Reward Preference Optimization for Implicit Model Fusion Fine-Tuning Language Models from Human Preferences

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:23.454765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:23.454765Z digest=sha256:0f2c3dceea3a8add4ea2aef4eb6fbe2bc9c71b7bfbc4d5a63ebe5e3b0d27d42b

Observation 9ffb35a9-2081-4577-9e53-59e0adfdc756 · outbound

This paper cites write newline.

Weighted-Reward Preference Optimization for Implicit Model Fusion write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:23.495146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:23.495146Z digest=sha256:e59aec6fbe15d3502905e84354ae2d56475309e3dde31d840f532dcf0ce16f16

Observation 16127d6e-d884-4d43-a4b4-7dfe77098d2b · outbound

This paper cites @esa (Ref.

Weighted-Reward Preference Optimization for Implicit Model Fusion @esa (Ref

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:23.526871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:23.526871Z digest=sha256:6de88a3e62a768bb2879b8640146e61a14e55259a2c47c65bf993d50e7148085

Observation 834b00d9-0ed4-4f10-a813-3308cfe5dc31 · outbound

This paper cites an unresolved cited work.

Weighted-Reward Preference Optimization for Implicit Model Fusion Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:23.574761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:23.574761Z digest=sha256:a0e5a4598665f8638ff7d330527c3c3a623c1a78755e369c25f72eeb74902f10

Observation 18d18745-b8f2-4f41-94bf-afcf3eaf9f6e · outbound

This paper cites an unresolved cited work.

Weighted-Reward Preference Optimization for Implicit Model Fusion Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:23.624843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:23.624843Z digest=sha256:08bd9d1645cabf1224ff049318ace3d78ab12c300a04cb40bc9ee0321ee0211e

Pith citing papers

Observation a0e3fefa-5e6f-441f-8143-9a1d15942330 · inbound

An Overview and Discussion on Using Large Language Models for Implementation Generation of Solutions to Open-Ended Problems cites this paper.

An Overview and Discussion on Using Large Language Models for Implementation Generation of Solutions to Open-Ended Problems Weighted-Reward Preference Optimization for Implicit Model Fusion

Reference 143

Resolution
unresolved
no resolver link, observed 2026-08-10T22:51:52.719416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:51:52.719416Z digest=sha256:70347003cca80a4baec6875b92fd96a4ab8a64d5feb52404890708e0bab922f3

Observation 048901a7-3233-4835-a467-f5e9c1f57323 · inbound

CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate cites this paper.

CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate Weighted-Reward Preference Optimization for Implicit Model Fusion

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:10.913347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:11:10.783132Z digest=sha256:789b02617744aff9c30c3e3c02817c55f6a90abaa475d955e92743fb3d8e2008