Pith. sign in

Paper Citation Record · LEDGER

WARP: On the Benefits of Weight Averaged Rewarded Policies

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2406.16768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.16768 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:53.839991Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:45.640428Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 428c4a58-44a2-4fa8-a3b3-dfb568d3269e · inbound

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities cites this paper.

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:16:04.840454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T22:16:04.386706Z digest=sha256:05c4123d095567060e35d0bc8a29ce274b5d5fff9242b5379f50296281f6a050

Observation f85ecb2d-b418-48fa-9bf8-4edc9e001be5 · inbound

If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs cites this paper.

If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T21:46:25.101061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:46:25.101061Z digest=sha256:a8e2b224847e3c76faf564d6ccb291e4c4fafdf5a80070e65629de6c30327fb0

Observation b6129d2b-060e-4258-b0ca-bb0d09a22cd2 · inbound

How to Merge Your Multimodal Models Over Time? cites this paper.

How to Merge Your Multimodal Models Over Time? WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:43.395373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:43.395373Z digest=sha256:4477d6356416631a9f93cc97090c28a13de2b780ffe1e1560d92f3a2f2275e61

Observation aead0e7d-5b5d-4a4e-9e6a-e63e1c18ff80 · inbound

Parameter-Efficient Interventions for Enhanced Model Merging cites this paper.

Parameter-Efficient Interventions for Enhanced Model Merging WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:06.636580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:06.636580Z digest=sha256:6ecdfd6829a64197200312f5a144128e4f4584a1acc51ae9a853d956474e26a7

Observation 86b9cb94-65cf-403b-af4a-d78036678a0b · inbound

Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? cites this paper.

Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T18:11:28.155863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:11:28.155863Z digest=sha256:6659657be8053a2d6e8bb70d059a94d86dfd486b5428ac150b35910c75353b8e

Observation 62c27199-f8e2-4e43-9735-7098902b0d17 · inbound

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning cites this paper.

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:25:45.301158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T23:24:43.556606Z digest=sha256:2b9d1e5a528d1c42c0bc72cf299b1ea3469b6dac401150617fc825576a61e060

Observation f45b4229-3401-43fd-821a-02d59d43cf5c · inbound

WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training cites this paper.

WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:49:40.309508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:49:40.309508Z digest=sha256:634ac7fcc9fdee0dac82267daf3a7b7db1b3cfcc6da31532f6b6dbf713a628e9

Observation 93163713-4833-40c2-ad31-24fd5d384d74 · inbound

Robust Reward Modeling for Large Language Models via Causal Decomposition cites this paper.

Robust Reward Modeling for Large Language Models via Causal Decomposition WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:29.610326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:50:47.902911Z digest=sha256:8e3820823c2fd7f3666c299ec00b1e8e538a3a21ed5bea336d54082882614ccf

Observation 019fa1e6-8d02-4490-9c74-c9e6e96ea15c · inbound

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors cites this paper.

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:44.118097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T19:05:51.423427Z digest=sha256:8096a0789ada9c0de0a7986a47f263180e2a775c91c19d1e8b2b2f6b9cacf5aa

Observation 2b0378f8-a329-4dc6-9fc7-2495ee8fec94 · inbound

Spectral Souping: A Unified Framework for Online Preference Alignment cites this paper.

Spectral Souping: A Unified Framework for Online Preference Alignment WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:59:50.966900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T07:54:56.356555Z digest=sha256:0bfd7dd7678a7c5e31cd29049680bb0715879cc4178f87cbe68b6c93feaef041

Observation e6ed6dcf-6dbc-4c62-8e34-598f32f03cf7 · inbound

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL cites this paper.

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:43:30.566219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-29T14:41:13.191919Z digest=sha256:dccb651733dbac61e56f8ede24109f645c453328e8139f85ab226104c1829c97

Observation 35c0e572-c422-4a72-b107-562b57def5ca · inbound

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games cites this paper.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:45.642601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:837e030e1dce5bdcb69a3f28fa299fe46e0dc13efd368f87e8b8779ebaa75bf8

Observation fa2ad06a-f529-4f59-ba2e-16cf000335e3 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 253

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:5c5eec091bf080a6490586f8d2c04989790793a0e98797cd82371c56414c3a98

Observation 1bb0dc9b-1b41-4640-afce-55a0a7b8c1cc · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 254

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:01.803367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:01.803367Z digest=sha256:9f5a518d1323adde31acf237b2baaedd6120f9be746785c57db44468310befc0

Observation df289768-6da3-4973-a71b-1b7e9b3e94bb · inbound

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation cites this paper.

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T14:00:00.388339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:00:00.388339Z digest=sha256:6d6ab627b093d48b03eee18fd7a109bd0860ef3473877c84e0a8c4033ef173bc

Observation 458571d0-4753-49b3-bc5e-b90ebef85fe3 · inbound

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation cites this paper.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.839991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.839991Z digest=sha256:dfdbbd200b9a577e8bd88ae65ba9854974b9546ee54030ff6c257eff253101fe