Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2406.16768.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:53.839991Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:49:45.640428Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 428c4a58-44a2-4fa8-a3b3-dfb568d3269e · inbound
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 177
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f85ecb2d-b418-48fa-9bf8-4edc9e001be5 · inbound
If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6129d2b-060e-4258-b0ca-bb0d09a22cd2 · inbound
How to Merge Your Multimodal Models Over Time? WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aead0e7d-5b5d-4a4e-9e6a-e63e1c18ff80 · inbound
Parameter-Efficient Interventions for Enhanced Model Merging WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86b9cb94-65cf-403b-af4a-d78036678a0b · inbound
Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62c27199-f8e2-4e43-9735-7098902b0d17 · inbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f45b4229-3401-43fd-821a-02d59d43cf5c · inbound
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93163713-4833-40c2-ad31-24fd5d384d74 · inbound
Robust Reward Modeling for Large Language Models via Causal Decomposition WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 019fa1e6-8d02-4490-9c74-c9e6e96ea15c · inbound
Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2b0378f8-a329-4dc6-9fc7-2495ee8fec94 · inbound
Spectral Souping: A Unified Framework for Online Preference Alignment WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e6ed6dcf-6dbc-4c62-8e34-598f32f03cf7 · inbound
Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 35c0e572-c422-4a72-b107-562b57def5ca · inbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fa2ad06a-f529-4f59-ba2e-16cf000335e3 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 253
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bb0dc9b-1b41-4640-afce-55a0a7b8c1cc · inbound
Multi-Turn On-Policy Distillation with Prefix Replay WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 254
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df289768-6da3-4973-a71b-1b7e9b3e94bb · inbound
REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 458571d0-4753-49b3-bc5e-b90ebef85fe3 · inbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.