Pith. sign in

Paper Citation Record · LEDGER

WARM: On the Benefits of Weight Averaged Reward Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2401.12187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.12187 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:41:10.519286Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b057dc0d-8de0-48b2-9089-8219d101ca8b · inbound

How to Merge Your Multimodal Models Over Time? cites this paper.

How to Merge Your Multimodal Models Over Time? WARM: On the Benefits of Weight Averaged Reward Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:43.411539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:43.411539Z digest=sha256:12f90865b5c29890eae3c234d12d29a0dd2d15572d70cf638d928805533556a1

Observation 0cc014de-00bb-47ba-90f6-8da9e69882ea · inbound

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment cites this paper.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment WARM: On the Benefits of Weight Averaged Reward Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.707005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.707005Z digest=sha256:ee6d90555a4598fb29cc97159f5ac87e402959a2bd1fb28ddf9fc554c2aa2edf

Observation cd98b722-028c-4980-8079-1ceae83db87f · inbound

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling cites this paper.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling WARM: On the Benefits of Weight Averaged Reward Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.257425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.257425Z digest=sha256:cf65440be76efb274f9977d85b284addee16e3e593b55c3a0e28db6004ef8a79

Observation a9917ce7-d6d8-45b9-a5cf-cfb28070d176 · inbound

On Teacher Hacking in Language Model Distillation cites this paper.

On Teacher Hacking in Language Model Distillation WARM: On the Benefits of Weight Averaged Reward Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T11:34:57.150732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:34:57.150732Z digest=sha256:efc7ce04715db1838cfae05891dedb6698b875dd8c23f8c8d80073703f33c552

Observation a2b21913-dd8a-401e-bf9e-3fe44b655998 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF WARM: On the Benefits of Weight Averaged Reward Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.649860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.649860Z digest=sha256:9ab16ace36d9074678af087b455d80da82a02a6545517205bf50e9157a3753f2

Observation 024add06-7af1-4b65-8d0d-cbf5366dfd3d · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling WARM: On the Benefits of Weight Averaged Reward Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.814125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:775a68f581aed5cadb686139886bc18e271bda1675a04f38a2ae98b8808b1892

Observation 608ab5de-8333-4751-8d8c-66a9fe2ffdbd · inbound

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary cites this paper.

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary WARM: On the Benefits of Weight Averaged Reward Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.316657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.316657Z digest=sha256:1c5b578ff3a990c323b09a8dc00081184bc076732b2b8d18fc32b5283ba5e955

Observation c96758dd-57fc-40a8-9c3f-1eaab651fd41 · inbound

Tiny Reward Models cites this paper.

Tiny Reward Models WARM: On the Benefits of Weight Averaged Reward Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.178654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.178654Z digest=sha256:f3f102a7289d969081dafe440376c1a80fa58c4bb7261384310fc9f2fe755f2f

Observation e2c4a7cc-89a7-40d8-941d-96d86ef9ce1e · inbound

Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback cites this paper.

Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback WARM: On the Benefits of Weight Averaged Reward Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:36:05.446280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:36:05.446280Z digest=sha256:0b2d56062990ec2232b3cd34fd00e744c18831dadf8da43f7b5eb7f822fdd247

Observation 5650bcbd-ea5c-4a08-9b87-388f4fe98063 · inbound

Towards Reliable, Uncertainty-Aware Alignment cites this paper.

Towards Reliable, Uncertainty-Aware Alignment WARM: On the Benefits of Weight Averaged Reward Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.470462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.470462Z digest=sha256:1c31aa0ed7e46c2ca5d09adc3a6dee63633eba05190b938f03f056ad1cf43ab1

Observation b34333c6-eb98-413a-9e3b-ea7b1f8d8977 · inbound

Mitigating Multimodal Hallucination via Phase-wise Self-reward cites this paper.

Mitigating Multimodal Hallucination via Phase-wise Self-reward WARM: On the Benefits of Weight Averaged Reward Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:43.298125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T05:10:45.144421Z digest=sha256:c24f407938c760490f7610b2c4e66d5d9857e30eca8ab3f0a1ad7e8d32016730

Observation 4dc419e5-df9d-4ad3-ac14-335e43c02c16 · inbound

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs cites this paper.

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs WARM: On the Benefits of Weight Averaged Reward Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:04:52.576970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T16:03:12.728352Z digest=sha256:116ce07964ceab4a28d796489a9e614e167b2c721b388bd710f380cba8a6dc56

Observation eb16ecb7-3933-47a1-9e97-7eaa7e4bb1dc · inbound

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity cites this paper.

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity WARM: On the Benefits of Weight Averaged Reward Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T17:31:06.966629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T17:26:17.072017Z digest=sha256:b61b40cdfbb63f52c1940a8a066ecc0ddc006836c82fa786d5682fbf044b9b2e

Observation 6fc38837-3c15-43a3-bcf9-70fdafaa011f · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization WARM: On the Benefits of Weight Averaged Reward Models

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.528193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:44fa79d3446614f69a38d0ee411511c7a6019e47e22bc2bdcdea747078d52f16

Observation 7de95927-6c1b-4e41-9256-d45d2ccd6eb2 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay WARM: On the Benefits of Weight Averaged Reward Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:c8cb2d7874105aa6108471806fcff43634d4cdd8b0abbc4a604ec83aaf69b3c2

Observation 42f096b6-45c0-42cb-8ad9-9578d5e75310 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay WARM: On the Benefits of Weight Averaged Reward Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:37.580679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:37.580679Z digest=sha256:f7fcbc4aea682b5321d75667af7a0f7b40752c5751161c84e86a8718695c2615

Observation d164ed71-f4cc-46b1-90dc-f45e60b3ff22 · inbound

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL cites this paper.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL WARM: On the Benefits of Weight Averaged Reward Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:58.293232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:58.293232Z digest=sha256:2089770780af7ac5dcec607b1dfa9a34c6b07c733eef1f5b00f1fa121d45fcdb

Observation 274a7b35-3326-4604-ad5d-4ad4d5a76151 · inbound

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees cites this paper.

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees WARM: On the Benefits of Weight Averaged Reward Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:41:10.519286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:41:10.519286Z digest=sha256:ae55dee00ce1b298e32cab3a086b311284f0185de748d317c84951db9517935a