Pith. sign in

Paper Citation Record · LEDGER

Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2401.16335.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.16335 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:54.207436Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:10:42.038007Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 88e81838-4d92-4525-9c47-3872bc995401 · inbound

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration cites this paper.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.430135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.430135Z digest=sha256:22cee17880789cd2b42fd0b8acbbf3eb43a416643347a0ff51bea2f2b92b9191

Observation b4131ff1-1cf1-4919-bf85-d5d5a7d31616 · inbound

When Can Proxies Improve the Sample Complexity of Preference Learning? cites this paper.

When Can Proxies Improve the Sample Complexity of Preference Learning? Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.109274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.109274Z digest=sha256:0015555dea13fc539634c5d94e014b92a72dc18b30c6639463c04ef1723f5e84

Observation 72dd7354-9ff1-4eef-b10d-c2f5168b8a3a · inbound

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models cites this paper.

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:00:49.045583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:00:49.045583Z digest=sha256:b5079227128002a7b09cc82e19daa0d2efdbb83a8b5f255018b1421931a5b1c9

Observation 0feb13fe-dc2c-486b-8f6c-179b65f3c4b8 · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T11:22:16.621648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:2f69270ab39ca1f2f39aec1669d75e3c07e05d76feb9f5703bb1f476094e8f2d

Observation 7448d546-462a-47b6-bb0a-3b9a27ab379e · inbound

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models cites this paper.

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:58.239660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:58.239660Z digest=sha256:0d093bfe429eb32712c0088042298c8cdce1f6d7fc13c36c8a4b7cc76e4899a9

Observation 99066ce0-bf66-45e2-bff9-babc6bf79869 · inbound

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary cites this paper.

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:32.278911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:32.278911Z digest=sha256:5a771b0fd1e597945832e99f7df62c297ff1b5ed0fe781ef4c259f7be1931174

Observation 8027a939-c023-4432-b829-2482dd1db6c2 · inbound

RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards cites this paper.

RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:10:42.039945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-21T22:09:47.649346Z digest=sha256:c6151a3a1a53772aea9c2c3ff423d5d08ff4c237a65a3b1631164b43a1caf258

Observation bab6fd4f-2ca2-4852-903e-283ac5acfd32 · inbound

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration cites this paper.

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:40:20.990503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-15T21:38:53.792221Z digest=sha256:5c94c379ed7dd2b0d29b133a53575ba301097b0be41edaadc0c85325b36c88f0

Observation ec45e4a4-4f48-4097-955a-686854b0ab13 · inbound

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization cites this paper.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:16:10.933497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:1eb63e2bd39f4e1b51ccd09e6adf5c8ad6d975788a12ddc796751b0af3d0508d

Observation d288a023-8ffe-4e7b-8cdd-cc4a31b371dc · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:46:00.102698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:4c2c9abf2749470f122391228e3a3db7fff0610ca9761829dcc0df9d0bec00f5

Observation b3d0f08e-3cf9-4443-9f84-41c7e45cb26e · inbound

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation cites this paper.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.207436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.207436Z digest=sha256:d0578d4495b68aec0dff9daaafcd46929fc8071c660778c401e77a8d20dcbac3