Pith. sign in

Paper Citation Record · LEDGER

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling

As of 9 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2506.08681.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08681 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:16:43.217090Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T03:41:52.859647Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T03:45:17.571390Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a0f685e8-a465-468b-a58f-4832787db942 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Scaling Laws for Reward Model Overoptimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.181106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.181106Z digest=sha256:567946709f77c724eb7df71e7af7f2feff94934e414ac21af6db400b9143cac8

Observation 258e424f-f7ff-40c9-a1f4-26cfa3be0f3d · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.185377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.185377Z digest=sha256:2e20f328a9d8d6abfe712e545a5e884b8d21b947b0568c2035c5045bb53509b2

Observation 4d66ed61-f61e-490b-ae51-adb4cca2b999 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.189540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.189540Z digest=sha256:18f491bb404869fda0b858ce4cfafd9ca597a183657af0229d75ea2b435dc2d3

Observation 2ecd451f-e590-4233-8ca6-48f70ef5035f · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.197365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.197365Z digest=sha256:5f3a654427454c535312561f2db18d1bbc1674f735de8fc2bbb95841ded2b1b1

Observation f2f2f166-723e-4ae7-9a6c-85d85b5b8c91 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.378995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:16:43.201513Z digest=sha256:c0b5110d627c02b7bc2e12541a7d3cdf65ad834a707d6b458669bb4b78b9d834

Observation 847a3684-3bc7-45c9-9457-2fc086b44a42 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.209489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.209489Z digest=sha256:40aef15ea23398757b0ccd3c9c9c58db65fafab1d8ecf9aae60c4a4d123b402d

Observation b43310e1-adb9-47f2-ba3b-74fa4eb505fe · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.353429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:16:43.213207Z digest=sha256:1572cce7dd7ab01b4254d20b80fabc84c9c07a0a40b5a08222b1995b1878d174

Observation 9c79d974-d0d4-4636-9e46-2afef4379b97 · outbound

This paper cites gpt-4o-mini.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling gpt-4o-mini

Reference 640

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:16:43.340209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:16:43.217090Z digest=sha256:09344b9b619a67364130b3dd5cebaab4a02e526ad913ee0e2957ce230c4c7873

Observation 6c2a0880-c598-4a3f-b7f5-1f9f2cf431ed · outbound

This paper cites cc/paper_files/paper/2017/file/d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling cc/paper_files/paper/2017/file/d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:16:43.405220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:16:43.171719Z digest=sha256:b5c5d6084a6875724b98c6c164bb7bba7149bb2bf48d9f252bf1130a274db4b7

Observation a77382f5-edcc-4ce8-af9b-2e9f9ac1bef2 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.391070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:16:43.193528Z digest=sha256:3af47be0170180c6788a45148ce0b3d05c6216f863f42d5557cb4a37114599d9

Observation 220cbbf9-df1c-41a9-9a11-0e648098282f · outbound

This paper cites doi.org/10.1002/9781118445112.stat08284.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling doi.org/10.1002/9781118445112.stat08284

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.176803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.176803Z digest=sha256:54db7f55e269a39388428c8eda0a121bfbaa65c0d1f44a6477823b967de62296

Observation a041404c-0989-4b13-b3a2-30e20740b42a · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.162886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.162886Z digest=sha256:b937a6255adeddd30d5e48eee4757259ddfa1d05892ded8bc3f600612ea9c156

Observation 4a290447-6663-444f-9ffa-cd2c74a00af0 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.365925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:16:43.205323Z digest=sha256:e1c93c5a45f6e4c390c997f2ea0406d10a315736700d7ea044e0e613d93d3c81

Observation e92b9e93-09fb-4fc7-a05a-db9669aceec6 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.157955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.157955Z digest=sha256:244429e9c4cbf29b7a71a9996372dfaed854bfabd44da949b43368da74b77c45

Observation d41e3211-0888-4820-835d-ee00b26f148a · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.418032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:16:43.167292Z digest=sha256:4df336654615864035534321140418a3e80c04ebb7eb8c107f905060815ae675

Pith citing papers

Observation e1f65f80-3589-45c1-a711-d6626b6e7187 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.573771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:e2109fdee2eda1128ddd87d735c1e4231bfc1b515605ff172369a63cd6acbc4a