Pith. sign in

Paper Citation Record · LEDGER

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling

As of 19 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2506.08681.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08681 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:16:43.217090Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T03:41:52.859647Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T03:45:17.571390Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a0f685e8-a465-468b-a58f-4832787db942 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Scaling Laws for Reward Model Overoptimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.181106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.181106Z digest=sha256:84ca62333ed6faaec7a2871df0fbd871fe21e8678fc13176b0b1dc89682bc550

Observation 258e424f-f7ff-40c9-a1f4-26cfa3be0f3d · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.185377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.185377Z digest=sha256:4be2b2e00e45b49be869b6447b9152296b775ad6143fa916c5f75262611d16b4

Observation 4d66ed61-f61e-490b-ae51-adb4cca2b999 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.189540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.189540Z digest=sha256:eb58e491728a1e958c69722ef3dae5512e88f13d1dba06bdfa70d4b5e5cd161c

Observation 2ecd451f-e590-4233-8ca6-48f70ef5035f · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.197365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.197365Z digest=sha256:1c60b0f0106cd01fbc8a8a08f2c189d971479a862b77a366cce7c7a4eb868b50

Observation f2f2f166-723e-4ae7-9a6c-85d85b5b8c91 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.378995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:16:43.201513Z digest=sha256:c466f71562954782aa86719a68a11a23db1159237cfd314ce0c965f76fbe361a

Observation 847a3684-3bc7-45c9-9457-2fc086b44a42 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.209489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.209489Z digest=sha256:0298ddca089316f28d0dcc6dd14142c0adf460a23261fea88de746929c4e1599

Observation b43310e1-adb9-47f2-ba3b-74fa4eb505fe · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.353429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:16:43.213207Z digest=sha256:1873b336890ae6f66b14d95b344d07b373ccdab08438627226445f05cbe01e94

Observation 9c79d974-d0d4-4636-9e46-2afef4379b97 · outbound

This paper cites gpt-4o-mini.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling gpt-4o-mini

Reference 640

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:16:43.340209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:16:43.217090Z digest=sha256:528ea046f5eb77ab62248f08480218e60a58e044a58c6b7342f988dad544f7e9

Observation 6c2a0880-c598-4a3f-b7f5-1f9f2cf431ed · outbound

This paper cites cc/paper_files/paper/2017/file/d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling cc/paper_files/paper/2017/file/d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:16:43.405220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:16:43.171719Z digest=sha256:ecd294f3a70741240a2907f975944923e4fd70e5e76edac4c0af717697ba3166

Observation a77382f5-edcc-4ce8-af9b-2e9f9ac1bef2 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.391070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:16:43.193528Z digest=sha256:17a920d4753585b7c82acd29fe77e2b9a92f08e37bc64b0e3f92b8ff8b394321

Observation 220cbbf9-df1c-41a9-9a11-0e648098282f · outbound

This paper cites doi.org/10.1002/9781118445112.stat08284.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling doi.org/10.1002/9781118445112.stat08284

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.176803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.176803Z digest=sha256:04b83bcce75444db511e959afeec35fef0a993d18bcc04022b8ec83d38b40f93

Observation a041404c-0989-4b13-b3a2-30e20740b42a · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.162886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.162886Z digest=sha256:423aedace5d4b445e03c0574af91529e853c843836bde6d2fb2cdcd989d04840

Observation 4a290447-6663-444f-9ffa-cd2c74a00af0 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.365925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:16:43.205323Z digest=sha256:7cf632df3e0582853dd8f295ec7a0c9ce9362b41f0a77686e0542f31b95cbacd

Observation e92b9e93-09fb-4fc7-a05a-db9669aceec6 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.157955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.157955Z digest=sha256:edf667aea95fe53f98e4699bcce13113ef6d0c056733e2fb0f037a4aceab248e

Observation d41e3211-0888-4820-835d-ee00b26f148a · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.418032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:16:43.167292Z digest=sha256:52f9b7cc5c1e1bebc11182f8a61a89c455715d9e5526f74077063d1dbd5c299c

Pith citing papers

Observation e1f65f80-3589-45c1-a711-d6626b6e7187 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.573771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:8556f67dd1cce69254893b9cd9e11821d9a22d65b25816b2cdabe8bf78d6b206