Pith. sign in

Paper Citation Record · LEDGER

Generalized Preference Optimization: A Unified Approach to Offline Alignment

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2402.05749.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.05749 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:17:45.991858Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T03:52:29.455696Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 47600c21-1a3d-4c32-8cd1-207366753daa · inbound

The Differences Between Direct Alignment Algorithms are a Blur cites this paper.

The Differences Between Direct Alignment Algorithms are a Blur Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:52:29.458777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-23T03:50:03.720389Z digest=sha256:049367c586af598541af9980609925d56660a089d93ba84ca952a0532b6fea8d

Observation 1708aeca-820c-4d29-93d4-736cf1eb84b1 · inbound

DPO-Shift: Shifting the Distribution of Direct Preference Optimization cites this paper.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.991858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.991858Z digest=sha256:26d30e483a56d3203176757792b83caa078c8cadf54ed829d5e7d5c77bd6ed18

Observation 35c96d81-5e34-4029-b468-edaa857d42f6 · inbound

Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance cites this paper.

Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T10:25:13.056866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:25:13.056866Z digest=sha256:42a7995d878c9aa0b57e0539e9da043955f0755711f6022b411a89b18910ce8d

Observation 400d8bdf-8572-4ceb-ba4e-f027564c672b · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.110662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.110662Z digest=sha256:36fd94ed578dbbc7e302bc6cfa1cd79fd07eea50424c7ce8282f131448a26b21

Observation 31f00f59-19a3-4c2b-a2f4-8491679d2c76 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.361604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.361604Z digest=sha256:ad4fe17365d505f972d10319da368ffa82995b5a423975c7e6cf378f9832a2bc

Observation 2e6ff718-c6fc-46f7-9aea-f0932bbcd6dd · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:46.360499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:46.360499Z digest=sha256:800978c95fdc097b2a13c0cceeea4aeaa4f8584e85a5ce5af7704da5d7980100

Observation ba1dbab5-2129-42cb-9ed6-8fe0fa6d62b7 · inbound

Learning Parametric Distributions from Samples and Preferences cites this paper.

Learning Parametric Distributions from Samples and Preferences Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:13.529395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:13.529395Z digest=sha256:9927b66f474550fc58d5f23120a895c9d88bad20d7dcf4338e713e27a7d05563

Observation 21d8ba4c-9c58-46cc-b33b-f6537323faba · inbound

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences cites this paper.

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:56.555683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:56.555683Z digest=sha256:aea5f0915634407679b7830abfdb572655f9844fd16cf3b55d7fb5f44741aad0

Observation 6bf9cade-a35f-49df-802a-f4f1d1c79197 · inbound

Explicit Preference Optimization: No Need for an Implicit Reward Model cites this paper.

Explicit Preference Optimization: No Need for an Implicit Reward Model Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.016343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.016343Z digest=sha256:e900a13da014a8221aaa43e1f30ec1ed52118e651a7c9dace4cbe33b334600e5

Observation 9fc43729-2309-44a1-b104-6dd6dd3a7ff4 · inbound

On a few pitfalls in KL divergence gradient estimation for RL cites this paper.

On a few pitfalls in KL divergence gradient estimation for RL Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:38.578236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:38.578236Z digest=sha256:094c9c095c65db1cfdc8375add2ea5295d8579af5ab6960e67bcdd2d97730d5a

Observation 5f28af2f-0126-4227-86bd-402bdf9fda61 · inbound

Curr-RLCER:Curriculum Reinforcement Learning For Coherence Explainable Recommendation cites this paper.

Curr-RLCER:Curriculum Reinforcement Learning For Coherence Explainable Recommendation Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:35:49.941183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:43:11.384787Z digest=sha256:77de076f17f60a9fad3cc06749f522f40797636e6741394c10c22d80c90ed9a9

Observation bfc0acf3-1e5c-4465-ba88-4d4026a4de31 · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 163

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:46:00.197978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:7f2074144d354a5745d6c4df80f12ba97aa353018163e8bd172d9853f982d28b

Observation 1590a79e-fe7d-455b-b731-62dfdfc85523 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 139

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.175932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:6fa29ce8ca196df49732f1949bb0718227f7afad5ad55c69840cfce43837399d

Observation 3b2bee7a-53ff-49ac-9411-5ca08a07456f · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 139

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.522644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:681934b1562fec07ce716d9b895b687f3f55a461992a01ccf2cb322758aefde9

Observation f2eb89b7-8cf7-4aaa-be79-a61664a91d56 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 190

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:6f66ef08329f0da044b9237fcdaa8a909da69095c4d9f7ff709d7018ecd25e26

Observation 0e16a44a-39cc-4b67-92e7-b47ce1f28de8 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:54.675387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:54.675387Z digest=sha256:7740557d945a40c2c57da1811b70a796a3ce1e8de14a4ee547b2b841ebbddfe6

Observation 9cfa9a43-235a-4937-86fe-7ae51fa1af33 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:58.919069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:58.919069Z digest=sha256:3cac0b68a70c644579f9c0d24774c1d9c071df50b8710980c7b5ee3f6b25374d

Observation d8a343d5-9c94-4a8b-b3da-b9f381382b9a · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:28.247565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:28.247565Z digest=sha256:2a3e6014d58b946f309468bd133a51e39ae65d74fb8b78baecc0ecacaca656a9