Pith. sign in

Paper Citation Record · LEDGER

Design Considerations in Offline Preference-based RL

As of 14 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2502.06861.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06861 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:40:42.080296Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T13:06:54.002248Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T13:10:10.404876Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 18182556-cba3-4140-bd73-4b789af87ceb · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Design Considerations in Offline Preference-based RL Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.977517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.977517Z digest=sha256:7fae64fc3ae91563e110f763b0998aa5ce1543cdcbfcc0dcd6e2e5f007f9b578

Observation 5c30b055-fb90-4bf3-8124-fbb6fabce50c · outbound

This paper cites an unresolved cited work.

Design Considerations in Offline Preference-based RL Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:40:42.415603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:40:42.074455Z digest=sha256:43319e70bccdccb3dd4b1e2ce728bda180dcb4178919ada402c1d9d8274aef31

Observation 89e4aabb-baae-45b5-af0f-9e2d9e9d6917 · outbound

This paper cites Robust Preference Optimization through Reward Model Distillation.

Design Considerations in Offline Preference-based RL Robust Preference Optimization through Reward Model Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.999377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.999377Z digest=sha256:3d5b4fce021443663eaa2e7823cc299095397973577e3c1660191739dd0d33f9

Observation 28b3bc5c-75f7-4b8e-bf00-8d71f0b107ff · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Design Considerations in Offline Preference-based RL Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.004209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.004209Z digest=sha256:ed143dc3fe949bc8ca0dd39145d4caa089912adec9b05d6c58f128b7e18891c2

Observation 8e476658-38ac-4c5e-8c3c-b49441490b4c · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Design Considerations in Offline Preference-based RL Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.009598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.009598Z digest=sha256:25f16874f88b825a974e9771a2cf3fef988bf4f70f705a5bb5887519a05c827a

Observation 9ba1fac7-9632-4932-a5e4-4a15ee9fb486 · outbound

This paper cites Nash Learning from Human Feedback.

Design Considerations in Offline Preference-based RL Nash Learning from Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.019294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.019294Z digest=sha256:6e5bd9b1fe9b5c8e43305ebd8be34663ab73ecd0be60f5be8a3a411869aa0eb0

Observation c94c4030-9263-4763-91fc-8508caf05beb · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Design Considerations in Offline Preference-based RL Disentangling Length from Quality in Direct Preference Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.029675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.029675Z digest=sha256:4b0a21cb3c515d7a28282f7d908f2475bf7f76bf4a6fdc3941a785822702f386

Observation 5f84b809-9c0b-464b-9ae6-7e16dff56c16 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Design Considerations in Offline Preference-based RL Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.034684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.034684Z digest=sha256:282b119a7f346097cd45b71e3ae30a99d340c967f9471549492a6b42d24a2b51

Observation 4639f06b-92d8-410e-aac7-ed4c6e12b946 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Design Considerations in Offline Preference-based RL Gemini: A Family of Highly Capable Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.049982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.049982Z digest=sha256:a645231da137ae3e1bbf3c9a55cb071ce254f3999cff2b62721a2c6e7de97610

Observation a64d54ee-d14d-4af9-a98e-368242dbd36e · outbound

This paper cites Is RLHF More Difficult than Standard RL?.

Design Considerations in Offline Preference-based RL Is RLHF More Difficult than Standard RL?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.054396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.054396Z digest=sha256:6bd0b6ea4633425dd9d7ce254e9e1cad16784145187fd3eb9b2246d174ccf0fe

Observation 3c00348f-ff9c-439e-a0ce-a04c3d03c991 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Design Considerations in Offline Preference-based RL SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.069597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.069597Z digest=sha256:1c0e1b54563ccb9d45796947d0ffccb0382fe9098a079b9be005d51da2e43fd7

Observation 5641bb44-ced0-44c8-8857-b9c0b6d82cc5 · outbound

This paper cites The optimizer used is Adafactor with learning rate that is constant with a linear warm-up for 2000 steps and a base rate of 1e −.

Design Considerations in Offline Preference-based RL The optimizer used is Adafactor with learning rate that is constant with a linear warm-up for 2000 steps and a base rate of 1e −

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:40:42.399596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:40:42.080296Z digest=sha256:5a5eedf3614a4f5b93f3a0d1360af3d95539fb6017add2d2abe6c79fbb1f1732

Observation 02e5e5a0-dffe-4b30-822d-13b599384c68 · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

Design Considerations in Offline Preference-based RL Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.989242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.989242Z digest=sha256:3f805bdf0ee7ff0adc41039e6743673ce9d2331f1182e2d38b8a162a2c7c481f

Observation e8739783-b312-4542-905a-f34c8373c5f4 · outbound

This paper cites Calibrating Sequence likelihood Improves Conditional Language Generation.

Design Considerations in Offline Preference-based RL Calibrating Sequence likelihood Improves Conditional Language Generation

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.064438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.064438Z digest=sha256:59f4aee44aa548ff442522e16dcf6d40da4f1cf74f954d2f7474069aaf56742f

Observation a37fcd4e-f981-4ab9-8bd2-ce2ed2866e01 · outbound

This paper cites On Regularization via Early Stopping for Least Squares Regression.

Design Considerations in Offline Preference-based RL On Regularization via Early Stopping for Least Squares Regression

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.039660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.039660Z digest=sha256:9abb240a30c1347837b8823da8cf9d4387dec493536789db69fe7d468866caee

Observation caabea31-3e40-41bf-b61e-2ab71ce38f96 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Design Considerations in Offline Preference-based RL SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.014455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.014455Z digest=sha256:669725f29b98326b8970a29d973d9902c68ef8c9f5b16de0473b3635f21a79be

Observation d202202d-2a25-428d-938b-5161964e17b7 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Design Considerations in Offline Preference-based RL KTO: Model Alignment as Prospect Theoretic Optimization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.994465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.994465Z digest=sha256:82655a7ef2e1db35573338eafecfa06f94ad99c462c255c658d586ff78202bca

Observation 1daba942-17dd-4a26-bfe4-8731dcd449fd · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

Design Considerations in Offline Preference-based RL A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.044924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.044924Z digest=sha256:d6df726eb6b10cea95280549e4c9ff0a3f9de2feabd9cf3e5e4c6371774434b6

Observation 7d1f0652-e606-4e17-8087-744fdbbf5409 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

Design Considerations in Offline Preference-based RL Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.059588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.059588Z digest=sha256:256e85172395de14cb76e39152338bb54a460e11b0a7936ed660512ef31a39f1

Observation 89d56662-a529-4de4-8ddc-6401e9a9a7d6 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Design Considerations in Offline Preference-based RL Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.024176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.024176Z digest=sha256:282fd0f650418e04e5bbbe0bd3205a71335d976dfd50c11c628f7e070529c223

Observation 09c3038c-a61c-4307-9b25-f209c322ef1c · outbound

This paper cites Direct Preference Optimization with an Offset.

Design Considerations in Offline Preference-based RL Direct Preference Optimization with an Offset

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.983451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.983451Z digest=sha256:fef6727fb66dabe2b412aaf24c49a747a421c9c6db57993a97871483d6d8a6a0

Pith citing papers

Observation 3fe1c548-1e61-48a0-9c5f-c0a9d67d8b9e · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Design Considerations in Offline Preference-based RL

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.516891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:17462f7ddcd186acbb807f516638fb3e93e7b6343492ac2284606914e0cd445f

Observation 6a5f2044-b3f1-4b49-9a79-dc6d9ad3e799 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Design Considerations in Offline Preference-based RL

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.406878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:f4cf8af217310408643f3072c065a22b22a1993a9843dd383172200b71f23727