Pith. sign in

Paper Citation Record · LEDGER

Direct Preference Optimization with an Offset

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2402.10571.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.10571 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:52:40.528456Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:24:23.529928Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b2d22c02-1478-414f-8744-afb40f444c3a · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Direct Preference Optimization with an Offset

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:59.493391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:8e1864a9aa0bd0000a161f007e42195e29b7ba56ec39611004fd32335e89a05c

Observation b416d928-f42d-4586-8c56-c34eb7bd0c49 · inbound

On Fairness of Unified Multimodal Large Language Model for Image Generation cites this paper.

On Fairness of Unified Multimodal Large Language Model for Image Generation Direct Preference Optimization with an Offset

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T04:52:40.528456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:52:40.528456Z digest=sha256:05e73cc9e8b67a4e923863e06c7662be3951d5715c3ec9ff9ccd51cd6ad3a854

Observation 09c3038c-a61c-4307-9b25-f209c322ef1c · inbound

Design Considerations in Offline Preference-based RL cites this paper.

Design Considerations in Offline Preference-based RL Direct Preference Optimization with an Offset

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.983451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.983451Z digest=sha256:d96c36be3b77680e32d8eaf6f6d453bca02f2542666bc16f2b12dab02b1cdfa7

Observation 4f9838e1-e980-48ff-95a1-a0e08ad56315 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Direct Preference Optimization with an Offset

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.539127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.539127Z digest=sha256:64fa30b8ed650d189d8ab52db377ec41b7c41ea0bb60d6ddf2e4fbdc9048cc8b

Observation 85257a83-460a-4538-aef6-ff869f385e16 · inbound

Proactive Guidance of Multi-Turn Conversation in Industrial Search cites this paper.

Proactive Guidance of Multi-Turn Conversation in Industrial Search Direct Preference Optimization with an Offset

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:14.904703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:14.904703Z digest=sha256:ea93b8fb2182541b7aa686fe2179bada58f639a656d90f13d4776ed276b8feb3

Observation e669799f-6b5e-4c36-a8fd-f7011e9720ec · inbound

AI Agent Behavioral Science cites this paper.

AI Agent Behavioral Science Direct Preference Optimization with an Offset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:00:53.302553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:00:53.302553Z digest=sha256:03e71f823e75cbf4b7ef81e4e0503b255982eda0405783d514da6c670cdc08d4

Observation 587bad29-1ff8-4a50-9e75-684abf84c47f · inbound

Explicit Preference Optimization: No Need for an Implicit Reward Model cites this paper.

Explicit Preference Optimization: No Need for an Implicit Reward Model Direct Preference Optimization with an Offset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.819727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.819727Z digest=sha256:df55e1ca6cb17729d62cbad5c47d367a3928f5fbf2d0f4e61b22ae6890278bd8

Observation 7e6f84a1-65fd-4f01-a787-58d2bebc5535 · inbound

Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks cites this paper.

Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks Direct Preference Optimization with an Offset

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:28.641196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:28.641196Z digest=sha256:346e70b866e02756b228c9ceccd7db2f30e662e79206776a8cb91792bf7fb033

Observation 8200f43e-7039-4015-b85e-48e6225b2dfc · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment Direct Preference Optimization with an Offset

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.531850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:9f52efe92dea990d9cf2de429775aeb6f0cf2fdd977f28b85c3ae321884a3251

Observation d26ae2f7-df07-4591-ab61-cb299ee78d85 · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences Direct Preference Optimization with an Offset

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:18.602366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:18.602366Z digest=sha256:8797806fb9b6d7e3bec93936e579ca3eab545af24b88192f6802be79693dad39

Observation 0b0e7a2d-45c4-4445-95f1-a1676ebdece4 · inbound

ARGUS: Policy-Adaptive Ad Governance via Evolving Reinforcement with Adversarial Umpiring cites this paper.

ARGUS: Policy-Adaptive Ad Governance via Evolving Reinforcement with Adversarial Umpiring Direct Preference Optimization with an Offset

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:21.958494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T19:36:05.119054Z digest=sha256:498c65d574e2294a8526150db1fa5a1afb876f70d3ad28f753f273630bbed1c8

Observation ff72c57d-92f4-4a47-8303-02b273dab3f7 · inbound

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models cites this paper.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Direct Preference Optimization with an Offset

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.350262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:c616abbc4d03af6294e9977b584d3c255c12f0f97295c798320d47e742682e25

Observation 11a4a659-dace-45ea-a427-ed31b495a799 · inbound

MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization cites this paper.

MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization Direct Preference Optimization with an Offset

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:24.730093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:14:37.374346Z digest=sha256:043b6f8e98a9daae9fa096638749c1a0be0aeeba580dd7886463e6dc1cba0f4e

Observation 104d1f22-4915-4274-ac35-a8d19d9c7e2b · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Direct Preference Optimization with an Offset

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:37:29.874766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:c8d2fd16582214f9b169066ff9e083f1e572f97cd1e588c69742e4508397a4a6

Observation 979f0979-22a4-4414-8841-c8a0c59959df · inbound

Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design cites this paper.

Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design Direct Preference Optimization with an Offset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T10:59:13.062985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:59:13.062985Z digest=sha256:d8a040f3d64067e1efc90f711e368b2e89fce7e3a9f895ab98bcabfc314c244a