Pith. sign in

Paper Citation Record · LEDGER

Direct Preference Optimization with an Offset

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2402.10571.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.10571 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:40:17.819727Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:24:23.529928Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b2d22c02-1478-414f-8744-afb40f444c3a · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Direct Preference Optimization with an Offset

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:59.493391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:5350cfc0fd587530bbc714dba4aa6d31e7d43db010a6c4e78589567d3259b2ba

Observation 587bad29-1ff8-4a50-9e75-684abf84c47f · inbound

Explicit Preference Optimization: No Need for an Implicit Reward Model cites this paper.

Explicit Preference Optimization: No Need for an Implicit Reward Model Direct Preference Optimization with an Offset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.819727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.819727Z digest=sha256:fa4050c038c016730e3eaed9f66654d025c4a883ed204425b15c2cf560d086dc

Observation 7e6f84a1-65fd-4f01-a787-58d2bebc5535 · inbound

Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks cites this paper.

Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks Direct Preference Optimization with an Offset

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:28.641196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:28.641196Z digest=sha256:ed62c59093a1a211d57e2bd4179208f3cc29d5b1a420991eacee64ee35bdfde6

Observation 8200f43e-7039-4015-b85e-48e6225b2dfc · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment Direct Preference Optimization with an Offset

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.531850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:563f65d2163c68c0c10801cee6e864b90c9e402dce239ca60ce2bfcc0dc2043b

Observation d26ae2f7-df07-4591-ab61-cb299ee78d85 · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences Direct Preference Optimization with an Offset

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:18.602366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:18.602366Z digest=sha256:30f5a0df3c1ea2b462528286dfb8e38f00c88d88518ba69b661557e58aed8441

Observation 0b0e7a2d-45c4-4445-95f1-a1676ebdece4 · inbound

ARGUS: Policy-Adaptive Ad Governance via Evolving Reinforcement with Adversarial Umpiring cites this paper.

ARGUS: Policy-Adaptive Ad Governance via Evolving Reinforcement with Adversarial Umpiring Direct Preference Optimization with an Offset

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:21.958494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T19:36:05.119054Z digest=sha256:f72704a3885a5718dc1c177d7a5e32d2bfc8ce96c24cc1970bf705bdeee794a5

Observation ff72c57d-92f4-4a47-8303-02b273dab3f7 · inbound

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models cites this paper.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Direct Preference Optimization with an Offset

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.350262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:0000fe434dc25c0ada75129ce07e0d415710bab0032e9218cffaba1a534977d6

Observation 11a4a659-dace-45ea-a427-ed31b495a799 · inbound

MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization cites this paper.

MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization Direct Preference Optimization with an Offset

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:24.730093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:14:37.374346Z digest=sha256:e3642daa7acad678c1aa966eaae51c262f9e4906c35db9bb53bc64bcd3999af9

Observation 104d1f22-4915-4274-ac35-a8d19d9c7e2b · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Direct Preference Optimization with an Offset

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:37:29.874766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:fe08109b494e667279681149db9eb72cdf9b7919149f2d7a59e41ca312a45317

Observation 979f0979-22a4-4414-8841-c8a0c59959df · inbound

Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design cites this paper.

Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design Direct Preference Optimization with an Offset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T10:59:13.062985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:59:13.062985Z digest=sha256:d8a040f3d64067e1efc90f711e368b2e89fce7e3a9f895ab98bcabfc314c244a