Pith. sign in

Paper Citation Record · LEDGER

A Survey of Direct Preference Optimization

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2503.11701.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.11701 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:09:44.129023Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e1c935f6-3a28-48d5-aa9a-b50a8efa91e9 · inbound

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning cites this paper.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning A Survey of Direct Preference Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:44.129023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:44.129023Z digest=sha256:79509e5e395ded9d27945263723b17f5ca6b44835a5488e1a1134a4bc134a6c3

Observation 99d4cb26-b296-413f-a0f2-d56a0e88cb02 · inbound

Intra-Trajectory Consistency for Reward Modeling cites this paper.

Intra-Trajectory Consistency for Reward Modeling A Survey of Direct Preference Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:34.053645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:34.053645Z digest=sha256:5bed71ed01e8479ab41d517f27be16484cc23b4a9d064bf9e1027d6d04c188a8

Observation 7a998fd7-a7b3-4929-9f05-f782f3d09b1b · inbound

A Novel Self-Evolution Framework for Large Language Models cites this paper.

A Novel Self-Evolution Framework for Large Language Models A Survey of Direct Preference Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:33.764607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:33.764607Z digest=sha256:d59045adf6d2271084be4ffbbbe99ca9fc06fb0dc6d5327b5ed2755a96ef33e1

Observation 8de070ee-5aae-42de-9bcb-a29fff768523 · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment A Survey of Direct Preference Optimization

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.437263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:3174d5187797fd99d88e5a27cb5676b8dacffc3f04274a81f20d2eefa89b3811

Observation ecd42d58-be7e-4676-946a-a1142d2fc2ae · inbound

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs cites this paper.

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs A Survey of Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T06:54:11.231474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:54:11.231474Z digest=sha256:4df3887ec2c4ff5e028d796bb6143b38022104a65c88bb0c7101068a9be96a82

Observation ee5a3204-80e7-427e-b045-6d032292e33a · inbound

VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models cites this paper.

VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models A Survey of Direct Preference Optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:39:54.350069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T09:38:32.977517Z digest=sha256:209187510477fc1e9ce7cb3dee807abf287fe111de2b2c4fda103fab24d8a298

Observation 1696a3ff-944e-435f-8c6c-275e5452166a · inbound

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning cites this paper.

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning A Survey of Direct Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:53.684688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:28:58.515666Z digest=sha256:e1e55b7573438b0f340dab681158097637583fe067a665789fcd605fd25de217

Observation fd377ff2-8a5f-44e1-bfc0-9fec693a5a7f · inbound

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization cites this paper.

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization A Survey of Direct Preference Optimization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:01.255353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:11:49.149334Z digest=sha256:98d8a85fc6fe200b0a443ccc28560f9a73625325937ee31524f07c3351949f0b

Observation 9c1d35da-0d6c-45d9-8dd0-cd8bd70ac6c5 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Survey of Direct Preference Optimization

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:16.671074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:939467d1d4ba9f6337f7f669ea9e2c9a984a0f27a0e8c254ef6e655362780900

Observation 6b978201-41a7-494c-915b-f4ea47e82fac · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Survey of Direct Preference Optimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:05.018076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:253e055bd01baea2f66193132afee7d3d739ec976adaaea66834ee426783579b

Observation 62e96638-a393-4a34-84f2-c538b10ce10c · inbound

TUX: Measuring Human--AI Tacit Understanding cites this paper.

TUX: Measuring Human--AI Tacit Understanding A Survey of Direct Preference Optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-28T21:22:38.748817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T21:18:37.366258Z digest=sha256:226709e3ff3286c5395e08b05939f0e5c9068d99b2318d5636e4e145a806cabb

Observation 20ac0933-1d06-433f-bea2-65123a72cac4 · inbound

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs cites this paper.

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs A Survey of Direct Preference Optimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.939868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T22:46:47.786929Z digest=sha256:5646366d90510ceae343e24f1dc34734f788ef715529f3ca1845b947ad32dd73

Observation 35783b72-aa3d-4470-8121-0eaf885c9031 · inbound

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text cites this paper.

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text A Survey of Direct Preference Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T13:36:56.375550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:36:56.375550Z digest=sha256:8739c51b38d07214663863b28bab271c51bbd81a7aa52f423eca5d8a11d192cb

Observation 6f2b483e-719d-4b54-a9a6-702675fde444 · inbound

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing cites this paper.

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing A Survey of Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T00:22:51.122740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:22:51.122740Z digest=sha256:9f58a078b853cdea48ed6515b09da962dfb7de00e6272f28a9f6f8f9b8ad68d1