Pith. sign in

Paper Citation Record · LEDGER

Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2503.12854.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12854 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:50:38.778526Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:02:33.817155Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation beacce03-b4a2-4061-b7a1-b2e2731349e0 · inbound

Mathesis: Towards Formal Theorem Proving from Natural Languages cites this paper.

Mathesis: Towards Formal Theorem Proving from Natural Languages Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:38.778526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:38.778526Z digest=sha256:339ccedf7c78f51ca73bd543b09e0db8262667c5144f6e9bb42b6020effaf54b

Observation 450bebd0-af38-4ece-ac96-ebf204566c01 · inbound

Enhancing Large Language Models through Structured Reasoning cites this paper.

Enhancing Large Language Models through Structured Reasoning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:59:47.417741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:59:47.417741Z digest=sha256:5860bfa85b4f863d476b7dbc965ee414267cccdebf0802a1ca8903e791ec657d

Observation e2d8a1c9-c418-4448-bdfa-652d2a27dfff · inbound

Sample-efficient LLM Optimization with Reset Replay cites this paper.

Sample-efficient LLM Optimization with Reset Replay Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:41:54.738087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T23:38:08.044450Z digest=sha256:ebb815bb1e4d646029f0b658ed428bd7b290eda8f09e4949afb1eb67a3f3c9dc

Observation 62801cd4-ce5a-4f1e-abfb-8c526f5297b7 · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:40.888746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:40.888746Z digest=sha256:47b5a52299c1970e3e45a9027c4fa71e85a23db33b25ae0d932c4b21ad064d9d

Observation 921189dd-55a7-47c7-8c07-88391aa3e3ff · inbound

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework cites this paper.

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:49.468251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:33:49.468251Z digest=sha256:bbc0224236a4e4b82dbc579fd73e66941fe700fbbbb654d3fa29553d05582353

Observation ae29fb2c-7c16-4ba5-b2d9-2e85298977f4 · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:25.124148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:25.124148Z digest=sha256:4e339fb7e6fc3633c1261c8f52be9ef097f0e2c3fd375bb9f520196141f5dc2f

Observation aa1fc4f7-5128-4fc3-9ee6-e1f725596d84 · inbound

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning cites this paper.

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:05.074027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:15:12.985803Z digest=sha256:859ca83ef72e85660d5cae7f5a1585ba8225d5e5747d116f64417f464c1f5bda

Observation 41a09afb-bd6c-4ae4-bfab-0ba069dbc723 · inbound

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models cites this paper.

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:25.953549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T06:06:53.959061Z digest=sha256:160d7076085621593a6c7b27b2428c92f24e5cfe54984c99b6c54667f07aeefe

Observation fd439143-50f9-432f-86e7-ec30bd916e06 · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:53.895403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:055b0224398f121a5a2f5b47aecf6746610cc2546b614d333f97946eb2bfbe77

Observation 8c53353d-3041-44a5-9005-e0d8f5f1affb · inbound

How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR cites this paper.

How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:24:00.555652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T06:22:55.286281Z digest=sha256:33f3894041646ac0245e6146600ab5588034cb3ebb7bf9758202f532ee81739c

Observation 173be850-c0a6-47df-9987-16f2e2589fb1 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.528307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:41c842805dc5a33e228b818329f5ae70e900342c6366dcaa34be8bbadfd4f079

Observation a6ce62f4-2f0c-435d-9727-edc4d7977589 · inbound

Enhancing LLM Metacognition via Cognitive Pairwise Training cites this paper.

Enhancing LLM Metacognition via Cognitive Pairwise Training Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:33.820741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T19:01:18.153145Z digest=sha256:f12f9c9e050788449033c5001b87950eae9b9d88e1ebed785bb872cb754de40b