Pith. sign in

Paper Citation Record · LEDGER

Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2503.12854.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12854 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:43:44.523299Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:02:33.817155Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1fd88e75-a1de-45c0-bb98-8b4defc56f8f · inbound

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models cites this paper.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.523299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.523299Z digest=sha256:5ac1aee0df9e9fb9f1a04f4f2c435316417f5afbd588ac8cb8144625c68bb6d3

Observation beacce03-b4a2-4061-b7a1-b2e2731349e0 · inbound

Mathesis: Towards Formal Theorem Proving from Natural Languages cites this paper.

Mathesis: Towards Formal Theorem Proving from Natural Languages Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:38.778526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:38.778526Z digest=sha256:bc3df4d12866ccc03a79ef8dc2f92d3bd68595ab496b825494be8a8ad48d7f90

Observation 450bebd0-af38-4ece-ac96-ebf204566c01 · inbound

Enhancing Large Language Models through Structured Reasoning cites this paper.

Enhancing Large Language Models through Structured Reasoning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:59:47.417741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:59:47.417741Z digest=sha256:003da5673bc3b3268abc0be9b13c2118e37b1988851e53a3273ad66008efe017

Observation e2d8a1c9-c418-4448-bdfa-652d2a27dfff · inbound

Sample-efficient LLM Optimization with Reset Replay cites this paper.

Sample-efficient LLM Optimization with Reset Replay Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:41:54.738087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T23:38:08.044450Z digest=sha256:cb7357536a46391a88a2f4921d0ab3fb62f0503892ad4d49ee6dbf5c8fbfbb49

Observation 62801cd4-ce5a-4f1e-abfb-8c526f5297b7 · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:40.888746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:40.888746Z digest=sha256:a9a688deaf8ef74e4b9a1a51b75ab9fc66579e741c1cb92713586d3921fd0a3f

Observation 921189dd-55a7-47c7-8c07-88391aa3e3ff · inbound

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework cites this paper.

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:49.468251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:33:49.468251Z digest=sha256:e633e8f9c4506c081a5318433374521bd31bdbdf13f05e6ae054ae15ff9efae7

Observation ae29fb2c-7c16-4ba5-b2d9-2e85298977f4 · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:25.124148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:25.124148Z digest=sha256:64079b34b931fdd46f62fae6afe42e0a49b4572876d6e844b8a09590e1514f5a

Observation aa1fc4f7-5128-4fc3-9ee6-e1f725596d84 · inbound

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning cites this paper.

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:05.074027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T01:15:12.985803Z digest=sha256:83913cfbe78f0181251c0c5cc4310d46ed0179c01836e04d9744e50e685c76eb

Observation 41a09afb-bd6c-4ae4-bfab-0ba069dbc723 · inbound

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models cites this paper.

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:25.953549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T06:06:53.959061Z digest=sha256:027c974d36890ec764d368b22a93e5279b200ede732bb1c1c9692a76e948eceb

Observation fd439143-50f9-432f-86e7-ec30bd916e06 · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:53.895403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:ea4045fc111a80baac52bb57240f4af10cd06cbf8aeb418ea8ed149f05e2be5a

Observation 8c53353d-3041-44a5-9005-e0d8f5f1affb · inbound

How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR cites this paper.

How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:24:00.555652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T06:22:55.286281Z digest=sha256:d0f569b0ee8c3a14050a1326602a033b91804de97ede7495e129293daa2a61cd

Observation 173be850-c0a6-47df-9987-16f2e2589fb1 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.528307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:2fca1d01c144832dbe11fef49ff70a1c6b4966b0968ec3e53756ce11a790cdfb

Observation a6ce62f4-2f0c-435d-9727-edc4d7977589 · inbound

Enhancing LLM Metacognition via Cognitive Pairwise Training cites this paper.

Enhancing LLM Metacognition via Cognitive Pairwise Training Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:33.820741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T19:01:18.153145Z digest=sha256:09a47cfba1dae3d01811941a47f3ad3f08ffbcfddd54a7c0deca86f6b3404cf4