Pith. sign in

Paper Citation Record · LEDGER

Iterative Reasoning Preference Optimization

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2404.19733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.19733 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:10:53.414173Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 169cde8f-8921-4e81-b7f9-13c9b775fde9 · inbound

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive cites this paper.

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Iterative Reasoning Preference Optimization

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T23:04:44.466816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T23:04:44.287660Z digest=sha256:b00529a78e24b85862ffbc9421b12b7ef14464a2e29fbb047d720a1419cac525

Observation ba3dc64d-5934-44c6-af64-4f35cbbaee7b · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Iterative Reasoning Preference Optimization

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:16:17.391561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:859fc81f8391bc63d06a0f56218d45163474d10b14f7c74d1aa7a4f0c93ff55e

Observation 8e87f634-90c2-4fe7-84c1-96679874929b · inbound

Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search cites this paper.

Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search Iterative Reasoning Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:07:30.594150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T04:06:23.521344Z digest=sha256:6928346641eb6176227680b4f1ee8a351c7f9eff38479fd94fe5eb84a8bac870

Observation 08114dfe-9127-43df-b5e1-9ec5f849d039 · inbound

LIMO: Less is More for Reasoning cites this paper.

LIMO: Less is More for Reasoning Iterative Reasoning Preference Optimization

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:37.170645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T02:11:36.932541Z digest=sha256:b109eee16a925bcce3e5786309b34a6b87f083b2e5165684b62f664957199fc3

Observation 9b4352dd-93a6-40ec-bb58-8b885156257f · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation Iterative Reasoning Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.414173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.414173Z digest=sha256:ba983cf8eb617d809bb7b885e22781d3bb492ee209c44bbcb73b5779c1ef8355

Observation 65119daf-0aba-4cb7-a45a-6001e6a64167 · inbound

DPO-Shift: Shifting the Distribution of Direct Preference Optimization cites this paper.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Iterative Reasoning Preference Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.953811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.953811Z digest=sha256:7fa9c8db1b78262c9c20135490e2a778a7d1dc67442811f8911aa9041d44505c

Observation 7de1d6b8-31cd-43dd-8c9e-2f83d7250299 · inbound

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models cites this paper.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Iterative Reasoning Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.193451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.193451Z digest=sha256:b4f095afdc18810fd4a924a44fa69c48f999ef829acdc3d8a206b0b6540e2637

Observation f2ec365d-a006-4fae-827b-a7f7d9981353 · inbound

SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models cites this paper.

SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models Iterative Reasoning Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T20:57:07.669510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:57:07.669510Z digest=sha256:4ba29e297d0f29038b00c5b6b3905291299c1e0d1a941510ee348c0940fe515c

Observation d098173b-6f47-467a-becf-6decefae2354 · inbound

Online Knowledge Distillation with Reward Guidance cites this paper.

Online Knowledge Distillation with Reward Guidance Iterative Reasoning Preference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.338434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.338434Z digest=sha256:d232cf50c8ddaa406f0274472649b073185554a8e2396f5e974beb80ec1ca3a2

Observation ef665640-8de3-49d2-8e93-6562c9d2f7a0 · inbound

Frictional Agent Alignment Framework: Slow Down and Don't Break Things cites this paper.

Frictional Agent Alignment Framework: Slow Down and Don't Break Things Iterative Reasoning Preference Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:01.534554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:01.534554Z digest=sha256:2a4267d8f7aff3cff76e8e0cf1049df99c5ef6175720971e21fd334593ffab20

Observation 6b8c74b3-c7d2-4fd9-9fa0-e88b8aaf2d78 · inbound

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning cites this paper.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Iterative Reasoning Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.881145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.881145Z digest=sha256:4d504788a546fd80e885eba1b68b58f5c24ca0414227c3c8177cff34313ff1ce

Observation f5605bbe-f396-4310-8d6b-4e4fd10e59b7 · inbound

Control-R: Towards controllable test-time scaling cites this paper.

Control-R: Towards controllable test-time scaling Iterative Reasoning Preference Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:20.784306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:14:20.784306Z digest=sha256:c6d6fd3d21807d26def0124c8fda24be634bbf65b5a96c49f5be829a6cfec29c

Observation 2e3e4396-081c-46b7-96e1-774db6aa8875 · inbound

PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization cites this paper.

PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization Iterative Reasoning Preference Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:47:50.045904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:47:50.045904Z digest=sha256:dabaec672a4af6b1c7788b26539daf6f4ec56ddce822f199276497e7e0547298

Observation 1c8bc8ab-dfff-4b38-b740-7d91927a1c8d · inbound

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment cites this paper.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Iterative Reasoning Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.682978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.682978Z digest=sha256:027e9bd155c693ff99a9bf72219a8e42f6a9a9a46bcc321ed14bbd30ccb1d480

Observation c3fc33cf-e7f8-400a-914a-28d445eab0cf · inbound

Optimising Language Models for Downstream Tasks: A Post-Training Perspective cites this paper.

Optimising Language Models for Downstream Tasks: A Post-Training Perspective Iterative Reasoning Preference Optimization

Reference 166

Resolution
unresolved
no resolver link, observed 2026-08-06T22:44:44.374791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:44:44.374791Z digest=sha256:cb20c8b5822d3cd1e6f71af526b9b4237648d93c4245e4bbb10fa99ee4b36524

Observation d2bfe77f-6232-485e-94f8-8b69dbbaa289 · inbound

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought cites this paper.

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought Iterative Reasoning Preference Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:37.514013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:37.514013Z digest=sha256:a844ff16533bf521bd90e3ca498009e08d3010becb052ec9d2b3f1cc15c91fa1

Observation 2e72cd4f-3d77-4dc0-bd7a-2c2772aac227 · inbound

Technical Report of TeleChat2, TeleChat2.5 and T1 cites this paper.

Technical Report of TeleChat2, TeleChat2.5 and T1 Iterative Reasoning Preference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:22.729159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:43:22.729159Z digest=sha256:d647bf0e65d6d4e1a21235ef0be3ba50d24fb77d53506f0d4ac5baed6eba16e8

Observation 32f8d4a6-ad75-4157-aa33-a7f4d287f706 · inbound

Learning to Configure Agentic AI Systems cites this paper.

Learning to Configure Agentic AI Systems Iterative Reasoning Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.393255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T13:06:56.207692Z digest=sha256:46b12cb07ed0d60012254d03a28e2d7ac6f52583bc1a7c005d42c2a20de15aa5

Observation 0a03a2dc-8c2b-4957-9d60-27c9f5ecf119 · inbound

Learning to Configure Agentic AI Systems cites this paper.

Learning to Configure Agentic AI Systems Iterative Reasoning Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:31:29.563857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T11:29:45.134060Z digest=sha256:77989bb0c274eda39f6ff88f7c60fa63d3cd52c04574a619fe0da1042b5f010d

Observation 225b7731-7a20-46dc-917b-d9c40ee5da45 · inbound

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models cites this paper.

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models Iterative Reasoning Preference Optimization

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:41:02.247061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:38:01.208136Z digest=sha256:a0c3c85db340d4011a5c5f7ed1334d639477677d9d79bf63ac7a141aee13cd83

Observation e2d574a8-4dd5-4852-b6cb-b2afba6acddc · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Iterative Reasoning Preference Optimization

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.375076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:2dbad178700411bcd855a0239342cdaf67ccd754eab4475baafcad13e4ef3781

Observation 590e9133-a022-4fef-9426-7f7586cede03 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Iterative Reasoning Preference Optimization

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.577200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:c0873ff8bc64608cf906eda35f92437f9022a10c525668f02d65ea354836c019

Observation a094ac28-9420-49a4-a90c-fe37fd65c469 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Iterative Reasoning Preference Optimization

Reference 186

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.562104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:d9c20ff1e4ae37718d11b0c524273c90d58fd4a1e5785e67eb5c01f4b3328606

Observation b1a7437f-f596-499d-ab0c-7fc41ea3c2cf · inbound

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training cites this paper.

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training Iterative Reasoning Preference Optimization

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:29:16.717946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:07:28.660119Z digest=sha256:ff1e225de96fa1f858f3075998293a8f3b4194e4e59818ee3c79d71fd5e41983

Observation 69c53d56-ac18-4f74-8d92-7bb60224c84a · inbound

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation cites this paper.

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation Iterative Reasoning Preference Optimization

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:34:26.755731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T08:30:37.016856Z digest=sha256:d1d780f77cbcaa780bd3997135eba0f65853d4d759664669355087146a3d206c

Observation c380513d-5864-4ce4-8214-e693e6d539a8 · inbound

Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement cites this paper.

Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement Iterative Reasoning Preference Optimization

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:04:28.460655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T07:55:28.254309Z digest=sha256:ecf5d7efa9248185010bbdc02cd7fc763bb63c30d8a109f815d44082729d82ea

Observation ca1d4301-fb79-49f9-a511-757e94e3061b · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Iterative Reasoning Preference Optimization

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:72cf69e969d733427aa6b1e3ff95b640be0d5992740261a9aed5495c0df3ffb4

Observation 511bdc20-4e3c-465d-951f-fbe531153738 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Iterative Reasoning Preference Optimization

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.278322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.278322Z digest=sha256:af3c554e933429bfc6b7e38f520fa580ad8479a58c5af03699bab1cc0f4a35c8

Observation a635ae33-147a-4384-a69a-afc62e4a058e · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Iterative Reasoning Preference Optimization

Reference 178

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:36.276989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:36.276989Z digest=sha256:9d900a4fa3c0d625bb680cab3ef0838da893b0dfeacac13ebdad34705d0f4a2a