Pith. sign in

Paper Citation Record · LEDGER

Token-level Direct Preference Optimization

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2404.11999.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.11999 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:43.539104Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:02.074671Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 078580d8-a1a5-4ca6-b1db-d8247483d942 · inbound

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types cites this paper.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Token-level Direct Preference Optimization

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.487999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:e9580fd396ed5f04afd0ff3a58539c28eaa17a39b4137ab0604680dfc9efcd54

Observation ba846ef6-2e07-4e29-a80a-bb825c1e5a78 · inbound

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd cites this paper.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Token-level Direct Preference Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.958038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.958038Z digest=sha256:0b65a81253724d89c9bc130c174153cdf1230f68f9e714358bbb1be9e68d2c0a

Observation 412575e6-f16c-4f16-b4a6-be6655f7d600 · inbound

Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability cites this paper.

Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability Token-level Direct Preference Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T05:43:32.080953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:43:32.080953Z digest=sha256:c5ff225528aae4251c9e00fe5447f83d0c400f4c0feb948ac42405ad0ef57537

Observation 44648541-cb76-49be-8203-d9cce74a5e55 · inbound

T-REG: Preference Optimization with Token-Level Reward Regularization cites this paper.

T-REG: Preference Optimization with Token-Level Reward Regularization Token-level Direct Preference Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:56.448673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:15:56.448673Z digest=sha256:73262dab6967078c655eaa20fbf3fb4d8a3eaa21c0a5f633513e611931454e7e

Observation df554fa1-4c4e-4e47-8839-3d33f48fea7a · inbound

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment cites this paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Token-level Direct Preference Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.898358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.898358Z digest=sha256:dd2212df453661787841dcce2058c57f0ba6da2ac5725d657205cbc5ffc4fc83

Observation 5614f044-9ccb-4a5a-902d-271b8ddd9c13 · inbound

SDPO: Segment-Level Direct Preference Optimization for Social Agents cites this paper.

SDPO: Segment-Level Direct Preference Optimization for Social Agents Token-level Direct Preference Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:28:56.807793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:28:56.807793Z digest=sha256:21a868068dbf9136e2409db2f9379bdffb0d43362ed617befa77d3d48be7895d

Observation 404700b7-9a2b-4d46-9a65-bf8dc02287f0 · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation Token-level Direct Preference Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.442993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.442993Z digest=sha256:e827b9a72bf5bfb3767ccd51e7010ffa7dbe92a0a8f69d4f5397cf50fd8c71f4

Observation 10e4deab-3929-4688-bbd8-84862fb74085 · inbound

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models cites this paper.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Token-level Direct Preference Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.539104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.539104Z digest=sha256:55cf2b296871712a1fd181b05e8e9358dffa2cb94660aeb79e3e41257f987a97

Observation 5cd4cba5-630c-4b7c-9577-6982d9e9e87a · inbound

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm cites this paper.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Token-level Direct Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.424127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.424127Z digest=sha256:2dff91565024c477552c674ee5742735f101227e9bd2b63673e7dd3a445e2105

Observation fd73f9b8-ffdf-4584-8b9a-66948967a9b3 · inbound

A Survey on Progress in LLM Alignment from the Perspective of Reward Design cites this paper.

A Survey on Progress in LLM Alignment from the Perspective of Reward Design Token-level Direct Preference Optimization

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:06.958529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:52:06.958529Z digest=sha256:7a4bcd8aac11e83fd01d18ac45b349ca2140afc84060007020342e73d96322bd

Observation c6c3cc89-936e-48cb-92b4-9bd3348890fb · inbound

Policy-labeled Preference Learning: Is Preference Enough for RLHF? cites this paper.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Token-level Direct Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.862316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.862316Z digest=sha256:d82c260695d2f9a9222bba526facf8fc63f2c182375d46b56eacdd0e3436fd92

Observation 113a36c2-1c1d-4bd1-90a2-8ee91dc66bb3 · inbound

SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment cites this paper.

SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment Token-level Direct Preference Optimization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:51.057004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:51.057004Z digest=sha256:156c223dc28067f0d48c7dcf641d490aefbadd0ab2fb99829d1ef5d7aba1b95e

Observation 8cc37db7-4a5c-40ce-99ae-84eaceba409b · inbound

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning cites this paper.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Token-level Direct Preference Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.646894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.646894Z digest=sha256:7458a4b57c3faad784d9b998a80592713dba2f2f9eb44b07cb01dfd06643dac2

Observation 74b81ec1-c15b-4fb3-a0bb-fb078d10d9bd · inbound

Learning Safety Constraints for Large Language Models cites this paper.

Learning Safety Constraints for Large Language Models Token-level Direct Preference Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.341645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.341645Z digest=sha256:16b8d855adb4ccbbf2b4a90b7c294faf6be10dd26406b6bbc84f7d3b1f39fb80

Observation afa98307-1de3-4953-b0fe-9cce2510cb6e · inbound

AI Agent Behavioral Science cites this paper.

AI Agent Behavioral Science Token-level Direct Preference Optimization

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-07T11:00:54.142086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:00:54.142086Z digest=sha256:414b24662450c03fb3310db793eda717979c5625a5b9c7aa6b2a6b7af8ec123d

Observation b868000e-f0a3-40b4-b6c6-1f723f5999fd · inbound

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations cites this paper.

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations Token-level Direct Preference Optimization

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T06:12:07.126000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-19T06:09:26.269452Z digest=sha256:47c953962625b5ee8b9ce3b5633404e23cf0a90d56c310cf5fa48751f44e5386

Observation c89f84c3-6e88-4ad9-9a0f-a6ab33c4fc8e · inbound

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap cites this paper.

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap Token-level Direct Preference Optimization

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:50:47.536613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T23:46:24.208438Z digest=sha256:99a6336ccdb2eacc44f9338869af079a349426a459624908414481c62634b56b

Observation d8f6c5cd-fbb1-44ae-8344-3539981e24fd · inbound

FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus cites this paper.

FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus Token-level Direct Preference Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T12:53:42.473198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:53:42.473198Z digest=sha256:28e1b2d49d6ad4c694cc7402df3979fc2a39e77ea299b7367d91d4d4d62b6669

Observation eb7e1abf-09c4-4a9e-9392-c24a0c7a4d6c · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment Token-level Direct Preference Optimization

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.532009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:da4980e13796b13c1e25628c99b4b0945b8fef15c5d6783a12631064d7066130

Observation 39310ff7-28d9-4fea-889c-7381e850af02 · inbound

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning cites this paper.

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning Token-level Direct Preference Optimization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T11:27:22.852709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:27:22.852709Z digest=sha256:2aa7cccb3d1bd85732c5458408e9ea66fd8d80e023e952490c473558831d8c72

Observation 107e3e7e-742e-4fc4-8ce9-5a2438b9cac2 · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Token-level Direct Preference Optimization

Reference 226

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:31:24.845889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T00:29:07.951709Z digest=sha256:14a424eff04ac8c1f7f37326e3db55cc5f5d6a1fd53a45eded11a8d77b515095

Observation 7edaff37-1b18-4831-ae58-90fd720dcab3 · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Token-level Direct Preference Optimization

Reference 226

Resolution
unresolved
no resolver link, observed 2026-08-03T18:19:31.598779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:19:31.598779Z digest=sha256:2311d49ada8dd34c1f0118f7370dbc63bb34acbab38ccaee486dc5c0778957f7

Observation e113064d-c96d-4e01-a4e1-845eef5436fa · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Token-level Direct Preference Optimization

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:37:28.588973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:1749ae09e50b7fb882e9ac089059c87f518cb6bd3017ef83c41fcdac670e3a9b

Observation b989e5c4-d7dc-4c55-9f54-8ec5a7a0c8d1 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Token-level Direct Preference Optimization

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:10:10.502504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:60e1a84762a827947883afcbad18cebebde3bcb58d9a10d9742262d87a263656

Observation 1004e4b8-4187-433f-a6bd-8bc92db8614c · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Token-level Direct Preference Optimization

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:14:10.968376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:e4a5fd36963188ef9d3d34bb497a0b710d728650dfe2006859fca4ab0d25e87a

Observation e8366876-ff4a-4a68-9242-88144359e01e · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Token-level Direct Preference Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:46.578112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:46.578112Z digest=sha256:925fde99ee1061b175a13ebecc04749c96e4770d7c1ab7a5b6f65f222d820984

Observation 05ac2df2-b9ff-45cf-9199-d8e5fcbf58e0 · inbound

RAD-DPO: Robust Adaptive Denoising Direct Preference Optimization for Generative Retrieval in E-commerce cites this paper.

RAD-DPO: Robust Adaptive Denoising Direct Preference Optimization for Generative Retrieval in E-commerce Token-level Direct Preference Optimization

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:06:30.691681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T19:05:00.826786Z digest=sha256:718508eab76711d007f2178b45ecaa9e4305a9b31ba566a4d66616273ac93d37

Observation 658ce455-d299-478c-9890-70c9f127eba5 · inbound

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization cites this paper.

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization Token-level Direct Preference Optimization

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:01.145338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:11:49.149334Z digest=sha256:2bfeadaf2bd6853d23eabdf0d79c8122ee30141b7484084b49e70ad70c94c144

Observation 3fd45605-c623-4937-a472-9f913ea078d2 · inbound

Step-level Denoising-time Diffusion Alignment with Multiple Objectives cites this paper.

Step-level Denoising-time Diffusion Alignment with Multiple Objectives Token-level Direct Preference Optimization

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:15:25.868966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:15:13.005136Z digest=sha256:105530fd0b04104d44bfd38aca40ae2e1ecfb87928b25d7b896eeacc397db1d5

Observation 1325efea-62bc-48c4-8707-c6fde731cd83 · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Token-level Direct Preference Optimization

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:41:44.376403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-09T19:26:57.596581Z digest=sha256:701a810016fd714460756ce7a64177145a517ec22c4b838ebc028ad6b3875aca

Observation f0ca4aaf-90d5-4706-bdf6-b147accde66f · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Token-level Direct Preference Optimization

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:55:56.600639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T02:07:21.806345Z digest=sha256:1e330f347d40a6afbccd70f082af2dd913d583d414d151d56390e0b030602715

Observation 57c9d52e-3b80-44bf-9d55-2c43f27631dc · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Token-level Direct Preference Optimization

Reference 192

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.310837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:257b66caedb416bedeeee346ffda5599a52418687c7fd3b6507216d3742a9aee

Observation 94827eba-847a-4d7b-9b89-3502ab126c78 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Token-level Direct Preference Optimization

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:05:47.253592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T22:15:53.639048Z digest=sha256:06be173378caa503db7ed98e315ad4cd303e53cb9d3abbf4938f3296e260dc9c

Observation a8e2813e-a815-4476-8a2a-e98ae7d2016c · inbound

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning cites this paper.

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning Token-level Direct Preference Optimization

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:47:23.097979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T20:08:29.208550Z digest=sha256:550098c7da6a03e7ed311f36b13f38a340f0db2953818731abad94f478082d14

Observation a5ba103b-130b-4b35-b5e4-2252069b86d8 · inbound

Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs cites this paper.

Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs Token-level Direct Preference Optimization

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.076164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T09:52:20.508013Z digest=sha256:4d8f163cf9b6916dfd8e06b99c204a163382f4ed1ea61cb2c8e6f88ed4b8742d

Observation ac02c51d-db5c-49a4-8a67-44dc879a8e09 · inbound

Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement cites this paper.

Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement Token-level Direct Preference Optimization

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:04:28.458045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T07:55:28.254309Z digest=sha256:f785a779441f863b5774c2f5fe456cfbc09c5bd93ae261a71187fbd89c1767ef

Observation e631884d-355e-47d0-9e7c-6c823cd3d4cd · inbound

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift cites this paper.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Token-level Direct Preference Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:42.921201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:42.921201Z digest=sha256:37e4edff87193c7073452f8bcea610c23335805622ec2dfcb196c5b91aacd56a

Observation 6cb38f9a-977e-4430-8c10-a4f24d11d9dc · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Token-level Direct Preference Optimization

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:36.606953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:36.606953Z digest=sha256:629246ce21f6f985ca02303397e841c80ec8e7c911a15d2e1a8f9651315b8f8a

Observation 54293756-9f1b-415b-aaa1-5be4700cd828 · inbound

Token-Level Credit Assignment Optimization for Generative Document Retrieval cites this paper.

Token-Level Credit Assignment Optimization for Generative Document Retrieval Token-level Direct Preference Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:23:45.924802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:23:45.924802Z digest=sha256:8493da15eededf4dd4713bed03da5c54563a7910056d20ddad14d856fb2efc2a