Pith. sign in

Paper Citation Record · LEDGER

Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.05534.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05534 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:27:11.518340Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:17:30.387482Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 14537d1c-3b73-45bc-a6d9-917caa9cc877 · inbound

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator cites this paper.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.518340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.518340Z digest=sha256:95158ca8c3ac50b35dd0e899ea0ad1ec74584054b08ec489ec90f2929650c838

Observation 433f1a7c-75d5-46e2-820b-a19450e9d175 · inbound

Salamandra Technical Report cites this paper.

Salamandra Technical Report Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-08T04:58:33.868928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:58:33.868928Z digest=sha256:3c3c79fd7228a64717de64a43c630a403af8fd2872fa5f599461c775086eae0d

Observation 708f84de-c075-4007-ac47-454f2cf4c6e4 · inbound

Do not Abstain! Identify and Solve the Uncertainty cites this paper.

Do not Abstain! Identify and Solve the Uncertainty Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:49.727328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:49.727328Z digest=sha256:679c50a5d3a2ae9b6554a844c92f72f0e17f387302e0330b3e05d8db624e8a2c

Observation 1b87ab13-8601-4918-857f-42e1daca9835 · inbound

Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities cites this paper.

Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:26.559009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:26.559009Z digest=sha256:b2312087f2efcab48c98f85b561befd3295cfdc2f0285a8bfac7861ba2e762f5

Observation fa140b0e-9a8c-4b3a-9fd7-c0f48c8c0439 · inbound

Phi-Ground Tech Report: Advancing Perception in GUI Grounding cites this paper.

Phi-Ground Tech Report: Advancing Perception in GUI Grounding Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T10:28:56.164873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:28:56.164873Z digest=sha256:62047b5ba3b0e1c735658e01931d37a88c6ed82371bdaac25a9fce2da0c711ff

Observation e8130233-b89f-4f30-ba84-b90345c91fb6 · inbound

Gene-R1: Reasoning with Data-Augmented Lightweight LLMs for Gene Set Analysis cites this paper.

Gene-R1: Reasoning with Data-Augmented Lightweight LLMs for Gene Set Analysis Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:57.769300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:57.769300Z digest=sha256:df59b52469d666b0546fc6337c37e4675e667557568a93a4547d8869cba34a8b

Observation e3ae8dbc-d63f-48d6-a867-279e600ec9ad · inbound

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning cites this paper.

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:53.842108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:28:58.515666Z digest=sha256:0f779c6f520e79efc6b352769add8da1fab72106599c937278aaae0943f9ba3b

Observation c8adf3e0-6aa8-4421-ba5c-f644bfdc0e22 · inbound

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization cites this paper.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:16:10.962351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:ca32e33e3a58f40b3f06d04f58b50d5ab5e5f3b984357d54c02d57bc468fe6d2

Observation 44fa5945-8ebe-40ea-a314-e28e646912c4 · inbound

RPO-PDT: Demonstrating Role-Play-Based Knowledge Adaptation for Student Support Dialogue (Demonstration System) cites this paper.

RPO-PDT: Demonstrating Role-Play-Based Knowledge Adaptation for Student Support Dialogue (Demonstration System) Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:30.389257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T16:45:27.352655Z digest=sha256:9e0bed175b6390a9c6ed072e2c318b5e3d90e21a53a2357b08039e3bc7e619e4

Observation 3bebc50f-93a1-44a5-b1fc-32a2ce506455 · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 179

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:36.362432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:36.362432Z digest=sha256:b20f0e0db738e567f1858133c640f2dd763fb1f4a055637b05d26eeb157b1701