Pith. sign in

Paper Citation Record · LEDGER

DPO Meets PPO: Reinforced Token Optimization for RLHF

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2404.18922.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.18922 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:10:53.450848Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:26:29.891495Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0247b814-5dab-4ed2-8f03-997e2df12d94 · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.450848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.450848Z digest=sha256:0b99dcf448ebf56510f1e1306bbe09cb61a933d5729d864ed5b308a3ae7da4b4

Observation 889509a3-c750-4706-85c5-c48ab90abe70 · inbound

Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning cites this paper.

Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:54.045722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:54.045722Z digest=sha256:1ab05ce1ba6905147e4fdf4db6ab706b584a5cc766e0c7f4229c2f4552275054

Observation d3bbb883-9740-4886-bf2b-91166959f7c4 · inbound

Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning cites this paper.

Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:36.509160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:36.509160Z digest=sha256:78af7669d6567ee08978f7109babbf8ced673c6013c2b84487febeadc85aee6f

Observation c4de7e60-b185-4ffc-a5e9-0029ea61649a · inbound

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models cites this paper.

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.199965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:32.199965Z digest=sha256:19ddf8da4a52e55a7f734167d583d4b7b8ac8645ba56e225b9bab365cf241e0c

Observation c31fd329-18ae-4c4a-af35-5ef3063ac485 · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.379167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.379167Z digest=sha256:613da14e69e611420ca478380fdbedd05e1e9fbfbcca136db06c322ebba17d02

Observation 4dc8e471-1f47-4447-9f08-d0e01cf6afef · inbound

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training cites this paper.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.770706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.770706Z digest=sha256:77207ee743eab13fe12a4a3d78264724ec682b4860328093fadad811806cfd45

Observation 99c591d9-b88d-4fbb-b681-f880534ff6fc · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.297970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.297970Z digest=sha256:779fe062e132eed7b01eb0069a4b970f37a673f2a816fa9fb07e2842024f3255

Observation 4b519ca3-8bb9-4684-b640-aa5f040cee17 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.171442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.171442Z digest=sha256:9830932da08e68c018babcf3d130875bca003b97c701bc95ea8cadd4a8e4180d

Observation 50ee1f6f-2450-4e74-b775-ef523dadca71 · inbound

Reinforced Language Models for Sequential Decision Making cites this paper.

Reinforced Language Models for Sequential Decision Making DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:56.616695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:56.616695Z digest=sha256:60372f191f36875ced2c4140e0f0d099d35372a92d560200a9abbc9d52110495

Observation 49e5d7b9-3262-4791-be36-9cff586dd82b · inbound

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training cites this paper.

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:30:35.448387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T20:30:17.581842Z digest=sha256:a27a6622b2907a14b5afc99cc58d70620ce57c6106a087c7888bff3e85f0a4a1

Observation 53e47d89-a31c-4e6e-af69-5298c870e248 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:45.181190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:45.181190Z digest=sha256:5726872ce7a88858b2e5561d071a632fdd4cf0732fcac0f476210645d4ebadfe

Observation 221083bd-c429-4d9a-9d64-47c01145b324 · inbound

Stabilizing Policy Optimization via Logits Convexity cites this paper.

Stabilizing Policy Optimization via Logits Convexity DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:09.115910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:09.115910Z digest=sha256:c6bff0039aa9ec9264c1eb3909f8bf34a5d6911cfc7b199128b2b38b97a17958

Observation c6edb85f-1d54-4e2e-b47c-4c321981fc4f · inbound

Data Agent: Learning to Select Data via End-to-End Dynamic Optimization cites this paper.

Data Agent: Learning to Select Data via End-to-End Dynamic Optimization DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:26:10.907667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T15:23:23.950152Z digest=sha256:a044837e7335daa07e4f6429eca57ceacb3f1e7f6f189fb35ba30fb19b0ef88f

Observation 67517349-28f0-43f7-bcc5-089ed583c2bc · inbound

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO cites this paper.

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:02.137107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:15:04.225059Z digest=sha256:76113d48e7a5db8710ca890133a48f282b43c952223d712ae360573d24733c6f

Observation 4f94fb72-2b09-4a9c-bbe7-9d8a3c520e93 · inbound

Leveraging RAG for Training-Free Alignment of LLMs cites this paper.

Leveraging RAG for Training-Free Alignment of LLMs DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:42:08.175821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T02:41:07.016045Z digest=sha256:631585dcba098d091fc22f9dcc16f5e5fe1c609056b17610d73aa279847baabb

Observation 5e8257b5-3f06-4443-a5d7-190b1c47cf5a · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:17.231849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:b22385cac3663205558ea1d323547787f9abf95c8aa00797f7cbae0e9a03778e

Observation 1d51c5b1-3850-4a7d-9b06-c6f90c87d07e · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:06.597442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:68a76ee55b38fd1c339317cf97ca2efe3a4c0f00ce581ab11f1df6efb03cb79d

Observation 3aaa981a-1563-4582-893a-02f8c5267fe0 · inbound

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition cites this paper.

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.011746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:44:14.878564Z digest=sha256:ecaa6d326c689ed91ce7bbc5692f813993da9910ce2fcb02eb2b13c16e5d0636

Observation bc0dafb7-0105-4645-91a7-0b0856f5d5c4 · inbound

Entropy Is Not Enough: Unlocking Effective Reinforcement Learning for Visual Reasoning via Vision-Anchored Token Selection cites this paper.

Entropy Is Not Enough: Unlocking Effective Reinforcement Learning for Visual Reasoning via Vision-Anchored Token Selection DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.893054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T09:58:34.557353Z digest=sha256:7e86252038f63af5ea55699c393c5d7afb5040bd30c67d68238148df17d4de42

Observation 697697b2-7bb0-4d44-972b-3bbb27b06c37 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 291

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:917368b496097e255526dea6ebcebe6607c6e1e9d7445ab5cbb656b91a5959ce

Observation 943e17f6-f6c4-493b-86dd-ea745a096419 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 292

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:06.696808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:06.696808Z digest=sha256:3860307f401f52bd5a5c511cd7fbae354a1f0e27fdfe58cbd70314f4475ea7c5

Observation ee81150d-a934-48b1-8599-f78542a401ea · inbound

Distilled Reinforcement Learning for LLM Post-training cites this paper.

Distilled Reinforcement Learning for LLM Post-training DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.185696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.185696Z digest=sha256:583914692c386027f8d47e14541fa441f65843c458fad5c3d8bdaa418176d8d4