Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:59:20.206994Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2506.14574.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:59:20.206994Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:12.210012Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T10:48:02.099479Z
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3ad100d9-c45f-4638-9378-5ed9748e2569 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Hindsight experience replay
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bbd70215-91af-4951-b798-b836fcaee5dd · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c9fa6ee3-4026-412a-8da2-bd38ac6842fd · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization On the weaknesses of reinforcement learning for neural machine translation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 78c8d9a9-286d-4134-ad09-942f9483a6d4 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Ultrafeedback: Boosting language models with scaled AI feedback
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 44c41176-263a-4e64-8e87-36373369656c · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Model alignment as prospect theoretic optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f0071c97-0b26-4991-92a9-674b2c2c068b · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64b30941-1df2-4f31-a4a2-c9c9e7725f52 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Direct Language Model Alignment from Online AI Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a540d8e-872c-4673-887f-cbc3c8e08264 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d123037-c166-44dc-b0d0-637531a95bbc · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization A., Choi, Y., and Hajishirzi, H
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fe5cdfaf-9896-4941-b5d5-01f5ff0bbc9b · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dd0ac76e-306c-4b3b-9352-4790887ba812 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization E., and Stoica, I
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7110caa9-8fb5-429d-8bf3-8fa3b2bd4db9 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fe53e676-c947-4b41-82fd-b97fbc51d34f · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization and Hutter, F
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde593bf-cd05-4d12-950e-4507bc753d8c · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Sim PO : Simple preference optimization with a reference-free reward
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a2bf27-688d-492d-84eb-f995dd1303ed · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Aligning CodeLLMs with Direct Preference Optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2ca7894-93c3-4ac1-8e17-1a2e6d2f3313 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization GPT-4 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cc2a536-2fbe-46f6-b235-e98933f2f72b · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization F., Leike, J., and Lowe, R
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6f390981-a170-46a6-9d8c-5abff0cba0ec · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Token-level Proximal Policy Optimization for Query Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5503b1c-4266-461c-b049-1f5c68aa6279 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Disentangling length from quality in direct preference optimization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d429dbc8-4ba8-4243-822c-5e29a3926c03 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization D., Ermon, S., and Finn, C
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cda84a96-082f-4e47-9b10-43126acd6890 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization From r to q^* : Your language model is secretly a q-function
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ec904eb3-fa10-4f52-bd62-24135df9f95e · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Proximal Policy Optimization Algorithms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7ef3a5c-3826-4a0a-b3ad-cb6aa5906d36 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Earlier tokens contribute more: Learning direct preference optimization from temporal decay perspective
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ed1b7731-cd1e-4065-8a89-1e86b47def95 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization V., Kostrikov, I., Su, Y., Yang, S., and Levine, S
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 61174b07-3329-4789-bbaf-2071270ecefb · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Gemini: A Family of Highly Capable Multimodal Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b3cf86-684d-4a55-b122-5f7dfc24864e · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Gemma 2: Improving Open Language Models at a Practical Size
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2bf55f2-6097-4376-ae9c-2fba496a0a3d · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization D., and Finn, C
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bb49bab6-7658-4085-94fe-bbcd71b8728c · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Interpretable preferences via multi-objective reward modeling and mixture-of-experts
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 51e64da4-8ebd-4d91-91a5-c7829650e1d4 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization A., Ostendorf, M., and Hajishirzi, H
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5102d338-e563-4f37-a5b3-c0f4b54cfe73 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Inverse- Q *: Token level reinforcement learning for aligning large language models without preference data
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ce39aa76-ffb9-4e6f-98a2-3b7aea2c304b · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Preference-grounded token-level guidance for language model fine-tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 48ac8dcc-c5fb-46f1-afa3-646180f2d26b · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization A dense reward view on aligning text-to-image diffusion with preference
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 87939929-964a-4e5a-88ea-1058304cd6c7 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73fd95e3-9db9-4fb9-96f6-ef65d94566d2 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Token-level direct preference optimization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 30e4808d-3d42-4cd7-9e0e-17f50d4987e8 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Judging LLM -as-a-judge with MT-Bench and Chatbot Arena
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f48c2fb8-b56e-43e0-8128-5d015c50ee54 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization DPO meets PPO : Reinforced token optimization for RLHF
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5d5d3081-9b40-43e8-8315-b047e71bbd0e · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Fine-Tuning Language Models from Human Preferences
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4428b8cf-84ce-460f-9649-6b7264d56101 · outbound
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization write newline
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84ec1dfe-f34e-4d75-8df3-c366e6e177a6 · inbound
Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 053fcd64-447e-461e-b378-fd25c27137ca · inbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.