Pith. sign in

Paper Citation Record · LEDGER

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization

As of 17 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2506.14574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14574 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:59:20.206994Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:12.210012Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:02.099479Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ad100d9-c45f-4638-9378-5ed9748e2569 · outbound

This paper cites Hindsight experience replay.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Hindsight experience replay

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.665968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.056903Z digest=sha256:85cf6fc8898246d53a89e95471fdd928293dd972d60a9ac70fc0cf9078ec329e

Observation bbd70215-91af-4951-b798-b836fcaee5dd · outbound

This paper cites an unresolved cited work.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:59:20.655381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.061791Z digest=sha256:28300168d92459d2951b5a5841922cead8e15e6bf0aafb1243f98804ef7d5a2a

Observation c9fa6ee3-4026-412a-8da2-bd38ac6842fd · outbound

This paper cites On the weaknesses of reinforcement learning for neural machine translation.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization On the weaknesses of reinforcement learning for neural machine translation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.644355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.066821Z digest=sha256:eac25b24e9a0c3a08b10014aa884bc6f4fcbb70ea9001c1aa519883598dd0fc0

Observation 78c8d9a9-286d-4134-ad09-942f9483a6d4 · outbound

This paper cites Ultrafeedback: Boosting language models with scaled AI feedback.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Ultrafeedback: Boosting language models with scaled AI feedback

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.633850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.071474Z digest=sha256:ed12de1215d40650f88d011fd0deb1c0d33b9816918a9b526bcfe917cfad858b

Observation 44c41176-263a-4e64-8e87-36373369656c · outbound

This paper cites Model alignment as prospect theoretic optimization.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Model alignment as prospect theoretic optimization

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.615533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.075557Z digest=sha256:3b7d1c9a77266a62fd66962ec27f593fa9fe7b172ff05d99824fe500d652dd6f

Observation f0071c97-0b26-4991-92a9-674b2c2c068b · outbound

This paper cites The Llama 3 Herd of Models.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.079682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.079682Z digest=sha256:628e3496753a7f8c8fa80496ff965b38349165524988ef1bc24f55991b18ff96

Observation 64b30941-1df2-4f31-a4a2-c9c9e7725f52 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.083973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.083973Z digest=sha256:5053f1e2d68ba41e333642491fb01f28ec65a9b006a667788ad4dbcc8e5e1b20

Observation 8a540d8e-872c-4673-887f-cbc3c8e08264 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.087957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.087957Z digest=sha256:61630dcb2e8227939458736a0b4c7a7ecb6f5948d3d352ecd9b84342bd7ac5da

Observation 6d123037-c166-44dc-b0d0-637531a95bbc · outbound

This paper cites A., Choi, Y., and Hajishirzi, H.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization A., Choi, Y., and Hajishirzi, H

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.602280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.092673Z digest=sha256:2a7b897912c0573b2ba787141547f58e34b1c1e5e88f36e520214825eed4c2dd

Observation fe5cdfaf-9896-4941-b5d5-01f5ff0bbc9b · outbound

This paper cites an unresolved cited work.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:59:20.591409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.096570Z digest=sha256:476a91397bb977ed0a709008f8119244c5418b5d6589a4f6b7d64b1b21d7c0c7

Observation dd0ac76e-306c-4b3b-9352-4790887ba812 · outbound

This paper cites E., and Stoica, I.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization E., and Stoica, I

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.578932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.101212Z digest=sha256:71bc85aeff71d08f70af35b7e41c4b7c2bf15ab6b00fe20e59bdf21eaee258ac

Observation 7110caa9-8fb5-429d-8bf3-8fa3b2bd4db9 · outbound

This paper cites an unresolved cited work.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:59:20.564851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.105042Z digest=sha256:049ab47719b1ab3efbc7cb83488f50f893372f079d9c0a7ea3e7e359c1413fff

Observation fe53e676-c947-4b41-82fd-b97fbc51d34f · outbound

This paper cites and Hutter, F.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization and Hutter, F

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.108793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.108793Z digest=sha256:6fed16cfd3b92120c5c540eb10cef8195e4c319f015b3b037c2ca1313c4e793e

Observation bde593bf-cd05-4d12-950e-4507bc753d8c · outbound

This paper cites Sim PO : Simple preference optimization with a reference-free reward.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Sim PO : Simple preference optimization with a reference-free reward

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.113317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.113317Z digest=sha256:ba85763ebddb19637e15978fbd27508d965387acdda1d6845048f17c9dda1e94

Observation 95a2bf27-688d-492d-84eb-f995dd1303ed · outbound

This paper cites Aligning CodeLLMs with Direct Preference Optimization.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Aligning CodeLLMs with Direct Preference Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.116800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.116800Z digest=sha256:feb3d6903773911da3f7d6e0c6252aa4d72f32d329052fe7d0fa9b092920f150

Observation a2ca7894-93c3-4ac1-8e17-1a2e6d2f3313 · outbound

This paper cites GPT-4 Technical Report.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization GPT-4 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.121968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.121968Z digest=sha256:a3ecdc6ce66add4f389b77230e5f5f2de461499abaf2848f802d62bb51ac1653

Observation 0cc2a536-2fbe-46f6-b235-e98933f2f72b · outbound

This paper cites F., Leike, J., and Lowe, R.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization F., Leike, J., and Lowe, R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.540547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.125966Z digest=sha256:5a31a5a11998bfaab8a9790fcfaaeb751ba57662be63a9c499a25cf5706da966

Observation 6f390981-a170-46a6-9d8c-5abff0cba0ec · outbound

This paper cites Token-level Proximal Policy Optimization for Query Generation.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Token-level Proximal Policy Optimization for Query Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.130052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.130052Z digest=sha256:d2a14e3d34b39e57eb80c9c37a56a1f839e9aa507ebe4ccd9a86e03b13fd16c1

Observation a5503b1c-4266-461c-b049-1f5c68aa6279 · outbound

This paper cites Disentangling length from quality in direct preference optimization.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Disentangling length from quality in direct preference optimization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.528127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.134128Z digest=sha256:694e45fd7403876d7ca2204c127ed9109004cda9aa80ba7e597e00575e0abadd

Observation d429dbc8-4ba8-4243-822c-5e29a3926c03 · outbound

This paper cites D., Ermon, S., and Finn, C.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization D., Ermon, S., and Finn, C

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.515215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.138652Z digest=sha256:d835c69a73677cbd5aa8bce2c1981feb9fde74237537799a831ed54e88e794b6

Observation cda84a96-082f-4e47-9b10-43126acd6890 · outbound

This paper cites From r to q^* : Your language model is secretly a q-function.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization From r to q^* : Your language model is secretly a q-function

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.504867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.142800Z digest=sha256:138afb2dc3822fb068ee30697d1f1d05e82222a9129529da669f2dc8ee3e5609

Observation ec904eb3-fa10-4f52-bd62-24135df9f95e · outbound

This paper cites Proximal Policy Optimization Algorithms.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.146238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.146238Z digest=sha256:eaaa76715c56f8b944dd7d254fa635cd51bb56b80a84e9b898bb7d440ddf0ce2

Observation d7ef3a5c-3826-4a0a-b3ad-cb6aa5906d36 · outbound

This paper cites Earlier tokens contribute more: Learning direct preference optimization from temporal decay perspective.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Earlier tokens contribute more: Learning direct preference optimization from temporal decay perspective

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.491433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.150241Z digest=sha256:3fa8681fe81dc1c6c18e33f878c30c6a7c0f2920ddda20e84acdc962e07dbdbe

Observation ed1b7731-cd1e-4065-8a89-1e86b47def95 · outbound

This paper cites V., Kostrikov, I., Su, Y., Yang, S., and Levine, S.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization V., Kostrikov, I., Su, Y., Yang, S., and Levine, S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.477722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.154039Z digest=sha256:b45299f28373959edb231142415857b0ad3615b0fce0b488e70908949e146f15

Observation 61174b07-3329-4789-bbaf-2071270ecefb · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Gemini: A Family of Highly Capable Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.157866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.157866Z digest=sha256:1e731da31855eda14be9bbf174391c51381411c11c076b4067a94d879062291a

Observation 17b3cf86-684d-4a55-b122-5f7dfc24864e · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Gemma 2: Improving Open Language Models at a Practical Size

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.161900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.161900Z digest=sha256:c39e175ef644668b7f6ef31eb443db801076264dbdf79e19b62a73485e69928e

Observation b2bf55f2-6097-4376-ae9c-2fba496a0a3d · outbound

This paper cites D., and Finn, C.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization D., and Finn, C

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.466316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.165866Z digest=sha256:d35409ab4c8ccc0983c59e8920f5e5daae34837fff17f0155e3a15d657ff2795

Observation bb49bab6-7658-4085-94fe-bbcd71b8728c · outbound

This paper cites Interpretable preferences via multi-objective reward modeling and mixture-of-experts.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Interpretable preferences via multi-objective reward modeling and mixture-of-experts

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.455391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.170856Z digest=sha256:fdc17dc43a9ee0c5f7a37e227e16bb07989845562c5dda25fcaedc6dcfcfc4d0

Observation 51e64da4-8ebd-4d91-91a5-c7829650e1d4 · outbound

This paper cites A., Ostendorf, M., and Hajishirzi, H.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization A., Ostendorf, M., and Hajishirzi, H

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.443704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.174094Z digest=sha256:5ceff2f050e6a1689c5037e4f29cd3f819528de67aad5786467160cac8d399ac

Observation 5102d338-e563-4f37-a5b3-c0f4b54cfe73 · outbound

This paper cites Inverse- Q *: Token level reinforcement learning for aligning large language models without preference data.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Inverse- Q *: Token level reinforcement learning for aligning large language models without preference data

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.431957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.177603Z digest=sha256:304e25b2bbf739fd2cd2d401b8a3631ee05df09499dcfb4bd40905c658420d24

Observation ce39aa76-ffb9-4e6f-98a2-3b7aea2c304b · outbound

This paper cites Preference-grounded token-level guidance for language model fine-tuning.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Preference-grounded token-level guidance for language model fine-tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.418301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.181731Z digest=sha256:33128c8d5a29d2a2a3e6dce18239be023bf82a2b2854bce859fa3d75b0dd7ddf

Observation 48ac8dcc-c5fb-46f1-afa3-646180f2d26b · outbound

This paper cites A dense reward view on aligning text-to-image diffusion with preference.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization A dense reward view on aligning text-to-image diffusion with preference

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.407009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.185336Z digest=sha256:88245502abe78e75da461faeb9ff4737cba984295aec1698fff62f444bac4b51

Observation 87939929-964a-4e5a-88ea-1058304cd6c7 · outbound

This paper cites Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.189439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.189439Z digest=sha256:8a8a9c58bb56ee28c3ab065a6b5568b5d8cdca430835e2979e9c708929db1bf4

Observation 73fd95e3-9db9-4fb9-96f6-ef65d94566d2 · outbound

This paper cites Token-level direct preference optimization.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Token-level direct preference optimization

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.396281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.193487Z digest=sha256:71865a43dd40feb82b584a3c33b01673dd97622fef953c61f8533f7bb9441811

Observation 30e4808d-3d42-4cd7-9e0e-17f50d4987e8 · outbound

This paper cites Judging LLM -as-a-judge with MT-Bench and Chatbot Arena.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Judging LLM -as-a-judge with MT-Bench and Chatbot Arena

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.385309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.196578Z digest=sha256:90f8fda9211d38ccda42ece8400aa499131e20fcbd01a483f5dc20397543f004

Observation f48c2fb8-b56e-43e0-8128-5d015c50ee54 · outbound

This paper cites DPO meets PPO : Reinforced token optimization for RLHF.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization DPO meets PPO : Reinforced token optimization for RLHF

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.374564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.199953Z digest=sha256:c577a7fd8caafac6a524e8aa7ba1b62198ec2638266303d80425cef9cac5a32f

Observation 5d5d3081-9b40-43e8-8315-b047e71bbd0e · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Fine-Tuning Language Models from Human Preferences

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.203188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.203188Z digest=sha256:8341fab37e343b50e2cacddd8ca2d12682f6ab7f519a886557e2f4be6e86698c

Observation 4428b8cf-84ce-460f-9649-6b7264d56101 · outbound

This paper cites write newline.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization write newline

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.206994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.206994Z digest=sha256:9bee25ca87735df05b392b68284b12120c9d7c61400aa9e01ddf1a16e96fbb59

Pith citing papers

Observation 84ec1dfe-f34e-4d75-8df3-c366e6e177a6 · inbound

Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs cites this paper.

Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:02.101463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:52:20.508013Z digest=sha256:be3734819e8abc6f1873648cb6b9ac38bbf041d268c5f5ea5f560b19dd433513

Observation 053fcd64-447e-461e-b378-fd25c27137ca · inbound

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization cites this paper.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.210012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.210012Z digest=sha256:65c74a2ed051666c3f62ce5f7cde98584124f5409cb1c4f36a6a4a8f559fbf4e