Pith. sign in

Paper Citation Record · LEDGER

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization

As of 16 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2506.14574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14574 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:59:20.206994Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:12.210012Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:02.099479Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ad100d9-c45f-4638-9378-5ed9748e2569 · outbound

This paper cites Hindsight experience replay.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Hindsight experience replay

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.665968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.056903Z digest=sha256:73061d8fc71bd4e8ea5f3e6468f47125a7ac5832ee335b347fb03b8e7c932fb3

Observation bbd70215-91af-4951-b798-b836fcaee5dd · outbound

This paper cites an unresolved cited work.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:59:20.655381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.061791Z digest=sha256:993d637518148dd4cc34fa5c0f7a8ffdc6e8dc84a07d22f77f252b6e0dd14237

Observation c9fa6ee3-4026-412a-8da2-bd38ac6842fd · outbound

This paper cites On the weaknesses of reinforcement learning for neural machine translation.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization On the weaknesses of reinforcement learning for neural machine translation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.644355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.066821Z digest=sha256:b04e0fbdac158d0808a7985966280fda713143469bac21db39a603c91a93fb0d

Observation 78c8d9a9-286d-4134-ad09-942f9483a6d4 · outbound

This paper cites Ultrafeedback: Boosting language models with scaled AI feedback.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Ultrafeedback: Boosting language models with scaled AI feedback

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.633850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.071474Z digest=sha256:e588fd5ddbe86522053727c358b2ec52d78592bb3f5b19306decd8b0ab011efc

Observation 44c41176-263a-4e64-8e87-36373369656c · outbound

This paper cites Model alignment as prospect theoretic optimization.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Model alignment as prospect theoretic optimization

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.615533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.075557Z digest=sha256:e4091fcba20b147fb57d899cfb929338926157bd01667c40935da0174bf7bdfa

Observation f0071c97-0b26-4991-92a9-674b2c2c068b · outbound

This paper cites The Llama 3 Herd of Models.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.079682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.079682Z digest=sha256:9ac6144c8103300582f1e5028c1d600c1d1f520f9a37c567b948516dac29167a

Observation 64b30941-1df2-4f31-a4a2-c9c9e7725f52 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.083973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.083973Z digest=sha256:8cddac8061bc65c7ffb2b998f9204e3c01da9af819d74c7365bf483633dfeb1b

Observation 8a540d8e-872c-4673-887f-cbc3c8e08264 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.087957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.087957Z digest=sha256:e816ec1bfb46a9f8ca70b300e1e74e8c018dbc5e3a1ef1e5e007a08ca9147de4

Observation 6d123037-c166-44dc-b0d0-637531a95bbc · outbound

This paper cites A., Choi, Y., and Hajishirzi, H.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization A., Choi, Y., and Hajishirzi, H

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.602280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.092673Z digest=sha256:c7cb1fe31be2cb1ed7737704152585b9654b1279e7ef984714b2792844c6a92b

Observation fe5cdfaf-9896-4941-b5d5-01f5ff0bbc9b · outbound

This paper cites an unresolved cited work.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:59:20.591409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.096570Z digest=sha256:0a75a4463ddac1b13e37cc393da29c28171f3cff17f69723e4b0338541235325

Observation dd0ac76e-306c-4b3b-9352-4790887ba812 · outbound

This paper cites E., and Stoica, I.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization E., and Stoica, I

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.578932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.101212Z digest=sha256:e6c498796619592e51def8ef1ca46f552e242e05ee0ef5b4604fb4ccd6d06f5c

Observation 7110caa9-8fb5-429d-8bf3-8fa3b2bd4db9 · outbound

This paper cites an unresolved cited work.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:59:20.564851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.105042Z digest=sha256:15c9a89de450c641210ca63d80a73ebf8a51c5aab080d372c5b3b4d4461989ac

Observation fe53e676-c947-4b41-82fd-b97fbc51d34f · outbound

This paper cites and Hutter, F.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization and Hutter, F

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.108793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.108793Z digest=sha256:d070ac662cca277be34cde90be3f0d4454538d66b97a349a7da4ecde1a39ae6d

Observation bde593bf-cd05-4d12-950e-4507bc753d8c · outbound

This paper cites Sim PO : Simple preference optimization with a reference-free reward.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Sim PO : Simple preference optimization with a reference-free reward

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.113317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.113317Z digest=sha256:d5354c8bf693306f285f080a4545ff33a1c398bd7ae526bfc7dcf8eed4716d30

Observation 95a2bf27-688d-492d-84eb-f995dd1303ed · outbound

This paper cites Aligning CodeLLMs with Direct Preference Optimization.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Aligning CodeLLMs with Direct Preference Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.116800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.116800Z digest=sha256:2e41df541bb4027546c97457403b3808bf73ee9e4641330fe0f5ad9aacbf994b

Observation a2ca7894-93c3-4ac1-8e17-1a2e6d2f3313 · outbound

This paper cites GPT-4 Technical Report.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization GPT-4 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.121968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.121968Z digest=sha256:fe150b69e8a446d04bc31c6c7ffea9acc331794a317f8c3ca26ac773ed8aaa92

Observation 0cc2a536-2fbe-46f6-b235-e98933f2f72b · outbound

This paper cites F., Leike, J., and Lowe, R.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization F., Leike, J., and Lowe, R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.540547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.125966Z digest=sha256:1ca5309a0b263039a675daa8b87961997db33ec332181794cc9447f210f0a747

Observation 6f390981-a170-46a6-9d8c-5abff0cba0ec · outbound

This paper cites Token-level Proximal Policy Optimization for Query Generation.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Token-level Proximal Policy Optimization for Query Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.130052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.130052Z digest=sha256:6a23265cf2f8f14e694ec5b90d32f9f6d29bc3f32cf3dec6373e4b9cc39a63d5

Observation a5503b1c-4266-461c-b049-1f5c68aa6279 · outbound

This paper cites Disentangling length from quality in direct preference optimization.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Disentangling length from quality in direct preference optimization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.528127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.134128Z digest=sha256:3b36bf51c5ee4f0d9cc85379cb80cbc44c881192977bbd712ebbf87bfdf2ac3e

Observation d429dbc8-4ba8-4243-822c-5e29a3926c03 · outbound

This paper cites D., Ermon, S., and Finn, C.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization D., Ermon, S., and Finn, C

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.515215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.138652Z digest=sha256:5eec42165f1b81670859aa72efcdbae13ffd7e6778f19e9545197022faf0418a

Observation cda84a96-082f-4e47-9b10-43126acd6890 · outbound

This paper cites From r to q^* : Your language model is secretly a q-function.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization From r to q^* : Your language model is secretly a q-function

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.504867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.142800Z digest=sha256:51a4ca72962b476384d8cfe279b749cc92bf5a69bf59750ffa7cce063d2f09b2

Observation ec904eb3-fa10-4f52-bd62-24135df9f95e · outbound

This paper cites Proximal Policy Optimization Algorithms.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.146238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.146238Z digest=sha256:348c8c9f60578699480587c62aa0c9489f7888fafe5573e59c28505327c52284

Observation d7ef3a5c-3826-4a0a-b3ad-cb6aa5906d36 · outbound

This paper cites Earlier tokens contribute more: Learning direct preference optimization from temporal decay perspective.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Earlier tokens contribute more: Learning direct preference optimization from temporal decay perspective

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.491433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.150241Z digest=sha256:da3209f19b2909a916bef319ac761c392c6b54be5582029c6a65d73a1d120b92

Observation ed1b7731-cd1e-4065-8a89-1e86b47def95 · outbound

This paper cites V., Kostrikov, I., Su, Y., Yang, S., and Levine, S.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization V., Kostrikov, I., Su, Y., Yang, S., and Levine, S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.477722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.154039Z digest=sha256:74681beacb7befc9ec86b481210e3752d2265b4f8027bd81b4e89d39e9b45bf7

Observation 61174b07-3329-4789-bbaf-2071270ecefb · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Gemini: A Family of Highly Capable Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.157866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.157866Z digest=sha256:5de42474700b79e5e2078b1ec8568caff2bdaa4055e6df95f7d7ca78fc00e6b6

Observation 17b3cf86-684d-4a55-b122-5f7dfc24864e · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Gemma 2: Improving Open Language Models at a Practical Size

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.161900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.161900Z digest=sha256:1037d361d74756ad19fd68e250ab9272b2243db6a1057996b18a0422088039c0

Observation b2bf55f2-6097-4376-ae9c-2fba496a0a3d · outbound

This paper cites D., and Finn, C.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization D., and Finn, C

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.466316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.165866Z digest=sha256:cc4a11ccbc331ef40e4bb295c38e7fe98f5ae73fb8a6e96ed048e8d948e93da4

Observation bb49bab6-7658-4085-94fe-bbcd71b8728c · outbound

This paper cites Interpretable preferences via multi-objective reward modeling and mixture-of-experts.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Interpretable preferences via multi-objective reward modeling and mixture-of-experts

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.455391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.170856Z digest=sha256:aec39ce3c70b096004ee9f8706a5a7a76e208f58e9ac7aa3f7f847a9e55bad85

Observation 51e64da4-8ebd-4d91-91a5-c7829650e1d4 · outbound

This paper cites A., Ostendorf, M., and Hajishirzi, H.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization A., Ostendorf, M., and Hajishirzi, H

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.443704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.174094Z digest=sha256:ff9becae7c2886fb04c84d5e4deb23d486bcfea24719cfaec54c4429266c9291

Observation 5102d338-e563-4f37-a5b3-c0f4b54cfe73 · outbound

This paper cites Inverse- Q *: Token level reinforcement learning for aligning large language models without preference data.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Inverse- Q *: Token level reinforcement learning for aligning large language models without preference data

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.431957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.177603Z digest=sha256:f666689c843e177ecd0c29af6094ce4f1808216c4c61bde1efe9df1397ec9e7e

Observation ce39aa76-ffb9-4e6f-98a2-3b7aea2c304b · outbound

This paper cites Preference-grounded token-level guidance for language model fine-tuning.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Preference-grounded token-level guidance for language model fine-tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.418301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.181731Z digest=sha256:3eb8e7cc8fd2a6c81a4d0dc574218a85234c431573efcf152b10b32309c5a7da

Observation 48ac8dcc-c5fb-46f1-afa3-646180f2d26b · outbound

This paper cites A dense reward view on aligning text-to-image diffusion with preference.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization A dense reward view on aligning text-to-image diffusion with preference

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.407009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.185336Z digest=sha256:d32682bee14598964701527c6c003ed1dc84bb665f78a61dba5b83fc6c5a2630

Observation 87939929-964a-4e5a-88ea-1058304cd6c7 · outbound

This paper cites Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.189439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.189439Z digest=sha256:9c91cf02c38694605baaf19bff987f65b9e4cd6051b119a2879993c86af3aeb3

Observation 73fd95e3-9db9-4fb9-96f6-ef65d94566d2 · outbound

This paper cites Token-level direct preference optimization.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Token-level direct preference optimization

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.396281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.193487Z digest=sha256:a24f3afe6f142f2a9bd3b1074cc693bfd33d2a8f702e9ead8c6561d3bcfdc2bd

Observation 30e4808d-3d42-4cd7-9e0e-17f50d4987e8 · outbound

This paper cites Judging LLM -as-a-judge with MT-Bench and Chatbot Arena.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Judging LLM -as-a-judge with MT-Bench and Chatbot Arena

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.385309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.196578Z digest=sha256:53c03f49ed5fa6581af481fc3a9dfb5edc29b12c35bffc2213ef49aaea0e4cee

Observation f48c2fb8-b56e-43e0-8128-5d015c50ee54 · outbound

This paper cites DPO meets PPO : Reinforced token optimization for RLHF.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization DPO meets PPO : Reinforced token optimization for RLHF

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:20.374564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:59:20.199953Z digest=sha256:ebeb8ec10e7d428d0414dd9c2fb11ca6ea2ee5c78695fa73e4e409e9a78f6340

Observation 5d5d3081-9b40-43e8-8315-b047e71bbd0e · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Fine-Tuning Language Models from Human Preferences

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.203188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.203188Z digest=sha256:5f05dddc0f10b9f2de97ec9e6691a14e348bf25e357396d22f782f45115fdad6

Observation 4428b8cf-84ce-460f-9649-6b7264d56101 · outbound

This paper cites write newline.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization write newline

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.206994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.206994Z digest=sha256:35ca650f95d1436037038780a2fd6b7ed86a7e022a99713350223611420b5ea9

Pith citing papers

Observation 84ec1dfe-f34e-4d75-8df3-c366e6e177a6 · inbound

Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs cites this paper.

Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:02.101463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T09:52:20.508013Z digest=sha256:6a267c40eb9f72b0a3aa880707b774cb89726d2271217310adf546c1d062a3e4

Observation 053fcd64-447e-461e-b378-fd25c27137ca · inbound

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization cites this paper.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.210012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.210012Z digest=sha256:79d6b99ec6adeabaa2826d413c45cc23420f7817726e781088eac249965face7