Pith. sign in

Paper Citation Record · LEDGER

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization

As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2608.09568.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09568 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:12.228176Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6fa8945-4983-4e15-b8ee-fb64a46a421f · outbound

This paper cites The Softplus activation ensures non-negative output.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization The Softplus activation ensures non-negative output

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:13.099952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:59:12.228176Z digest=sha256:04efdb12879145dcfebf44efdd837b34bd39ca5780d43ab4665e867e42499e87

Observation 8a8df1ca-40a5-474e-8c68-72426770d7b4 · outbound

This paper cites TAB-PO: Preference Optimization with a Token-Level Adaptive Barrier for Token-Critical Structured Generation.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization TAB-PO: Preference Optimization with a Token-Level Adaptive Barrier for Token-Critical Structured Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.089794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.089794Z digest=sha256:474890f5e9d0c8eae59e00b3235762b7f2db34d8808385b961a2d18e444b0dee

Observation 33b45022-8242-4122-9b16-6b43985a1c3a · outbound

This paper cites The Llama 3 Herd of Models.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.102438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.102438Z digest=sha256:eb2ed2cd8986de75b0ec4a0b9ac367cf48bf47d9d1ec2c37ab579d17b29765f1

Observation 70d24357-23bb-4e09-8381-07e1d75c7ae3 · outbound

This paper cites Adaptive batch-wise sample scheduling for direct preference optimization.arXiv preprint arXiv:2506.17252,.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Adaptive batch-wise sample scheduling for direct preference optimization.arXiv preprint arXiv:2506.17252,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.108854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.108854Z digest=sha256:f1b5c2a8bad1147cbf9896b5f89c41fa790dea2aaee80805384864c6742e509f

Observation 13b60a3a-61fa-46b4-872d-bc34ccacddfb · outbound

This paper cites Kl penalty control via perturbation for direct preference optimization.arXiv preprint arXiv:2502.13177,.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Kl penalty control via perturbation for direct preference optimization.arXiv preprint arXiv:2502.13177,

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:59:12.785116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:59:12.120007Z digest=sha256:79e660ca9d0eb4bb5c33d5b0581ef4061914772d831ac941978f1e0e34af894f

Observation 9b1089d8-0b76-4953-8b22-cc2c7b53c939 · outbound

This paper cites AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:59:12.690701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:59:12.125194Z digest=sha256:9e60d661ad9a6f42e6b077a5d79e8e53a00d622a3eff3c0b988caac297533994

Observation 121b761d-a040-4c1f-8fb9-9f2acc0b40d4 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.131103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.131103Z digest=sha256:48c8ceec38eaeef7bb31dfc40cf394c858ea19852fce7952ca47e0d9d23cf73f

Observation 49016e83-99d3-461e-9614-038d2905fce0 · outbound

This paper cites TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.137307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.137307Z digest=sha256:8467a58e4ef2d3fcf868bc8de26484e79dd30f576cb6a159c59d0d7bf963a4b2

Observation fa33753e-55f5-45cd-9414-fc31134649ba · outbound

This paper cites Autoregressive Direct Preference Optimization.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Autoregressive Direct Preference Optimization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:59:12.616020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:59:12.142774Z digest=sha256:3244a6aff94ea3e501a1cc0bf3a066df40c6e13605bb1c5f273a49d2db49d0f1

Observation 56850233-3fec-4c8e-8ff8-4b7147d966d8 · outbound

This paper cites Small-margin preferences still matter—if you train them right.arXiv preprint arXiv:2602.00954,.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Small-margin preferences still matter—if you train them right.arXiv preprint arXiv:2602.00954,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.148038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.148038Z digest=sha256:0b5bbd4deec38b452f216e232654f856719533be9261addd74cdfb90af714791

Observation 45ea035a-bcc7-47b1-a214-9cf9a7f8fc28 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.152668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.152668Z digest=sha256:a34338579f662790788f337bf6433af7862ad8c9992a166f917ed6cf23fe84d4

Observation a99d5f1a-a66f-4978-839a-6f23484af853 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Proximal Policy Optimization Algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.157429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.157429Z digest=sha256:abc537f0b495a4cd106b294d5db15a067ba1441db006192e36ab13d61d6b93b1

Observation 71bbb7a6-f4c8-4a2c-bf4c-ed864b7bcea5 · outbound

This paper cites Interpretable preferences via multi-objective reward modeling and mixture-of-experts.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Interpretable preferences via multi-objective reward modeling and mixture-of-experts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:13.194401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:59:12.166894Z digest=sha256:fcfd881ce24cee5636b5350e0e18888bb33e537f3f934c6e2ba10123fbd19a2c

Observation 63b16c52-ce48-4e97-8a1e-7bbafd9a170b · outbound

This paper cites Explore the reasoning capability of LLMs in the chess testbed.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Explore the reasoning capability of LLMs in the chess testbed

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:13.173052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:59:12.172751Z digest=sha256:bf83cc2d65f22df4b4c4e27d07d501418bb4b5aa43b189868de445089b5cf000

Observation 065bbcbc-0ea8-4de3-9c63-5b860f4f90ab · outbound

This paper cites URLhttps://aclanthology.org/2025.naacl-short.52/.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization URLhttps://aclanthology.org/2025.naacl-short.52/

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:13.154557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:59:12.180368Z digest=sha256:def75ceebc09cfa57d726ec46d3c49ab089fc333d417e9e6aa73d878f133c6d9

Observation a0d28849-a59f-498a-bcf0-7315c6904fc0 · outbound

This paper cites Se- lective preference optimization via token-level reward function estimation.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Se- lective preference optimization via token-level reward function estimation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.186010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.186010Z digest=sha256:add764c4274ef7febfeb537dcfb0b55ea291e276bd8cf9e1acbbd07df852e0df

Observation c911f9fa-1abc-46b8-9a46-443e0f913cc4 · outbound

This paper cites RUBRIC: Realism--Utility Balanced Ranking for Imbalanced Classification.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization RUBRIC: Realism--Utility Balanced Ranking for Imbalanced Classification

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:59:12.357820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:59:12.192104Z digest=sha256:23f59fa460f9981bbc5638c843a7a0d487f75a78bbef00a909abbd2a2e3d6d2c

Observation 90cbbf2e-91f9-40bf-9088-e69e9fdbd730 · outbound

This paper cites A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:59:12.329770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:59:12.197591Z digest=sha256:fb7c5f89cf77e39f22e2e910711b7d6fa0685e301abba7c04a4735c411c24d7f

Observation 3ae436d4-1233-42b5-b5d3-5b999e4641c1 · outbound

This paper cites Wpo: Enhancing rlhf with weighted preference optimization.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Wpo: Enhancing rlhf with weighted preference optimization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:13.136581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:59:12.202844Z digest=sha256:6c3459f304617dadc59b10b1c268b993755d813b7adaac3891a43137b11a3747

Observation 053fcd64-447e-461e-b378-fd25c27137ca · outbound

This paper cites TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.210012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.210012Z digest=sha256:2d3ea24b5eb946a93a0cc52483f33203c99313265278fb1bdbc60777c309f9a2

Observation d324eaf6-ebbf-4695-8e22-4b4855768640 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Fine-Tuning Language Models from Human Preferences

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.216222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.216222Z digest=sha256:c0ca7d3e7a07fbd29be99a40a575ba4ccf58c271182b30e3cc04563e039503d0

Observation fee99bea-86ff-4e65-b828-eb08c74dd85a · outbound

This paper cites Sparsepo: Controlling preference alignment of llms via sparse token masks.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Sparsepo: Controlling preference alignment of llms via sparse token masks

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.076346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.076346Z digest=sha256:6e19f0f75a3161da66688f9640f11ad6c224b7e54905ecf30ada0321d7ead5bd

Observation 4a452e08-4ea5-4a91-80dc-99e6096b74e7 · outbound

This paper cites Under independent noise, Var[ˆ∆] =∑t c2 t σ2 t.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Under independent noise, Var[ˆ∆] =∑t c2 t σ2 t

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:13.117690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:59:12.222521Z digest=sha256:ddd764f980bcee1d6c5d904f72450d9682aedf03384b15401fd0451ed5211fb8

Observation 564c92a3-b4ae-40ee-ae6c-fa07150f4530 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.161810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.161810Z digest=sha256:3ef646c256d060891606f15d60d46937c61891237b689876365b64170ac6e379

Observation 959f098a-23bb-411c-9ad1-170ba632a6f2 · outbound

This paper cites Discriminative Policy Optimization for Token-Level Reward Models.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Discriminative Policy Optimization for Token-Level Reward Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.070055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.070055Z digest=sha256:92d44a818ada470a433def28d6e9144416ec049e534ede3a5d2cb44a9948b377

Observation e050620a-34d5-46d5-b4d6-852f9bc765c1 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.114561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.114561Z digest=sha256:bd7e6e37e7c26ba7c84e51d58ef35e02617fcce7ce95360b272b5df1b4ee7e75

Observation aa42e25b-f3b6-4a78-bc79-2cee2585bf2e · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.063432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.063432Z digest=sha256:aff8648e18136374ea75fddf7e62d6940cf229f6997554a1b658548d42ac82f6

Observation f4e325dd-45d5-4f10-88c2-f9f7f3e02579 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Process Reinforcement through Implicit Rewards

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.083599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.083599Z digest=sha256:639addc1d3827b050dc7aa870617d93f7716ddeacb07718735c909abc346ac5c

Observation de9a80bd-8553-44e9-b2f7-41399341a89c · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Gemma 2: Improving Open Language Models at a Practical Size

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.096578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.096578Z digest=sha256:4a8b034a6e2a8a01ed1df6a9e41d2be235a385f4860d6ce762256bce79a2862c

Pith citing papers

No inbound Pith citation observations are available.