Pith. sign in

Paper Citation Record · LEDGER

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner

As of 12 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2604.18239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.18239 v4

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T15:57:22.808390Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed02e5cb-4f79-4b94-b74c-7d36adc85df9 · outbound

This paper cites an unresolved cited work.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:20.675979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:20.675979Z digest=sha256:bb7f35a8a02931adf89af02b992367c3f008c93b6bbee2b70060deacdd7f864a

Observation 8c3dd996-5c70-4b40-934f-2fb3d16de745 · outbound

This paper cites Li, J., Chen, W., Liu, Y ., Yang, J., Zeng, D., and Zhou, Z.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Li, J., Chen, W., Liu, Y ., Yang, J., Zeng, D., and Zhou, Z

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:21.079138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:21.079138Z digest=sha256:a16e5a0081a73929dc9c7bce12b0571045a4db0cc751b9f29de4dc2a91c3fe00

Observation 6ff9588e-e5c6-46a1-b969-5f7f158cfa14 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:21.269332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:21.269332Z digest=sha256:7d8306c195b6b175433bedf22c4b53cb2e01893ddd44cb6b731302a86970eb3f

Observation 383d8080-aede-45ec-b06c-084084af2df4 · outbound

This paper cites Razin, N., Malladi, S., Bhaskar, A., Chen, D., Arora, S., and Hanin, B.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Razin, N., Malladi, S., Bhaskar, A., Chen, D., Arora, S., and Hanin, B

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:21.500647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:21.500647Z digest=sha256:6362c9257f464855526d3444c5b43dd308de55f716bba8cf6fd87b8ab2ef5921

Observation 3f1fcaca-4824-4e75-bd9d-7b4e4ca0c61b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:21.707722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:21.707722Z digest=sha256:2a693b0b6960132f6b858e80d95924b4737152f05f0163aa9bf126cd545e5943

Observation 0f8cf12b-f05d-45d5-9a0d-4440303b316c · outbound

This paper cites org/CorpusID:274859421.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner org/CorpusID:274859421

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:22.566238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:22.566238Z digest=sha256:3396f675969bbeb487fd275b2100122434bcf11d23a0a3b9b63381700f7ad13b

Observation 4ae23a2a-babf-42dc-84f5-065a128dbc3a · outbound

This paper cites Yuan, H., Yuan, Z., Tan, C., Wang, W., Huang, S., and Huang, F.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Yuan, H., Yuan, Z., Tan, C., Wang, W., Huang, S., and Huang, F

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:22.644269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:22.644269Z digest=sha256:66d5a9996b29f0cbdceff354ef4ebd569e512c84b26e8ec3c7e3e919cd6547c6

Observation 3c43a103-a3d0-4bdc-928e-da3b95fd21fc · outbound

This paper cites Zeng, Y ., Liu, G., Ma, W., Yang, N., Zhang, H., and Wang, J.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Zeng, Y ., Liu, G., Ma, W., Yang, N., Zhang, H., and Wang, J

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:22.746322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:22.746322Z digest=sha256:a0c8078e02f6d80e771fe1a8d6658ca83660c8c4a73a4c2a21475cbd4e1dca47

Observation 2c3cf895-0603-4e0c-bd38-353f1448397c · outbound

This paper cites Semantic-Aware Logical Reasoning via a Semiotic Framework.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Semantic-Aware Logical Reasoning via a Semiotic Framework

Reference 17

Resolution
malformed identifier
no resolver link, observed 2026-08-02T15:57:22.808390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:22.808390Z digest=sha256:9e282ceb34e751b12728885662c27fee78db57083f14482c776300cefd8b6b0d

Observation 8a5ce3ed-3eaa-478d-a5fd-1f96f64a4f1d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:20.848345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:20.848345Z digest=sha256:072d477a5ed536b2acb258a109e074fa08b8a881c6b2c3920cb73bbd45d68cf2

Observation 97ac27b5-c67e-475f-bfed-2669eb1af9d6 · outbound

This paper cites org/CorpusID:44112860.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner org/CorpusID:44112860

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:21.944889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:21.944889Z digest=sha256:10404efc7ed40f8b24602580bf0596b63a16eea73939ec4b08b4574d5412476d

Observation c3b268b7-6730-49c6-b7a8-8003c6578aac · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:22.150586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:22.150586Z digest=sha256:8e1695d0d5ba3f7ba43262eab47dc3cc79d7ee03be833c908259dd3cd3a7d62d

Observation 397a7f76-438b-4f00-b3c1-3d5e9d3a3029 · outbound

This paper cites org/CorpusID:244478113.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner org/CorpusID:244478113

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:20.055196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:20.055196Z digest=sha256:4a24e27cc216284eab1d61ec4e1ac9f6f258d2c73eae998ae6cfd6ba1e647bf4

Observation 2ab998f5-6a8a-4fd7-aa39-5d6206574981 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:20.515549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:20.515549Z digest=sha256:7a8a8489b61b6f668f85ddbd3202a76beb3247727ffdb4cf9888c240fd9ca37b

Observation 563c051b-2318-443e-9735-177bf98dcc7e · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:20.248367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:20.248367Z digest=sha256:e4edadf340d16b310d4dd47a2884c71d7ab1acb347b09ce7f422d526ec89f395

Observation 25e3e511-d3fb-433c-ad6c-2e30300deb58 · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:19.944569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:19.944569Z digest=sha256:c2e6ddf1054a42a2332629ce98a49811fd49f7980e34648a034bc0ef607d3ab3

Observation 7c0e464d-ba3c-448d-b2a2-294b1bcf9ce7 · outbound

This paper cites Qwen2.5 Technical Report.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner Qwen2.5 Technical Report

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:22.408332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:22.408332Z digest=sha256:1ecf273c5d8642f31cc0644817ad85cf98a2b15eb4fb22a4bb5b90d7b19fc9d5

Pith citing papers

No inbound Pith citation observations are available.