Pith. sign in

Paper Citation Record · LEDGER

Length Desensitization in Direct Preference Optimization

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2409.06411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.06411 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:37:34.210531Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.691954Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3aa25225-c77a-477a-9353-1cb33690dc55 · inbound

Hansel: Output Length Controlling Framework for Large Language Models cites this paper.

Hansel: Output Length Controlling Framework for Large Language Models Length Desensitization in Direct Preference Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:37:34.210531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:37:34.210531Z digest=sha256:d1c3d8df9a68ec0e17a9843ef585202339b9058661bc2659611286b326ef3a97

Observation 5f33d3a0-bab8-4df3-8ffb-e0ad92964bbe · inbound

A Technical Survey of Reinforcement Learning Techniques for Large Language Models cites this paper.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models Length Desensitization in Direct Preference Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:31.732769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:31.732769Z digest=sha256:6dede8d2624be3f5ff2238a437a8b58bc70a130968db0aebe64202dfece69c2e

Observation abf57df3-9381-41e3-8389-0b164cbcf2f0 · inbound

TiCo: Time-Controllable Spoken Dialogue Model cites this paper.

TiCo: Time-Controllable Spoken Dialogue Model Length Desensitization in Direct Preference Optimization

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.752476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T00:38:52.182973Z digest=sha256:56b3e396b6f3d6c0e5393c80466401d6c9f7234ad73d79241446dbfee7f374d5

Observation 340fa0a7-083c-4f74-9c30-7a9ce73f343e · inbound

GroupDPO: Memory efficient Group-wise Direct Preference Optimization cites this paper.

GroupDPO: Memory efficient Group-wise Direct Preference Optimization Length Desensitization in Direct Preference Optimization

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:48.709683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T09:43:18.432084Z digest=sha256:ca58c661f1f06cc08a8a28650b694f8089eaaa213be14368504433cced144a4b

Observation 0b5c5264-4c30-4508-b3c8-3df38d12a197 · inbound

Learning to Control Summaries with Score Ranking cites this paper.

Learning to Control Summaries with Score Ranking Length Desensitization in Direct Preference Optimization

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:36.366133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T06:41:27.954392Z digest=sha256:f2d4ff1af566f60b8ed281a6662749f141ba75b53fdebc0a48cd392f84c12984

Observation df62599c-a6da-476b-8033-298497b91682 · inbound

From Coarse to Fine: Self-Adaptive Hierarchical Planning for LLM Agents cites this paper.

From Coarse to Fine: Self-Adaptive Hierarchical Planning for LLM Agents Length Desensitization in Direct Preference Optimization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:10.387921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-08T08:10:36.579810Z digest=sha256:f2d892cbf3abbd61b3548863878b4909530bf2888e814cbc6cfccb5825a5a81e

Observation bdb75aa9-7689-4f60-aefa-d4d84e10b311 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Length Desensitization in Direct Preference Optimization

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.693757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:deae50c6034bbbf12ae9f6491b3885fc9a278044c93c6ee1bcdd7f975d55dc65