Pith. sign in

Paper Citation Record · LEDGER

LiPO: Listwise Preference Optimization through Learning-to-Rank

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2402.01878.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.01878 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:12.662588Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T21:23:27.413513Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2ebd1d91-ecbf-45fb-bfd3-e5688bb9c815 · inbound

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types cites this paper.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.417555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:14f5ed62959ec807f47c955f3c89bb446569a9f4a921d98d696c3445b8a9769d

Observation 0bb642c8-3734-47cb-8b39-7bd02e952bdb · inbound

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd cites this paper.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.880969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.880969Z digest=sha256:0a248c6b7ee810abec6f8e24feff6886ecf8aea8e96210bbedfe6dad6065589c

Observation ecebbdea-9c34-4716-8e01-db7c56e17d37 · inbound

ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning cites this paper.

ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T05:13:41.491585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:13:41.491585Z digest=sha256:26bc948a7070976ea1797b5d5b74603ba911b2408c58d4c162f6ab49801599c4

Observation 52f1e475-40f3-437a-a27d-c611f82978de · inbound

Controllable Protein Sequence Generation with LLM Preference Optimization cites this paper.

Controllable Protein Sequence Generation with LLM Preference Optimization LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:51.815631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:51.815631Z digest=sha256:64a319713f952f8f044b516ed0d3d452f3a6a8e677ed85eba2711acbe73b32b7

Observation 17aeefa1-009b-4909-aa4d-091165f021e5 · inbound

The Differences Between Direct Alignment Algorithms are a Blur cites this paper.

The Differences Between Direct Alignment Algorithms are a Blur LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:52:29.529197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T03:50:03.720389Z digest=sha256:24ee79637963cbbd6f92a005825a5e38fddb8b1f7aa3f5986d74a6a0a4a026c2

Observation 8b1e524f-e854-4e74-b820-d5a26e47e936 · inbound

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective cites this paper.

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T04:08:51.724809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:08:51.724809Z digest=sha256:90526aea8f6a94650708102c13242e8fac39c4128f88591b2a17585269b174db

Observation f00dd238-8ddf-480e-a080-9759f83a6233 · inbound

PerPO: Perceptual Preference Optimization via Discriminative Rewarding cites this paper.

PerPO: Perceptual Preference Optimization via Discriminative Rewarding LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T06:01:13.306315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T06:01:13.306315Z digest=sha256:c9cc5f84d89bb53f9ed83e265230c4a13f67e628665eb7d6dde6aae8edfe63e2

Observation f7881d25-80d0-457c-a75b-53286b7bad60 · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 221

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:12.662588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:12.662588Z digest=sha256:954c75d5911b9cf2f8a77079d908da040750125af951b744effe670201c2d4d1

Observation ac0a7142-37ff-468f-868a-eb0e608aa41a · inbound

Advancing LLM Safe Alignment with Safety Representation Ranking cites this paper.

Advancing LLM Safe Alignment with Safety Representation Ranking LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.276005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.276005Z digest=sha256:1b9c6fea1ff0d17a742cbada832a5df6e9cfa3d3fe1365e8d136696a2103b35f

Observation 6852c46f-e6fb-4d63-972d-aac640c962ee · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:28.962243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:28.962243Z digest=sha256:7b822fcd6aaf8c5f520326d32d1e58be85bf8469ef96655a99d83424b7786fdc

Observation 8d7eba4a-3ff9-4c93-ad87-bddebd038360 · inbound

LPOI: Listwise Preference Optimization for Vision Language Models cites this paper.

LPOI: Listwise Preference Optimization for Vision Language Models LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:55.816724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:55.816724Z digest=sha256:30ec6af57cf44ea07f0ce97b44b52c283a6e4a51e59058e224c591b64b06893d

Observation c104ef4b-df02-4545-b873-a02691f147ea · inbound

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models cites this paper.

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:00:48.923951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:00:48.923951Z digest=sha256:37e750c037845e12c5e5a288ea31a7dc5a9b53dea66c92f5cb246c081c6e7f70

Observation 98d5ad29-7c6a-41bb-900b-d36ab5f2d551 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:48.212275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:48.212275Z digest=sha256:7281abc8e2a1043c73b43d1e1bbfbafb0c1e744d1cf6ab4fd72d398ff55449a2

Observation 2fe16626-45d2-4e09-a520-a493d26b41fe · inbound

Reward Models in Deep Reinforcement Learning: A Survey cites this paper.

Reward Models in Deep Reinforcement Learning: A Survey LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.674235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.674235Z digest=sha256:e2ea2788e463e2ba58bd94fd49496d1dbd1e2fe75e1ed5a8ad949a605c98c89c

Observation 9bfe8b37-794b-4f8e-8281-d7588622099f · inbound

HAEPO: History-Aggregated Exploratory Policy Optimization cites this paper.

HAEPO: History-Aggregated Exploratory Policy Optimization LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T16:13:08.310472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:13:08.310472Z digest=sha256:ea0c48d094fd028529a171ae165067d4fa87dfb66cc771e7ab157ae2c9e06bb8

Observation 582509b3-ae83-4d8b-ad9a-5603ce94dccc · inbound

Threshold-Guided Optimization for Visual Generative Models cites this paper.

Threshold-Guided Optimization for Visual Generative Models LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:46:07.941014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T17:14:36.632493Z digest=sha256:2e8091fbebb846dcf93f79d96ddd53424c7cf9fb12a083278ca688acd551ed45

Observation 1838de3e-0a41-455a-b996-f0c3cb00713a · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 160

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:46:00.020414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:c7663055fb5ef899f58e15c0956e62c850472e1c925138a5a21b6649ae0c8261

Observation c99c18e0-2630-4461-99d9-37a51bcdf627 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 231

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:de6524e43c9d842faf8109fd76979448b5df26ec002e79c96f2386b6e009af6b

Observation b3380034-14bd-4832-8393-397a7c40bb0f · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:4fb53ec6b995515a1619db2278db38fe79beb91e8fd0968d0ff3f6c5f7aaa988

Observation 910aa229-e412-462f-b930-696349ee7d06 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:41.165786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:41.165786Z digest=sha256:2af5026c4737d1ed99147907a238c771002a21e8438f9bf42f6f3104e3cb5c0e