Pith. sign in

Paper Citation Record · LEDGER

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples

As of 21 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2412.15244.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15244 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:21:35.492844Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T04:14:37.374346Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T06:31:24.733435Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 97938775-c86a-4530-a403-76c05dd628f5 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.363100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.363100Z digest=sha256:ddb6b57a9833e050df2cd36a589f56c59f87c7ff697fafb9b0ad46db1fe12c4e

Observation e177efc5-4444-4654-9f74-e2481b5db51a · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples KTO: Model Alignment as Prospect Theoretic Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.388264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.388264Z digest=sha256:9803e11c286bb796c609bf78641a457fbae9162f2ed9bf8af3f2f5d05b1e2655

Observation fb4e39e4-66e8-410e-a234-af809a403d63 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples ORPO: Monolithic Preference Optimization without Reference Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.396942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.396942Z digest=sha256:5e4c4a1e4c309aaa7a3bb834ab5eb83a4871f636bbb8534fee881e3b81097302

Observation 22304670-4771-4e4c-b3e7-b02ca9d3d064 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.407893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.407893Z digest=sha256:a10e3f90f9b316b1137679749f089f7855c46592408c567a7ecbeb6e7ef98e25

Observation 4f7cc2a7-9a45-42bb-9af7-c779cc913570 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.421940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.421940Z digest=sha256:04dcc287fc6f9f3192ada11f46ee0df4074ace59966d33256c573689986c6d42

Observation 137be3cb-7081-4eec-947a-93a7b8abd905 · outbound

This paper cites gradient descent.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples gradient descent

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:21:36.050712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T16:21:35.435417Z digest=sha256:69968c3e37200ae9b8193c2c48c0026b16c18e2ac23b2f3c8aea2927d9355df6

Observation bfd35f52-86e3-424c-99c5-6b7fc15b386c · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.447841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.447841Z digest=sha256:c6fe0ba814e62a4ac8b5eaeea27faeb94cd1bf41e890a3c9c0b56004988a0aae

Observation 772c3859-bb46-4d48-871a-6071ffccf479 · outbound

This paper cites Proximal Policy Optimization Algorithms.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.457164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.457164Z digest=sha256:1683cae96b376abb6f0fa970ba51c0acf21badc7b9af593aacac29a2d4e8b2c2

Observation 3036090b-1a64-403f-ae07-08697bc4ff2c · outbound

This paper cites Preference Ranking Optimization for Human Alignment.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Preference Ranking Optimization for Human Alignment

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.463801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.463801Z digest=sha256:365a1a03e8ddb43a311263d8e2421beb580cb27f5d4a82dca02397a67ee7cce7

Observation 3143a5c9-f2f4-44da-8125-ab7d5e5383f7 · outbound

This paper cites Fine-tuning Language Models for Factuality.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Fine-tuning Language Models for Factuality

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.470037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.470037Z digest=sha256:ba51eb0c73ed8ba63c1820eabdbcace91271fef8998faebe8832834a6948077e

Observation 7927400a-778e-46be-b9c0-a5a542b649af · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Zephyr: Direct Distillation of LM Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.477729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.477729Z digest=sha256:c72c2095a0305dc5733aad85a52acb8db4ba4d90550c813f0ddca027d4732ceb

Observation b44864e9-f666-4875-a8b8-13dd0bb47821 · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.486443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.486443Z digest=sha256:8e61ce25bde3703d0b9a6d3aee7c97e63284d151d572399622deb419f09b615a

Observation 731a6462-5443-4126-aed9-de4b0c4c6a0c · outbound

This paper cites Deep reinforcement learning from human preferences.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Deep reinforcement learning from human preferences

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.380643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.380643Z digest=sha256:14f4a89b3875b636cd4352dd7c0478ea133d6d51b911d57408b33a18025d58df

Observation 9ce35d6d-ddfa-4af4-b12a-8cd66b1c7922 · outbound

This paper cites Language Models are Few-Shot Learners.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Language Models are Few-Shot Learners

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.354232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.354232Z digest=sha256:2f17a4a423a68967892acbdac0d29e2939c30044ac0eae2d9185fd3119dfa53d

Observation e1bf28cb-c1e2-442a-b97b-a79d5388adc5 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Finetuned Language Models Are Zero-Shot Learners

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.492844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.492844Z digest=sha256:a9dcf77b590ea16a74bde71d8575f14761e2c5355551cd0c6c1e9969ca756c22

Observation 49b1ecca-c134-48d6-850d-34a91fc6f83a · outbound

This paper cites Training language models to follow instructions with human feedback.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.414681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.414681Z digest=sha256:a477c407b7b1f9bbd2df7464d8b63d74553b71e215908b1b31aece4dd9e85784

Observation 1c811978-cd77-4be0-b743-9c738c9f8055 · outbound

This paper cites The Falcon Series of Open Language Models.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples The Falcon Series of Open Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.346624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.346624Z digest=sha256:76cc876f0d1a2a978e944fbe533b2d1ff2d55eacff5fad7fa9c664970f532250

Observation 42d25986-8008-4009-b677-f91ee9a3dcac · outbound

This paper cites Can AI Assistants Know What They Don't Know?.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Can AI Assistants Know What They Don't Know?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.371583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.371583Z digest=sha256:3513b8dbc0acdfc98f51dba701978d1e2293e8dc15fce95f438a99e7a775b9cd

Pith citing papers

Observation ab0ab55d-1da6-4289-847f-ed009875cb0c · inbound

MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization cites this paper.

MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:24.736800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:14:37.374346Z digest=sha256:a5bc341bfba9b4fe884577422bfba238aad5887b3c235894c4dc02f1594277e7