Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:21:35.492844Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2412.15244.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:21:35.492844Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-12T04:14:37.374346Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T06:31:24.733435Z
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 97938775-c86a-4530-a403-76c05dd628f5 · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e177efc5-4444-4654-9f74-e2481b5db51a · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples KTO: Model Alignment as Prospect Theoretic Optimization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb4e39e4-66e8-410e-a234-af809a403d63 · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples ORPO: Monolithic Preference Optimization without Reference Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22304670-4771-4e4c-b3e7-b02ca9d3d064 · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f7cc2a7-9a45-42bb-9af7-c779cc913570 · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 137be3cb-7081-4eec-947a-93a7b8abd905 · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples gradient descent
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bfd35f52-86e3-424c-99c5-6b7fc15b386c · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 772c3859-bb46-4d48-871a-6071ffccf479 · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Proximal Policy Optimization Algorithms
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3036090b-1a64-403f-ae07-08697bc4ff2c · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Preference Ranking Optimization for Human Alignment
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3143a5c9-f2f4-44da-8125-ab7d5e5383f7 · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Fine-tuning Language Models for Factuality
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7927400a-778e-46be-b9c0-a5a542b649af · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Zephyr: Direct Distillation of LM Alignment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b44864e9-f666-4875-a8b8-13dd0bb47821 · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 731a6462-5443-4126-aed9-de4b0c4c6a0c · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Deep reinforcement learning from human preferences
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ce35d6d-ddfa-4af4-b12a-8cd66b1c7922 · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Language Models are Few-Shot Learners
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1bf28cb-c1e2-442a-b97b-a79d5388adc5 · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Finetuned Language Models Are Zero-Shot Learners
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49b1ecca-c134-48d6-850d-34a91fc6f83a · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Training language models to follow instructions with human feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c811978-cd77-4be0-b743-9c738c9f8055 · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples The Falcon Series of Open Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42d25986-8008-4009-b677-f91ee9a3dcac · outbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Can AI Assistants Know What They Don't Know?
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab0ab55d-1da6-4289-847f-ed009875cb0c · inbound
MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.