Pith. sign in

Paper Citation Record · LEDGER

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models

As of 23 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 3 inbound Pith citation observations for arXiv:2505.10892.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10892 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:08:56.977697Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T08:52:34.039089Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 105f9312-5888-4084-a3d8-a1e01985d193 · outbound

This paper cites GPT-4 Technical Report.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.695415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.695415Z digest=sha256:9f5fb17cdcc7904ec9a32f386b768fa34816ec0779d0075c7a2ec92271ddc71b

Observation 83c854ab-e810-460a-9b1b-59be6dc88438 · outbound

This paper cites Agnihotri, R.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Agnihotri, R

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.849923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.700924Z digest=sha256:af7ba1dbc8b6288a82894f1ad26d63f988a79b439b1c0fdbf7a4a7cace9cd1c9

Observation e8de9cd9-9b15-4cce-8d72-2b7cc90129f4 · outbound

This paper cites Agnihotri, R.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Agnihotri, R

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.836979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.705282Z digest=sha256:e1b453f88978596410e054f9012ec3adc3164289c55a3f36c00984791dec4180

Observation 78ffbc57-53dc-4415-a587-4c2f3dd468d9 · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.822286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.708943Z digest=sha256:84b58fde053a731728fa1a9c3d5a48953ded36d8dd42266d60ef8b608376b5d9

Observation 3926a990-920f-41bc-91c1-fcda019f91bd · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.713803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.713803Z digest=sha256:f2be1dee1127991fcce747f6efd15091aa235221e3ceb265f51e3eb9be61b6ba

Observation 7e23d1bb-d8a5-431c-881c-0952b41b01e1 · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.717922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.717922Z digest=sha256:f0375b419ff537dffbc69464f2641605e8ebb9f55f9016b7ccd295c6602c0039

Observation 2b2e651f-988d-4485-b9d1-b048739ab88b · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.721917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.721917Z digest=sha256:0b3a5f412cd5874d3bdef4cc250013c4ea8c30cb890ccaeeca209c0985215ee9

Observation 06055c03-1ca4-4c4b-acf6-91ef81907a33 · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.726366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.726366Z digest=sha256:feff1177219798085b7944764f795397274431cdff2f936bff5e382343c43798

Observation 9c5ce716-0b9e-4e6d-a9ba-75bc9456c04e · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.784676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.730063Z digest=sha256:64243f28c6820b3dd2d42e4cbaba4945c355a76a28b7a4deed4061249363fdcf

Observation fb1562da-c785-4105-9329-f3f2b0623316 · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.766504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.733853Z digest=sha256:1087f835270be8d851c21861c59b167a144225c8dbb38666ec39c3e3c0051b77

Observation 0b2bd1b1-6df3-413d-a873-b1808da60bbd · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.749022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.738966Z digest=sha256:51e7f076b87c2c143156f409d810b3e89f7204838fb5341894feaf68a6d00a9c

Observation 58d45108-64b3-4cce-8227-6193b3d3e271 · outbound

This paper cites Efron and R.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Efron and R

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.736085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.743148Z digest=sha256:b05ba8ebc8d6f4859de4be9aa43386315a1f7b4064217ec33e7bd64145af7144

Observation 28faa723-bfc5-4951-84b7-726389ab41d3 · outbound

This paper cites Gheshlaghi Azar, Z.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Gheshlaghi Azar, Z

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.722260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.746931Z digest=sha256:307bdad610f014e44832163f223ff173c3b610b3ee25a113cf7e679784056664

Observation 766d6ed0-76c0-4859-aade-e7ab2841a7e4 · outbound

This paper cites The Llama 3 Herd of Models.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.751323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.751323Z digest=sha256:acbb1c9a9e3e4565acc8c2d814200998075b29852e580f23c8c9f66c1ae41e2c

Observation bdbb9e00-543d-4ec9-ac7b-2017fbd23f8e · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.755215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.755215Z digest=sha256:6c47c71889877cb177f0cd9343dbe818fc95d2db0b0f0bba1719bf97f1a727e2

Observation 2a0676ce-b45c-4989-bd44-a8b01b82fddc · outbound

This paper cites Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.759061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.759061Z digest=sha256:b7eb735764848ccad1829eb078b25ac1845ae9849227feee1cefebfa541da3c6

Observation fc99bbc1-4db5-44f1-8a80-f7d11356335a · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.701031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.763800Z digest=sha256:bf4df5836fc61c37589b40f827ec63b58930079fb5fd5131d1a744e5c9c17922

Observation 8cd5aaee-f7ae-4d6e-943e-4f70dbc71bb2 · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.686860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.767733Z digest=sha256:29fa561e679e97da09a7091bc9aefda21071f1f54673cfb4fba3719eea68d0b9

Observation f27c4b67-d78f-46d7-aca5-8187db0c70a6 · outbound

This paper cites Aligning Language Models with Offline Learning from Human Feedback.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Aligning Language Models with Offline Learning from Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.773942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.773942Z digest=sha256:931a175241b120179ed5aef2b71874a6310e7509bcd762099187c88d217b291b

Observation 580db4c9-4cb3-4d74-831b-9c2edd48f800 · outbound

This paper cites Huang, A.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Huang, A

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.673694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.779255Z digest=sha256:9ecc6684dec10de51ec73629e84c5f4c748bfa4451c83d8d07c578f2e9d75e73

Observation ac85fcf5-3744-4e26-89a1-59bb0f23865e · outbound

This paper cites Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.784337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.784337Z digest=sha256:7296e6b4825236aac637bbd9a840460445b3cd06afcc51d8d3aeadc6496d44e3

Observation cad7586f-a052-4093-af39-af5e3c51c5ac · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.659469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.788896Z digest=sha256:6daaf96a610a1820df3b4c4b16edb3af68140fc9e1ad6a37802c35e6476fdc7c

Observation 00a2dd24-dc43-498f-84ca-0a124e823fa5 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Adam: A Method for Stochastic Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.793145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.793145Z digest=sha256:1c0581eebf9f238bdebc182dff9abf93c734bf81bf1acedb5a56d466320c3d42

Observation 5f76ffc1-0cc8-4b63-bcff-535e3fd580ac · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.797792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.797792Z digest=sha256:05f4cf34d5a77e575e5cdd61c1b2f61b3af0b2e580059ac1e8483a54b42c2e10

Observation f839d209-68a9-4d12-bd70-3086b445ef82 · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.646083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.801762Z digest=sha256:5c6cef1ef4f6b85afa66cbf423fe79fb6608265496a15c1f335fefb9f95cb369

Observation 8dae53ae-623d-4beb-8510-8949d7117130 · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.806276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.806276Z digest=sha256:7de586e35a84c0c2062e7c92afc0ead85976083b2b773e052ac0155db8c6481d

Observation 73a526ea-6440-4f5c-9389-5392865a2fa3 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Textbooks Are All You Need II: phi-1.5 technical report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.811426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.811426Z digest=sha256:3f9f6beb1bda1a09b4f8f02a292cb4a2a1da20e86be0f34443308a7817520250

Observation 9219ab43-7c5e-4305-969b-4431df926d3b · outbound

This paper cites C-MORL: Multi-Objective Reinforcement Learning through Efficient Discovery of Pareto Front.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models C-MORL: Multi-Objective Reinforcement Learning through Efficient Discovery of Pareto Front

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.817094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.817094Z digest=sha256:1d50bc1f556f3c193fc0063347459cfa8a6cf2be196fb09c5825d0a11c477061

Observation 8f537bb4-cba6-4aae-aa08-89f635c58bf8 · outbound

This paper cites Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.821866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.821866Z digest=sha256:646ed0a6ba7402dfb13800fbcd15794bb3db14e7f3e6fc13768f1a25877906f7

Observation 5b424fa6-521f-40bd-8ebd-71fd51511d39 · outbound

This paper cites Miettinen.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Miettinen

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.826965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.826965Z digest=sha256:86cb6f604a2992b5565b9a1e759a33bc777fcbc8a2df3893b0ebeb45dbf700d9

Observation b3dd2f31-c10e-454a-80d8-943543c47e6b · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Playing Atari with Deep Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.831351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.831351Z digest=sha256:fe64be928828571182e6a8b72a59086384157f8ef74c2276572254d1a7d8e7e2

Observation f5bbc9b6-96cf-43ae-bbab-10479171ff50 · outbound

This paper cites Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.835885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.835885Z digest=sha256:e092a332fc0bba86638a578621f4d8fdbc00c2ef84ae91bd6eb0a8e8c01aa306

Observation 4a02459c-9a13-4c0b-bd04-d4fe2ce87cc6 · outbound

This paper cites Nash Learning from Human Feedback.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Nash Learning from Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.842003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.842003Z digest=sha256:36473182b38275f231462fe06d9efb7a67aa3211c25c707401131ea82c3770d6

Observation fac20a13-0c40-406d-a9b7-4a1e03449bed · outbound

This paper cites Nesterov.Lectures on Convex Optimization.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Nesterov.Lectures on Convex Optimization

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.623299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.846442Z digest=sha256:5b8e8da28efb1b98faebf59d6a72e398709df6a6f340379bf55c9058f3b2b966

Observation c2e6e737-dd3e-4f36-8a1f-4b68a7768097 · outbound

This paper cites Ouyang, J.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Ouyang, J

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.850761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.850761Z digest=sha256:339d6a7fc39031ede29539a076dfffca0b3c9af6c5432b912988e3f2d714b97f

Observation 105e4be2-4440-4070-a3b7-8e5af5cf5807 · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.600262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.856185Z digest=sha256:6ca1797cd4e908b2e0981b23b908f4b6f5cc480b7aad24041d147bb8eaa8aea4

Observation f6a8772d-65a1-4d8c-b3b7-20b6c374bccb · outbound

This paper cites Instruction Tuning with GPT-4.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Instruction Tuning with GPT-4

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.861883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.861883Z digest=sha256:824a682cebfbc96a56c18f3b73128e0301ec2c2d13052930399f72bf4ad44b32

Observation 3967cbdb-e400-4cc1-ba9c-428c0dc91d7f · outbound

This paper cites Rafailov, A.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Rafailov, A

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.585329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.867351Z digest=sha256:a69fbe0b58aecf876314230c2066fb135f188bfa1f3a5410eeb42e470fc66f62

Observation 2291ff78-929f-4d2e-8784-b2e19075bac4 · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.572840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.871803Z digest=sha256:7184d7d1257f27bf50a247ab311f1d98f1542f897eb8d49690a616351a11c5e4

Observation 6be56b49-96bb-4a5f-8214-c123907126df · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.558779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.876996Z digest=sha256:950312a4b14ac950aed7add631b90f4bb928cbdfdfc225133afee16b6f192d0e

Observation e5cda545-9fa3-4340-b923-cc82302bcf30 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Proximal Policy Optimization Algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.881623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.881623Z digest=sha256:cd3cadb32ae178f477b7f0a33bc1265fa0a51b62272539d1b047fe288be1d73d

Observation afd30265-c73c-4dce-989c-88dd644cec04 · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.546219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.886824Z digest=sha256:64cf55c5b40c7b3ffc6764e1fc42e75829122eda3217f4b2ac9a93b7720aa5f7

Observation 51222244-0ed7-4f36-8d8c-ce489914eb3b · outbound

This paper cites Stiennon, L.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Stiennon, L

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.893071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.893071Z digest=sha256:aa3859bf1512137acfb541b748ff85cfa0ebdc1e41e80dbf4c01a5c8b774fd8b

Observation c2a6a508-17ed-428c-8f7d-aed6243c3c26 · outbound

This paper cites Gemma 3 Technical Report.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Gemma 3 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.896996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.896996Z digest=sha256:a1a7bc3c0bcc85305a56b44cd69cf71c472bc229409a1c4cd5cd9a8039b16f22

Observation 94072b2a-6de9-48cb-b716-30cbd0683cb7 · outbound

This paper cites Van Moffaert and A.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Van Moffaert and A

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.524477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.901994Z digest=sha256:d6c4fac6a5a0c0884edf51a0d22dc6284174bd7b482d73bddfe634381d09a045

Observation 6b764e25-b99b-4bd2-8694-1900509d9bc4 · outbound

This paper cites von Werra, Y.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models von Werra, Y

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.511798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.906906Z digest=sha256:247ae783f3226830ce60513e3102ed1563e60aafcd539cc46a5ffcb53933a3b5

Observation 042a9d0a-0f3c-4566-a3e3-91228f7fd0a9 · outbound

This paper cites von Werra, Y.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models von Werra, Y

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.499491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.910695Z digest=sha256:ef58525667f00e4be59e2e08b0fd6878de6d0b2e43cc4108550d6c886d52123b

Observation 31b027a1-7302-4f0f-9d53-6502de4fb9bd · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.488356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.914460Z digest=sha256:c0bc554a05f422ca78cf139c728db828b32cb8d6734c4377101e6f176c2c2f1c

Observation 3aee28d5-a3a3-47f9-bc4f-ad5443d8e10b · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.475939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.919179Z digest=sha256:2e28404907b5a6819ec807e879394e05b6f1c6eda54b2bf2316f81bf18332efa

Observation 63933835-d9dd-44a5-b42f-57788a1a3070 · outbound

This paper cites Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.923655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.923655Z digest=sha256:0a1d825cf37f8fb9dfad28c7ce41ac37a10fe88168a3ca0faa8efc39671d5052

Observation b4da37b9-4d56-488c-8df6-daeba4423801 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.928698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.928698Z digest=sha256:f180d30b51d36785b6a3d0cc2e7cad1eed9fac3fc8c311ade9a500e80b5fe747

Observation 35c71cd1-2d5f-4f56-8bd6-3f102cdcf0ec · outbound

This paper cites Wortsman, G.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Wortsman, G

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.463254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.932835Z digest=sha256:62606df5fa741b149557d176d5692e867016951b3a16431e6c6e80e9c2d4b706

Observation 4a7fa947-b0db-4897-8338-4ba6e0f838a1 · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.937018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.937018Z digest=sha256:8738ace37a3cbf6d5ee10b7cf704d827df136078f708d4ad24d573c2117f7d6d

Observation 33380517-f687-4950-98ab-c8daa1119659 · outbound

This paper cites an unresolved cited work.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:57.443958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.941580Z digest=sha256:8a2e5edf260b7263035e447cb6b2ac7560843beff4b920ec613a89c928f31d79

Observation 8a0f6e7d-2ea4-48ed-8701-b3a5af2080f3 · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.946275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.946275Z digest=sha256:53864d09b8702506de69bdfb4052d694ec652202f722eda53952707e7287e454

Observation 9974e1cf-4a27-4df4-9986-444f91439db5 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.950817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.950817Z digest=sha256:e8d606a0e4fc09e814d0b9f5736e2baa46b46c4f4eeb0f8aa36b961154c5080c

Observation da7eb7c4-a63a-408c-972c-234aaa23fa80 · outbound

This paper cites Zhang, S.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Zhang, S

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.432643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.955070Z digest=sha256:ea450e575169fd82cb61f62347b0a0d877bc4272b143686d94535440b5a0e7d2

Observation f60cd730-7c06-47ff-ba87-002bce33f06b · outbound

This paper cites Zhong, C.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Zhong, C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.419888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.958865Z digest=sha256:31359892f519443c9bd02415faefa991b0a45a0e7a56232c5f705973f5f969cb

Observation 5c9d276b-a7b8-433e-ae9b-1d89198140ca · outbound

This paper cites Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.964024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.964024Z digest=sha256:c050053447ff7ec3c526a07fd2747da687bf740f14328616fcce213591e97816

Observation a712e480-ee0c-4e85-a4b2-1c3ea06ad12f · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Fine-Tuning Language Models from Human Preferences

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.968309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.968309Z digest=sha256:60ef595c3304091d39c16037a116fdc583c12489b876f96e9742c5dc8425b92c

Observation faa9016d-dcc0-421d-a4e5-f2b6e5e7d29c · outbound

This paper cites KX i=1 wiRi(y) # =p +q +r +s. (18) While by the selection ofk, we have E y∼πA,w.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models KX i=1 wiRi(y) # =p +q +r +s. (18) While by the selection ofk, we have E y∼πA,w

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.407048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.973004Z digest=sha256:569116a8deb5a6547f4e5a2a807b610d25bb43fd0daa5d6ed87e5f16042bc0f2

Observation d1343081-3ca8-4c45-8cf3-e47cb7e7f967 · outbound

This paper cites reverse_kld.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models reverse_kld

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:57.394592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:08:56.977697Z digest=sha256:a0c5776713d7fea9284b4d88c3042e1cce94cbb4eff8f1b68195be41a64dae2e

Pith citing papers

Observation 2d5844c4-8adb-4f7a-8a6c-5936bd1060f1 · inbound

MGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM Alignment cites this paper.

MGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM Alignment Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-06-08T02:03:42.868631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T01:38:49.892824Z digest=sha256:f225bc790592917704c26090b5ee5b409d4c33056f5eec1455410cf8ecf38403

Observation eccd3dea-1ffc-4b32-947f-b1be0bdb3176 · inbound

Multi-Objective Exploration and Preference Optimization via Mutual Information cites this paper.

Multi-Objective Exploration and Preference Optimization via Mutual Information Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models

Reference 101

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T21:18:57.948481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-03T21:17:46.551850Z digest=sha256:8f3564c308db0261f50ef08b61a5fc94efb16f5263424f7aeffa813aac875c94

Observation ff0edeb1-b97d-4968-97f2-3708657bd0e7 · inbound

Multi-Objective Exploration and Preference Optimization via Mutual Information cites this paper.

Multi-Objective Exploration and Preference Optimization via Mutual Information Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T08:52:34.039089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:52:34.039089Z digest=sha256:4f7fd16648102a6264be7c640bd6894e92236ff65f981261cedc4de4ff7a3f45