Pith. sign in

Paper Citation Record · LEDGER

InfoPO: On Mutual Information Maximization for Large Language Model Alignment

As of 23 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2505.08507.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08507 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:01:30.169786Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 853756a2-cb3f-41c9-a2ed-471a2f0f7948 · outbound

This paper cites URL: " 'urlintro :=.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.910923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.910923Z digest=sha256:44bf3acfb6f0ff2670103a2e50f38f2e4e2470e6028a405dd3a63e2e38ddb7b2

Observation 3bfc6b3f-6fcf-40c6-adf9-d6e882849614 · outbound

This paper cites write newline.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.916104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.916104Z digest=sha256:4be5c81c317dbe38557cf04ca0f6d0a5e9c7a029555f0017258ecf6132ebb241

Observation 8bab2e73-a639-4005-a585-ef0d2c974344 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.999530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:29.920988Z digest=sha256:1d66e9b1634e1de267df8892e6dfae4c12b183697ed0736d57208c2c014ea035

Observation 8c749b8c-beaf-4c4f-bf2a-54467204a8a3 · outbound

This paper cites Program Synthesis with Large Language Models.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.925276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.925276Z digest=sha256:9e4cfe69a6697d166b9176a182a119fd4414898d0e8cdc98459340178859289b

Observation 27bfdd50-6ab8-4a53-9ce1-5d65395375c8 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.929616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.929616Z digest=sha256:fb6cd7c46cb02c3502bc30a3414cdcbdae4561aa9a3f25e023d1d52b59d0a868

Observation 02b95cc4-86ed-4004-80f9-ec17e8151620 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.978666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:29.933643Z digest=sha256:e8b3d4c4daf8a10632ada16b3682c8eaaefcd4269630b80091144f060f73d031

Observation c225d0ac-622e-4e90-96ba-7e40c2504415 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.937797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.937797Z digest=sha256:693313b8e3d9cfa34e52bc2358d157e132ae6431cbc1a7ae19782d91d2a2cb2d

Observation 963bdade-a1cd-41ea-9026-4265a94aee23 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.965732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:29.942419Z digest=sha256:a3619dade570c93f64e85c5912b9bbc2353a2d57cdd076525d6a7526e638a68a

Observation bfd77ebc-9fec-48b0-a3f6-f2b591677b94 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.952830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:29.946785Z digest=sha256:93fe6d189d7dbe2d200b8093ffd00aaf573ebe55cc7c8cd97f02348fe2ee0b5e

Observation 15e7a395-9685-469b-91de-623e75d0b860 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.950804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.950804Z digest=sha256:8cd5412060aea878970829dd34feec2584d9e9cb3d8c2e9f62cd94b0d2cf9c48

Observation 26742ea3-0049-4310-a1c8-395adcd07b2d · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.931636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:29.955072Z digest=sha256:37d46e1f49487677c3d154ae400755e283e029b1877bc3a97cdc53a8fa27f10e

Observation a325e666-b571-44ca-bee0-c1f652982af7 · outbound

This paper cites Noise Contrastive Alignment of Language Models with Explicit Rewards.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Noise Contrastive Alignment of Language Models with Explicit Rewards

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.958940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.958940Z digest=sha256:6f80d4173f153d022e38dc61bbbd128c9fcd92abb098a739e05e3cc7e7f4fe03

Observation 2a9ab82c-44bd-469b-a058-7ddb4ed2bb14 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Evaluating Large Language Models Trained on Code

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.963072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.963072Z digest=sha256:9e637118c83e51bdd3dd0ae7c11b5bf34b07c5dd0be74894bfd907c231896676

Observation 6d0bc3c3-f682-4764-b7bb-d663e45670a0 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.967280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.967280Z digest=sha256:deb63e805b9bc629020badc13c1affe6addf239b9c7bdfedf6d4a15c3d27878f

Observation 6aec295b-c6ec-42da-a3c1-07f5f202412d · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.971268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.971268Z digest=sha256:1a41f93fb6be6caf0688824e39bd7278f9739aa1e46417fcb56d8b9e1b633e8e

Observation 7f617a0b-208a-4516-8e88-cd641025a7c3 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.975433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.975433Z digest=sha256:132bc1d4082b5b0e6ca6c117a35a7c57c97be0d34fd66f1143751779e1ff0db9

Observation cf92b869-7082-413c-bbaf-d803c7d9a0cd · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.902559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:29.979776Z digest=sha256:c2134f31b3c9c22ef79f831ad3c7316401c08b7cdf137ccd2182b96a44752d9c

Observation 59e5b8b6-137d-4104-a5a0-18150ad54181 · outbound

This paper cites The Llama 3 Herd of Models.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.983638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.983638Z digest=sha256:e58bbc7643bfde1a48ec6146e296a3b734f0196011db8d8a1d2def355f633f3b

Observation c0f95fb3-d458-4f9c-b11c-26206f7b2e5f · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.889861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:29.987813Z digest=sha256:3edf6b0bbaa94eaeaffa852963d119a4214c4be0eb6aa4d2e02a3e589f51cb23

Observation 1c5ee20f-8b5a-48fc-9666-890439d09f8d · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment KTO: Model Alignment as Prospect Theoretic Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.991766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.991766Z digest=sha256:971a5635534e53e238eb524c141803d096d8dacbf70f954c197888b5c61fa5a8

Observation 121172a4-d12c-42e8-965e-9f6a9286b9f7 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.875948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:29.995808Z digest=sha256:15d89dddb08d26c9745f78328620026dc8cffc0c06e9b37223dc1e44e8833b6c

Observation f0cb6f24-07e8-4a6d-95c8-75e401b93a70 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:29.999619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:29.999619Z digest=sha256:33464fb27abe71588cce5e042f83dc33e7abdfde27e0b0bfef0fe4af56b40bfa

Observation 5650d78d-a847-42cc-8221-fc0c13ce9ded · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Direct Language Model Alignment from Online AI Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.003661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.003661Z digest=sha256:dc8d3c4312f8d019e64797768fb951788f3fc3ea4f62d305c405ea93344da520

Observation 8546ea88-9230-40d5-a710-81ab1d0c384f · outbound

This paper cites Mistral 7B.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.007659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.007659Z digest=sha256:b4c950db96ad983f372ff07ad6763fc1253b68dcd7b2f6e164704c4ffaedb74d

Observation 2a71715f-59f3-463c-b74e-c608740154ba · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.011597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.011597Z digest=sha256:80c33e65830a5808d64e89290061d484b67514f1b6bdad82ab924364dda7a0f2

Observation 89cc55cf-13a4-43f1-bede-b113fe9a3ef6 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Adam: A Method for Stochastic Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.015824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.015824Z digest=sha256:5c408e49455eeb2eab81f4fb3bf70d2c5dc49b93281035aa3b584040517f1067

Observation 263df5ce-1722-45d3-af34-0bdefee097a3 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.863370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:30.019748Z digest=sha256:66e83a81fdbc3af846f11d78fbdab6a80ccb8104043d0103db6627649b348812

Observation 32991a22-e661-43fb-8392-6e1d5b1e03c5 · outbound

This paper cites Conditional Contrastive Learning for Improving Fairness in Self-Supervised Learning.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Conditional Contrastive Learning for Improving Fairness in Self-Supervised Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.023561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.023561Z digest=sha256:da9c74d03e309a43aff9d7957744d9e7c3bb9c546a7800bc08e4807484d1e13c

Observation dd1d3041-7ae2-49bc-8346-b91dad26bed4 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.027666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.027666Z digest=sha256:58e27f594c3fcef910c08cf626f3f39f4676c8ba749021ccc950ca9f5dc9d617

Observation a66cb64b-1fc4-43dc-8c11-9fce68a0a370 · outbound

This paper cites Nash Learning from Human Feedback.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Nash Learning from Human Feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.031609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.031609Z digest=sha256:810777d0c5ed021cfef90708f4513c7b2746387a0039ec192594739a6133715f

Observation e96d362d-19f3-486d-a58c-8fa83520d088 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.850799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:30.035663Z digest=sha256:86a03f1efd31fc734105cf9f868f7fc490a0d8f0d11a7ef9bb80321387eef19f

Observation 85f69cbc-9499-43ed-803e-bbed47271c8e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Representation Learning with Contrastive Predictive Coding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.039509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.039509Z digest=sha256:88fc9cd70526fd4791df42b10d488e037eadb3f608cf168bd7271bf08c12e049

Observation ef33d5bc-1499-4d85-950b-a4d1f9ad92e7 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.838073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:30.043616Z digest=sha256:018ff7b9940f5bc8878b91aeead3dfbc38b7b2b02b7d8c6bceb712f1e67d7237

Observation 646779c3-7538-45e3-a613-4a3a5dc9ba8c · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.047460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.047460Z digest=sha256:4dae5814801709b8deed195428afddae26413b386fb98ab83a7be43b227a3c16

Observation 55c5d506-89c0-4514-8191-f7d6f9d01ae2 · outbound

This paper cites Iterative Reasoning Preference Optimization.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Iterative Reasoning Preference Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.051783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.051783Z digest=sha256:c0a32b8fcfdbddf0076de6485c9fdd23fc1313c592d3f68764f1f0a1e84282b1

Observation b9af7862-07bc-4ea4-9094-1a5eba98fd14 · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Disentangling Length from Quality in Direct Preference Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.055880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.055880Z digest=sha256:40b124431eaa6545bd630a58f5d8ed4b082eb1be0bb36fa269c31bd8c11d49ce

Observation ddc650b3-2547-43c7-a5bc-df3506484bed · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.825400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:30.060027Z digest=sha256:80b1a2e2f13cfd4f5d488496f153ca7d23319a30a47aebcd8bd14e82f42a37ee

Observation 51132a8b-a32e-40f4-bd44-c6b1801c1a4d · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.063686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.063686Z digest=sha256:36fd342d38551ea3ae5814447ca78f0e3ad67ed4d7b3930a0be180d1b586d18e

Observation 27a16a56-8388-416c-94ff-08dc03f11475 · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.067470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.067470Z digest=sha256:9a3b6d3bf937cf7a097c4bd428d32022de5a7b31d23b61e5ca9e9f33fe029088

Observation 6f1c94b2-353a-4f2a-b76f-6944bb8ccbf4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Proximal Policy Optimization Algorithms

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.071471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.071471Z digest=sha256:b4f98cb53e3d24a6e531a08dbfc01140367e1d7fed53e4c02da3c1fb40a233c5

Observation bbb7a90f-d231-4161-8ab7-323a002690b2 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.804224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:30.075358Z digest=sha256:6af8f3c2e29290c6feecd944e9611cd8bf3bd9408c52feeeff1e1d7cac057c2c

Observation 9720ebef-292f-4cf2-90ae-1d34025d9f4b · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.079064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.079064Z digest=sha256:b02036e119929580df7e67b3c0280b12aff141d1811088c0406fd51dcb2098e2

Observation 2eb8e6dd-4863-4eab-ad9f-10620b29cf7c · outbound

This paper cites Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.083231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.083231Z digest=sha256:68fd740213be6d33fb79e3c846efa54056c65581e28a4c0342f41e06e75c24d8

Observation e1a42b13-7698-4f3a-a3e8-d98314293d05 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.087528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.087528Z digest=sha256:68d70d1c6b467363d40640775d331b9365019b1e11fac2753393c4be4028cfbe

Observation 358c5337-4b1c-429a-bcef-dff33df4ad90 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.783161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:30.091647Z digest=sha256:d074b13b6e806d1fcd37324993d27277aeedd2516490328b53a02b4b1d0d9414

Observation 46fc1427-e61b-4968-9948-69d86abfabbb · outbound

This paper cites On Mutual Information Maximization for Representation Learning.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment On Mutual Information Maximization for Representation Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.095660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.095660Z digest=sha256:e4f427f9a67bca83bccae1a2718d71d6b6db02300a7a508e38fdc5a3878807c5

Observation b23bdb74-7636-4e0a-9422-aa4a67bf5dec · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Zephyr: Direct Distillation of LM Alignment

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.099954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.099954Z digest=sha256:e02a3163043704d5b442e2851e8e12cc3c43d431163a50f7f3e8dab1361d4ab8

Observation 1964ddcf-7864-4d77-a235-c7055fa0edad · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.104470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.104470Z digest=sha256:d3bc13b820c116f1b493da3ed7c392c1c81b5590e260a7819e331910ec07a0ae

Observation 9aa5b613-2996-4847-a5d9-d919e2b97e44 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.761207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:30.108562Z digest=sha256:8736431dc98e6ce15d0235d8a432867a85f79e89c5f0e099d495232f070d06bc

Observation 9d69d374-c9ae-4a8a-beca-57798c86d61a · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.747119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:30.112738Z digest=sha256:64de5313b5e4ffc1d5dc319b22cdbd47c9a1a768284443a4ed59cb78572cacc4

Observation 77e56128-ddb8-41bc-a98e-8db3d5e14aa2 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.733249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:30.116989Z digest=sha256:c27bda8731405a68c9993ce90f821f3c61308ff1cd09ce60df37d581935c01b0

Observation 72e51d6b-cbf6-4ec3-b7a6-0cdc1ea58549 · outbound

This paper cites SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.121138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.121138Z digest=sha256:9b572c98871d8a76b876bbabe478443e19c5132bdebc88cf3147b19d14de5e58

Observation cb486f40-3385-4352-a47b-572b1bacb018 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.720409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:30.125471Z digest=sha256:4da45b4789184235788ba1a783fbd78d62201ae1a35aaa80bec04d05b6e142ba

Observation 5d5d4293-279b-4ff4-810c-6722d49ffc12 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.129520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.129520Z digest=sha256:d173a99bed2c6bd5d21c5b903f26a24a842aebf6f491ea992250443fcefa0f97

Observation 966b6536-ee11-494a-8920-914433fc1d58 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:01:30.707494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T22:01:30.133963Z digest=sha256:e551e29e9109d9edee4141e41f63b242cbd9d5a67b2f70f7c635b8bdfe23a847

Observation 56f4439f-686d-447f-bd31-9b07ce9e441e · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.138334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.138334Z digest=sha256:86a158959d19fe9a74bbf36040306bc54cf1aca26cc67368e98359c2e3619545

Observation b527fa5f-5611-4f88-811b-e9b95c6ec9f1 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.142617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.142617Z digest=sha256:d0c4be08e146dbdff8a5a21908258975033060fc6c508e219340e52432f5280b

Observation b0971951-d3cb-4414-b510-8a0031962c4b · outbound

This paper cites Self-Rewarding Language Models.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Self-Rewarding Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.147175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.147175Z digest=sha256:bf9087f66392b3af56b87d5885a8f3a9971dc8d528967ae95d8b931807e3cacc

Observation fd360663-2ac3-415c-af6b-3ab7597baae8 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.152786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.152786Z digest=sha256:3b7ac72bfefb0a8d05539502cd7c3b176fcced0119f04857211dc5893de09a45

Observation d663b943-b965-4288-bcb8-73c011e67336 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Secrets of RLHF in Large Language Models Part I: PPO

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.157129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.157129Z digest=sha256:d7a362eff8f2a04d3323a7a7b2cd5d592bf27424fe1e3946fcd58694dcf59c70

Observation d937ad8b-5700-4532-8594-bd385cd4b38a · outbound

This paper cites @esa (Ref.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment @esa (Ref

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.161337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.161337Z digest=sha256:95ad4a74b44a1991fb66abcd97b04823e94864f4b3559151fe2a7e6a7a16ae87

Observation c301a447-869c-49a7-8d37-ee14c95a6878 · outbound

This paper cites an unresolved cited work.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.165599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.165599Z digest=sha256:4c3fcc78187254ebe8abcead812bacbe8f7bd91c20bb700a293931f79432ed8a

Observation 5ad0620b-fd3e-4269-aed9-65167606532b · outbound

This paper cites accepted.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment accepted

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.169786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.169786Z digest=sha256:dc4bd5044edbd378a64b3b9ea54a3d2744d32bbe69027b530b7520dd7c04a6ab

Pith citing papers

No inbound Pith citation observations are available.