Pith. sign in

Paper Citation Record · LEDGER

Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2404.03715.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.03715 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:33.077519Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:49:37.721117Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ff11211d-bc73-4838-908e-26f990d76380 · inbound

KTO: Model Alignment as Prospect Theoretic Optimization cites this paper.

KTO: Model Alignment as Prospect Theoretic Optimization Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:17:53.564060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T12:17:53.478052Z digest=sha256:c497bef64cc7498847b6887822bf49fee0785be9ab3260a619332358c94e15e3

Observation 35f6eb9c-ac9b-4b3f-858d-507259d3fdbb · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:04:10.572479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:e74f1e7f23b8fcac020f7380af57eab9b3f1b7cffd27b26f582ec637071c43a9

Observation 05481b0e-c6d2-4d15-9327-5f77272944cb · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:23:31.113118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:adeea0b57763f219d33813c74145c65fa4046b5c89fe72be0f2f9adf6927e179

Observation 3a9b81d3-fad4-400f-8a5b-9c988348fc5d · inbound

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers cites this paper.

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-07T22:50:33.077519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:50:33.077519Z digest=sha256:800353952fd7fb04a935c147152454bd72985e71cf37be43040bf60ccc14b25c

Observation b9384154-a853-48d5-8e9e-5e44e46c2217 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 205

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.386647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:569e58ad466de8a730bd4b273ed20439a4b718a322eb6935456cbf791ff14fee

Observation c8df08d7-fa7c-40af-90c7-3558c8025d87 · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:30.769929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:30.769929Z digest=sha256:bd4c5652aec163c46f50d0093f3f3beccdb895e84efc032fea4f6e430a9049e3

Observation cf583308-af18-4836-b38e-fa24ff026969 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.734257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.734257Z digest=sha256:6b8ab81e704b4ea47feeb238b10f725a9e9ec0108a6ab7f3065e0ba3dfffa98c

Observation 2c612a59-241f-4f9b-a974-7f75e4c5f031 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:59.736916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:59.736916Z digest=sha256:85fc461b42164b0ea1bd85a97614cc885ffc1bc6f646872bf4cd65e9d74b5b57

Observation 1b6b2cad-3767-4e73-9981-da4ba7c71a92 · inbound

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment cites this paper.

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.872410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.872410Z digest=sha256:cd8e15031c50f434a0b1aa60dc21a9cb2ebf912a654a42e820eadebcd0b85986

Observation e7f15eb2-1f01-4a41-b31c-5f4fe31cdcbc · inbound

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function cites this paper.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.109193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.109193Z digest=sha256:729859cbe728d14deee32f8998844af7a442973d4f224ad309678d8995f20c1e

Observation 0fc2bec0-5495-40c1-a5e9-18cfb280fb3c · inbound

Debiasing Online Preference Learning via Preference Feature Preservation cites this paper.

Debiasing Online Preference Learning via Preference Feature Preservation Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.672178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.672178Z digest=sha256:502789769870a667c367151c6bac678a659ec391ee655cc78e5cfbf228e10b95

Observation 4bbf2b40-5962-4839-b7d1-094cdec1c94d · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.944650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:1b691482eff81beebc3dffda376b79bef51fed1134b4891bee593b5d33f77a53

Observation 375fda13-12d7-495a-ac33-4357f018d7f5 · inbound

Improved Bounds for Private and Robust Alignment cites this paper.

Improved Bounds for Private and Robust Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T13:42:06.117000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:42:06.117000Z digest=sha256:bbf22934fda33dfba0b6fb9a7cf2c133ff39a173dc5a880f3f4ea41174fb93c3

Observation f59d078f-612b-4b29-96c8-2b6e0079f15a · inbound

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective cites this paper.

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:23.730311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:23.730311Z digest=sha256:1eb3931b3c2b6dee6edf4270f86ae4714e7baf06552f84fe75273327b48f6a21

Observation 13f26489-4421-4416-a739-9397259589c6 · inbound

CoAct: Co-Active LLM Preference Learning with Human-AI Synergy cites this paper.

CoAct: Co-Active LLM Preference Learning with Human-AI Synergy Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.584601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:26:18.163720Z digest=sha256:52ec31014b5a004e0a05f93f46b2d18b3158dc9cb17349dd4394ba20b4d8f1d0

Observation b200f40d-5da8-47c7-b62e-11bdda45921e · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:30.887352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:de39b1ca0615afb01a5e51ba671fe7214f8dd6e3d8663a7b36a6c3634b1a5d57

Observation 5d79cd9f-3635-4db2-96a7-9fdd7e45ab3b · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 154

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.354453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:28f80fa2a9b5343bafa5b8a9e1e3ddc40b4765b94509f12881c760360e1b11ed

Observation a65fb4dd-ff9b-4306-8680-9a32351a9c0b · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 154

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.654565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:9d2a953c5427da5e775812e2dc7459cf7a9a2b94b048d8b4e02ad574c5bf89b1

Observation 454872d7-679a-46f8-baa3-63fe12fbeb02 · inbound

Common-agency Games for Multi-Objective Test-Time Alignment cites this paper.

Common-agency Games for Multi-Objective Test-Time Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:15:06.588171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T06:14:53.685486Z digest=sha256:fb001fdf3a6b8d8ed70810f70aea14df93cb51b36ef7f658fc9a38332643f82c

Observation 160f2efb-8f9e-48a8-9cb4-0241ff246eda · inbound

Learning from Language Feedback via Variational Policy Distillation cites this paper.

Learning from Language Feedback via Variational Policy Distillation Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:39:00.287412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T20:34:36.764090Z digest=sha256:aa8f157ac1cc80fea3be91098eba52b66cf611c7ef31b008a850b4ed9d6f43fa

Observation eed979e2-bcfc-4804-a09e-539e2e4fa576 · inbound

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment cites this paper.

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.293961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T14:26:06.428076Z digest=sha256:d18f635bc12a3d1de06f073291d0b4bbcaf7c47edc8f15cfa0966fcc927d7cb3

Observation 39657248-29cf-426c-9954-0c4cfe4a1083 · inbound

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization cites this paper.

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.406664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:58:33.167891Z digest=sha256:f0315b54912b48dad29705b62e2d3a3dfc4857c387690b79a6af547542b567a6

Observation 8a86bec4-31c5-4500-aca5-98e5479967fc · inbound

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs cites this paper.

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:46:33.121702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T09:40:35.258818Z digest=sha256:4a0492b080b933cb5fc5567d96092291c8cf8d40c38df7e7be99458ef0386ee9

Observation 8af116e8-2e4a-4849-8af6-03779b44affc · inbound

AI Alignment From Social Choice Perspectives cites this paper.

AI Alignment From Social Choice Perspectives Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:37.723384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T14:12:36.892697Z digest=sha256:c1672ab97532d8c4428168b5df31806fbefce2dd1823ee113c437c6e51693b3a