Pith. sign in

Paper Citation Record · LEDGER

Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2407.00617.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.00617 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:14.166568Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T06:15:06.516739Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7d609393-a485-48e9-831d-0d2108ca7be1 · inbound

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function cites this paper.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.166568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.166568Z digest=sha256:e7d60b2d8ef59ed796237cbcb491c2dfec90f49318476296626df674d7bcaa21

Observation 78f2aad9-cdc8-4194-835f-960b014972e0 · inbound

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization cites this paper.

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:58.431783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:40:58.431783Z digest=sha256:dae07266b79fae570234fc128145a09901b18a44787a99c2f4c270253a170fbd

Observation 0bea63a5-2892-41ef-8df9-08596b5ac2cb · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:25.524406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:25.524406Z digest=sha256:701466f385c535994e52c9bc813d8c7935b2a2daea65e7f4c813704c1a44644c

Observation e3367dc2-ea39-4b96-bdba-d8f844872624 · inbound

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective cites this paper.

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:24.161428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:24.161428Z digest=sha256:b1e3442f50b84308488d6369f6d5a97d574249bf787119909f7ff7fafce2c0fd

Observation 12b22110-0364-4920-8647-55160f8ecb2a · inbound

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium cites this paper.

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:06.862801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T16:54:58.732444Z digest=sha256:90bed532449acd76737a82240e1059c828b855062143494604692df5bd582558

Observation 3de88b08-e481-495b-9ba6-eb8aae212352 · inbound

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective cites this paper.

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:55.373741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T01:19:58.448344Z digest=sha256:95a9f63c8b279ba16d2a052bf4e60f7655608848719df467f16fc36039974af9

Observation 494ba96e-1b3c-4bb4-8a3a-7784ce071891 · inbound

Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions cites this paper.

Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:31.596943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T03:46:35.908522Z digest=sha256:eb8bc0fc34ce020a00c13380e593c8ac5722202cc1962176d63fe8d6a1446855

Observation 4b125e06-5f14-43d4-85bb-79ce9038f6cb · inbound

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning cites this paper.

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:26.511794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:15:24.919355Z digest=sha256:ee0a96b7f741aa71b852466bad8e751aab1adad752974848eff3a6b721714511

Observation 0b3fc829-0a8d-4059-8640-cab73bdfc385 · inbound

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning cites this paper.

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:05.102397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T22:03:05.102274Z digest=sha256:27c95ef32e7ef2d4f20ee4acc5664f37153579ff00b37c6c6e37758a2bf101a9

Observation fd179071-170e-4843-bc7f-41dad63fcca3 · inbound

Common-agency Games for Multi-Objective Test-Time Alignment cites this paper.

Common-agency Games for Multi-Objective Test-Time Alignment Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:15:06.520722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T06:14:53.685486Z digest=sha256:b0b8ff74fac55abeb080fb5f0c179d0f6f8d67d543103bbb80b2fd0fdc8f678e