Pith. sign in

Paper Citation Record · LEDGER

Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2502.16852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.16852 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:49:14.854778Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T19:32:01.290815Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 46aa73ec-3aab-4942-ba24-86c20a2b6047 · inbound

SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence cites this paper.

SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T23:49:14.854778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:49:14.854778Z digest=sha256:def4c0c64eb3c0a632c3eb9d83985a723a46df23d30c43116ed523776d6ab8af

Observation 1bc552a7-eb6f-4ba2-be9c-c225b57e6375 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.293338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:809b446cc7a58b460f7c2c7152959372dc13723635cdb1d167d08bff55099b14

Observation d8df3fe8-36b7-4734-82a2-68f7d0244696 · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.022728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:7632b438ab3fccb87dbbe4007d4518082369f0351bfdfae6ba774999ed827234

Observation 644bbec5-543a-4555-819a-4c4ca0d16576 · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:25.789502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:25.789502Z digest=sha256:f216df9ea289a0264d48beaaee13209d37c67112a1b82cff947d6abacd927162

Observation 564865ed-5513-407d-87d0-01759a8b9402 · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:25.670135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:25.670135Z digest=sha256:ffdc4ac0d0feb61863d3fcb52a4eddcd257a5679c8fde097cc6a2c09a891a7c0

Observation 7a2fc9aa-b845-4e20-82e8-52e78af90a38 · inbound

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium cites this paper.

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:06.812051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T16:54:58.732444Z digest=sha256:c9524e84301897deaa96d10f4e2e7ab361e4b18ca84420fc9c63a9df2a0d3397

Observation b3a700e2-9e5a-403b-8dac-0574b3b2c91d · inbound

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective cites this paper.

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:55.154910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:19:58.448344Z digest=sha256:bf7a39e8466f094606d9adb4b42b7dc6c370869f1efeb46cea6c06875c96a767

Observation a50f61f9-77eb-403f-b166-1f1aa1d90a84 · inbound

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment cites this paper.

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.277911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T14:26:06.428076Z digest=sha256:77981cd4a0984ebfda4fe054fada64334aa5bf9686b6e9ef9f8b25c82b3cd156

Observation c42f0dc0-1a24-4b14-b95b-906cda524044 · inbound

Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion cites this paper.

Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T02:24:42.953753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:24:42.953753Z digest=sha256:06339163e665e4cf4fba910d46b83ccc7317531d8c300826c0190fb167f7e338