Pith. sign in

Paper Citation Record · LEDGER

Human Alignment of Large Language Models through Online Preference Optimisation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2403.08635.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.08635 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:32:20.963680Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T13:11:24.028811Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e57a0052-665a-4d6a-a9bf-e207e09a7be3 · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models Human Alignment of Large Language Models through Online Preference Optimisation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:58:16.898151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:3b6f34751d0ecd67401284a0f597960e76a2d6dca763a46a5b403a49dda1bb80

Observation 9f07e96b-2159-4740-a1d1-875239dda663 · inbound

Preference learning made easy: Everything should be understood through win rate cites this paper.

Preference learning made easy: Everything should be understood through win rate Human Alignment of Large Language Models through Online Preference Optimisation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.963680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.963680Z digest=sha256:ad7ef22dca8128dc3f2b2994a59c2054038a32aa4932cf1a4d54ee0046e78555

Observation 15b86385-82d0-4368-b077-1c142b2587e1 · inbound

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought cites this paper.

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought Human Alignment of Large Language Models through Online Preference Optimisation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:37.284453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:37.284453Z digest=sha256:5a37e06ab9aeab7921d082edbbba34bda15d87537935a0fb06fda9c94cbc03ee

Observation 90faa925-f9ee-4c16-8373-4b82a5ff6928 · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Human Alignment of Large Language Models through Online Preference Optimisation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.031523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:82a22ad62b9b308dfaf0f37eeef63285cabcd1e6346ad29da0bef8c2e7d6a9a6

Observation 95ae58a7-83aa-40ae-bf8c-5e961fcbf0ce · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Human Alignment of Large Language Models through Online Preference Optimisation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:24.182186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:24.182186Z digest=sha256:1b0db5e5d0a1a4bfb841c1be30b4a90d5cedf6581670037d8006d51333357d60

Observation 7e042000-3c1a-486f-a564-7ec1890f3e78 · inbound

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning cites this paper.

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning Human Alignment of Large Language Models through Online Preference Optimisation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:26.503032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:15:24.919355Z digest=sha256:8e8b3e53fcfb22d147cddc7ee8daccb5d4e59f5f5e4275a9407410355aca54ed

Observation 0f7ab43a-c49f-4647-a192-6e4db481869c · inbound

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning cites this paper.

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning Human Alignment of Large Language Models through Online Preference Optimisation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:05.096769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T22:03:05.102274Z digest=sha256:cdabc826305b4624d83f3830d2a1cc6217f2a7cb9011735b76348b8300010c53