Pith. sign in

Paper Citation Record · LEDGER

Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2402.07314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.07314 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:33:57.516039Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:06:55.872592Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1db72c5e-c4f9-4a0f-a0a4-ca77ce1a4faf · inbound

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs cites this paper.

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:05:15.209576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:05:15.209576Z digest=sha256:f897faf4ccaf6e0e9a8dba8879767d4367622b95eeab04250b2786fe81ee609b

Observation e98b4f2e-9bb9-48cf-aa18-cc9c55bd8095 · inbound

Self-Improvement in Language Models: The Sharpening Mechanism cites this paper.

Self-Improvement in Language Models: The Sharpening Mechanism Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.827757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.827757Z digest=sha256:e562afc36145c7903901afba57a589ccc412c47c00f9394eab697f44e26d2e67

Observation 800496e3-91e4-41d9-88cc-260b1ff6e4e5 · inbound

Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems cites this paper.

Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:35:31.494836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T00:35:31.375020Z digest=sha256:ee7564dd8da5e7a5d904682d5caae69b266707f684e3dd7e884f172b5c85a2b8

Observation f56f94a3-d0d3-4a03-a3a8-816047500123 · inbound

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration cites this paper.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.419053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.419053Z digest=sha256:81217fdfbba91b0a7121594341dd6ae2ca07499d7b19777af1ce3e105abad4f6

Observation 4a78435b-0abb-422c-95b9-2722526d7a4a · inbound

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? cites this paper.

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:46.376259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:46.376259Z digest=sha256:92a770ec4e768cb552aab875ccf935eeb93686851964dc264bc100825367893f

Observation 62417c3d-95fc-47b8-9e9b-f5ba558409e3 · inbound

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits cites this paper.

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:43.394782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:10:43.394782Z digest=sha256:44dbddb940918308d3da67de7b74ec568ed2481d03e9b453e65c6e93d6d9aa2c

Observation 683f5707-3ded-4ea2-9359-4f792525dc08 · inbound

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis cites this paper.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:59.337652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:59.337652Z digest=sha256:934cb69522cfa8909342979a795d45548ef6a0319bf6da1696874f4d68850c5c

Observation 681a11cf-4fc8-4558-aaed-db5ba133eae9 · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.027663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:b75fb22221750c2da5f584a9259a7f909b40abad19a3977cc78059fcd4211736

Observation 15262199-a874-49e4-8ca7-17e4457fbee4 · inbound

Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions cites this paper.

Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.615988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T03:46:35.908522Z digest=sha256:f293c6e05cfff00d270ac6fd8b08a275ee2337f96a817c3a96cbe4e18dea5220

Observation c02ab51e-a296-4c6f-ab53-b9a3347391f8 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.873915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T02:31:11.200818Z digest=sha256:1fcd7cfa4f80c7a8a31b2bbd8e8859a4a9dc17b710204b89557c20b632b9c491

Observation d4255a95-1077-4dd9-8b92-ceef7e6f0dd4 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T18:23:21.126826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:23:21.126826Z digest=sha256:12b815914cb22feac9b0ebde8290027475e92e4dffb78a735368793b878414ec

Observation 02081b23-56b9-48b2-9542-1b49feb4a21a · inbound

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration cites this paper.

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 161

Resolution
unresolved
no resolver link, observed 2026-08-15T14:33:57.516039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:33:57.516039Z digest=sha256:38519b4fb2e2988ba7770570ca6640c7bceaecccc608fd94a98265c2c9b49c3b