Pith. sign in

Paper Citation Record · LEDGER

Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2312.11456.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.11456 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:50.431221Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T08:07:45.294768Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation be79fbfa-d555-494a-8c4a-1fae5e1094ea · inbound

Self-Rewarding Language Models cites this paper.

Self-Rewarding Language Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.513253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:2e5e938428117483a0f9e1d37ad1d388866d260a8432b23f3fdc95fed3b0fc1d

Observation cb628c24-4853-460c-a3d4-056fa4cffffd · inbound

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning cites this paper.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 4344

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.431221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.431221Z digest=sha256:7290ac7e93d30684987e54f38f73c4203d56f1217dca04f4776c86e6e435fc00

Observation c925f773-402e-4005-aad0-6c67f4ba00c6 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.783481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.783481Z digest=sha256:b73ef0cd961ab043d1ce1943bc7c3107ec860602d348597137a6d2fd6c380f85

Observation e26b3256-2698-4ae6-8ffa-1f03b099f100 · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:46.467622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:46.467622Z digest=sha256:10d39673737d3d9a66655a2f6667d35f5349f50f304da5930fd4355beaccaf3f

Observation 2fd6201e-da41-4d52-bed5-5e89a29b156c · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:30.488898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:30.488898Z digest=sha256:472647a99575d16c09e0e97f1c12a81c2cd0a41ef8ef1f6e0f155d5c17b54327

Observation d633d1f9-2b01-4aa8-929c-e316d7ca4f71 · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.435213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.435213Z digest=sha256:e1053ef039de9bde3922c6a72888c44f96ec29a60f10e421fa0969ed3c5f4c3d

Observation 738674b3-804c-466f-8cb8-2176cf94dc03 · inbound

Aligning Large Language Models with Implicit Preferences from User-Generated Content cites this paper.

Aligning Large Language Models with Implicit Preferences from User-Generated Content Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 12

Resolution
malformed identifier
no resolver link, observed 2026-08-07T10:50:52.162823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.162823Z digest=sha256:eeaa50fdd0c6d94dbe2d62ba9d32ae87a69cccd9d779e5a3fe1445e1ea130e55

Observation 2cbcdcb0-37db-4fba-b52b-61e85f8acf78 · inbound

Boosting LLM Reasoning via Spontaneous Self-Correction cites this paper.

Boosting LLM Reasoning via Spontaneous Self-Correction Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:30.736946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:30.736946Z digest=sha256:b92d509ba85cfc838ce8d6317157e40218f03946e62b081beecde0ffdfc71abf

Observation aadb1a06-6ba7-44f1-8ad0-d2b140535341 · inbound

Bridging Offline and Online Reinforcement Learning for LLMs cites this paper.

Bridging Offline and Online Reinforcement Learning for LLMs Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.397544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.397544Z digest=sha256:821de703da81dd838bdfef3de3fed0f43c11d40df120f0b389253b61d2f06908

Observation 7a5f7fda-8506-4f53-83f0-03e83d0c13f5 · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.794667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:54aaaf19633d930dc6fd0750180fe2acde5aa0f19f34a62bde6edb87dd292c61

Observation caa7a167-a786-40b8-ad11-8e33e6c466dd · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.249776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.249776Z digest=sha256:9d17b7fa5966368a59bf635b16471f77b329afa306f8c93074b370061662d1d5

Observation fce8968b-b7ec-4828-97f6-67d3d1edf231 · inbound

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention cites this paper.

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:56:51.807191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T20:54:30.449792Z digest=sha256:66b815253ea6a077363c5e6595d419db6c5f941aad04923cb3849dba46ce58a8

Observation e9f47699-f286-41a3-83c0-869043ed15e1 · inbound

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning cites this paper.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.556600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:64dc0654b8260f5ddb3a04d6ce502253a9bbbf292c659b4a5a1f7e8d6557eb00

Observation ca620d9e-6f09-4fc7-bfbe-5a69903cbd8f · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.577593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.577593Z digest=sha256:67a4eea17f4d9c2b9fa05bb304d13c2917c1fe4b2b3bf2ade809e6cfb5c1fe42

Observation f46296f8-9f42-4f61-bc8a-4a54339c294d · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 207

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:44.508870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:44.508870Z digest=sha256:b416f8f3e1ebb875eb0e7d9da7b22c7bd079db647ba06df8fa51a717c4ad6e1b

Observation 54e2a028-9ce8-437c-82c9-b871889ae3d1 · inbound

T-TAMER: Provably Taming Trade-offs in ML Serving cites this paper.

T-TAMER: Provably Taming Trade-offs in ML Serving Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T14:57:00.278197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:57:00.278197Z digest=sha256:fa57bdaf223f3ba931ea23fde301a1685974a9c048e41caef9e888450cb47788

Observation a187e3c4-ed2a-4fd9-bc9c-25ad2c792f07 · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.027237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:2775e666f5cfda0070f267f1fa35ba8c7c086cce755de9826e39fa282a2fbacb

Observation 29f80c5d-e9d3-4b69-a6ca-ff32d85aa70e · inbound

Improved Bounds for Private and Robust Alignment cites this paper.

Improved Bounds for Private and Robust Alignment Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-03T13:42:06.121305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:42:06.121305Z digest=sha256:fde157ca1a84bc353f713ad6232af5488b48f209c31050bc3bda53255f818fcb

Observation 9f384787-e9db-4a2a-bf85-b581734215fd · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:11.017458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:865b3f336610ded296bc28b5e9d7aa6e3e48a9100e37d2e48640b027d069c5e3

Observation 51151045-dabe-41db-85a2-582cede2776b · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:45.088960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:45.088960Z digest=sha256:38861e5d61a232ab42fb2e34bcf3e97e024024c127468d1993e6b1b49dd387c0

Observation 11128f59-79a4-4dc7-90d6-57e1d37c1fe3 · inbound

Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference Learning cites this paper.

Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference Learning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.664074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T14:09:06.139625Z digest=sha256:374eba6f2c46acc8f67b3efe617f425da781f9b5fa74fd04c1d1fe4006991670

Observation 40f61a43-74ac-4b71-8338-b529d1583785 · inbound

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning cites this paper.

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:05.005307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:15:12.985803Z digest=sha256:853860d4750b3b5215f4d9c113a2ae9e3261003c84fd9a7772ae916f1a787422

Observation 110138e7-d015-4b0d-a398-f9f63980401c · inbound

RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences cites this paper.

RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:21:07.087022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T17:27:05.533755Z digest=sha256:780d8c71cce9bc1d025a2488f4d278ef01aee0ce24a2d4447d89f79a62a9796a

Observation 3ef6e794-1ad0-4748-a798-b34efe831837 · inbound

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses cites this paper.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:58.665306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:13:29.292351Z digest=sha256:c2c4653ac26595f078efef94e158aab2413bc9cdba916c09fb1ff3482c58db4c

Observation 4ba08c5f-0a38-41ec-8355-2b6b0715ab24 · inbound

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training cites this paper.

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:32:24.270894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T06:30:51.812541Z digest=sha256:90d32ea57b10320a74868654311eec7e3dfc9288280044b3fa9f94f7ba590158

Observation eb3dd61e-1eb5-4ff0-841d-65edf791a5a6 · inbound

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models cites this paper.

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:26.783878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:32:16.166724Z digest=sha256:fa14812a8235fb737b22bf4507fa9d5878387c549a5385f5acabc09ae9021da8

Observation 07b58e96-e25d-472f-80e9-06479c7bbd52 · inbound

The Power of Test-Time Training for Approximate Sampling cites this paper.

The Power of Test-Time Training for Approximate Sampling Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:07:45.296622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T11:07:47.543592Z digest=sha256:142647914bd884be9324fdd5dbc5483f6e4e04a14aabd3572670f8398d7bc987

Observation cec4c83d-8ba5-45b8-b569-d375ceb5c8f1 · inbound

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon cites this paper.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:44:21.719763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:8474944e695ffe85d3ab58a3d6f363f925d50f58b7acde867cff7fd634df26e5

Observation 05b50ba8-3813-405f-8765-7a6df5cff1c2 · inbound

Subjective Risk Decomposition: A New View for Uncertainty Quantification cites this paper.

Subjective Risk Decomposition: A New View for Uncertainty Quantification Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:59:41.739998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:59:41.739998Z digest=sha256:8f65f9d5497a7ceec9e2ed90a563256acaaf6bc86659d67d7733d9d7f46275cd

Observation ab20eedc-4fd4-4669-abb6-94f61bdc20e0 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:59.330889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:59.330889Z digest=sha256:f1c998cdeaab42e9d9f13dc4fdf698f71724b67d37fed7086aa157a3df2421b9