Pith. sign in

Paper Citation Record · LEDGER

Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2312.11456.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.11456 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:05:15.202021Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T08:07:45.294768Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation be79fbfa-d555-494a-8c4a-1fae5e1094ea · inbound

Self-Rewarding Language Models cites this paper.

Self-Rewarding Language Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.513253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:6f3bc4e834aa1e9a1816436d8ec5d43b52c17c700765c6ee8b94dd99baa56046

Observation 53ddb564-ab81-432f-997c-92e62b299710 · inbound

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs cites this paper.

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:05:15.202021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:05:15.202021Z digest=sha256:eee0b81cd919b93e986af3117678d4c5b3068eb4429d38fbe594d40135fbfe85

Observation 182c56cc-2c66-43db-b674-7b15889833d9 · inbound

Self-Improvement in Language Models: The Sharpening Mechanism cites this paper.

Self-Improvement in Language Models: The Sharpening Mechanism Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.818560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.818560Z digest=sha256:5efe4f8bd959216dac5afe0264c038c4fd893b3c6cec3f504324b5648a1ce31f

Observation c38c503f-75ed-414e-a33c-f363807bb01e · inbound

T-REG: Preference Optimization with Token-Level Reward Regularization cites this paper.

T-REG: Preference Optimization with Token-Level Reward Regularization Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:56.417911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:15:56.417911Z digest=sha256:67a3f21f9dbf99b7c9ab9dbbf51af1378526e3c3ef9151821e822fdd6fa57c44

Observation 0762ebe0-d00d-443c-ba4f-58f67230247e · inbound

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment cites this paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.851281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.851281Z digest=sha256:d51b7ed2008e68a93a9d43be09a095aaa4406e0d13f7716f90d829e2c8847629

Observation 68877fec-7230-405e-b79d-0453e484ba40 · inbound

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment cites this paper.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.645871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.645871Z digest=sha256:4aca094fecf909a04151b23f02bd33fa4c2c53881415db46d97246ff654e5065

Observation 5d3c12fb-96dc-4564-94f4-9bdabdee110e · inbound

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs cites this paper.

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:32:47.964012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:32:47.964012Z digest=sha256:97fa4b8ee9eaa900dfb89c25d8de639f64ff985ae70539718e9bc45239c03c66

Observation cb628c24-4853-460c-a3d4-056fa4cffffd · inbound

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning cites this paper.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 4344

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.431221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.431221Z digest=sha256:cab6d89a98af97365fa6ee3c1868e41e19917b65d847f0531de0fbd106d47b5f

Observation c925f773-402e-4005-aad0-6c67f4ba00c6 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.783481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.783481Z digest=sha256:5027f428d07346878f85b3662bd584d49887f437cbff6d910efe184768a86884

Observation e26b3256-2698-4ae6-8ffa-1f03b099f100 · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:46.467622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:46.467622Z digest=sha256:aa88a6ea57cfe731c231c89cb19a27223c296ccae1bb707af87fcb882ce9a9b3

Observation 2fd6201e-da41-4d52-bed5-5e89a29b156c · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:30.488898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:30.488898Z digest=sha256:f131092334ee154278e226aef77c4000d626644ccc9eb7225f8e2ce4d24a5268

Observation d633d1f9-2b01-4aa8-929c-e316d7ca4f71 · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.435213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.435213Z digest=sha256:81c6c94780976537f0a0aaf5468d9571f19f4db133a85d96179860133255aa29

Observation 738674b3-804c-466f-8cb8-2176cf94dc03 · inbound

Aligning Large Language Models with Implicit Preferences from User-Generated Content cites this paper.

Aligning Large Language Models with Implicit Preferences from User-Generated Content Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 12

Resolution
malformed identifier
no resolver link, observed 2026-08-07T10:50:52.162823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.162823Z digest=sha256:dd3ecabb235517205fabd343bbefd132152340261c70255f5c0699ef6c70f3ed

Observation 2cbcdcb0-37db-4fba-b52b-61e85f8acf78 · inbound

Boosting LLM Reasoning via Spontaneous Self-Correction cites this paper.

Boosting LLM Reasoning via Spontaneous Self-Correction Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:30.736946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:30.736946Z digest=sha256:6887c813c98d44f8fd77d14d151f7c5ea0ac85edc77882b02baef2cbb794ad57

Observation aadb1a06-6ba7-44f1-8ad0-d2b140535341 · inbound

Bridging Offline and Online Reinforcement Learning for LLMs cites this paper.

Bridging Offline and Online Reinforcement Learning for LLMs Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.397544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.397544Z digest=sha256:2ec9e39428794219af76275d18c2de4adc419de895368313a614d7a011899d03

Observation 7a5f7fda-8506-4f53-83f0-03e83d0c13f5 · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.794667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:c49098e22df2d5d26886de5d72b7adf1bc3f4ff8737b89ff748455e6527da2f9

Observation caa7a167-a786-40b8-ad11-8e33e6c466dd · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.249776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.249776Z digest=sha256:9c99d736f3abe8b60f2920a90776f39266be07e29c0858f70077073b1ac94eb3

Observation fce8968b-b7ec-4828-97f6-67d3d1edf231 · inbound

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention cites this paper.

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:56:51.807191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T20:54:30.449792Z digest=sha256:54eea8f9565c49a087d328df49d68c923ed66df6ac10bbfe001ec4e9775b886d

Observation e9f47699-f286-41a3-83c0-869043ed15e1 · inbound

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning cites this paper.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.556600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:26ce0b8e5addd2db91b9b528ccdb43b3292777a81c0baf5c9edf634944f72ac7

Observation ca620d9e-6f09-4fc7-bfbe-5a69903cbd8f · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.577593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.577593Z digest=sha256:b0df782a573101b78b814cc1ce907294b291bc33daf7a3d71fecba72c7547ab6

Observation f46296f8-9f42-4f61-bc8a-4a54339c294d · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 207

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:44.508870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:44.508870Z digest=sha256:c83b19e0e7997c7699474bcb56132e7a3859bf141ca2cbb7082fc6b9577d6807

Observation 54e2a028-9ce8-437c-82c9-b871889ae3d1 · inbound

T-TAMER: Provably Taming Trade-offs in ML Serving cites this paper.

T-TAMER: Provably Taming Trade-offs in ML Serving Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T14:57:00.278197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:57:00.278197Z digest=sha256:2f11eaf0750d1c011156725b3f660aaad1080a70e3e47aed0b52b6b1831797ba

Observation a187e3c4-ed2a-4fd9-bc9c-25ad2c792f07 · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.027237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:eee9589513266a48fc09d1f6d7b75077c00a55b5cb8b485b9a58170c5facfbc4

Observation 29f80c5d-e9d3-4b69-a6ca-ff32d85aa70e · inbound

Improved Bounds for Private and Robust Alignment cites this paper.

Improved Bounds for Private and Robust Alignment Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-03T13:42:06.121305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:42:06.121305Z digest=sha256:3b104b23f7fd540c06db0c43d44ac8167dc17d9995df3576167e19f34eca98f3

Observation 9f384787-e9db-4a2a-bf85-b581734215fd · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:11.017458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:10c5700a0b37eb27f64756a897cd58e21e140d3e36a5e8614bcada5c831abba2

Observation 51151045-dabe-41db-85a2-582cede2776b · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:45.088960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:45.088960Z digest=sha256:2e739c3d19a28478d822f4eed29f38702bd1248fd083f0963c1e6bb110363bcd

Observation 11128f59-79a4-4dc7-90d6-57e1d37c1fe3 · inbound

Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference Learning cites this paper.

Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference Learning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.664074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T14:09:06.139625Z digest=sha256:d60e40de9fc4518c17088241d730bacda59f5c99bf3de8c9ab1602106a861812

Observation 40f61a43-74ac-4b71-8338-b529d1583785 · inbound

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning cites this paper.

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:05.005307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T01:15:12.985803Z digest=sha256:f2dcb982680ea65cb6a789eec1919bcc1080050b07d17a7d1a03e02a4b76750c

Observation 110138e7-d015-4b0d-a398-f9f63980401c · inbound

RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences cites this paper.

RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:21:07.087022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T17:27:05.533755Z digest=sha256:68b2ef55d3ae7e778df57c187bd06b993c28e255178cf1b10b3208bad4ebf1a3

Observation 3ef6e794-1ad0-4748-a798-b34efe831837 · inbound

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses cites this paper.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:58.665306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:13:29.292351Z digest=sha256:acc50b1aadd5f62c231b30612d2b03eb8222b29dd949bbd75cf88d095fec5f49

Observation 4ba08c5f-0a38-41ec-8355-2b6b0715ab24 · inbound

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training cites this paper.

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:32:24.270894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T06:30:51.812541Z digest=sha256:9191b800de7f58fd7a8978e1a8b0c5bbd71e70a98c186aa513702d547c5acd8f

Observation eb3dd61e-1eb5-4ff0-841d-65edf791a5a6 · inbound

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models cites this paper.

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:26.783878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T11:32:16.166724Z digest=sha256:4e941dfcf5b544fa159256877fd3b2d4c7721221fd6d15a1abd14ba77b4c10ee

Observation 07b58e96-e25d-472f-80e9-06479c7bbd52 · inbound

The Power of Test-Time Training for Approximate Sampling cites this paper.

The Power of Test-Time Training for Approximate Sampling Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:07:45.296622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T11:07:47.543592Z digest=sha256:38c3e1d4a7fef5601b11e4c89ca98b5f591eaace193f9b4efb17a26015988401

Observation cec4c83d-8ba5-45b8-b569-d375ceb5c8f1 · inbound

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon cites this paper.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:44:21.719763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:96db8f1b94e68d36a5fd166aa42adddc251148e8393f5a8be83b8f49c313f599

Observation 05b50ba8-3813-405f-8765-7a6df5cff1c2 · inbound

Subjective Risk Decomposition: A New View for Uncertainty Quantification cites this paper.

Subjective Risk Decomposition: A New View for Uncertainty Quantification Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:59:41.739998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:59:41.739998Z digest=sha256:b8a814af2fbb9c1fafa25597da1dc0a3f7ccf5ebf558ed02ae432ed02d836d65

Observation ab20eedc-4fd4-4669-abb6-94f61bdc20e0 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:59.330889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:59.330889Z digest=sha256:7112e8f462bd2242a20a146870f2e3e612cca5e882fb8e9f9046302b4ccbced3

Observation 801be598-9771-4bfb-a8e4-51aa7ce27033 · inbound

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation cites this paper.

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:03:47.550403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:03:47.550403Z digest=sha256:f2d70edaf020b83fd4f303cc8e458416b8fd12c9639e099df29f1a96efc670f3