Pith. sign in

Paper Citation Record · LEDGER

Provable Offline Preference-Based Reinforcement Learning

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2305.14816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.14816 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:20:13.002840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:29:02.979336Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ba8ba420-fd18-4f34-b324-9b965813f310 · inbound

Combinatorial Reinforcement Learning with Preference Feedback cites this paper.

Combinatorial Reinforcement Learning with Preference Feedback Provable Offline Preference-Based Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.002840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.002840Z digest=sha256:9eb2fc177cb46cd05662c92f52a422310db8ab02c4aefa0cc02b6280704e71bc

Observation 4c1e7082-13de-40a7-a7a9-ffe05b95f5db · inbound

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO cites this paper.

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Provable Offline Preference-Based Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:09.500623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:09.500623Z digest=sha256:62ec6d2a357b0fd20083fc897dac1ec20bd413c2342baf47040fa30459da61b6

Observation 16c687fe-e2a0-4cef-b0f2-4cef349d565e · inbound

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits cites this paper.

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Provable Offline Preference-Based Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:43.538251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:10:43.538251Z digest=sha256:138aebe1424969cbb714b7b554086199152541197193a0ddbe3347d5fce93ec5

Observation a3074639-539a-49a7-97b8-496c5ef4f467 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Provable Offline Preference-Based Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.936809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.936809Z digest=sha256:253327a8c0847bf5dc4257806243f63fc9ebecae40493dc67916128918209214

Observation be121500-6594-443b-9b65-15b395ac386f · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Provable Offline Preference-Based Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.030374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.030374Z digest=sha256:e49336bf06e7c1e41983ca91ffb61548bf806828ff377426a877de3d03be06bb

Observation faa4c54d-4083-4620-95f4-cb0d669efcad · inbound

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis cites this paper.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Provable Offline Preference-Based Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:59.418783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:59.418783Z digest=sha256:ef49ca8c53caee8036e987d97aa18c4637cba27682f7c3219a7f923986686ab4

Observation 73973510-7d79-4868-92e3-2ab89a0ef4cb · inbound

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model cites this paper.

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model Provable Offline Preference-Based Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T14:03:36.586968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:03:36.586968Z digest=sha256:fadbb46c918eeb2bab6e4c6d62965f166002da44a29ef37256d615d2c6086538

Observation 0c8d5c55-211b-4cc2-88ee-685fe1b9f111 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Provable Offline Preference-Based Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.562069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:98acb7e8a646aba99be9e5e61eb08df48c1f0a9ac7567311b06781fec40ae0e3

Observation 072c8af7-3370-41b3-a0c1-ca10dfdcd607 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Provable Offline Preference-Based Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.495290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:c78e99c65ff667b1a9b59260d536cca9538687ceb8532369c94d0efa2028f6a9

Observation efb1a5a2-635e-408b-865b-236fb3a78783 · inbound

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback cites this paper.

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback Provable Offline Preference-Based Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:29.063328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:03:48.813600Z digest=sha256:9944d10a3021d7d53cfc9070def1d2cae7851c56b5e91cf633b2386b2d77708f

Observation cf716446-1873-4fa2-9096-1d4cc83eddd8 · inbound

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration cites this paper.

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Provable Offline Preference-Based Reinforcement Learning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:40:20.967674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-15T21:38:53.792221Z digest=sha256:c5e440edae1b06b962c68837d7f1570b846b01e4b74ef95e5c7422bdbac2cc03

Observation ad244cf2-7e9a-4b06-b231-43d295016ddc · inbound

Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback cites this paper.

Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback Provable Offline Preference-Based Reinforcement Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:06:03.890094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T02:23:20.208976Z digest=sha256:e0c88c12fb31a66c251a4eacd634c144e49852bf0548ae61c9dab9603a8ebe4b

Observation f8da3cfa-a9c5-4e30-8936-b6b158725598 · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Provable Offline Preference-Based Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:31.090069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:9c4c8206f8363bfec22a58cf4db0c5225e3d214d19b302eda0d8f0fdfa09334a

Observation a6624355-e3cf-4a8d-ae53-9721c458c5cb · inbound

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? cites this paper.

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? Provable Offline Preference-Based Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:29:02.980864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T22:06:24.412259Z digest=sha256:4d1c9d3eea3c8a1fbd2bcf86d64fbecbd4dbcdf75e7b466e007b029edf1b6ea7

Observation 4f46e08b-8347-4ca8-8c67-e650cc688904 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Provable Offline Preference-Based Reinforcement Learning

Reference 228

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:bc13c072151dbe4ae61d180695d1f681ae36a1e87bed148d8b10f1cd15299e66

Observation 454061c9-42bb-493b-b17a-5224417792c2 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Provable Offline Preference-Based Reinforcement Learning

Reference 229

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:58.612845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:58.612845Z digest=sha256:71f1b3cdc9490c6f48388da16e31e1e92e56e689d806991ee7fd91d73ef20a8a