Pith. sign in

Paper Citation Record · LEDGER

Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2411.04625.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04625 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:24:30.557757Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:06:55.879807Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a5c9c86d-4bf6-4af7-9c22-e8a3cc37626d · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:30.557757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:30.557757Z digest=sha256:41d888cb21d18d132a57ff1b317db34428ee942b82855b10754fc4e57f8bc17a

Observation c246fe7b-c98b-49c4-ab0f-333531e32621 · inbound

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis cites this paper.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:59.510021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:59.510021Z digest=sha256:670d192d59488072d1b2e9f60d3d609144afa642c24aeb0041e95f8c642c47f4

Observation e5c294c0-4fdc-43c3-b1bc-f90b10593d33 · inbound

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses cites this paper.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:58.743351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:13:29.292351Z digest=sha256:08d0111df2130ef6ec497cb122eed9faa48777ea5dafb6c8d4d1905640be8289

Observation df772fe2-d199-4d0a-a663-157a83450cb4 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.881073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T02:31:11.200818Z digest=sha256:44b913447b8e5eba0fe290012822e42542e5bd38d74cfa143ad66fe482d5ad97

Observation 5c771a33-d4db-4b2d-a6e8-86f3d462dbcb · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T18:23:21.126826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:23:21.126826Z digest=sha256:65026f7979aa5ef760a92ea6a871acefac5d916cc11ff585ae2c7df8052d46ef

Observation 26fdd42b-e885-4dd0-a223-ff9cfa071e81 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T18:23:21.126826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:23:21.126826Z digest=sha256:05835db7984979b237c666fbea4364615c7f905bc98b85070b9171fa0641ac81

Observation 5a1d4ea9-bd30-4d74-a74e-896919596d63 · inbound

Graph Dimensionality Reduction for Contextual Bandits: Structure-Specific Regret Bounds under Approximate Smoothness and Noisy Eigenspaces cites this paper.

Graph Dimensionality Reduction for Contextual Bandits: Structure-Specific Regret Bounds under Approximate Smoothness and Noisy Eigenspaces Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:55:51.230113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T04:24:32.483334Z digest=sha256:9b6b1a1fc40fd11e5cba71f187c0783d210d2af85d51ba4f18f15649936a63bd