Pith. sign in

Paper Citation Record · LEDGER

Provable Benefits of Policy Learning from Human Preferences in Contextual Bandit Problems

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 3 inbound Pith citation observations for arXiv:2307.12975.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.12975 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 3 of 3 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:26:27.949991Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T14:10:46.616356Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation badb5ea3-3f00-44ec-a56a-facf1f0a8407 · inbound

Fusing Reward and Dueling Feedback in Stochastic Bandits cites this paper.

Fusing Reward and Dueling Feedback in Stochastic Bandits Provable Benefits of Policy Learning from Human Preferences in Contextual Bandit Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.949991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.949991Z digest=sha256:15b0d3b69a6617200dbe9647f868ada4e50e03428d1d78d7492effbefedb99b7

Observation c1ddcf5e-e2d6-4180-9c18-9b1d98fc76a2 · inbound

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration cites this paper.

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration Provable Benefits of Policy Learning from Human Preferences in Contextual Bandit Problems

Reference 235

Resolution
unresolved
no resolver link, observed 2026-08-15T14:33:57.861566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:33:57.861566Z digest=sha256:ad4dcb7f18b667961f1037557f291121a4207544a449e272bc498499c77bb74d

Observation 392f1c28-c7cc-4b29-8c7e-e03f1565dd43 · inbound

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling cites this paper.

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling Provable Benefits of Policy Learning from Human Preferences in Contextual Bandit Problems

Reference 264

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T14:10:46.622427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T14:10:46.230891Z digest=sha256:e182f7c2d6758ca2107ac9ea51e3c56a5409f77f3503ebecf94f426155783cef