Pith. sign in

Paper Citation Record · LEDGER

Dueling Posterior Sampling for Preference-Based Reinforcement Learning

As of 17 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 2 inbound Pith citation observations for arXiv:1908.01289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.01289 v4

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T15:26:01.391615Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:43:48.515394Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:51.302063Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 152a63c1-09b9-4c0a-879e-08bcdeb43aff · outbound

This paper cites Then, asymptotically bound the one-sided regret rate forπi2 (Appendix A.2).

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, asymptotically bound the one-sided regret rate forπi2 (Appendix A.2)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.506566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:26:01.357974Z digest=sha256:b2e92e2ef36d925ca0e64cc4b86b63ab5db1474af44c1ce63de78382f2ba7a3e

Observation 61d8394d-7f07-4896-91e1-c922e3c34acf · outbound

This paper cites Then, asymptotically bound the one-sided regret rate forπi2 (Appendix A.3).

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, asymptotically bound the one-sided regret rate forπi2 (Appendix A.3)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.497607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:26:01.361240Z digest=sha256:7477537aa488797fae5b61f0fac39bc45cdccf47bd2c2ea1f8ad955622547837

Observation d21d7a6a-3d17-4c37-897c-12656ee1cf04 · outbound

This paper cites Then, the sequence ˆr1, ˆr2,.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, the sequence ˆr1, ˆr2,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.488366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:26:01.364296Z digest=sha256:caba6fd029113c6e91860430b2cc3dd148b3ef55f94842fa0e9327a068cff70e

Observation eafaab60-048d-4698-84fd-0596ea717907 · outbound

This paper cites an unresolved cited work.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:26:01.478425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:26:01.368113Z digest=sha256:0e3d31875c6824bf4b8426468d34b7172bbd27e969016cbf26c4eb46848b1f77

Observation 87587dc9-8fb9-4c0b-893d-2970a49e6efc · outbound

This paper cites an unresolved cited work.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:26:01.469433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:26:01.370873Z digest=sha256:5e7ea7430606d8bf6d67a9728919e3e36e13a6639fa35422ee91fffa055fe39b

Observation ec22b845-b717-46ec-8736-d62a7780ee71 · outbound

This paper cites To inferr, we approximate eachr(τi) with its preference labelyi.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning To inferr, we approximate eachr(τi) with its preference labelyi

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.459969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:26:01.374474Z digest=sha256:7c133981af3525a8543762f9d00dff7de0aaf6c7e964c281dcc3059fd446f9c4

Observation de74fe6d-ea40-4db9-88c6-017d4a1aba64 · outbound

This paper cites an unresolved cited work.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:26:01.515394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:26:01.377899Z digest=sha256:16066784f713b2519930b2eb992cb1a19cd0c57559349dded46a11f7584c8c74

Observation 86fd71c8-f962-49d7-b9b5-0694f2fc9104 · outbound

This paper cites Then, asymptotically bound the one-sided regret forπi2 (Appendix A.2).

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, asymptotically bound the one-sided regret forπi2 (Appendix A.2)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.450647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:26:01.380876Z digest=sha256:5bd47fbf3a2da432875ef9675f0f5fc5078c9ad630d86ef9f823e473da0dbb9c

Observation e90baac9-c09d-4834-be39-06f1efc27bcf · outbound

This paper cites Then, asymptotically bound the one-sided regret forπi2 (Appendix A.3).

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, asymptotically bound the one-sided regret forπi2 (Appendix A.3)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.441397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:26:01.384064Z digest=sha256:5c4a2372bf43afcbe3578d9a3a5c5a7d11665a31d1479f2eec15d17d76cc8f5e

Observation 4d8afffe-5e44-4c74-8147-cc7ac2248b99 · outbound

This paper cites The information-theoretic perspective used to prove 2) and 3) likely applies to a wide class of credit assignment models.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning The information-theoretic perspective used to prove 2) and 3) likely applies to a wide class of credit assignment models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.431118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:26:01.387446Z digest=sha256:21ababba36c6ed02cad650726aab56268132a47907ccc9eb5f7ce00aeb3a2141

Observation 35329daa-3728-431b-9b9e-cd449b87295f · outbound

This paper cites Under this assumption, one can show that applying the martingale techniques yields the following variant of (65): n∑ i=1 2∑ j=1 xT ij ( λI + i−1∑ s=1 xsxT s )−1 xij.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Under this assumption, one can show that applying the martingale techniques yields the following variant of (65): n∑ i=1 2∑ j=1 xT ij ( λI + i−1∑ s=1 xsxT s )−1 xij

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.421469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:26:01.391615Z digest=sha256:81e6e5d9cefb70115243fcbbf8d101e61ed5d16794e4b5bec2158fe825db5c07

Observation 5e9faa1d-c807-422c-9abf-98f47acad9c5 · outbound

This paper cites an unresolved cited work.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Unresolved cited work

Reference 248

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:26:01.523773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:26:01.349082Z digest=sha256:2fcdf55989463e532a5d2e8f1de4759946cb803c45a8f4638fd6c8d316deac08

Pith citing papers

Observation 99e536a6-804c-4e8e-9d7d-6a8133fa2060 · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Dueling Posterior Sampling for Preference-Based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.515394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.515394Z digest=sha256:4c45ccd386247b80c68ae244c9057027ef5e767c14fe76a61295aec271df8045

Observation 2c4d08ec-b37c-43e1-b5d5-8834b378f028 · inbound

Finding Stationary Points by Comparisons cites this paper.

Finding Stationary Points by Comparisons Dueling Posterior Sampling for Preference-Based Reinforcement Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:51.303516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T05:26:07.720218Z digest=sha256:ddc57b29111ac073753a4ea0561dbc0aedf75641a09d8500830dc35d99b70015