Pith. sign in

Paper Citation Record · LEDGER

Dueling Posterior Sampling for Preference-Based Reinforcement Learning

As of 17 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 2 inbound Pith citation observations for arXiv:1908.01289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.01289 v4

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T15:26:01.391615Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:43:48.515394Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:51.302063Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 152a63c1-09b9-4c0a-879e-08bcdeb43aff · outbound

This paper cites Then, asymptotically bound the one-sided regret rate forπi2 (Appendix A.2).

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, asymptotically bound the one-sided regret rate forπi2 (Appendix A.2)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.506566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:26:01.357974Z digest=sha256:45f6c6d11f2fde5fb2b1e857c7b339f31de55e001b5bd86005580890e41d5ae1

Observation 61d8394d-7f07-4896-91e1-c922e3c34acf · outbound

This paper cites Then, asymptotically bound the one-sided regret rate forπi2 (Appendix A.3).

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, asymptotically bound the one-sided regret rate forπi2 (Appendix A.3)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.497607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:26:01.361240Z digest=sha256:e96930b59f92693c190c527622986e87da44272de8b6633ff3dcab77f7624479

Observation d21d7a6a-3d17-4c37-897c-12656ee1cf04 · outbound

This paper cites Then, the sequence ˆr1, ˆr2,.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, the sequence ˆr1, ˆr2,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.488366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:26:01.364296Z digest=sha256:309e01cf4b99dc26968cad7bb1ba239cc1b8bf1db830a5b3f1e9eb750f1e4d7f

Observation eafaab60-048d-4698-84fd-0596ea717907 · outbound

This paper cites an unresolved cited work.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:26:01.478425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:26:01.368113Z digest=sha256:a622b65f2a09b61f2fb1fb108f61b85301cace3c8d094f9a7203da51e009c77a

Observation 87587dc9-8fb9-4c0b-893d-2970a49e6efc · outbound

This paper cites an unresolved cited work.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:26:01.469433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:26:01.370873Z digest=sha256:6880d2656774cec3cf858721dd3ff22aad44a552f1ae6e6710361072decbb912

Observation ec22b845-b717-46ec-8736-d62a7780ee71 · outbound

This paper cites To inferr, we approximate eachr(τi) with its preference labelyi.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning To inferr, we approximate eachr(τi) with its preference labelyi

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.459969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:26:01.374474Z digest=sha256:c3ecf3fad72ffaff4ebe6135199ef7c563ffc1810de4fc2642de423e90fc1d29

Observation de74fe6d-ea40-4db9-88c6-017d4a1aba64 · outbound

This paper cites an unresolved cited work.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:26:01.515394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:26:01.377899Z digest=sha256:d8c5a6fd89fbb1ac4e0c668cd9de777e499c65facdd7fbc32a8824b344fdb2d6

Observation 86fd71c8-f962-49d7-b9b5-0694f2fc9104 · outbound

This paper cites Then, asymptotically bound the one-sided regret forπi2 (Appendix A.2).

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, asymptotically bound the one-sided regret forπi2 (Appendix A.2)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.450647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:26:01.380876Z digest=sha256:1e87d05befecddc7145b82beab68bb8f901f59f4e7dd16851daffddcbc7db95b

Observation e90baac9-c09d-4834-be39-06f1efc27bcf · outbound

This paper cites Then, asymptotically bound the one-sided regret forπi2 (Appendix A.3).

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, asymptotically bound the one-sided regret forπi2 (Appendix A.3)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.441397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:26:01.384064Z digest=sha256:f5a5b3238dcf8aa0ef9c16e915fe7fd5b00abc3f15c6f8f99d4338a5fc3645a0

Observation 4d8afffe-5e44-4c74-8147-cc7ac2248b99 · outbound

This paper cites The information-theoretic perspective used to prove 2) and 3) likely applies to a wide class of credit assignment models.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning The information-theoretic perspective used to prove 2) and 3) likely applies to a wide class of credit assignment models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.431118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:26:01.387446Z digest=sha256:beff732e5a6b68edf20101b91fa3eb0ca9b5fa52d66ace0d9d7bfda3a1bc8f9f

Observation 35329daa-3728-431b-9b9e-cd449b87295f · outbound

This paper cites Under this assumption, one can show that applying the martingale techniques yields the following variant of (65): n∑ i=1 2∑ j=1 xT ij ( λI + i−1∑ s=1 xsxT s )−1 xij.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Under this assumption, one can show that applying the martingale techniques yields the following variant of (65): n∑ i=1 2∑ j=1 xT ij ( λI + i−1∑ s=1 xsxT s )−1 xij

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:26:01.421469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:26:01.391615Z digest=sha256:b5f2a8be646805d81425882cd911c3c77af9a49f6e5800d1308b4a507e13a536

Observation 5e9faa1d-c807-422c-9abf-98f47acad9c5 · outbound

This paper cites an unresolved cited work.

Dueling Posterior Sampling for Preference-Based Reinforcement Learning Unresolved cited work

Reference 248

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:26:01.523773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:26:01.349082Z digest=sha256:236161799555efc03742097ac8106897bdcc938f181a110911bd27dfedeb496d

Pith citing papers

Observation 99e536a6-804c-4e8e-9d7d-6a8133fa2060 · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Dueling Posterior Sampling for Preference-Based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.515394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.515394Z digest=sha256:4c45ccd386247b80c68ae244c9057027ef5e767c14fe76a61295aec271df8045

Observation 2c4d08ec-b37c-43e1-b5d5-8834b378f028 · inbound

Finding Stationary Points by Comparisons cites this paper.

Finding Stationary Points by Comparisons Dueling Posterior Sampling for Preference-Based Reinforcement Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:51.303516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T05:26:07.720218Z digest=sha256:f34d1d2f7a94aa4fec567c6036391948732291d65df847288f42eaa8f323cbc4