Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T15:26:01.391615Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 2 inbound Pith citation observations for arXiv:1908.01289.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T15:26:01.391615Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:43:48.515394Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:09:51.302063Z
12 of 12 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 152a63c1-09b9-4c0a-879e-08bcdeb43aff · outbound
Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, asymptotically bound the one-sided regret rate forπi2 (Appendix A.2)
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 61d8394d-7f07-4896-91e1-c922e3c34acf · outbound
Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, asymptotically bound the one-sided regret rate forπi2 (Appendix A.3)
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d21d7a6a-3d17-4c37-897c-12656ee1cf04 · outbound
Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, the sequence ˆr1, ˆr2,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eafaab60-048d-4698-84fd-0596ea717907 · outbound
Dueling Posterior Sampling for Preference-Based Reinforcement Learning Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 87587dc9-8fb9-4c0b-893d-2970a49e6efc · outbound
Dueling Posterior Sampling for Preference-Based Reinforcement Learning Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ec22b845-b717-46ec-8736-d62a7780ee71 · outbound
Dueling Posterior Sampling for Preference-Based Reinforcement Learning To inferr, we approximate eachr(τi) with its preference labelyi
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation de74fe6d-ea40-4db9-88c6-017d4a1aba64 · outbound
Dueling Posterior Sampling for Preference-Based Reinforcement Learning Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 86fd71c8-f962-49d7-b9b5-0694f2fc9104 · outbound
Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, asymptotically bound the one-sided regret forπi2 (Appendix A.2)
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e90baac9-c09d-4834-be39-06f1efc27bcf · outbound
Dueling Posterior Sampling for Preference-Based Reinforcement Learning Then, asymptotically bound the one-sided regret forπi2 (Appendix A.3)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4d8afffe-5e44-4c74-8147-cc7ac2248b99 · outbound
Dueling Posterior Sampling for Preference-Based Reinforcement Learning The information-theoretic perspective used to prove 2) and 3) likely applies to a wide class of credit assignment models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 35329daa-3728-431b-9b9e-cd449b87295f · outbound
Dueling Posterior Sampling for Preference-Based Reinforcement Learning Under this assumption, one can show that applying the martingale techniques yields the following variant of (65): n∑ i=1 2∑ j=1 xT ij ( λI + i−1∑ s=1 xsxT s )−1 xij
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5e9faa1d-c807-422c-9abf-98f47acad9c5 · outbound
Dueling Posterior Sampling for Preference-Based Reinforcement Learning Unresolved cited work
Reference 248
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 99e536a6-804c-4e8e-9d7d-6a8133fa2060 · inbound
Thompson Sampling in Online RLHF with General Function Approximation Dueling Posterior Sampling for Preference-Based Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c4d08ec-b37c-43e1-b5d5-8834b378f028 · inbound
Finding Stationary Points by Comparisons Dueling Posterior Sampling for Preference-Based Reinforcement Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.