Pith. sign in

Paper Citation Record · LEDGER

Thompson Sampling in Online RLHF with General Function Approximation

As of 17 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.23927.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23927 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:43:50.639923Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact3
  • verified fuzzy31
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4249897-9122-4b45-8413-d0da2bfbad51 · outbound

This paper cites Analysis of thompson sampling for the multi-armed bandit problem.

Thompson Sampling in Online RLHF with General Function Approximation Analysis of thompson sampling for the multi-armed bandit problem

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:57.353799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:47.472924Z digest=sha256:606cc2edff3aa3c9bbdb7eb42d480d4ca2cfe968784fc7df359ceef6a256cf5b

Observation 20d9d1ba-99c3-4248-8eb1-27b873063592 · outbound

This paper cites Thompson sampling for contextual bandits with linear payoffs.

Thompson Sampling in Online RLHF with General Function Approximation Thompson sampling for contextual bandits with linear payoffs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:57.169067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:47.503626Z digest=sha256:f24f1044d316cf992a297e3e77ffb889a5552e95a81f6dd4e2f824da384991e7

Observation 123a3047-6d7c-4587-b00f-bc2f1b16ed18 · outbound

This paper cites Near-optimal regret bounds for thompson sampling.J.

Thompson Sampling in Online RLHF with General Function Approximation Near-optimal regret bounds for thompson sampling.J

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:57.050843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:47.527002Z digest=sha256:d03187c89559ea6e418abfc6c071f73b4b9b85cf5010152e0ec4f17af7245474

Observation ffa5cb5e-80ae-48f0-a949-f154c1aeaa4c · outbound

This paper cites Preference-based online learning with dueling bandits: a survey.J.

Thompson Sampling in Online RLHF with General Function Approximation Preference-based online learning with dueling bandits: a survey.J

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.921778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:47.557153Z digest=sha256:d8f2a78058f114dd6f31ba8b9f9c9ba57a6fe73d2dd2885babd32a7f2d62655a

Observation 068d17e4-c7b3-451c-b836-a752a0dd49bd · outbound

This paper cites an unresolved cited work.

Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:56.839597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:47.588635Z digest=sha256:a2828834c7c76002da72d58cdf802ede1ed7c47b97377d0f9f446a66681b4c11

Observation 1c0b73b2-aa5c-49b1-990e-9906ac6bb7ed · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Thompson Sampling in Online RLHF with General Function Approximation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.614974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.614974Z digest=sha256:3461fa933c33deff9d65fc94a0b21b55fa6315e94d313951a07057c7b0a36619

Observation bf60b172-58a5-4b1c-a562-17d5db2b54ab · outbound

This paper cites Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation.

Thompson Sampling in Online RLHF with General Function Approximation Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.660403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.660403Z digest=sha256:ddbc0152d9b3ac228322346dd3b43fa4de88fa2ff41faa2d9839f773a9744b24

Observation df94777c-9a6c-42ee-a834-a528c133d4c5 · outbound

This paper cites On the Weaknesses of Reinforcement Learning for Neural Machine Translation.

Thompson Sampling in Online RLHF with General Function Approximation On the Weaknesses of Reinforcement Learning for Neural Machine Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.697268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.697268Z digest=sha256:dbccbacede379702ca5c68a10b42aa682ac36cdab138b7b704c54a604b6638b7

Observation ec97d09e-1422-4c8e-beb9-24e7242a2e90 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Thompson Sampling in Online RLHF with General Function Approximation Christiano, Jan Leike, Tom B

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.773058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:47.725692Z digest=sha256:f6b8895c4fbbbcb08d7af9da50f7a0c75ecca4402f673734b3e057e30a61b1b5

Observation c3b6e898-2535-472e-bb8e-1f0b4efb17ce · outbound

This paper cites RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023.

Thompson Sampling in Online RLHF with General Function Approximation RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.691296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:47.755339Z digest=sha256:602f81b1255aa70f66c9ad20b7f01f50d9ea4bba5b0e26837b533a9da3fc22af

Observation 9b05ccd0-1758-4e94-898e-5e4cf8905e2f · outbound

This paper cites Schapire, Aleksandrs Slivkins, and Masrour Zoghi.

Thompson Sampling in Online RLHF with General Function Approximation Schapire, Aleksandrs Slivkins, and Masrour Zoghi

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.598480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:47.799854Z digest=sha256:cadca3256035b7ce679823552300ccf061d8c910c048c6dffe86bb0a371a7961

Observation 19324a54-6076-42a7-9974-d420050cd73f · outbound

This paper cites Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO.

Thompson Sampling in Online RLHF with General Function Approximation Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.833667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.833667Z digest=sha256:2df74b1dfa9b09e491a83c2eb456240b5f18e7e34ffde784dcbc87e7682585f1

Observation 3b0a6cb4-d055-45c1-963e-d178d68d1d07 · outbound

This paper cites Foster and Alexander Rakhlin.

Thompson Sampling in Online RLHF with General Function Approximation Foster and Alexander Rakhlin

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.515634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:47.872516Z digest=sha256:0b4a6e05ad736f3e8c8c7093a30bc9400b1e50f008312aeddef95d533edec653

Observation 7e17df10-3712-4376-ba25-69acb9567738 · outbound

This paper cites Scaling laws for reward model overoptimization.

Thompson Sampling in Online RLHF with General Function Approximation Scaling laws for reward model overoptimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.445967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:47.899602Z digest=sha256:39702a03dda566337c0e0de6d1e6e10153cffb30aa733bd73aa785ab5b91f215

Observation f93d914f-0c4c-473a-82e8-2c013645e3ae · outbound

This paper cites A General Theoretical Paradigm to Understand Learning from Human Preferences.

Thompson Sampling in Online RLHF with General Function Approximation A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.923648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.923648Z digest=sha256:404b15ee213bc2b17597a6c296e9be121bc235a20af64b9819f53d1c4f638015

Observation e83e8437-ccdd-42d5-b84d-f82d4a42da7e · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Thompson Sampling in Online RLHF with General Function Approximation Reinforced Self-Training (ReST) for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.949387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.949387Z digest=sha256:51be037779be91fee053494cd2160ddb4cc2b2885205aea704720c066e93a21b

Observation 577cacf1-c0d0-4a3d-81be-5207f551a71d · outbound

This paper cites Randomized Exploration for Reinforcement Learning with General Value Function Approximation.

Thompson Sampling in Online RLHF with General Function Approximation Randomized Exploration for Reinforcement Learning with General Value Function Approximation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:43:51.698037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:47.982915Z digest=sha256:2e26a823039d905967bf190e06db55262f7c0799638d754a66a9edbe5c830f0f

Observation f297a784-efaf-4ffe-89e3-7dcd0598597c · outbound

This paper cites Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo.

Thompson Sampling in Online RLHF with General Function Approximation Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.010070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.010070Z digest=sha256:3952c96dd23472ff99cca2dd2341f3e4ddb71ec1f6b4aae57484e15c77c58c93

Observation c42b36f4-36c4-48c0-91b7-2a61f39d6dd5 · outbound

This paper cites Learning trajectory preferences for manipulators via iterative improvement.

Thompson Sampling in Online RLHF with General Function Approximation Learning trajectory preferences for manipulators via iterative improvement

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.372101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.044855Z digest=sha256:e78e7cd41e8290a0a58ff043dad0f8f3fa0555e7714cbdc69d763f088fb77703

Observation 8fcfabfa-f0f1-4928-b963-d64cda7e98fd · outbound

This paper cites Schapire.

Thompson Sampling in Online RLHF with General Function Approximation Schapire

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.305869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.086054Z digest=sha256:5986baf5272f8ccbf550d8bc93c2f4473ac97191e76b008f1618e26ac74b507c

Observation 8d7c2828-6114-47ee-9b1c-75a721be65d2 · outbound

This paper cites Bellman eluder dimension: new rich classes of rl problems, and sample-efficient algorithms.

Thompson Sampling in Online RLHF with General Function Approximation Bellman eluder dimension: new rich classes of rl problems, and sample-efficient algorithms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.160284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.220473Z digest=sha256:caad36d32f7da89b4aa5a48a00d667a41ab4aab80ba16fa89213ae21abe9ce76

Observation 0c2c094a-4035-4108-932e-73172d5f9e74 · outbound

This paper cites An Introduction to Variational Autoencoders.

Thompson Sampling in Online RLHF with General Function Approximation An Introduction to Variational Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.246384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.246384Z digest=sha256:8a99d366c4acb72efc9c7caf0f6329bd29dda2eaba4e5ecdbc5f02fce79c84c6

Observation 42031251-04a4-4d60-98f6-1d92937d10ec · outbound

This paper cites Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism.

Thompson Sampling in Online RLHF with General Function Approximation Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.277156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.277156Z digest=sha256:6a746efc8cc050e410a66b0017802b9c8541379a1122c687609879fa135ed388

Observation f0776b47-d595-408d-8199-0e20cbb831ad · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Thompson Sampling in Online RLHF with General Function Approximation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.349648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.349648Z digest=sha256:ab512f9a24646a43935fdd5f322d32c9edca409600a51de18b712d8283f8aca0

Observation 5e9a7c7f-6d8b-48bb-a13c-f3dd10c93cb9 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Thompson Sampling in Online RLHF with General Function Approximation Statistical Rejection Sampling Improves Preference Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.431444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.431444Z digest=sha256:0bc53c77abfeebd62d121de825becdcf2c6e9a7bac38c4a421b675d29c7ea803

Observation 2d07115f-7ad9-4934-859e-b3cfd268c94f · outbound

This paper cites Roberts, Matthew E.

Thompson Sampling in Online RLHF with General Function Approximation Roberts, Matthew E

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.983442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.478015Z digest=sha256:9afd9d98f3f4f9052dffae265b7dc2484af79f50fe64e564bbd0df32a3106d04

Observation 99e536a6-804c-4e8e-9d7d-6a8133fa2060 · outbound

This paper cites Dueling Posterior Sampling for Preference-Based Reinforcement Learning.

Thompson Sampling in Online RLHF with General Function Approximation Dueling Posterior Sampling for Preference-Based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.515394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.515394Z digest=sha256:4fa7e9a23278a708fbfc4cf3f473587f5cddc1d9276c5cb28218b9f8fdb2922f

Observation d2dfc404-e7b3-4272-a6ce-8e481e070dfe · outbound

This paper cites GPT-4 Technical Report.

Thompson Sampling in Online RLHF with General Function Approximation GPT-4 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.557964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.557964Z digest=sha256:83dce908e297df014c331ebe148b9d5081916b503456d67ad44c8f20f42565b2

Observation d119f7ce-7e96-44a5-8ae4-32de28a7a17b · outbound

This paper cites Randomized prior functions for deep reinforcement learning.

Thompson Sampling in Online RLHF with General Function Approximation Randomized prior functions for deep reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.724220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.595572Z digest=sha256:fb11d0586ae4749d118347dbd794e19c91b8207a5ae6b2fb5b2d93109982e0a6

Observation 49fef49c-a677-43d7-aabc-9d0303095e41 · outbound

This paper cites Approximate thompson sampling via epistemic neural networks.

Thompson Sampling in Online RLHF with General Function Approximation Approximate thompson sampling via epistemic neural networks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.458395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.687662Z digest=sha256:1214773500275f064a68c01dc39660326b6334e3a0ef6399dbd749c82fca0251

Observation a4149ebe-e20f-44e2-98f4-d8e5ea133972 · outbound

This paper cites an unresolved cited work.

Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:55.265065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.737226Z digest=sha256:429d7e5df6f19772fa68fdbd020850d41e59485007d94708e6b9b4ab0044a080

Observation 6f9f557f-4443-4c58-8032-dc79c85f99d1 · outbound

This paper cites Training language models to follow instructions with human feedback.

Thompson Sampling in Online RLHF with General Function Approximation Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.769424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.769424Z digest=sha256:1faf0324736347029c10548f7cd134974ea6340a7aa366d2ef2f1a09267a931b

Observation 861b40f8-ed0f-40de-9e17-2b78ec60c871 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Thompson Sampling in Online RLHF with General Function Approximation Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.796716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.796716Z digest=sha256:43a5ca3ef84057f3eb6056f6e39ae6180e2e4a12a5b73a82356bcb69814d7610

Observation 211392fb-78f6-416f-88c2-2596ae83e015 · outbound

This paper cites Worst-Case Regret Bounds for Exploration via Randomized Value Functions.

Thompson Sampling in Online RLHF with General Function Approximation Worst-Case Regret Bounds for Exploration via Randomized Value Functions

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:43:51.321069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.814525Z digest=sha256:cf03abf08980282001a308deaa0856533a415c92199672996fb75ccf9124d169

Observation 55c69687-57ec-4809-8c77-83e041a8c371 · outbound

This paper cites Learning to optimize via posterior sampling.Math.

Thompson Sampling in Online RLHF with General Function Approximation Learning to optimize via posterior sampling.Math

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.197918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.838269Z digest=sha256:709ad7fbf05d743245e8875173894763782a1ca835958b9dfd9408b625eb98f0

Observation bc4131b1-1e3a-4dc0-9db7-ff14eb6c686a · outbound

This paper cites Optimal algorithms for stochastic contextual preference bandits.

Thompson Sampling in Online RLHF with General Function Approximation Optimal algorithms for stochastic contextual preference bandits

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.096880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.861004Z digest=sha256:767ff3ccd0da09a544580b767cf738b62441472660038afb7a932922abca6894

Observation d14f95a0-6435-44a0-ad5c-650065971df0 · outbound

This paper cites Efficient and optimal algorithms for contextual dueling bandits under realizability.

Thompson Sampling in Online RLHF with General Function Approximation Efficient and optimal algorithms for contextual dueling bandits under realizability

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:54.840829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.890800Z digest=sha256:f2717be876707a19f0da9daf4130e5f80529441dd310d3e8923e42e496f86e73

Observation 96c7ef00-ab1d-4d6b-afc8-59b1df7108b3 · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Thompson Sampling in Online RLHF with General Function Approximation Dueling rl: Reinforcement learning with trajectory preferences

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:54.465692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.921797Z digest=sha256:0e071dcbb68b667cd40ff2b3e22fc6fbd94d7e536c0f428c633954096d8d183f

Observation 49191887-b0d3-43ff-95e7-cb122383c61b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Thompson Sampling in Online RLHF with General Function Approximation Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.940005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.940005Z digest=sha256:c64849989366a63d1b8c74d7ee4b120cbf03fc759f72caeca21d830f53a1d913

Observation 279c74ef-69af-4a5a-80c5-1c6e4b99451a · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano.

Thompson Sampling in Online RLHF with General Function Approximation Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:54.117813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.969661Z digest=sha256:975e3ce42518e15c54101210c28da89e5c636f74391aec405602c34feac02ba3

Observation fa5d1bcb-9c9a-48b3-8d23-0f6bb19a2314 · outbound

This paper cites Thompson.

Thompson Sampling in Online RLHF with General Function Approximation Thompson

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.933023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:48.992418Z digest=sha256:e0702f4c2fd8709d1bc917e47ab54a2c9728d574dabe4162ef5b5587476a7ea4

Observation d0e6e6ee-2a2d-4ea5-9d0c-675fddd8b452 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Thompson Sampling in Online RLHF with General Function Approximation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.018024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.018024Z digest=sha256:7c0206ecfbf5e99726fd55a4ce977a3f43124d6f088fb56716e4fbf414787619

Observation 5ade00be-bf41-4888-ae26-23ff03ca7444 · outbound

This paper cites van de Geer.

Thompson Sampling in Online RLHF with General Function Approximation van de Geer

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.773838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:49.052190Z digest=sha256:7570b3d4be216901a2aab70c17ef128ae805356af316abb30fbf642ce74d76d0

Observation 76980b50-eb90-47f5-9a42-dc7bc440db34 · outbound

This paper cites Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints.

Thompson Sampling in Online RLHF with General Function Approximation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.083529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.083529Z digest=sha256:8e031fb292351758b9415cf9778e8757a13bb9ba3685941caea8358a5144ae3d

Observation 7f75d0cc-fa25-45bd-bef3-aa7b602e3ef7 · outbound

This paper cites Thompson sampling for combinatorial semi-bandits.

Thompson Sampling in Online RLHF with General Function Approximation Thompson sampling for combinatorial semi-bandits

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.720079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:49.108031Z digest=sha256:ac052712db44caa9a16a96805ad830df4dd99fe3267500751cebb901a840dbef

Observation 66e2e15c-29c5-4fef-b039-035424f36808 · outbound

This paper cites Is rlhf more difficult than standard rl? a theoretical perspective.

Thompson Sampling in Online RLHF with General Function Approximation Is rlhf more difficult than standard rl? a theoretical perspective

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.586179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:49.117791Z digest=sha256:ddb156edf63824ca897c518e99f4a6ee80775ad6886b4175c2f033273fa64452

Observation 32e371bb-2530-41ee-91aa-d4f60570624f · outbound

This paper cites A survey of preference-based reinforcement learning methods.Journal of Machine Learning Research, 18(136):1–46, 2017.

Thompson Sampling in Online RLHF with General Function Approximation A survey of preference-based reinforcement learning methods.Journal of Machine Learning Research, 18(136):1–46, 2017

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.466104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:49.197952Z digest=sha256:9cd06af5b71c6fb16a4a39f1817825492b10f6c3fb0759ab87eda560b8fa1dd5

Observation aaaeabf1-1a17-4283-b3f8-afbfa8d4821b · outbound

This paper cites Making RL with Preference-based Feedback Efficient via Randomization.

Thompson Sampling in Online RLHF with General Function Approximation Making RL with Preference-based Feedback Efficient via Randomization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.305402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.305402Z digest=sha256:1bc0a3e848c5cd82209c6dc0098f1f72e2b4632d41db3e6d8fc413a386731581

Observation 3bff1bf3-7f85-42d6-9c5b-5fa3c5a8987b · outbound

This paper cites Borda regret minimization for generalized linear dueling bandits.

Thompson Sampling in Online RLHF with General Function Approximation Borda regret minimization for generalized linear dueling bandits

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.309730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:49.376621Z digest=sha256:853cd7d0c049d048f9ee6cd58baaa22c983748d4e4acf601782ea7c0461694a2

Observation d633d1f9-2b01-4aa8-929c-e316d7ca4f71 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Thompson Sampling in Online RLHF with General Function Approximation Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.435213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.435213Z digest=sha256:e3d9a45bb849c937617f0b7114d6a5a496eb6d34699ac6874846cc139b540143

Observation 50e372fb-8206-41ca-be88-96913864589f · outbound

This paper cites Near-optimal randomized exploration for tabular markov decision processes.

Thompson Sampling in Online RLHF with General Function Approximation Near-optimal randomized exploration for tabular markov decision processes

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.149868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:49.519282Z digest=sha256:e84d7c13ed7531f827ba0ca045855a838c796fa8b78ba66019e3be40a33e187a

Observation 358a7f11-bc21-4f36-ae13-e03d2f64e846 · outbound

This paper cites Yang, Aarti Singh, and Artur Dubrawski.

Thompson Sampling in Online RLHF with General Function Approximation Yang, Aarti Singh, and Artur Dubrawski

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.943796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:49.630519Z digest=sha256:c646517edf91dfc26032bc980d1f4af1ff2daa0bb4b474d275dcd8d35e50d186

Observation 2224a913-2e3b-4ce0-89b8-bb0c530ca86a · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

Thompson Sampling in Online RLHF with General Function Approximation RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.730231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.730231Z digest=sha256:5076b0efe975e1f350a123c906f2fbac8bd7a53e623cc1f405c5ac753713019b

Observation efdf6433-3dac-4f16-8619-95bc291194bb · outbound

This paper cites The k-armed dueling bandits problem.Journal of Computer and System Sciences, 78(5):1538–1556, 2012.

Thompson Sampling in Online RLHF with General Function Approximation The k-armed dueling bandits problem.Journal of Computer and System Sciences, 78(5):1538–1556, 2012

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.725910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:49.832227Z digest=sha256:7b92b0d5ad824e7b4346ac57c428e15bdfb8940cf906bc7bdfb777ed68e75dc8

Observation 74185e95-b077-4d1a-8ed4-77e557bb356e · outbound

This paper cites Frequentist Regret Bounds for Randomized Least-Squares Value Iteration.

Thompson Sampling in Online RLHF with General Function Approximation Frequentist Regret Bounds for Randomized Least-Squares Value Iteration

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:43:50.963293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:49.926426Z digest=sha256:d125478cf040c001f534f8564f6153d69605acaf87290ad51c89f6014ac88f99

Observation be121500-6594-443b-9b65-15b395ac386f · outbound

This paper cites Provable Offline Preference-Based Reinforcement Learning.

Thompson Sampling in Online RLHF with General Function Approximation Provable Offline Preference-Based Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.030374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.030374Z digest=sha256:fbaccc8aed0415acccd1fbe3ee4551d9b6ca3ee7552c58be9451f7c5e8f2bf13

Observation 52573980-9b4b-4c97-91d0-417248f344a7 · outbound

This paper cites Lee, and Wen Sun.

Thompson Sampling in Online RLHF with General Function Approximation Lee, and Wen Sun

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.502747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:50.138027Z digest=sha256:568c45a7ae4f69913a232d1cdf1c03f45d38b75316381d5cd448adabfca8ccfc

Observation 0b8a432e-e368-4c56-b05f-05153a48c8e2 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Thompson Sampling in Online RLHF with General Function Approximation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.251391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.251391Z digest=sha256:85c82317b39b23e5d8bbb3f58e167096fbd49b8fb233eb00bfa1b490e635a6eb

Observation 03c377c5-e3ab-436d-8210-6ccd562be404 · outbound

This paper cites Fine-Tuning Language Models with Advantage-Induced Policy Alignment.

Thompson Sampling in Online RLHF with General Function Approximation Fine-Tuning Language Models with Advantage-Induced Policy Alignment

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.418201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.418201Z digest=sha256:600606768b04e466c313385b6f20e8cbc22129d0685d9201967cfb8e90034d3c

Observation 77e137d6-8164-40ce-abb3-10a5008f955d · outbound

This paper cites Efficient active learning with abstention.

Thompson Sampling in Online RLHF with General Function Approximation Efficient active learning with abstention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.301764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:50.546951Z digest=sha256:c82ef92718bd1f0a6d9fe53c36f8cda94339750b28fba5802950857230a4b383

Observation b14aa731-9812-46c9-a59b-4c5cdea52644 · outbound

This paper cites an unresolved cited work.

Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:52.029084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:43:50.639923Z digest=sha256:fcc44b8c2ac5d59ae241a80b0d9cbdbfbc9c3a5b2ff50d5ffa298db744c62b6d

Pith citing papers

No inbound Pith citation observations are available.