Pith. sign in

Paper Citation Record · LEDGER

Thompson Sampling in Online RLHF with General Function Approximation

As of 16 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.23927.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23927 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:43:50.639923Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact3
  • verified fuzzy31
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4249897-9122-4b45-8413-d0da2bfbad51 · outbound

This paper cites Analysis of thompson sampling for the multi-armed bandit problem.

Thompson Sampling in Online RLHF with General Function Approximation Analysis of thompson sampling for the multi-armed bandit problem

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:57.353799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:47.472924Z digest=sha256:a1b3bae1ccb50db95f121ab20464a50428ace9929491388f9abaadbd9090d5b0

Observation 20d9d1ba-99c3-4248-8eb1-27b873063592 · outbound

This paper cites Thompson sampling for contextual bandits with linear payoffs.

Thompson Sampling in Online RLHF with General Function Approximation Thompson sampling for contextual bandits with linear payoffs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:57.169067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:47.503626Z digest=sha256:168faaa2f8b510cd91c0f2f256e54ae0fb160f0b03c402dcfbd36405a7f7f4a2

Observation 123a3047-6d7c-4587-b00f-bc2f1b16ed18 · outbound

This paper cites Near-optimal regret bounds for thompson sampling.J.

Thompson Sampling in Online RLHF with General Function Approximation Near-optimal regret bounds for thompson sampling.J

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:57.050843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:47.527002Z digest=sha256:3c4038a07a0483967ccd142054f1d06a6b0c50a0cac9250ef875dfca7bad42a9

Observation ffa5cb5e-80ae-48f0-a949-f154c1aeaa4c · outbound

This paper cites Preference-based online learning with dueling bandits: a survey.J.

Thompson Sampling in Online RLHF with General Function Approximation Preference-based online learning with dueling bandits: a survey.J

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.921778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:47.557153Z digest=sha256:615af3efe3e22ac3e1537dfb6d55b8627a8284a408ccd490e561ac69d7c34128

Observation 068d17e4-c7b3-451c-b836-a752a0dd49bd · outbound

This paper cites an unresolved cited work.

Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:56.839597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:47.588635Z digest=sha256:21674e5cbc6a3fdc10a72bc17232dc411df18eec268053134e8e0c5e9f43ff7e

Observation 1c0b73b2-aa5c-49b1-990e-9906ac6bb7ed · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Thompson Sampling in Online RLHF with General Function Approximation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.614974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.614974Z digest=sha256:01b4a7b7b22d4a9fe7046517ad3ddec89f98034b478e5718c31b038a21dd62a9

Observation bf60b172-58a5-4b1c-a562-17d5db2b54ab · outbound

This paper cites Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation.

Thompson Sampling in Online RLHF with General Function Approximation Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.660403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.660403Z digest=sha256:194325b6091b1cb3b6b867cad1807edfcf977f5db99e1695efd9afad0630123c

Observation df94777c-9a6c-42ee-a834-a528c133d4c5 · outbound

This paper cites On the Weaknesses of Reinforcement Learning for Neural Machine Translation.

Thompson Sampling in Online RLHF with General Function Approximation On the Weaknesses of Reinforcement Learning for Neural Machine Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.697268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.697268Z digest=sha256:b55a69e6deacf267034198bfdcdb34530121f9e314ea845c8ffd5642c5ef8e07

Observation ec97d09e-1422-4c8e-beb9-24e7242a2e90 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Thompson Sampling in Online RLHF with General Function Approximation Christiano, Jan Leike, Tom B

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.773058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:47.725692Z digest=sha256:2fe87662cce289617573e664251e153537e70abbcb03a7d13c393d5a0b1d6f98

Observation c3b6e898-2535-472e-bb8e-1f0b4efb17ce · outbound

This paper cites RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023.

Thompson Sampling in Online RLHF with General Function Approximation RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.691296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:47.755339Z digest=sha256:7b9c082e7f0793ec590f62556d5f5d41472af44ab23628f4e615e88518fb8ff5

Observation 9b05ccd0-1758-4e94-898e-5e4cf8905e2f · outbound

This paper cites Schapire, Aleksandrs Slivkins, and Masrour Zoghi.

Thompson Sampling in Online RLHF with General Function Approximation Schapire, Aleksandrs Slivkins, and Masrour Zoghi

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.598480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:47.799854Z digest=sha256:75e20e170095ddd0cbaa874ef1a6ab2875b41edae0c8c27fa5efd83ca0703779

Observation 19324a54-6076-42a7-9974-d420050cd73f · outbound

This paper cites Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO.

Thompson Sampling in Online RLHF with General Function Approximation Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.833667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.833667Z digest=sha256:5578f536d509147251bf6079ea8880e73369436415b53848eb1638d535e2c17c

Observation 3b0a6cb4-d055-45c1-963e-d178d68d1d07 · outbound

This paper cites Foster and Alexander Rakhlin.

Thompson Sampling in Online RLHF with General Function Approximation Foster and Alexander Rakhlin

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.515634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:47.872516Z digest=sha256:a85a8c68a340543621d6b977e9f414a1c7c422904ad24d3858fde16265a5cf67

Observation 7e17df10-3712-4376-ba25-69acb9567738 · outbound

This paper cites Scaling laws for reward model overoptimization.

Thompson Sampling in Online RLHF with General Function Approximation Scaling laws for reward model overoptimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.445967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:47.899602Z digest=sha256:f4fc105093b8501d8d4a7d9bf2b0fc9e3a57a56a5063179e72bc0d205c954355

Observation f93d914f-0c4c-473a-82e8-2c013645e3ae · outbound

This paper cites A General Theoretical Paradigm to Understand Learning from Human Preferences.

Thompson Sampling in Online RLHF with General Function Approximation A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.923648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.923648Z digest=sha256:9e415a1c6be0defee66ba158518e9078ac863d4e896c535426faaadda01babcd

Observation e83e8437-ccdd-42d5-b84d-f82d4a42da7e · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Thompson Sampling in Online RLHF with General Function Approximation Reinforced Self-Training (ReST) for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.949387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.949387Z digest=sha256:db4425357154bc783640623c5ddada5084f8610a5ada3e6a7caf23e281af1a1c

Observation 577cacf1-c0d0-4a3d-81be-5207f551a71d · outbound

This paper cites Randomized Exploration for Reinforcement Learning with General Value Function Approximation.

Thompson Sampling in Online RLHF with General Function Approximation Randomized Exploration for Reinforcement Learning with General Value Function Approximation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:43:51.698037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:47.982915Z digest=sha256:7deb01c763131d5f37c5b98fb64235b69b68d66890af6b52648ac6a0f59df3ff

Observation f297a784-efaf-4ffe-89e3-7dcd0598597c · outbound

This paper cites Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo.

Thompson Sampling in Online RLHF with General Function Approximation Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.010070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.010070Z digest=sha256:fb48881c5ef56b2b4ee0158acd57ef30f04e0f7db3651f59b530567a281f9ce0

Observation c42b36f4-36c4-48c0-91b7-2a61f39d6dd5 · outbound

This paper cites Learning trajectory preferences for manipulators via iterative improvement.

Thompson Sampling in Online RLHF with General Function Approximation Learning trajectory preferences for manipulators via iterative improvement

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.372101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.044855Z digest=sha256:fb770bb18223912d081fbe4d07966a66bc051022c39583ada6774dcd1b202d05

Observation 8fcfabfa-f0f1-4928-b963-d64cda7e98fd · outbound

This paper cites Schapire.

Thompson Sampling in Online RLHF with General Function Approximation Schapire

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.305869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.086054Z digest=sha256:a5ae87b4862b67be90e8199f047aa97e5cb8642b764d3e9e437e43fbb36c1154

Observation 8d7c2828-6114-47ee-9b1c-75a721be65d2 · outbound

This paper cites Bellman eluder dimension: new rich classes of rl problems, and sample-efficient algorithms.

Thompson Sampling in Online RLHF with General Function Approximation Bellman eluder dimension: new rich classes of rl problems, and sample-efficient algorithms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.160284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.220473Z digest=sha256:80123a38968efa928b98a6852d17bb8bfea4e0ae039e465ec6d597ff37bed3f1

Observation 0c2c094a-4035-4108-932e-73172d5f9e74 · outbound

This paper cites An Introduction to Variational Autoencoders.

Thompson Sampling in Online RLHF with General Function Approximation An Introduction to Variational Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.246384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.246384Z digest=sha256:97ebc2b577043a91d13ccc4cc2c8b3b11333341beb64619331ce055e3efe77d0

Observation 42031251-04a4-4d60-98f6-1d92937d10ec · outbound

This paper cites Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism.

Thompson Sampling in Online RLHF with General Function Approximation Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.277156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.277156Z digest=sha256:40187065fa320c8db801703df2f2e45c5d9dbcafb98b06e364d13f8e147d0e3c

Observation f0776b47-d595-408d-8199-0e20cbb831ad · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Thompson Sampling in Online RLHF with General Function Approximation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.349648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.349648Z digest=sha256:0fe3163202d8399fe04c74736dc324127526bd312f43252a2ce7008828c971d9

Observation 5e9a7c7f-6d8b-48bb-a13c-f3dd10c93cb9 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Thompson Sampling in Online RLHF with General Function Approximation Statistical Rejection Sampling Improves Preference Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.431444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.431444Z digest=sha256:ea29c8a16ad241b0f313ef99ea432ca84c8f6989046c59ac893c8d11ae44fd97

Observation 2d07115f-7ad9-4934-859e-b3cfd268c94f · outbound

This paper cites Roberts, Matthew E.

Thompson Sampling in Online RLHF with General Function Approximation Roberts, Matthew E

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.983442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.478015Z digest=sha256:a4edb7eee1676963474749674ef3aa28142f295cd84929554a7cfaa75f47f668

Observation 99e536a6-804c-4e8e-9d7d-6a8133fa2060 · outbound

This paper cites Dueling Posterior Sampling for Preference-Based Reinforcement Learning.

Thompson Sampling in Online RLHF with General Function Approximation Dueling Posterior Sampling for Preference-Based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.515394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.515394Z digest=sha256:4c45ccd386247b80c68ae244c9057027ef5e767c14fe76a61295aec271df8045

Observation d2dfc404-e7b3-4272-a6ce-8e481e070dfe · outbound

This paper cites GPT-4 Technical Report.

Thompson Sampling in Online RLHF with General Function Approximation GPT-4 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.557964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.557964Z digest=sha256:b4acaa66203d8bd9631bc083b4f8a6a9bc5239afb128f676ed8e3844c4729afb

Observation d119f7ce-7e96-44a5-8ae4-32de28a7a17b · outbound

This paper cites Randomized prior functions for deep reinforcement learning.

Thompson Sampling in Online RLHF with General Function Approximation Randomized prior functions for deep reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.724220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.595572Z digest=sha256:d613900d9ce92861b8c81b2c95fa92772c850b0a7233c3ed52ad104c58c84160

Observation 49fef49c-a677-43d7-aabc-9d0303095e41 · outbound

This paper cites Approximate thompson sampling via epistemic neural networks.

Thompson Sampling in Online RLHF with General Function Approximation Approximate thompson sampling via epistemic neural networks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.458395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.687662Z digest=sha256:1b2b6147430c96d9f4a5657960afe9eb0cfce938fc884a321acfac070422e2a4

Observation a4149ebe-e20f-44e2-98f4-d8e5ea133972 · outbound

This paper cites an unresolved cited work.

Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:55.265065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.737226Z digest=sha256:ad594c574a071dd6fc0a6330d74f2e9083040234ac0a733fe6a728685f8cd557

Observation 6f9f557f-4443-4c58-8032-dc79c85f99d1 · outbound

This paper cites Training language models to follow instructions with human feedback.

Thompson Sampling in Online RLHF with General Function Approximation Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.769424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.769424Z digest=sha256:5a8a3de268a13c01d67fa4a503f3711ea4317138fc0cdddc3af73ac0112f7b8d

Observation 861b40f8-ed0f-40de-9e17-2b78ec60c871 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Thompson Sampling in Online RLHF with General Function Approximation Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.796716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.796716Z digest=sha256:e03b181b110aa101af14b927d1229646525e26bcb73610999087ceb9bbacd718

Observation 211392fb-78f6-416f-88c2-2596ae83e015 · outbound

This paper cites Worst-Case Regret Bounds for Exploration via Randomized Value Functions.

Thompson Sampling in Online RLHF with General Function Approximation Worst-Case Regret Bounds for Exploration via Randomized Value Functions

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:43:51.321069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.814525Z digest=sha256:429274221d18ee22ad5fb2e7a43b415352365792047f3ed6411f66dc1c2918c5

Observation 55c69687-57ec-4809-8c77-83e041a8c371 · outbound

This paper cites Learning to optimize via posterior sampling.Math.

Thompson Sampling in Online RLHF with General Function Approximation Learning to optimize via posterior sampling.Math

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.197918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.838269Z digest=sha256:8f8076167a57c526e34e9f527e62e71d98e309de5b40626618f8333094554d4c

Observation bc4131b1-1e3a-4dc0-9db7-ff14eb6c686a · outbound

This paper cites Optimal algorithms for stochastic contextual preference bandits.

Thompson Sampling in Online RLHF with General Function Approximation Optimal algorithms for stochastic contextual preference bandits

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.096880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.861004Z digest=sha256:415c72f0b3d3ac9957d7cf5cd24c5e907b10854a4ef917548bdeab9b86124983

Observation d14f95a0-6435-44a0-ad5c-650065971df0 · outbound

This paper cites Efficient and optimal algorithms for contextual dueling bandits under realizability.

Thompson Sampling in Online RLHF with General Function Approximation Efficient and optimal algorithms for contextual dueling bandits under realizability

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:54.840829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.890800Z digest=sha256:0a5a7e1bc4bba5685bcbfcd6d361bd4531999d6104c944df4ea1c2156fce043c

Observation 96c7ef00-ab1d-4d6b-afc8-59b1df7108b3 · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Thompson Sampling in Online RLHF with General Function Approximation Dueling rl: Reinforcement learning with trajectory preferences

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:54.465692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.921797Z digest=sha256:4b2086db2a070a46e9999698e935060e3e94d40f5790f87b4438ac07803e672c

Observation 49191887-b0d3-43ff-95e7-cb122383c61b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Thompson Sampling in Online RLHF with General Function Approximation Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.940005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.940005Z digest=sha256:5702d4593d5f4e50cfff57b243bc6efca56fd67fd037c0311b66e12060b032bd

Observation 279c74ef-69af-4a5a-80c5-1c6e4b99451a · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano.

Thompson Sampling in Online RLHF with General Function Approximation Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:54.117813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.969661Z digest=sha256:3b099a1ce2039ae299e72f4c62bd1954f40e0cec22bc0b4fcab9920f0e0cb964

Observation fa5d1bcb-9c9a-48b3-8d23-0f6bb19a2314 · outbound

This paper cites Thompson.

Thompson Sampling in Online RLHF with General Function Approximation Thompson

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.933023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:48.992418Z digest=sha256:0db1d2cbb85b4b60b980df4516f7d2d098be1e0380c4f718364f8a76af4a730b

Observation d0e6e6ee-2a2d-4ea5-9d0c-675fddd8b452 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Thompson Sampling in Online RLHF with General Function Approximation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.018024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.018024Z digest=sha256:6119c9ce74264d544b603cd39ae0d50707389fa6b5e8eec4ba50e1427e30a377

Observation 5ade00be-bf41-4888-ae26-23ff03ca7444 · outbound

This paper cites van de Geer.

Thompson Sampling in Online RLHF with General Function Approximation van de Geer

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.773838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:49.052190Z digest=sha256:a1af7a7eb6f60cd0594ef4f26699f66d2febb2b98f78f34d8f90e088c3cb3897

Observation 76980b50-eb90-47f5-9a42-dc7bc440db34 · outbound

This paper cites Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints.

Thompson Sampling in Online RLHF with General Function Approximation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.083529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.083529Z digest=sha256:59e65dbad9a2cc4e18be7245d2f611a00d827101f44576b9a6470da7c801cddb

Observation 7f75d0cc-fa25-45bd-bef3-aa7b602e3ef7 · outbound

This paper cites Thompson sampling for combinatorial semi-bandits.

Thompson Sampling in Online RLHF with General Function Approximation Thompson sampling for combinatorial semi-bandits

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.720079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:49.108031Z digest=sha256:312bc357ce1b90ae96f63b03dfe093472d53655f1e58be39d321190b46af54e3

Observation 66e2e15c-29c5-4fef-b039-035424f36808 · outbound

This paper cites Is rlhf more difficult than standard rl? a theoretical perspective.

Thompson Sampling in Online RLHF with General Function Approximation Is rlhf more difficult than standard rl? a theoretical perspective

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.586179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:49.117791Z digest=sha256:f37730183a86e08430df3f19e6efad133b7c39f464927babe3aa5797a07877a6

Observation 32e371bb-2530-41ee-91aa-d4f60570624f · outbound

This paper cites A survey of preference-based reinforcement learning methods.Journal of Machine Learning Research, 18(136):1–46, 2017.

Thompson Sampling in Online RLHF with General Function Approximation A survey of preference-based reinforcement learning methods.Journal of Machine Learning Research, 18(136):1–46, 2017

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.466104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:49.197952Z digest=sha256:9a6af0527007a0cff0cbbea8d8bcf25f21779a3a6955ebb2ff7afb3c179e7733

Observation aaaeabf1-1a17-4283-b3f8-afbfa8d4821b · outbound

This paper cites Making RL with Preference-based Feedback Efficient via Randomization.

Thompson Sampling in Online RLHF with General Function Approximation Making RL with Preference-based Feedback Efficient via Randomization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.305402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.305402Z digest=sha256:f43b6716aa0f4951189acfa0db4d2c92ddab30a404d23a5d8c63b83bf572c6c5

Observation 3bff1bf3-7f85-42d6-9c5b-5fa3c5a8987b · outbound

This paper cites Borda regret minimization for generalized linear dueling bandits.

Thompson Sampling in Online RLHF with General Function Approximation Borda regret minimization for generalized linear dueling bandits

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.309730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:49.376621Z digest=sha256:a272d7cc180cce1b84651a44f6a14445ab44018a9a202be45eb9b5b56b814e24

Observation d633d1f9-2b01-4aa8-929c-e316d7ca4f71 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Thompson Sampling in Online RLHF with General Function Approximation Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.435213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.435213Z digest=sha256:f12035fce3eb47f6f307ad5050b914eb29391c4a3f54e76a67dcc8d720fc8636

Observation 50e372fb-8206-41ca-be88-96913864589f · outbound

This paper cites Near-optimal randomized exploration for tabular markov decision processes.

Thompson Sampling in Online RLHF with General Function Approximation Near-optimal randomized exploration for tabular markov decision processes

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.149868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:49.519282Z digest=sha256:22ba952a5fdd1ce85561a6567b381d67b5f12ce56ac48a2a1439f41c09cc73ba

Observation 358a7f11-bc21-4f36-ae13-e03d2f64e846 · outbound

This paper cites Yang, Aarti Singh, and Artur Dubrawski.

Thompson Sampling in Online RLHF with General Function Approximation Yang, Aarti Singh, and Artur Dubrawski

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.943796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:49.630519Z digest=sha256:97105e20609ffa5a72333c98ab8e21f2482e31475be6786f2527f9dc5c7a081b

Observation 2224a913-2e3b-4ce0-89b8-bb0c530ca86a · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

Thompson Sampling in Online RLHF with General Function Approximation RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.730231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.730231Z digest=sha256:bd2c6a94ab001f45b613eeca79ae615f69d33978009bdb66945394c3431e1c01

Observation efdf6433-3dac-4f16-8619-95bc291194bb · outbound

This paper cites The k-armed dueling bandits problem.Journal of Computer and System Sciences, 78(5):1538–1556, 2012.

Thompson Sampling in Online RLHF with General Function Approximation The k-armed dueling bandits problem.Journal of Computer and System Sciences, 78(5):1538–1556, 2012

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.725910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:49.832227Z digest=sha256:6ebd123aaa8dd8596df1970ff2da9cf5dd89b07847c57bbd7bca70a485525333

Observation 74185e95-b077-4d1a-8ed4-77e557bb356e · outbound

This paper cites Frequentist Regret Bounds for Randomized Least-Squares Value Iteration.

Thompson Sampling in Online RLHF with General Function Approximation Frequentist Regret Bounds for Randomized Least-Squares Value Iteration

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:43:50.963293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:49.926426Z digest=sha256:19aea1e45e4092b8280d54ca5edc283361a57284ec6601bad2a2ec4dee1a8382

Observation be121500-6594-443b-9b65-15b395ac386f · outbound

This paper cites Provable Offline Preference-Based Reinforcement Learning.

Thompson Sampling in Online RLHF with General Function Approximation Provable Offline Preference-Based Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.030374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.030374Z digest=sha256:e095d90588fbffb910d1b5de7aada49f3c17ee96afc048dd2ef148fd4570f531

Observation 52573980-9b4b-4c97-91d0-417248f344a7 · outbound

This paper cites Lee, and Wen Sun.

Thompson Sampling in Online RLHF with General Function Approximation Lee, and Wen Sun

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.502747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:50.138027Z digest=sha256:f9cd21564fda77e03b0015704490bdac371404d39fe78d8bad15fea8cd3ff0e0

Observation 0b8a432e-e368-4c56-b05f-05153a48c8e2 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Thompson Sampling in Online RLHF with General Function Approximation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.251391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.251391Z digest=sha256:29c4a6f2ab636a9396289e900e1f78ee9b928e6a6eeeb87b3c94150c23e50751

Observation 03c377c5-e3ab-436d-8210-6ccd562be404 · outbound

This paper cites Fine-Tuning Language Models with Advantage-Induced Policy Alignment.

Thompson Sampling in Online RLHF with General Function Approximation Fine-Tuning Language Models with Advantage-Induced Policy Alignment

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.418201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.418201Z digest=sha256:7c563eb098a475b0cfe19fca7626175607c1f2d8812a65b55c3f4709befe3b9d

Observation 77e137d6-8164-40ce-abb3-10a5008f955d · outbound

This paper cites Efficient active learning with abstention.

Thompson Sampling in Online RLHF with General Function Approximation Efficient active learning with abstention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.301764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:50.546951Z digest=sha256:405b2e5288dfbd1fde86105d97b0d763cf3509ef1a0b6ca2b0fa364693f3413b

Observation b14aa731-9812-46c9-a59b-4c5cdea52644 · outbound

This paper cites an unresolved cited work.

Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:52.029084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:43:50.639923Z digest=sha256:a13546a6fe461b65c4afeb14962d9b8dd2a9075c7a2e7101610f99eb7fc28cc1

Pith citing papers

No inbound Pith citation observations are available.