Pith. sign in

REVIEW 5 cited by

Foundations of Reinforcement Learning and Interactive Decision Making

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.16730 v1 pith:5END7BQI submitted 2023-12-27 cs.LG math.OCmath.STstat.MLstat.TH

classification cs.LGmath.OCmath.STstat.MLstat.TH
keywords learningdecisionmakingreinforcementbanditsfoundationsinteractiveaddressing
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

These lecture notes give a statistical perspective on the foundations of reinforcement learning and interactive decision making. We present a unifying framework for addressing the exploration-exploitation dilemma using frequentist and Bayesian approaches, with connections and parallels between supervised learning/estimation and decision making as an overarching theme. Special attention is paid to function approximation and flexible model classes such as neural networks. Topics covered include multi-armed and contextual bandits, structured bandits, and reinforcement learning with high-dimensional feedback.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimizing the Preconditioner: A Black-box Online-to-Nonconvex Conversion with Static Regret Minimization Oracles

    cs.LG 2026-07 conditional novelty 7.0 of 10

    An OCO algorithm with only O(√T) static regret, pluggable as a preconditioner selector, recovers the classical O(1/√T) stationarity rate on smooth stochastic nonconvex problems and the O(T^{-2/7}) rate on nonsmooth ones.

  2. A Jointly Efficient and Optimal Algorithm for Heteroskedastic Generalized Linear Bandits with Adversarial Corruptions

    cs.LG 2026-02 conditional novelty 7.0 of 10

    A per-round O(1) algorithm for generalized linear bandits achieves near-optimal regret with time-varying dispersion and adversarial corruptions, up to a κ factor.

  3. Self-Improvement in Language Models: The Sharpening Mechanism

    cs.AI 2024-12 conditional novelty 7.0 of 10

    Self-improvement in language models can be understood as amortizing best-of-N inference-time selection, with minimax-optimal guarantees for SFT and provable coverage-free benefits for RL with exploration.

  4. Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A hybrid online-plus-offline preference optimization algorithm, HPO, provably needs fewer samples than pure online or offline RLHF in linear MDP settings.

  5. A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

    cs.AI 2024-11 conditional novelty 2.0 of 10

    A comprehensive but flawed survey of RL algorithms that catalogs many methods and applications without rigorous comparative analysis.

Pith tools