Pith. sign in

REVIEW 1 cited by

Hypermodels for Exploration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.07464 v1 pith:WXIYULFZ submitted 2020-06-12 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords hypermodelssamplingensemblesexplorationthompsonapproximateconsiderelements
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the use of hypermodels to represent epistemic uncertainty and guide exploration. This generalizes and extends the use of ensembles to approximate Thompson sampling. The computational cost of training an ensemble grows with its size, and as such, prior work has typically been limited to ensembles with tens of elements. We show that alternative hypermodels can enjoy dramatic efficiency gains, enabling behavior that would otherwise require hundreds or thousands of elements, and even succeed in situations where ensemble methods fail to learn regardless of size. This allows more accurate approximation of Thompson sampling as well as use of more sophisticated exploration schemes. In particular, we consider an approximate form of information-directed sampling and demonstrate performance gains relative to Thompson sampling. As alternatives to ensembles, we consider linear and neural network hypermodels, also known as hypernetworks. We prove that, with neural network base models, a linear hypermodel can represent essentially any distribution over functions, and as such, hypernetworks are no more expressive.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration

    cs.LG 2025-01 reject novelty 5.0 of 10

    Concurrent RLSVI with aggregated states is shown to have worst-case regret O~(K H^(5/2) Γ √N) and per-agent regret 1/√N, with an analogous infinite-horizon bound.

Pith tools