Pith. sign in

REVIEW 1 cited by

Pathologies in priors and inference for Bayesian transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.04020 v2 pith:INC3MH7A submitted 2021-10-08 cs.LG stat.ML

classification cs.LGstat.ML
keywords bayesianfindinferencetransformertransformersapplicationsfunction-spacelearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, the transformer has established itself as a workhorse in many applications ranging from natural language processing to reinforcement learning. Similarly, Bayesian deep learning has become the gold-standard for uncertainty estimation in safety-critical applications, where robustness and calibration are crucial. Surprisingly, no successful attempts to improve transformer models in terms of predictive uncertainty using Bayesian inference exist. In this work, we study this curiously underpopulated area of Bayesian transformers. We find that weight-space inference in transformers does not work well, regardless of the approximate posterior. We also find that the prior is at least partially at fault, but that it is very hard to find well-specified weight priors for these models. We hypothesize that these problems stem from the complexity of obtaining a meaningful mapping from weight-space to function-space distributions in the transformer. Therefore, moving closer to function-space, we propose a novel method based on the implicit reparameterization of the Dirichlet distribution to apply variational inference directly to the attention weights. We find that this proposed method performs competitively with our baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Singular Bayesian Neural Networks

    stat.ML 2026-01 unverdicted novelty 7.0 of 10

    Low-rank weight factorization creates singular posteriors in Bayesian neural networks that scale as sqrt(r(m+n)) in complexity and use up to 33x fewer parameters than ensembles.

Pith tools