Pith. sign in

REVIEW 3 cited by

Generalized Variational Inference: Three arguments for deriving new Posteriors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.02063 v4 pith:MSKBRUNU submitted 2019-04-03 stat.ML cs.AIcs.LG

classification stat.MLcs.AIcs.LG
keywords inferencebayesianposteriorsstandardthreevariationalargumentsincluding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We advocate an optimization-centric view on and introduce a novel generalization of Bayesian inference. Our inspiration is the representation of Bayes' rule as infinite-dimensional optimization problem (Csiszar, 1975; Donsker and Varadhan; 1975, Zellner; 1988). First, we use it to prove an optimality result of standard Variational Inference (VI): Under the proposed view, the standard Evidence Lower Bound (ELBO) maximizing VI posterior is preferable to alternative approximations of the Bayesian posterior. Next, we argue for generalizing standard Bayesian inference. The need for this arises in situations of severe misalignment between reality and three assumptions underlying standard Bayesian inference: (1) Well-specified priors, (2) well-specified likelihoods, (3) the availability of infinite computing power. Our generalization addresses these shortcomings with three arguments and is called the Rule of Three (RoT). We derive it axiomatically and recover existing posteriors as special cases, including the Bayesian posterior and its approximation by standard VI. In contrast, approximations based on alternative ELBO-like objectives violate the axioms. Finally, we study a special case of the RoT that we call Generalized Variational Inference (GVI). GVI posteriors are a large and tractable family of belief distributions specified by three arguments: A loss, a divergence and a variational family. GVI posteriors have appealing properties, including consistency and an interpretation as approximate ELBO. The last part of the paper explores some attractive applications of GVI in popular machine learning models, including robustness and more appropriate marginals. After deriving black box inference schemes for GVI posteriors, their predictive performance is investigated on Bayesian Neural Networks and Deep Gaussian Processes, where GVI can comprehensively improve upon existing methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Calibrating Wireless AI via Meta-Learned Context-Dependent Conformal Prediction

    eess.SP 2025-01 conditional novelty 6.0 of 10

    ML-WCP meta-learns a context-dependent likelihood ratio and uses it inside weighted conformal prediction to calibrate wireless AI with zero runtime data.

  2. Rethinking Likelihood distributions: Student's t Likelihood Boosts Bayesian Neural Network Performance

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Student's t likelihood (ν=5) is a robust default for VI-trained BNNs, improving CRPS in most tested settings while occasionally losing on MSE to Gaussian under lognormal noise.

  3. Uncertainty in Physics and AI: Taxonomy, Quantification, and Validation

    stat.ML 2026-05 conditional novelty 4.0 of 10

    A unified taxonomy of uncertainty in ML for physics is introduced together with validation tools such as coverage, calibration, and proper scoring rules, illustrated on regression and classification tasks.

Pith tools