Pith. sign in

REVIEW 1 cited by

Gaussian Stochastic Weight Averaging for Bayesian Low-Rank Adaptation of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.03425 v2 pith:6J44XPZC submitted 2024-05-06 cs.CL

classification cs.CL
keywords bayesianlanguagellmsadaptationaveragingcalibrationfine-tunedgaussian
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Fine-tuned Large Language Models (LLMs) often suffer from overconfidence and poor calibration, particularly when fine-tuned on small datasets. To address these challenges, we propose a simple combination of Low-Rank Adaptation (LoRA) with Gaussian Stochastic Weight Averaging (SWAG), facilitating approximate Bayesian inference in LLMs. Through extensive testing across several Natural Language Processing (NLP) benchmarks, we demonstrate that our straightforward and computationally efficient approach improves model generalization and calibration competitively with comparable, more sophisticated methods for Bayesian inference in LLMs. We further show that our method exhibits greater robustness against distribution shift, as reflected in its improved performance on out-of-distribution tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Promoting Ensemble Diversity with Interactive Bayesian Distributional Robustness for Fine-tuning Foundation Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    IBDR couples Bayesian LoRA fine-tuning with a diversity-promoting divergence loss and Wasserstein distributional robustness, improving average ensemble accuracy on VTAB-1K and commonsense reasoning benchmarks.

Pith tools