Pith. sign in

REVIEW 1 cited by

Transformers are Minimax Optimal Nonparametric In-Context Learners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.12186 v2 pith:FFCSP4PZ submitted 2024-08-22 stat.ML cs.LG

classification stat.MLcs.LG
keywords in-contextlearningboundsemphexamplesgeneralizationminimaxnonparametric
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In-context learning (ICL) of large language models has proven to be a surprisingly effective method of learning a new task from only a few demonstrative examples. In this paper, we study the efficacy of ICL from the viewpoint of statistical learning theory. We develop approximation and generalization error bounds for a transformer composed of a deep neural network and one linear attention layer, pretrained on nonparametric regression tasks sampled from general function spaces including the Besov space and piecewise $\gamma$-smooth class. We show that sufficiently trained transformers can achieve -- and even improve upon -- the minimax optimal estimation risk in context by encoding the most relevant basis representations during pretraining. Our analysis extends to high-dimensional or sequential data and distinguishes the \emph{pretraining} and \emph{in-context} generalization gaps. Furthermore, we establish information-theoretic lower bounds for meta-learners w.r.t. both the number of tasks and in-context examples. These findings shed light on the roles of task diversity and representation learning for ICL.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Privacy in Decentralized Min-Max Optimization: A Differentially Private Approach

    cs.LG 2025-08 reject novelty 6.0 of 10

    DPMixSGD injects calibrated Gaussian noise into local gradient estimates to make decentralized nonconvex-strongly-concave min-max optimization differentially private, while claiming to preserve the STORM convergence rate.

Pith tools