Pith. sign in

REVIEW 1 cited by

How Infinitely Wide Neural Networks Can Benefit from Multi-task Learning -- an Exact Macroscopic Characterization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.15577 v4 pith:VJTSK4VD submitted 2021-12-31 cs.LG stat.ML

classification cs.LGstat.ML
keywords learningmulti-taskneuralinfinite-widthnetworkswidebenefitcharacterization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In practice, multi-task learning (through learning features shared among tasks) is an essential property of deep neural networks (NNs). While infinite-width limits of NNs can provide good intuition for their generalization behavior, the well-known infinite-width limits of NNs in the literature (e.g., neural tangent kernels) assume specific settings in which wide ReLU-NNs behave like shallow Gaussian Processes with a fixed kernel. Consequently, in such settings, these NNs lose their ability to benefit from multi-task learning in the infinite-width limit. In contrast, we prove that optimizing wide ReLU neural networks with at least one hidden layer using L2-regularization on the parameters promotes multi-task learning due to representation-learning - also in the limiting regime where the network width tends to infinity. We present an exact quantitative characterization of this infinite width limit in an appropriate function space that neatly describes multi-task learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Nonparametric Filtering, Estimation and Classification using Neural Jump ODEs

    stat.ML 2024-12 conditional novelty 6.0 of 10

    An input-output variant of Neural Jump ODEs is proven to converge to the L2-optimal conditional expectation for online filtering and classification with irregularly sampled, partially observed data.

Pith tools