REVIEW 2 major objections 5 minor 56 references
Meta-learning of shared linear representations beyond well-specified linear regression
T0 review · 2 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read The paper establishes that shared low-dimensional linear structure across tasks can be recovered for any convex objective satisfying Hessian and noise concentration, with sample complexity matching the lower bound known from linear…
desk verdict The general-convex framework and Theorems 1-2 are solid and worth a referee, but the one-sample-per-task Theorem 3 has a backwards monotonicity inequality in the proof and is not established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument's load-bearing object is the pair of concentration assumptions: Assumption 3 says that gradients of each loss at its optimum are $\sigma_*^2$-subexponential along every direction, and Assumption 4 says the same for every quadratic form $\langle w_1,\nabla^2F_i(w,\xi)w_2\rangle$ of the Hessian with scale $\sigma^2$. These assumptions make the excess loss decomposable through Bregman divergences: the empirical risk's regret is bounded by a noise term controlled by Assumption 3 and a Hessian-concentration term controlled by Assumption 4, with uniform bounds obtained by covering the low-rank constraint set by $\varepsilon$-nets of metric entropy $O(r(d+n)\ln(1/\varepsilon))$. For the one-sample-per-task regime the machinery changes to admissibility: an orthonormal matrix $U$ is admissible if every single sample equation $y_i=\langle Uv_i,x_i\rangle$ can be fitted with bounded coefficients; the paper bounds the probability that a bad $U$ is admissible, which yields the exponential dependence on $r$.
What would settle it
Take a logistic-regression or other GLM task with heavy-tailed features and measure, over a fine $\varepsilon$-net of directions $w,w_1,w_2$, the subexponential norm of $\langle w_1,\nabla^2F_i(w,\xi)w_2\rangle$: if that norm grows with the net resolution, Assumption 4 fails and the paper's rates should not hold, so the empirical excess risk versus total samples $nm$ should show a worse dependence than $\sqrt{r(d+n)/(mn)}$.
Extended reading notes
Core claim
The paper's central claim is that shared low-rank or clustered structure among task optimizers can be recovered in general convex stochastic optimization, provided the noise at each optimum and the Hessians of the individual losses concentrate as subexponential random variables. The main generalization bound states that with high probability the rank-constrained estimator satisfies $f(\widehat{W}_{\mathrm{lowrank}})-f(W^*)=\widetilde{O}(B^2(\sigma^2+\sigma_*^2)\sqrt{r(d+n)/(mn)})$, and, under strong convexity, $\frac1n\|\widehat{W}_{\mathrm{lowrank}}-W^*\|_F^2=\widetilde{O}(B^2(\sigma^4+\sigma_*^4)r(d+n)/(\mu^2 mn))$. This reaches the same $nm\sim rd$ sample complexity that was previously known only for linear regression. The paper further claims that with one sample per task, the subspace is still identifiable by the rank-constrained estimator, but only once the number of tasks grows exponentially in the subspace dimension $r$; below that threshold, bad minimizers orthogonal to the true subspace exist. Finally, a nuclear-norm relaxation is claimed to be polynomial-time computable at the cost of requiring $m\gtrsim\sqrt{rd}$ samples per task, the geometric mean of the no-collaboration and ideal collaborative rates.
Load-bearing premise
The whole sample-complexity story rests on Assumption 4: for every task and every pair of unit directions, the quadratic form of the empirical Hessian must be subexponential with a common scale; if Hessians concentrate more slowly, say because losses have unbounded second derivatives or features are heavy-tailed, the stated rates are not guaranteed.
Editorial extensions
If this is right
- The rank-constrained estimator reaches optimal total sample complexity $nm\sim rd$ for general convex objectives, so collaborative gains once reserved for linear regression transfer to GLMs and classification.
- In the strongly convex case the same estimator recovers each task's optimizer accurately enough to reconstruct the shared subspace, and a new task then learns with only $m_1\gtrsim r$ samples once the representation is known.
- The clustered estimator needs only $m\gtrsim 1$ samples per task, bypassing the $m\gtrsim r$ constraint of the low-rank estimator, at the price of a $\ln r / m$ term in the rate.
- With one sample per task, subspace recovery remains possible but requires $n$ exponential in $r$; the paper also gives a matching failure regime where bad minimizers orthogonal to the truth appear.
- The nuclear-norm relaxation is convex and polynomial-time, and its sample requirement $m\gtrsim\sqrt{rd}$ is the geometric mean of the no-collaboration rate $m\gtrsim d$ and the idealized rate $m\gtrsim r$, quantifying the cost of convex relaxation.
Reading between the lines
- Editorially, the same proof template should extend to non-parametric feature learning when the shared object is a function class rather than a matrix, though the covering numbers and concentration arguments would need reworking for function spaces.
- The exponential-in-$r$ task count for $m=1$ suggests a fundamental price for dropping the linear-regression likelihood: moment-based estimators can exploit second-order statistics, whereas the rank-constrained estimator sees only the raw convex losses, so a general-convex method may need exponentially many tasks.
- A testable consequence is that heavy-tailed features or losses with unbounded second derivatives should degrade the rates predictably: the effective scale in the bounds becomes the subexponential norm of the Hessian quadratic forms, so the degradation is governed by that constant rather than by the ambient dimension alone.
- The $\sqrt{rd}$ sample cost of the nuclear-norm relaxation indicates a statistical-computational gap for convex surrogates; whether a polynomial-time algorithm can close the gap to the rank-constrained $r$ rate is a question the paper leaves open.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies meta-learning of shared linear representations under general convex objectives, going beyond the well-specified linear regression setting that dominates prior work. The authors introduce two structural assumptions: noise concentration at the optimum (Assumption 3) and a uniform subexponential condition on Hessian quadratic forms (Assumption 4). Under these assumptions, they analyze rank-constrained estimators (Theorem 1), clustered estimators (Theorem 2), and a nuclear-norm-regularized estimator (Theorem 4), proving excess-risk and parameter-recovery rates. The paper's main novelty is Section 3.3, which treats the m=1 setting (one sample per task) and claims in Theorem 3 that the rank-constrained estimator recovers the subspace U* provided the number of tasks scales exponentially in the subspace dimension r. The proofs rely on an admissibility argument over an epsilon-net of the Grassmannian, formalized in Proposition 3, together with a lower bound in Proposition 2.
Significance. If the results hold, this would be a meaningful extension of representation-learning theory from quadratic regression to general convex losses, with sample complexity nm ~ rd for the rank-constrained estimator and a polynomial-time nuclear-norm relaxation with sample complexity interpolating between the optimal and no-collaboration rates. The detailed proofs with explicit constants and high-probability bounds are a strength. However, the m=1 subspace-recovery theorem is a highlighted contribution, and the section containing it is not rigorously established: the key admissibility-probability estimate in Proposition 3 uses an incorrect monotonicity direction, and the final failure probability in Theorem 3 does not follow from the union bound as written.
major comments (2)
- [Appendix B, Proposition 3] The displayed inequality P(U is δ-admissible) ≤ (1−e^{−C1 r/(λε)})^{λε n/2} follows from the preceding sentence by the wrong monotonicity direction. The function f(α)=1−e^{−cr/α} is decreasing in α; hence the established bound α_i ≤ λε/4 on at least λε n/2 indices actually gives f(α_i) ≥ f(λε/4)=1−e^{−4cr/(λε)}, which is a lower bound on the product, not the advertised upper bound. Obtaining the claimed upper bound would require α_i ≥ λε/4, the opposite of what the averaging argument supplies. Since this exponential-tail estimate is the only quantitative control of admissibility used in the net argument for Theorem 3, the proof of Theorem 3 is not valid as written.
- [Appendix B, proof of Theorem 3] The union bound over the η-net of the Grassmannian gives failure probability at most exp(crd ln(1/η) − c1 n e^{−c2 r/(λε)}). With the stated sample condition n ≥ (e^{c2 r/(λε)}/c1) rd(1+c ln(1/δ)), this expression is at most exp(−rd(1+c ln(1/δ)) + crd ln(1/η)), i.e., at most e^{−Ω(rd)} after choosing η, not e^{−rdn}. Reaching e^{−rdn} would require c1 e^{−c2 r/(λε)} ≥ rd, which fails for all sufficiently large r. Thus the failure probability claimed in Theorem 3 is not a consequence of the submitted argument; this issue is independent of the monotonicity problem in Proposition 3.
minor comments (5)
- [Appendix B] There are two consecutive subsections both titled 'Proof of Proposition 2' that appear to be duplicates; the redundant copy should be removed.
- [Theorem 2] The statement says 'Assume that Assumption 2 holds (underlying low rank assumption)', but Assumption 2 is the clustered-clients assumption; please correct the parenthetical.
- [Section 2.1, Example 2] Assumption 4 is said to hold for GLMs only when the second derivative ℓ2 of the loss is bounded; this restriction should be stated explicitly in the main text, since it excludes common settings such as unbounded losses or heavy-tailed features and tempers the claimed degree of generality beyond linear regression.
- [Lemma 5] The stated probability '1 − 2e^{r(d+n)} − 2e^{nmd}' appears to have missing minus signs in the exponents; compare with the theorem statements, which use '1 − 4e^{−r(d+n)} − 4e^{−nmd}'. Please check the typesetting.
- [Section 2] There are unresolved placeholder references '??' for the definition of the principal angle distance and in Lemma 1; these should be filled in.
Circularity Check
No circularity: the generalization and subspace-recovery guarantees are derived from explicit distributional assumptions and concentration lemmas, not from the conclusions they target.
full rationale
The paper's derivation chain is self-contained rather than circular. Theorem 1 follows from Lemma 2 (a Bregman-divergence decomposition) plus Lemmas 5 and 6, which are uniform concentration bounds proved from Assumptions 3 and 4 using subexponential tail bounds and metric entropy arguments; the target excess-risk and parameter-error quantities do not appear as assumptions. Assumptions 3 and 4 are stated as problem-dependent conditions on the data distributions, not as fitted parameters or as renamed versions of the conclusions. Theorem 3's admissibility argument computes probabilities for fixed U from Gaussian chi-square tails and then takes a union bound over a net; it does not assume that the estimated subspace is close to U*. The lower bound of Tripuraneni et al. is used only as an external benchmark to claim optimality, not to derive the upper bound. The self-citation to Even and Massoulie (2021) appears only in a parenthetical remark about replacing d with an effective dimension and is not load-bearing for any theorem. Whatever the merits of the Proposition 3 monotonicity step flagged by the skeptic, that is a mathematical correctness concern, not a circularity: it does not involve the paper assuming its conclusion or fitting a parameter and renaming it a prediction. No step was found in which a quoted equation reduces to its own input by construction.
Assumptions & free parameters
assumptions (6)
- domain assumption Assumption 1: there exist minimizers w_i* of each f_i such that the matrix W* = (w_1*|...|w_n*) has rank at most r.
- domain assumption Assumption 3: for each i and each unit vector w, the random variable <grad F_i(w_i*, xi_i), w> is sigma_*^2-subexponential.
- domain assumption Assumption 4: for all unit vectors w, w1, w2 and all i, the quadratic form <w1, grad^2 F_i(w, xi_i) w2> is sigma^2-subexponential.
- domain assumption Section 3.3 restricts to noiseless Gaussian linear regression with equal-norm heads and the spanning condition 1/n sum_i w_i* w_i*^T >= lambda B^2/r U* U*^T.
- standard math Hanson-Wright inequality and subexponential tail bounds (cited as Hanson and Wright 1971, Rudelson and Vershynin 2013).
- standard math Metric entropy bound for rank-r matrices with bounded columns, O(r(d+n) ln(B/epsilon)) (cited from Candes and Plan).
Cite this review
Pith. "Pith review of Meta-learning of shared linear representations beyond well-specified linear regression." pith.science (2026). https://pith.science/paper/OKJOF5ZP
@misc{pith2026250118975,
author = {Pith},
title = {Pith review of: Meta-learning of shared linear representations beyond well-specified linear regression},
year = {2026},
howpublished = {\url{https://pith.science/paper/OKJOF5ZP}},
note = {Machine review of arXiv:2501.18975}
}
read the original abstract
Motivated by multi-task and meta-learning approaches, we consider the problem of learning structure shared by tasks or users, such as shared low-rank representations or clustered structures. While all previous works focus on well-specified linear regression, we consider more general convex objectives, where the structural low-rank and cluster assumptions are expressed on the optima of each function. We show that under mild assumptions such as \textit{Hessian concentration} and \textit{noise concentration at the optimum}, rank and clustered regularized estimators recover such structure, provided the number of samples per task and the number of tasks are large enough. We then study the problem of recovering the subspace in which all the solutions lie, in the setting where there is only a single sample per task: we show that in that case, the rank-constrained estimator can recover the subspace, but that the number of tasks needs to scale exponentially large with the dimension of the subspace. Finally, we provide a polynomial-time algorithm via nuclear norm constraints for learning a shared linear representation in the context of convex learning objectives.
Reference graph
Works this paper leans on
-
[1]
Alekh Agarwal, Sahand Negahban, and Martin J. Wainwright. Fast global convergence of gradient methods for high-dimensional statistical recovery . The Annals of Statistics, 40 0 (5): 0 2452 -- 2482, 2012. doi:10.1214/12-AOS1032. URL https://doi.org/10.1214/12-AOS1032
-
[2]
Andreas Argyriou, Theodoros Evgeniou, and Massimiliano Pontil. Multi-task feature learning. In B. Sch\" o lkopf, J. Platt, and T. Hoffman, editors, Advances in Neural Information Processing Systems, volume 19. MIT Press, 2006. URL https://proceedings.neurips.cc/paper_files/paper/2006/file/0afa92fc0f8a9cf051bf2961b06ac56b-Paper.pdf
work page 2006
-
[3]
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives. IEEE Trans. Pattern Anal. Mach. Intell., 35 0 (8): 0 1798–1828, aug 2013. ISSN 0162-8828. doi:10.1109/TPAMI.2013.50. URL https://doi.org/10.1109/TPAMI.2013.50
-
[4]
Interpretable Meta-Learning of Physical Systems
Matthieu Blanke and Marc Lelarge. Interpretable Meta-Learning of Physical Systems . In ICLR 2024 - The Twelfth International Conference on Learning Representations , Vienne, Austria, May 2024. URL https://hal.science/hal-04513216
work page 2024
-
[5]
Thomas Blumensath and Mike E. Davies. Iterative hard thresholding for compressed sensing. Applied and Computational Harmonic Analysis, 27 0 (3): 0 265--274, 2009. ISSN 1063-5203. doi:https://doi.org/10.1016/j.acha.2009.04.002. URL https://www.sciencedirect.com/science/article/pii/S1063520309000384
-
[6]
Trace norm regularization for multi-task learning with scarce data
Etienne Boursier, Mikhail Konobeev, and Nicolas Flammarion. Trace norm regularization for multi-task learning with scarce data. In Po-Ling Loh and Maxim Raginsky, editors, Proceedings of Thirty Fifth Conference on Learning Theory, volume 178 of Proceedings of Machine Learning Research, pages 1303--1327. PMLR, 02--05 Jul 2022. URL https://proceedings.mlr.p...
work page 2022
-
[7]
Convex optimization: Algorithms and complexity
S\' e bastien Bubeck. Convex optimization: Algorithms and complexity. Foundations and Trends in Machine Learning, 8 0 (3-4): 0 231--357, 2015
work page 2015
-
[8]
Generalize across tasks: Efficient algorithms for linear representation learning
Brian Bullins, Elad Hazan, Adam Kalai, and Roi Livni. Generalize across tasks: Efficient algorithms for linear representation learning. In Aurélien Garivier and Satyen Kale, editors, Proceedings of the 30th International Conference on Algorithmic Learning Theory, volume 98 of Proceedings of Machine Learning Research, pages 235--246. PMLR, 22--24 Mar 2019....
work page 2019
Show all 56 references
-
[9]
Monteiro
Samuel Burer and Renato D.C. Monteiro. A nonlinear programming algorithm for solving semidefinite programs via low-rank factorization. Mathematical Programming, 95 0 (2): 0 329–357, February 2003. ISSN 1436-4646. doi:10.1007/s10107-002-0352-8. URL http://dx.doi.org/10.1007/s10...
2003 doi
-
[10]
Monteiro
Samuel Burer and Renato D.C. Monteiro. Local minima and convergence in low-rank semidefinite programming. Mathematical Programming, 103 0 (3): 0 427–444, December 2004. ISSN 1436-4646. doi:10.1007/s10107-004-0564-1. URL http://dx.doi.org/10.1007/s10107-004-0564-1
2004 doi
-
[11]
Candès and Yaniv Plan
Emmanuel J. Candès and Yaniv Plan. Tight oracle inequalities for low-rank matrix recovery from a minimal number of noisy random measurements. IEEE Transactions on Information Theory, 57 0 (4): 0 2342--2359, 2011. doi:10.1109/TIT.2011.2111771
2011
-
[12]
Multi-task learning
Rich Caruana. Multi-task learning. Machine Learning, 28 0 (1): 0 41–75, 1997. ISSN 0885-6125. doi:10.1023/a:1007379606734. URL http://dx.doi.org/10.1023/A:1007379606734
1997 doi
-
[13]
Multitask online mirror descent
Nicol \`o Cesa-Bianchi, Pierre Laforgue, Andrea Paudice, and massimiliano pontil. Multitask online mirror descent. Transactions on Machine Learning Research, 2022. ISSN 2835-8856. URL https://openreview.net/forum?id=zwRX9kkKzj
2022
-
[14]
Convex learning of multiple tasks and their structure
Carlo Ciliberto, Youssef Mroueh, Tomaso Poggio, and Lorenzo Rosasco. Convex learning of multiple tasks and their structure. In Francis Bach and David Blei, editors, Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learni...
2015
-
[15]
Exploiting shared representations for personalized federated learning
Liam Collins, Hamed Hassani, Aryan Mokhtari, and Sanjay Shakkottai. Exploiting shared representations for personalized federated learning. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings o...
2021
-
[16]
Fedavg with fine tuning: Local updates lead to representation learning
Liam Collins, Hamed Hassani, Aryan Mokhtari, and Sanjay Shakkottai. Fedavg with fine tuning: Local updates lead to representation learning. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022 a ....
2022
-
[17]
MAML and ANIL provably learn representations
Liam Collins, Aryan Mokhtari, Sewoong Oh, and Sanjay Shakkottai. MAML and ANIL provably learn representations. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, editors, Proceedings of the 39th International Conference on Machine L...
-
[18]
Provable multi-task representation learning by two-layer relu neural networks, 2023
Liam Collins, Hamed Hassani, Mahdi Soltanolkotabi, Aryan Mokhtari, and Sanjay Shakkottai. Provable multi-task representation learning by two-layer relu neural networks, 2023
2023
-
[19]
Learning-to-learn stochastic gradient descent with biased regularization
Giulia Denevi, Carlo Ciliberto, Riccardo Grazzi, and Massimiliano Pontil. Learning-to-learn stochastic gradient descent with biased regularization. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, vo...
2019
-
[20]
Online-within-online meta-learning
Giulia Denevi, Dimitris Stamos, Carlo Ciliberto, and Massimiliano Pontil. Online-within-online meta-learning. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alch\' e -Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 32. Curran...
2019
-
[21]
Distributed personalized empirical risk minimization, 2023
Yuyang Deng, Mohammad Mahdi Kamani, Pouria Mahdavinia, and Mehrdad Mahdavi. Distributed personalized empirical risk minimization, 2023
2023
-
[22]
Collaborative learning by detecting collaboration partners
Shu Ding and Wei Wang. Collaborative learning by detecting collaboration partners. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=fkiFqG-muu
2022
-
[23]
Du, Wei Hu, Sham M
Simon S. Du, Wei Hu, Sham M. Kakade, Jason D. Lee, and Qi Lei. Few-shot learning via learning the representation, provably, 2021
2021
-
[24]
Efficient projections onto the l1-ball for learning in high dimensions
John Duchi, Shai Shalev-Shwartz, Yoram Singer, and Tushar Chandra. Efficient projections onto the l1-ball for learning in high dimensions. In Proceedings of the 25th International Conference on Machine Learning, ICML '08, page 272–279, New York, NY, USA, 2008. Association for ...
2008
-
[25]
Subspace recovery from heterogeneous data with non-isotropic noise
John Duchi, Vitaly Feldman, Lunjia Hu, and Kunal Talwar. Subspace recovery from heterogeneous data with non-isotropic noise. Advances in Neural Information Processing Systems, 2022
2022
-
[26]
Metric entropy of some classes of sets with differentiable boundaries
R.M Dudley. Metric entropy of some classes of sets with differentiable boundaries. Journal of Approximation Theory, 10 0 (3): 0 227--236, 1974. ISSN 0021-9045. doi:https://doi.org/10.1016/0021-9045(74)90120-8. URL https://www.sciencedirect.com/science/article/pii/0021904574901208
1974
-
[27]
Concentration of non-isotropic random tensors with applications to learning and empirical risk minimization
Mathieu Even and Laurent Massoulie. Concentration of non-isotropic random tensors with applications to learning and empirical risk minimization. In Mikhail Belkin and Samory Kpotufe, editors, Proceedings of Thirty Fourth Conference on Learning Theory, volume 134 of Proceedings...
2021
-
[28]
On sample optimality in personalized collaborative and federated learning
Mathieu Even, Laurent Massouli \'e , and Kevin Scaman. On sample optimality in personalized collaborative and federated learning. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022. URL https://...
2022
-
[29]
Nonparametric linear feature learning in regression through regularisation, 2024
Bertille Follain and Francis Bach. Nonparametric linear feature learning in regression through regularisation, 2024. URL https://arxiv.org/abs/2307.12754
2024 arXiv
-
[30]
Iterative hard thresholding for low-rank recovery from rank-one projections, 2018
Simon Foucart and Srinivas Subramanian. Iterative hard thresholding for low-rank recovery from rank-one projections, 2018
2018
-
[31]
An efficient framework for clustered federated learning
Avishek Ghosh, Jichan Chung, Dong Yin, and Kannan Ramchandran. An efficient framework for clustered federated learning. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS'20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9...
2020
-
[32]
D. L. Hanson and F. T. Wright. A Bound on Tail Probabilities for Quadratic Forms in Independent Random Variables . The Annals of Mathematical Statistics, 42 0 (3): 0 1079 -- 1083, 1971. doi:10.1214/aoms/1177693335. URL https://doi.org/10.1214/aoms/1177693335
1971
-
[33]
Lower Bounds and Optimal Algorithms for Personalized Federated Learning
Filip Hanzely, Slavomír Hanzely, Samuel Horváth, and Peter Richtarik. Lower Bounds and Optimal Algorithms for Personalized Federated Learning . In Advances in Neural Information Processing Systems , volume 33, pages 2304--2315. Curran Associates, Inc., 2020
2020
-
[34]
Hospedales, A
T. Hospedales, A. Antoniou, P. Micaelli, and A. Storkey. Meta-learning in neural networks: A survey. IEEE Transactions on Pattern Analysis; Machine Intelligence, 44 0 (09): 0 5149--5169, sep 2022. ISSN 1939-3539. doi:10.1109/TPAMI.2021.3079209
2022
-
[35]
Clustered multi-task learning: A convex formulation
Laurent Jacob, Jean-philippe Vert, and Francis Bach. Clustered multi-task learning: A convex formulation. In D. Koller, D. Schuurmans, Y. Bengio, and L. Bottou, editors, Advances in Neural Information Processing Systems, volume 21. Curran Associates, Inc., 2008. URL https://pr...
2008
-
[36]
Revisiting Frank-Wolfe : Projection-free sparse convex optimization
Martin Jaggi. Revisiting Frank-Wolfe : Projection-free sparse convex optimization. In Sanjoy Dasgupta and David McAllester, editors, Proceedings of the 30th International Conference on Machine Learning, volume 28 of Proceedings of Machine Learning Research, pages 427--435, Atl...
2013
-
[37]
Adaptive gradient-based meta-learning methods
Mikhail Khodak, Maria-Florina Balcan, and Ameet Talwalkar. Adaptive gradient-based meta-learning methods. Curran Associates Inc., Red Hook, NY, USA, 2019 a
2019
-
[38]
Provable guarantees for gradient-based meta-learning, 2019 b
Mikhail Khodak, Maria-Florina Balcan, and Ameet Talwalkar. Provable guarantees for gradient-based meta-learning, 2019 b
2019
-
[39]
Gradient-based meta-learning with learned layerwise metric and subspace
Yoonho Lee and Seungjin Choi. Gradient-based meta-learning with learned layerwise metric and subspace. In Jennifer Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages...
2018
-
[40]
Hierarchical clustering multi-task learning for joint human action grouping and recognition
An-An Liu, Yu-Ting Su, Wei-Zhi Nie, and Mohan Kankanhalli. Hierarchical clustering multi-task learning for joint human action grouping and recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39 0 (1): 0 102--114, 2017. doi:10.1109/TPAMI.2016.2537337
2017
-
[41]
G. G. Lorentz. Metric entropy and approximation . Bulletin of the American Mathematical Society, 72 0 (6): 0 903 -- 937, 1966
1966
-
[42]
Taking advantage of sparsity in multi-task learning
Karim Lounici, Massimiliano Pontil, AB Tsybakov, and SA Geer. Taking advantage of sparsity in multi-task learning. Proceedings of the 22nd Conference on Information Theory, 12 2009
2009
-
[43]
Three Approaches for Personalization with Applications to Federated Learning
Yishay Mansour, Mehryar Mohri, Jae Ro, and Ananda Theertha Suresh. Three Approaches for Personalization with Applications to Federated Learning . arXiv:2002.10619 [cs, stat], July 2020. arXiv: 2002.10619
2002 arXiv
-
[44]
Bounds for linear multi-task learning
Andreas Maurer. Bounds for linear multi-task learning. Journal of Machine Learning Research, 7 0 (5): 0 117--139, 2006. URL http://jmlr.org/papers/v7/maurer06a.html
2006
-
[45]
The benefit of multitask representation learning
Andreas Maurer, Massimiliano Pontil, and Bernardino Romera-Paredes. The benefit of multitask representation learning. J. Mach. Learn. Res., 17 0 (1): 0 2853–2884, January 2016. ISSN 1532-4435
2016
-
[46]
Communication-Efficient Learning of Deep Networks from Decentralized Data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communication-Efficient Learning of Deep Networks from Decentralized Data . In Aarti Singh and Jerry Zhu, editors, Proceedings of the 20th International Conference on Artificial Intelligence ...
2017
-
[47]
Collaborative learning with shared linear representations: Statistical rates and optimal algorithms, 2024
Xiaochun Niu, Lili Su, Jiaming Xu, and Pengkun Yang. Collaborative learning with shared linear representations: Statistical rates and optimal algorithms, 2024. URL https://arxiv.org/abs/2409.04919
2024
-
[48]
Excess risk bounds for multitask learning with trace norm regularization
Massimiliano Pontil and Andreas Maurer. Excess risk bounds for multitask learning with trace norm regularization. In Shai Shalev-Shwartz and Ingo Steinwart, editors, Proceedings of the 26th Annual Conference on Learning Theory, volume 30 of Proceedings of Machine Learning Rese...
2013
-
[49]
Tsybakov
Angelika Rohde and Alexandre B. Tsybakov. Estimation of high-dimensional low-rank matrices. The Annals of Statistics, 39 0 (2): 0 887--930, 2011
2011
-
[50]
Hanson-Wright inequality and sub-gaussian concentration
Mark Rudelson and Roman Vershynin. Hanson-Wright inequality and sub-gaussian concentration . Electronic Communications in Probability, 18 0 (none): 0 1 -- 9, 2013. doi:10.1214/ECP.v18-2865. URL https://doi.org/10.1214/ECP.v18-2865
2013 doi
-
[51]
Clustered federated learning: Model-agnostic distributed multitask optimization under privacy constraints
Felix Sattler, Klaus-Robert Müller, and Wojciech Samek. Clustered federated learning: Model-agnostic distributed multitask optimization under privacy constraints. IEEE Transactions on Neural Networks and Learning Systems, PP: 0 1--13, 08 2020. doi:10.1109/TNNLS.2020.3015958
2020
-
[52]
Statistically and computationally efficient linear meta-representation learning
Kiran Koshy Thekumparampil, Prateek Jain, Praneeth Netrapalli, and Sewoong Oh. Statistically and computationally efficient linear meta-representation learning. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing S...
2021
-
[53]
Provable meta-learning of linear representations
Nilesh Tripuraneni, Chi Jin, and Michael Jordan. Provable meta-learning of linear representations. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 10434...
2021
-
[54]
Wainwright
Martin J. Wainwright. High-Dimensional Statistics: A Non-Asymptotic Viewpoint. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 2019. doi:10.1017/9781108627771
2019 doi
-
[55]
First-order ANIL provably learns representations despite overparametrisation
O g uz Y \"u ksel, Etienne Boursier, and Nicolas Flammarion. First-order ANIL provably learns representations despite overparametrisation. In NeurIPS 2023 Workshop on Mathematics of Modern Machine Learning, 2023. URL https://openreview.net/forum?id=89fKHVtsMR
2023
-
[56]
Clustered multi-task learning via alternating structure optimization
Jiayu Zhou, Jianhui Chen, and Jieping Ye. Clustered multi-task learning via alternating structure optimization. In Proceedings of the 24th International Conference on Neural Information Processing Systems, NIPS'11, page 702–710, Red Hook, NY, USA, 2011. Curran Associates Inc. ...
2011
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.