Pith. sign in

REVIEW 3 cited by

MGDA Converges under Generalized Smoothness, Provably

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19440 v5 pith:TJQWHJIA submitted 2024-05-29 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords epsilonmgdadirectiongeneralizedmathcalsmooththeyalgorithms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Multi-objective optimization (MOO) is receiving more attention in various fields such as multi-task learning. Recent works provide some effective algorithms with theoretical analysis but they are limited by the standard $L$-smooth or bounded-gradient assumptions, which typically do not hold for neural networks, such as Long short-term memory (LSTM) models and Transformers. In this paper, we study a more general and realistic class of generalized $\ell$-smooth loss functions, where $\ell$ is a general non-decreasing function of gradient norm. We revisit and analyze the fundamental multiple gradient descent algorithm (MGDA) and its stochastic version with double sampling for solving the generalized $\ell$-smooth MOO problems, which approximate the conflict-avoidant (CA) direction that maximizes the minimum improvement among objectives. We provide a comprehensive convergence analysis of these algorithms and show that they converge to an $\epsilon$-accurate Pareto stationary point with a guaranteed $\epsilon$-level average CA distance (i.e., the gap between the updating direction and the CA direction) over all iterations, where totally $\mathcal{O}(\epsilon^{-2})$ and $\mathcal{O}(\epsilon^{-4})$ samples are needed for deterministic and stochastic settings, respectively. We prove that they can also guarantee a tighter $\epsilon$-level CA distance in each iteration using more samples. Moreover, we analyze an efficient variant of MGDA named MGDA-FA using only $\mathcal{O}(1)$ time and space, while achieving the same performance guarantee as MGDA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sign Operator for Coping with Heavy-Tailed Noise in Non-Convex Optimization: High Probability Bounds Under $(L_0, L_1)$-Smoothness

    math.OC 2025-02 conditional novelty 7.0 of 10

    First high-probability bounds for SignSGD with batching or majority voting under (L0, L1)-smoothness and heavy-tailed noise, with near-optimal epsilon-dependencies.

  2. Revisiting Convergence: Shuffling Complexity Beyond Lipschitz Smoothness

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Shuffling gradient methods converge without Lipschitz smoothness under a sub-quadratic ℓ-smoothness condition, matching Lipschitz-case rates when ℓ is constant.

  3. Multiple Wasserstein Gradient Descent Algorithm for Multi-Objective Distributional Optimization

    cs.LG 2025-05 conditional novelty 5.0 of 10

    MWGraD aggregates multiple Wasserstein gradients with dynamically updated weights to find Pareto-stationary distributions, with convergence guarantees and improved multi-task accuracy.

Pith tools