Pith. sign in

REVIEW 7 cited by

Reasonable Effectiveness of Random Weighting: A Litmus Test for Multi-Task Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.10603 v2 pith:NB3G7NLC submitted 2021-11-20 cs.LG

classification cs.LG
keywords methodsrandomweightingachieveeffectivenessgradientlossbaselines
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-Task Learning (MTL) has achieved success in various fields. However, how to balance different tasks to achieve good performance is a key problem. To achieve the task balancing, there are many works to carefully design dynamical loss/gradient weighting strategies but the basic random experiments are ignored to examine their effectiveness. In this paper, we propose the Random Weighting (RW) methods, including Random Loss Weighting (RLW) and Random Gradient Weighting (RGW), where an MTL model is trained with random loss/gradient weights sampled from a distribution. To show the effectiveness and necessity of RW methods, theoretically we analyze the convergence of RW and reveal that RW has a higher probability to escape local minima, resulting in better generalization ability. Empirically, we extensively evaluate the proposed RW methods to compare with twelve state-of-the-art methods on five image datasets and two multilingual problems from the XTREME benchmark to show RW methods can achieve comparable performance with state-of-the-art baselines. Therefore, we think that the RW methods are important baselines for MTL and should attract more attentions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring Line Bundle Standard Models with Transformers

    hep-th 2026-06 unverdicted novelty 7.0 of 10

    A Transformer trained by reinforcement learning generates heterotic line-bundle sums that satisfy anomaly-cancellation, stability, and chirality constraints, and its policy transfers usefully across Calabi-Yau geometries.

  2. Flatness and Gradient Alignment Are Both Necessary: Spectral-Aware Gradient-Aligned Exploration for Multi-Distribution Learning

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    SAGE, an optimizer combining spectral polar-factor perturbation with gradient-agreement-scaled noise, reports 78.9% average on DomainBed.

  3. AutoScale: Linear Scalarization Guided by Multi-Task Optimization Metrics

    cs.LG 2025-08 conditional novelty 6.0 of 10

    AutoScale selects fixed linear-scalarization weights by optimizing multi-task optimization metrics during a short exploration phase, matching grid-searched performance without search.

  4. FastCAR: Fast Classification And Regression for Task Consolidation in Multi-Task Learning to Model a Continuous Property Variable of Detected Object Class

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A single regression network trained on class-shifted hardness labels can jointly classify steel microstructures and predict hardness, outperforming multi-task baselines on the authors' own dataset.

  5. FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A multi-objective fine-tuning method with Minimum Potential Delay fairness improves hand and face quality in human image generation while maintaining global quality.

  6. Controlled Data Rebalancing in Multi-Task Learning for Real-World Image Super-Resolution

    cs.CV 2025-06 conditional novelty 5.0 of 10

    The paper proposes splitting real-world degradation into four tasks and adaptively rebalancing each task's training data volume, reporting consistent gains over prior super-resolution methods.

  7. Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning

    cs.LG 2026-07 conditional novelty 4.0 of 10

    On a large World of Tanks dataset, a shared multi-task model with equal weighting or PCGrad outperforms single-task models on average, and task/map pre-training helps most in low-data regimes.

Pith tools