Pith. sign in

REVIEW 14 cited by

Multi-Task Learning as a Bargaining Game

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.01017 v2 pith:QPX2RUUI submitted 2022-02-02 cs.LG cs.GT

classification cs.LGcs.GT
keywords jointbargaininggradientslearningmulti-tasktasksdirectiongame
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In Multi-task learning (MTL), a joint model is trained to simultaneously make predictions for several tasks. Joint training reduces computation costs and improves data efficiency; however, since the gradients of these different tasks may conflict, training a joint model for MTL often yields lower performance than its corresponding single-task counterparts. A common method for alleviating this issue is to combine per-task gradients into a joint update direction using a particular heuristic. In this paper, we propose viewing the gradients combination step as a bargaining game, where tasks negotiate to reach an agreement on a joint direction of parameter update. Under certain assumptions, the bargaining problem has a unique solution, known as the Nash Bargaining Solution, which we propose to use as a principled approach to multi-task learning. We describe a new MTL optimization procedure, Nash-MTL, and derive theoretical guarantees for its convergence. Empirically, we show that Nash-MTL achieves state-of-the-art results on multiple MTL benchmarks in various domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DanceOPD: On-Policy Generative Field Distillation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Hard-routed, single low-noise on-policy velocity matching composes conflicting image-generation capabilities into one flow student better than joint training, merging, or dense OPD baselines.

  2. S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF

    cs.AI 2026-05 conditional novelty 6.0 of 10

    S2T-RLHF splits each response-level RLHF reward into sentence shares and then token shares, via bargaining and Dirichlet weighting, yielding steadier training with competitive preference alignment.

  3. GreenRFM: Learning a resource-efficient radiology vision-language foundation model via supervision-centric pre-training

    cs.CV 2026-03 conditional novelty 6.0 of 10

    MUST supervision—LLM-distilled diagnostic labels plus two-stage ubiquitous training—lets a 33M-parameter 3D ResNet-18 reach 84.8 zero-shot AUC on CT-RATE in 24 GPU-hours and transfer across institutions and MRI.

  4. Simple Optimizers for Convex Aligned Multi-Objective Optimization

    cs.LG 2025-09 reject novelty 6.0 of 10

    Convex AMOO is analyzed under Lipschitz and smooth assumptions with a maximum-gap metric, giving simple gradient methods with rates independent of the number of objectives, plus a flawed equal-weights lower bound.

  5. AutoScale: Linear Scalarization Guided by Multi-Task Optimization Metrics

    cs.LG 2025-08 conditional novelty 6.0 of 10

    AutoScale selects fixed linear-scalarization weights by optimizing multi-task optimization metrics during a short exploration phase, matching grid-searched performance without search.

  6. Synchronizing Task Behavior: Aligning Multiple Tasks during Test-Time Training

    cs.LG 2025-07 conditional novelty 6.0 of 10

    S4T synchronizes multi-task test-time adaptation by learning cross-task relations on the source domain and using them to align task predictions on the target domain.

  7. Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A token-space SVD-based method that separately resolves gradient conflicts in the range and null spaces of transformer tokens improves multi-task learning performance with minimal extra parameters.

  8. Interaction-Merged Motion Planning: Effectively Leveraging Diverse Motion Datasets for Robust Planning

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A two-step model-merging method transfers interaction knowledge from multiple motion datasets to a target domain, outperforming ensembling and domain adaptation at the same inference cost.

  9. FastCAR: Fast Classification And Regression for Task Consolidation in Multi-Task Learning to Model a Continuous Property Variable of Detected Object Class

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A single regression network trained on class-shifted hardness labels can jointly classify steel microstructures and predict hardness, outperforming multi-task baselines on the authors' own dataset.

  10. GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining

    cs.LG 2025-05 conditional novelty 6.0 of 10

    GRAPE uses a minimax group-DRO scheme to reweight both source domains and target tasks during pretraining, improving multi-task reasoning and low-resource language modeling.

  11. FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A multi-objective fine-tuning method with Minimum Potential Delay fairness improves hand and face quality in human image generation while maintaining global quality.

  12. MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Grouping instruction-tuning datasets by redundancy, uniqueness, or synergy of text-image interaction improves vision-language model accuracy over single-task and unselective multi-task tuning.

  13. AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs

    cs.LG 2025-05 conditional novelty 5.0 of 10

    AutoMixAlign adaptively reweights or resamples task data during DPO training to match specialist-model losses, improving average performance on helpfulness, coding, and safety benchmarks compared to standard DPO and m...

  14. Expert Merging in Sparse Mixture of Experts with Nash Bargaining

    cs.LG 2025-10 conditional novelty 4.0 of 10

    Expert merging in MoE models can be improved by setting per-expert weights with the Nash bargaining solution instead of simple averaging.

Pith tools