REVIEW 3 cited by
Adaptive Scheduling for Multi-Task Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
To train neural machine translation models simultaneously on multiple tasks (languages), it is common to sample each task uniformly or in proportion to dataset sizes. As these methods offer little control over performance trade-offs, we explore different task scheduling approaches. We first consider existing non-adaptive techniques, then move on to adaptive schedules that over-sample tasks with poorer results compared to their respective baseline. As explicit schedules can be inefficient, especially if one task is highly over-sampled, we also consider implicit schedules, learning to scale learning rates or gradients of individual tasks instead. These techniques allow training multilingual models that perform better for low-resource language pairs (tasks with small amount of data), while minimizing negative effects on high-resource tasks.
Forward citations
Cited by 3 Pith papers
-
Learning in Deep Networks under Dale's Constraint
An on-off two-channel network with fixed-sign synapses and local Hebbian learning is claimed to recover backpropagation exactly under symmetric weights and to beat comparable vanilla networks on Tiny ImageNet.
-
Efficient Task Grouping Through Samplewise Optimisation Landscape Analysis
Task groupings are inferred from one-step gradient distances between tasks at initialization and clustered with graph attention networks, matching TAG/HOA performance at lower cost.
-
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
Selectively fine-tuning only 40% of a multimodal LLM's layers and neurons can match or slightly beat full fine-tuning on multilingual image-to-text translation benchmarks, though the measured gains are marginal.
Discussion (0). Continue with ORCID to comment.