REVIEW 2 cited by
Model Merging by Uncertainty-Based Gradient Matching
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Models trained on different datasets can be merged by a weighted-averaging of their parameters, but why does it work and when can it fail? Here, we connect the inaccuracy of weighted-averaging to mismatches in the gradients and propose a new uncertainty-based scheme to improve the performance by reducing the mismatch. The connection also reveals implicit assumptions in other schemes such as averaging, task arithmetic, and Fisher-weighted averaging. Our new method gives consistent improvements for large language models and vision transformers, both in terms of performance and robustness to hyperparameters. Code available here.
Forward citations
Cited by 2 Pith papers
-
SAFE-Merge: Data-Free Continual Model Merging with General Knowledge Preservation
SAFE-Merge masks risk-prone parameter updates and recovers lost task information with a constrained low-rank correction, achieving the best H-score in data-free continual model merging benchmarks.
-
STAR: Spectral Truncation and Rescale for Model Merging
STAR merges fine-tuned models by truncating small singular values of task vectors and rescaling to restore the nuclear norm, improving multi-task merging performance.
Discussion (0). Continue with ORCID to comment.