Pith. sign in

REVIEW 9 cited by

Fusing finetuned models for better pretraining

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.03044 v1 pith:ACWHURWL submitted 2022-04-06 cs.CL cs.CVcs.LG

classification cs.CLcs.CVcs.LG
keywords fusingmodelsbetterintertrainingmodelpretrainedpretrainingapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pretrained models are the standard starting point for training. This approach consistently outperforms the use of a random initialization. However, pretraining is a costly endeavour that few can undertake. In this paper, we create better base models at hardly any cost, by fusing multiple existing fine tuned models into one. Specifically, we fuse by averaging the weights of these models. We show that the fused model results surpass the pretrained model ones. We also show that fusing is often better than intertraining. We find that fusing is less dependent on the target task. Furthermore, weight decay nullifies intertraining effects but not those of fusing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Gradient Descent Simulate Prompting?

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A MAML-style meta-training objective makes a single gradient step on new text recover part of the performance that prompting achieves, on reversal-curse and passage-QA tasks.

  2. DA-MergeLoRA: Hypernetwork-Based LoRA Merging for Few-Shot Test-Time Domain Adaptation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hypernetwork generates per-column merging weights to combine source LoRA modules on CLIP, achieving state-of-the-art few-shot test-time domain adaptation.

  3. The Appeal and Reality of Recycling LoRAs with Adaptive Merging

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Adaptive merging of recycled LoRAs gives little benefit over training a target-task LoRA, and randomly initialized LoRAs work as well as real ones once the target LoRA is in the pool.

  4. Interaction-Merged Motion Planning: Effectively Leveraging Diverse Motion Datasets for Robust Planning

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A two-step model-merging method transfers interaction knowledge from multiple motion datasets to a target domain, outperforming ensembling and domain adaptation at the same inference cost.

  5. Update Your Transformer to the Latest Release: Re-Basin of Task Vectors

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A two-level permutation strategy over attention heads and their inner units transfers task vectors from an old pretrained Transformer to a new one, data-free.

  6. Intrinsic Strain-Driven Topological Evolution in SrRuO3 via Flexural Strain Engineering

    cond-mat.mtrl-sci 2025-08 unverdicted novelty 5.0 of 10

    The abstract reports a 21% anomalous Hall conductivity increase in flexurally strained SrRuO3, but the submitted full text belongs to a different machine learning paper.

  7. EpiCoDe: Boosting Model Performance Beyond Training with Extrapolation and Contrastive Decoding

    cs.CL 2025-06 conditional novelty 5.0 of 10

    EpiCoDe builds an extrapolated checkpoint from early and late finetuned models, then subtracts the late model's logits from the extrapolated model's logits during decoding to boost accuracy.

  8. Continual Learning in Vision-Language Models via Aligned Model Merging

    cs.CV 2025-05 conditional novelty 5.0 of 10

    PAM merges a task-specific LoRA into a global LoRA and re-initializes sign-conflicting weights during training, reducing catastrophic forgetting in continual VLM learning.

  9. Neural Parameter Search for Slimmer Fine-Tuned Models and Better Transfer

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Neural Parameter Search (NPS) prunes fine-tuned models by evolutionary reweighting of magnitude-based task vector subspaces, improving transfer, fusion, and compression.

Pith tools