Pith. sign in

REVIEW 4 major objections 4 minor 68 references

Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation

T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Fine-tuning with a small subset of LoRA blocks active can match full-LoRA target accuracy while preserving zero-shot and out-of-distribution knowledge.

desk verdict A broad, mostly consistent empirical study of gated LoRA for CLIP/DINO, but the headline CLIP claim needs a random-selection control and the unreported threshold tau makes the efficiency numbers unverifiable. read the letter →

arxiv 2501.15377 v1 pith:4UKPXFMQ submitted 2025-01-26 cs.CV

classification cs.CV
keywords catastrophicforgettingparameter-efficientfine-tuninglow-rankadaptationindicatorfunctionout-of-distributiongeneralizationzero-shotretentionselectiveblockactivationCLIP
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that fine-tuning pretrained vision and vision-language models does not require updating all low-rank blocks: a learned gate can switch most of them off. It adds a per-block score $s_\ell$ and a straight-through indicator $I_\tau(s_\ell)$ in front of each LoRA residual, with an $\ell^1$ penalty on the scores to push most gates to zero. On CLIP fine-tuned on ImageNet-1K at rank 128, activating only 6.25% of blocks matches LoRA's 81.77% accuracy at 81.84%, while zero-shot classification improves from 51.44 to 61.48 and zero-shot retrieval loss stays under 5.73% instead of about 28% for full fine-tuning. The same gating works when applied to DoRA and to DINO-ViT backbones, keeping target accuracy comparable while retaining far more of the pretrained knowledge. A reader should care because it suggests a cheap, model-agnostic way to adapt foundation models without erasing what they already know.

What carries the argument

The central object is a per-block gate built from a learnable scalar score $s_\ell$ and a straight-through estimator (STE) indicator $I_\tau(s_\ell)$ that multiplies each low-rank residual $A_\ell B_\ell$, trained with an $\ell^1$ sparsity penalty on the scores. The gate makes the set of active blocks a trainable, sparse choice: the scores decide which low-rank blocks stay on for the target task, and the $\ell^1$ term is what produces the tiny active-block percentages (1.39–6.25% on CLIP, 2.7–44% on DINO depending on $\lambda$, rank, and dataset). The straight-through estimator lets gradients flow through the binary decision so the selection itself is learned, not set by hand.

What would settle it

Repeat the rank-128 CLIP fine-tuning with the paper's equations but sweep the threshold $\tau$ across a wide range, say 0, 0.01, 0.1, and 0.5, while keeping $\lambda=1$; the claim fails if any run either leaves nearly all blocks active, drops ImageNet accuracy below LoRA's 81.77%, or loses the zero-shot retention advantage, since that would show the result depends on the unspecified threshold rather than on the gating mechanism itself.

Watch

Extended reading notes

Core claim

The central discovery is that the low-rank updates learned during fine-tuning are highly redundant: a small, task-dependent subset of blocks carries almost all of the adaptation signal, and the rest can be left off without losing target-task accuracy while preserving the pretrained feature space. Formally, the paper updates each block as $W_\ell = W_{0,\ell} + I_\tau(s_\ell)A_\ell B_\ell$, where $I_\tau(s_\ell)=1$ if $s_\ell\ge\tau$ and $0$ otherwise, and regularizes the gate scores with $\lambda\sum_\ell |s_\ell|$. With this update, LoRA at rank 128 on CLIP reaches 81.84% ImageNet-1K accuracy (versus 81.77% for full LoRA) using 6.25% of the blocks; zero-shot classification average rises to 61.48 versus 51.44 for LoRA, zero-shot retrieval loss falls to at most 5.73% versus about 28% for FLYP, and unmerged inference becomes up to 2.9x faster for LoRA and 5x faster for DoRA at rank 256. On DINO-ViT, activating as few as 2.7% of blocks keeps target accuracy comparable to LoRA while source-domain accuracy on ImageNet-100 is retained substantially better than under full fine-tuning.

Load-bearing premise

The load-bearing premise is that the sparsity penalty and an unspecified threshold $\tau$ together yield the reported small active-block counts; if $\tau$ is set differently or depends on the scale of the learned scores, the headline efficiency and forgetting numbers would not reproduce.

Editorial extensions

If this is right

  • Fine-tuning a foundation model on a new dataset can be done with only a small fraction of its low-rank blocks active, cutting inference FLOPs and memory proportionally; the paper measures up to 2.9x (LoRA) and 5x (DoRA) faster unmerged inference at rank 256.
  • Because the gate applies to any LoRA-variant method (the paper demonstrates DoRA), it is a plug-in that can reduce catastrophic forgetting across PEFT techniques.
  • Higher ranks, which normally cause stronger forgetting, become usable: at rank 256 the gated LoRA reaches 82.31% ImageNet accuracy with 5.56% of blocks active, while plain LoRA's out-of-distribution mean drops to 53.81.
  • The retained zero-shot abilities mean a fine-tuned CLIP can still serve as a general image-text retriever and zero-shot classifier, not just as a specialist on the target classes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: the layer-activation maps, which concentrate in feedforward MLP blocks and in later CLIP vision layers, suggest that a fixed or pretrained subset of blocks could be enough, which would remove the need to learn gate scores at all.
  • Inference: the reported preservation suggests low-rank updates are redundant across transformer residuals generally, implying selective LoRA could slot into continual-learning pipelines without rehearsal buffers.
  • Inference: because $\tau$ is unspecified, a normalized gate such as a hard sigmoid over scores divided by the maximum score would make the method reproducible and remove dependence on the score scale.
  • Inference: if sparsity is what preserves knowledge, then the method's benefit should be testable against an ablation that selects the same number of blocks by gradient magnitude or by the paper's block-activation heatmaps rather than training scores.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a PEFT method that extends Task Adaptive Parameter Sharing (TAPS) by adding an indicator function I_tau(s_i) that selectively activates LoRA blocks, with an L1 penalty on the block scores s_i. The central claim is that this gating mechanism allows fine-tuned models to match full-LoRA in-distribution accuracy while using only a small fraction of active blocks, thereby retaining zero-shot and OOD knowledge and reducing inference cost. Experiments are reported on DINO ViT-S/B-16 fine-tuned on six target datasets and on CLIP fine-tuned on ImageNet-1K, evaluated on OOD, zero-shot classification, and zero-shot retrieval benchmarks. The method is also applied to DoRA. The headline result is that CLIP at rank 128 with 6.25% active blocks reaches 81.84% ImageNet accuracy versus 81.77% for full LoRA, while zero-shot classification improves from 51.44 to 61.48 and retrieval retention is substantially better than full LoRA.

Significance. If the selective-activation mechanism is genuinely responsible for the observed preservation of pre-trained knowledge, the method is a simple, practical extension of LoRA-style PEFT with broad applicability. The paper’s strengths are its extensive empirical coverage: multiple backbones (DINO ViT-S/B, CLIP), multiple PEFT variants (LoRA and DoRA), many transfer datasets, several ranks and regularization strengths, and separate evaluations of OOD robustness, zero-shot classification, and zero-shot retrieval. The inference-FLOPs analysis is a useful practical contribution. However, the central scientific claim—that learned selection, rather than merely reducing the number of adapted parameters, yields the retention gains—is not fully supported for the main CLIP experiment, and the reproducibility of the reported active-block percentages is compromised by the unspecified threshold tau. The paper also overstates its contribution in the title and abstract relative to the explicit limitation stated in Section 5.

major comments (4)
  1. [Sec. 3.2, Eq. (3)] This is a load-bearing issue: the claimed efficiency gains and the claimed trade-off curves depend on where the threshold sits.
  2. [Sec. 4.2, Fig. 6 and Table 10] The central claim 'selective activation preserves knowledge' is not distinguishable from 'fewer adapted parameters preserve knowledge' without this control.
  3. [Sec. 5, Conclusion] The current framing overclaims the contribution relative to what is measured in the experiments.
  4. [Tables 10-15] This concern is particularly relevant to the DINO random-selection comparison in Table 1, where the claimed advantage is about 2%.
minor comments (4)
  1. [Supplementary Table 4] The Linear row in Table 4 lists CIFAR-100=72.07, IN-100=87.20, Mean=88.32, but the mean of 72.07 and 87.20 is 79.64. This appears to be a copy-paste error from Table 3 (where the mean 88.45 is consistent with the CIFAR-10 numbers). Please correct.
  2. [Sec. 2, Related Works] Given that AdaLoRA is cited as an adaptive rank-allocation method, a direct empirical comparison with AdaLoRA (or a brief explanation of why it is not compared) would help position the contribution. Currently the method is compared only with full LoRA/DoRA and with full fine-tuning/linear probing.
  3. [Sec. 4.3, Fig. 9] The FLOPs comparison is presented for the unmerged setting where inactive blocks are skipped. It would be helpful to state explicitly whether the reported speedups assume that active LoRA blocks are not merged into the base weights, since merging would eliminate the inference benefit; the current text implies but does not state this.
  4. [Sec. 4.4, Fig. 10] The figure caption says 'across 8 runs' for CLIP, but the text describes 7 runs at one lambda plus 6 runs at other lambdas. Please reconcile the count and clarify whether the shown activations are pooled over all runs or averaged per configuration.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the central LoRA-plus-indicator result is measured against held-out OOD and zero-shot benchmarks, and the self-citations only supply setup and inspiration, not the reported numbers.

full rationale

The paper's derivation chain is empirical rather than definitional. Equation (2) defines the adapted weight as W_l = W_0,l + I_tau(s_i) A_l B_l, and Equation (3) defines the indicator via an unspecified threshold tau, but the headline claims (e.g., 6.25% active blocks matching LoRA on ImageNet, zero-shot retention within 5.73%) are measured outcomes on held-out benchmarks, not quantities derived from the definition. There is no fitted parameter that is later renamed as a prediction: the reported active-block percentages are post-training counts, and the ID/OOD/zero-shot accuracies come from external evaluation datasets, so the central comparisons are not forced by construction. The paper does cite the authors' own prior work: reference [3] (same first and last authors) is used for the problem setup, namely 'Following [3]', and reference [53] (TAPS, with overlapping author Avinash Ravichandran) is the stated inspiration for the indicator function. These self-citations provide framing and a borrowed mechanism, but they do not by themselves imply the reported accuracy or retention numbers; the method is additionally supported by a random-selection control (Table 1) on DINO/CIFAR-100, even though no such control is given for the headline CLIP experiment. The conclusion's admission that the method 'does not directly address catastrophic forgetting' is an internal inconsistency with the title and abstract, but inconsistency is not circularity. The missing threshold tau is a reproducibility concern, not a circularity concern: it affects whether the exact 6.25% figure can be audited, but it does not make any prediction equal to an input by definition. Overall, the central claim retains independent empirical content, so the circularity score is low.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central mechanism has no fitted free parameters in a derived sense, but the block selection depends on manually set tau, lambda, rank, and K-NN K. The paper provides no tau value, and lambda is tuned per dataset. Axioms are standard domain assumptions about pretrained representations and evaluation protocols, plus an assumption that unmerged LoRA is the right inference regime.

free parameters (4)
  • threshold tau = not reported
    Eq. (3) defines the indicator I_tau(s_i)=1 if s_i>=tau, but no value is given in Appendix B; all block-activation percentages depend on it.
  • regularization lambda = 0.1, 0.5, 1.0, 1.5, chosen per dataset/rank
    Eq. (4) controls sparsity; the paper selects lambda per DINO dataset (e.g., 0.5 for Flowers, 1.0 for DTD), which affects the reported active-block percentage.
  • LoRA rank = 4, 8, 16, 32, 64, 128, 256
    Rank is a hand-chosen hyperparameter; higher ranks increase ID accuracy but increase forgetting, so rank choice interacts with the main claim.
  • K-NN K = 20
    Used to measure source-domain retention for DINO; the forgetting numbers depend on this evaluation choice.
assumptions (4)
  • standard math Straight-through estimator provides usable gradients for the discrete indicator
    Training in Eq. (2)-(3) relies on STE through I_tau(s_i); the paper cites [62] and does not analyze bias.
  • domain assumption Pretrained CLIP and DINO features are the knowledge to preserve; adapting a few LoRA blocks retains them
    The method freezes the backbone and only trains LoRA and scores; retention is then measured on pretraining-related data.
  • domain assumption K-NN on ImageNet-100 (DINO) and zero-shot retrieval on COCO/Flickr (CLIP) measure catastrophic forgetting
    Section 4.1 and 4.2 use these protocols; if they are not faithful measures of retained pretraining knowledge, the forgetting claims weaken.
  • domain assumption Inference efficiency is assessed with LoRA/DoRA adapters left unmerged
    Section 4.3 computes FLOPs and speedups assuming adapters are not merged; merged LoRA would add no inference FLOPs regardless of rank.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation." pith.science (2026). https://pith.science/paper/4UKPXFMQ

@misc{pith2026250115377,
  author       = {Pith},
  title        = {Pith review of: Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4UKPXFMQ}},
  note         = {Machine review of arXiv:2501.15377}
}
read the original abstract

Adapting deep learning models to new domains often requires computationally intensive retraining and risks catastrophic forgetting. While fine-tuning enables domain-specific adaptation, it can reduce robustness to distribution shifts, impacting out-of-distribution (OOD) performance. Pre-trained zero-shot models like CLIP offer strong generalization but may suffer degraded robustness after fine-tuning. Building on Task Adaptive Parameter Sharing (TAPS), we propose a simple yet effective extension as a parameter-efficient fine-tuning (PEFT) method, using an indicator function to selectively activate Low-Rank Adaptation (LoRA) blocks. Our approach minimizes knowledge loss, retains its generalization strengths under domain shifts, and significantly reduces computational costs compared to traditional fine-tuning. We demonstrate that effective fine-tuning can be achieved with as few as 5\% of active blocks, substantially improving efficiency. Evaluations on pre-trained models such as CLIP and DINO-ViT demonstrate our method's broad applicability and effectiveness in maintaining performance and knowledge retention.

Figures

Figures reproduced from arXiv: 2501.15377 by the authors.

Figure 1
Figure 1. Comparison of pretrained CLIP, fully fine-tuned CLIP [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of our proposed method incorporating LoRA [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The figure compares the top-1 accuracy of pretrained and fine-tuned DINO ViT/S-16 models on CIFAR-10 and CIFAR-100 using [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: The figure compares top-1 accuracy of fine-tuned DINO ViT/S-16 models across transfer datasets and ImageNet-100. It demon [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Comparison of top-1 accuracy for fine-tuned DINO [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: We compare in-distribution, out-of-distribution, zero-shot classification, and retrieval performance of the fine-tuned CLIP model [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Comparison of in-distribution, out-of-distribution, zero-shot classification, and zero-shot retrieval performance of the fine-tuned [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 8
Figure 8. Figure 8: The y-axis represents the percentage of activated blocks [PITH_FULL_IMAGE:figures/full_fig_p006_8.png]
Figure 10
Figure 10. Figure 10: Visualization of the normalized count of activated [PITH_FULL_IMAGE:figures/full_fig_p007_10.png]
Figure 11
Figure 11. Figure 11: Normalized count of activated blocks across 15 runs [PITH_FULL_IMAGE:figures/full_fig_p007_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

68 extracted references · 43 canonical work pages

  1. [1]

    Agiza, Marina Neseem, and Sherief Reda

    Ahmed A. Agiza, Marina Neseem, and Sherief Reda. Mt- lora: A low-rank adaptation approach for efficient multi-task learning. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16196–16205, 2024. 2

  2. [2]

    data2vec: A general framework for self-supervised learning in speech, vision and language

    Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. data2vec: A general framework for self-supervised learning in speech, vision and language. ArXiv, abs/2202.03555, 2022. 2

  3. [3]

    Parameter efficient fine-tuning of self-supervised vits without catastrophic forgetting

    Reza Akbarian Bafghi, Nidhin Harilal, Claire Monteleoni, and Maziar Raissi. Parameter efficient fine-tuning of self-supervised vits without catastrophic forgetting. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 3679–3684, 2024. 2, 3, 4

  4. [4]

    Tenenbaum, and Boris Katz

    Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Joshua B. Tenenbaum, and Boris Katz. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In Neural Information Processing Systems, 2019. 6

  5. [5]

    Cunningham

    Dan Biderman, Jose Gonzalez Ortiz, Jacob Portes, Mansheej Paul, Philip Greengard, Connor Jennings, Daniel King, Sam Havens, Vitaliy Chiley, Jonathan Frankle, Cody Blakeney, and John P. Cunningham. Lora learns less and forgets less. ArXiv, abs/2405.09673, 2024. 2

  6. [6]

    Food-101 - mining discriminative components with random forests

    Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101 - mining discriminative components with random forests. In European Conference on Computer Vision, 2014. 4, 6

  7. [7]

    Emerg- ing properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 2, 4

  8. [8]

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Ge- offrey E. Hinton. A simple framework for contrastive learn- ing of visual representations. ArXiv, abs/2002.05709, 2020. 2

Show all 68 references
  1. [9]

    Describing textures in the wild

    Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. 2014 IEEE Conference on Computer Vision and Pat- tern Recognition, pages 3606–3613, 2013. 4, 6

  2. [10]

    Ng, and Honglak Lee

    Adam Coates, A. Ng, and Honglak Lee. An analysis of single-layer networks in unsupervised feature learning. InIn- ternational Conference on Artificial Intelligence and Statis- tics, 2011. 6

  3. [11]

    Compacter: Efficient low-rank hypercomplex adapter layers

    Joe Davison. Compacter: Efficient low-rank hypercomplex adapter layers. In Neural Information Processing Systems ,

  4. [12]

    Li, and Li Fei-Fei

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, K. Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009. 5

  5. [13]

    The mnist database of handwritten digit images for machine learning research [best of the web]

    Li Deng. The mnist database of handwritten digit images for machine learning research [best of the web]. IEEE Signal Processing Magazine, 29:141–142, 2012. 6

  6. [14]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...

  7. [15]

    Data determines distributional robustness in contrastive language image pre-training (clip)

    Alex Fang, Gabriel Ilharco, Mitchell Wortsman, Yu Wan, Vaishaal Shankar, Achal Dave, and Ludwig Schmidt. Data determines distributional robustness in contrastive language image pre-training (clip). In International Conference on Machine Learning, 2022. 2

  8. [16]

    Mixture-of-loras: An efficient multitask tuning method for large language models

    Wenfeng Feng, Chuzhan Hao, Yuewei Zhang, Yu Han, and Hao Wang. Mixture-of-loras: An efficient multitask tuning method for large language models. ArXiv, abs/2403.03432,

  9. [17]

    Finetune like you pretrain: Im- proved finetuning of zero-shot vision models

    Sachin Goyal, Ananya Kumar, Sankalp Garg, Zico Kolter, and Aditi Raghunathan. Finetune like you pretrain: Im- proved finetuning of zero-shot vision models. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19338–19347, 2022. 2, 3

  10. [18]

    Richemond, Elena Buchatskaya, Carl Do- ersch, Bernardo ´Avila Pires, Zhaohan Daniel Guo, Mo- hammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, R´emi Munos, and Michal Valko

    Jean-Bastien Grill, Florian Strub, Florent Altch’e, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Do- ersch, Bernardo ´Avila Pires, Zhaohan Daniel Guo, Mo- hammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, R´emi Munos, and Michal Valko. Bootstrap your own ...

  11. [19]

    Anchor- based robust finetuning of vision-language models

    Jinwei Han, Zhiwen Lin, Zhongyi Sun, Yingguo Gao, Ke Yan, Shouhong Ding, Yuan Gao, and Gui-Song Xia. Anchor- based robust finetuning of vision-language models. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26909–26918, 2024. 1, 3

  12. [20]

    Towards a unified view of parameter-efficient transfer learning

    Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg- Kirkpatrick, and Graham Neubig. Towards a unified view of parameter-efficient transfer learning. ArXiv, abs/2110.04366, 2021. 2

  13. [21]

    Zhang, Shaoqing Ren, and Jian Sun

    Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. 2016 IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) , pages 770–778, 2015. 1

  14. [22]

    Natural adversarial exam- ples

    Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Stein- hardt, and Dawn Xiaodong Song. Natural adversarial exam- ples. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15257–15266, 2019. 5

  15. [23]

    The many faces of robust- ness: A critical analysis of out-of-distribution generalization

    Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kada- vath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Lixuan Zhu, Samyak Parajuli, Mike Guo, Dawn Xiaodong Song, Ja- cob Steinhardt, and Justin Gilmer. The many faces of robust- ness: A critical analysis of out-of-distribution...

  16. [24]

    Parameter-efficient transfer learning for nlp

    Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for nlp. ArXiv, abs/1902.00751, 2019. 2

  17. [25]

    Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen

    J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. ArXiv, abs/2106.09685, 2021. 1, 2, 3

  18. [26]

    Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V . Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Con- ference on Machine Learning, 2021. 2

  19. [27]

    Belongie, Bharath Hariharan, and Ser Nam Lim

    Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge J. Belongie, Bharath Hariharan, and Ser Nam Lim. Vi- sual prompt tuning. ArXiv, abs/2203.12119, 2022. 2

  20. [28]

    Mora: High- rank updating for parameter-efficient fine-tuning

    Ting Jiang, Shaohan Huang, Shengyue Luo, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang, Deqing Wang, and Fuzhen Zhuang. Mora: High- rank updating for parameter-efficient fine-tuning. ArXiv, abs/2405.12130, 2024. 2

  21. [29]

    Vera: Vector-based random matrix adaptation.ArXiv, abs/2310.11454, 2023

    Dawid Jan Kopiczko, Tijmen Blankevoort, and Yuki Markus Asano. Vera: Vector-based random matrix adaptation.ArXiv, abs/2310.11454, 2023. 2

  22. [30]

    Learning multiple layers of features from tiny images

    Alex Krizhevsky. Learning multiple layers of features from tiny images. 2009. 4, 6

  23. [31]

    Fine-tuning can distort pretrained features and underperform out-of-distribution

    Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. Fine-tuning can distort pretrained features and underperform out-of-distribution. ArXiv, abs/2202.10054, 2022. 1, 3

  24. [32]

    The power of scale for parameter-efficient prompt tuning

    Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In Confer- ence on Empirical Methods in Natural Language Processing,

  25. [33]

    Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

    Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, 2014. 6

  26. [34]

    Grounding dino: Marrying dino with grounded pre-training for open-set object detec- tion

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chun yue Li, Jianwei Yang, Hang Su, Jun-Juan Zhu, and Lei Zhang. Grounding dino: Marrying dino with grounded pre-training for open-set object detec- tion. ArXiv, abs/2303.05499, 2023. 3

  27. [35]

    Blaschko, and Andrea Vedaldi

    Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew B. Blaschko, and Andrea Vedaldi. Fine-grained visual classi- fication of aircraft. ArXiv, abs/1306.5151, 2013. 6

  28. [36]

    Saft: Towards out-of-distribution generalization in fine-tuning

    Bac Nguyen, Stefan Uhlich, Fabien Cardinaux, Lukas Mauch, Marzieh Edraki, and Aaron Courville. Saft: Towards out-of-distribution generalization in fine-tuning. ArXiv, abs/2407.03036, 2024. 3

  29. [37]

    Continual vision-language representation learning with off-diagonal information

    Zixuan Ni, Longhui Wei, Siliang Tang, Yueting Zhuang, and Qi Tian. Continual vision-language representation learning with off-diagonal information. In International Conference on Machine Learning, 2023. 3

  30. [38]

    Automated flower classification over a large number of classes

    Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing, pages 722–729, 2008. 4, 6

  31. [39]

    Parkhi, Andrea Vedaldi, Andrew Zisserman, and C

    Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V . Jawahar. Cats and dogs. 2012 IEEE Conference on Computer Vision and Pattern Recognition , pages 3498– 3505, 2012. 4, 6

  32. [40]

    Girshick, Piotr Doll ´ar, Trevor Dar- rell, and Bharath Hariharan

    Deepak Pathak, Ross B. Girshick, Piotr Doll ´ar, Trevor Dar- rell, and Bharath Hariharan. Learning features by watching objects move. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 6024–6033, 2016. 3

  33. [41]

    Plummer, Liwei Wang, Christopher M

    Bryan A. Plummer, Liwei Wang, Christopher M. Cervantes, Juan C. Caicedo, J. Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspon- dences for richer image-to-sentence models. International Journal of Computer Vision, 123:74 – 93, 2015. 6

  34. [42]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In International Con...

  35. [43]

    Residual prompt tuning: Improving prompt tuning with residual reparameterization

    Anastasia Razdaibiedina, Yuning Mao, Rui Hou, Madian Khabsa, Mike Lewis, Jimmy Ba, and Amjad Almahairi. Residual prompt tuning: Improving prompt tuning with residual reparameterization. In Annual Meeting of the As- sociation for Computational Linguistics, 2023. 2

  36. [44]

    Do imagenet classifiers generalize to im- agenet? In International Conference on Machine Learning,

    Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to im- agenet? In International Conference on Machine Learning,

  37. [45]

    Tied-lora: Enhancing parameter efficiency of lora with weight tying

    Adithya Renduchintala, Tugrul Konuk, and Oleksii Kuchaiev. Tied-lora: Enhancing parameter efficiency of lora with weight tying. In North American Chapter of the Association for Computational Linguistics, 2023. 2

  38. [46]

    Laion-400m: Open dataset of clip-filtered 400 million image-text pairs

    Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. ArXiv, abs/2111.02114, 2021. 2

  39. [47]

    Ziplora: Any subject in any style by effectively merging loras.ArXiv, abs/2311.13600, 2023

    Viraj Shah, Nataniel Ruiz, Forrester Cole, Erika Lu, Svet- lana Lazebnik, Yuanzhen Li, and Varun Jampani. Ziplora: Any subject in any style by effectively merging loras.ArXiv, abs/2311.13600, 2023. 2

  40. [48]

    Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

    Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning. In An- nual Meeting of the Association for Computational Linguis- tics, 2018. 2

  41. [49]

    Improved deep metric learning with multi- class n-pair loss objective

    Kihyuk Sohn. Improved deep metric learning with multi- class n-pair loss objective. InNeural Information Processing Systems, 2016. 2

  42. [50]

    A closer look at the robustness of contrastive language-image pre-training (clip)

    Weijie Tu, Weijian Deng, and Tom Gedeon. A closer look at the robustness of contrastive language-image pre-training (clip). ArXiv, abs/2402.07410, 2024. 1

  43. [51]

    van de Ven, Nicholas Soures, and Dhireesha Ku- dithipudi

    Gido M. van de Ven, Nicholas Soures, and Dhireesha Ku- dithipudi. Continual learning and catastrophic forgetting. ArXiv, abs/2403.05175, 2024. 1

  44. [52]

    Repre- sentation learning with contrastive predictive coding

    A ¨aron van den Oord, Yazhe Li, and Oriol Vinyals. Repre- sentation learning with contrastive predictive coding. ArXiv, abs/1807.03748, 2018. 2

  45. [53]

    Fowlkes, Rahul Bhotika, and Ste- fan 0 Soatto

    Matthew Wallingford, Hao Li, Alessandro Achille, Avinash Ravichandran, Charless C. Fowlkes, Rahul Bhotika, and Ste- fan 0 Soatto. Task adaptive parameter sharing for multi-task learning. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 7551–75...

  46. [54]

    Xing, and Zachary Chase Lipton

    Haohan Wang, Songwei Ge, Eric P. Xing, and Zachary Chase Lipton. Learning robust global representations by penalizing local predictive power. In Neural Information Processing Systems, 2019. 6

  47. [55]

    Do clips always generalize better than imagenet models? ArXiv, abs/2403.11497, 2024

    Qizhou Wang, Yong Lin, Yongqiang Chen, Ludwig Schmidt, Bo Han, and Tong Zhang. Do clips always generalize better than imagenet models? ArXiv, abs/2403.11497, 2024. 3

  48. [56]

    Robust fine-tuning of zero-shot models

    Mitchell Wortsman, Gabriel Ilharco, Mike Li, Jong Wook Kim, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, and Ludwig Schmidt. Robust fine-tuning of zero-shot models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 7949–7961, 2021. 1, 2, 3

  49. [57]

    A practi- cal contrastive learning framework for single-image super- resolution

    Gang Wu, Junjun Jiang, and Xianming Liu. A practi- cal contrastive learning framework for single-image super- resolution. IEEE Transactions on Neural Networks and Learning Systems, 2023. 3

  50. [58]

    Ehinger, Aude Oliva, and Antonio Torralba

    Jianxiong Xiao, James Hays, Krista A. Ehinger, Aude Oliva, and Antonio Torralba. Sun database: Large-scale scene recognition from abbey to zoo. 2010 IEEE Computer Soci- ety Conference on Computer Vision and Pattern Recognition, pages 3485–3492, 2010. 6

  51. [59]

    Detco: Unsuper- vised contrastive learning for object detection

    Enze Xie, Jian Ding, Wenhai Wang, Xiaohang Zhan, Hang Xu, Peize Sun, Zhenguo Li, and Ping Luo. Detco: Unsuper- vised contrastive learning for object detection. In Proceed- ings of the IEEE/CVF international conference on computer vision, pages 8392–8401, 2021. 3

  52. [60]

    In- stance localization for self-supervised detection pretraining

    Ceyuan Yang, Zhirong Wu, Bolei Zhou, and Stephen Lin. In- stance localization for self-supervised detection pretraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3987–3996, 2021. 3

  53. [61]

    Dora: Weight-decomposed low-rank adaptation

    Shih yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, and Min-Hung Chen. Dora: Weight-decomposed low-rank adaptation. ArXiv, abs/2402.09353, 2024. 1, 2

  54. [62]

    Lyu, Shuai Zhang, S

    Penghang Yin, J. Lyu, Shuai Zhang, S. Osher, Yingyong Qi, and Jack Xin. Understanding straight-through esti- mator in training activation quantized neural nets. ArXiv, abs/1903.05662, 2019. 2

  55. [63]

    Low-rank few-shot adaptation of vision-language models.2024 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition Work- shops (CVPRW), pages 1593–1603, 2024

    Maxime Zanella and Ismail Ben Ayed. Low-rank few-shot adaptation of vision-language models.2024 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition Work- shops (CVPRW), pages 1593–1603, 2024. 1

  56. [64]

    Adalora: Adaptive budget allocation for parameter-efficient fine-tuning

    Qingru Zhang, Minshuo Chen, Alexander Bukharin, Nikos Karampatziakis, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. Adalora: Adaptive budget allocation for parameter-efficient fine-tuning. 2023. 2

  57. [65]

    Preventing zero-shot transfer degradation in continual learning of vision-language mod- els

    Zangwei Zheng, Mingyu Ma, Kai Wang, Ziheng Qin, Xi- angyu Yue, and Yang You. Preventing zero-shot transfer degradation in continual learning of vision-language mod- els. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 19068–19079, 2023. 1

  58. [66]

    Multi-lora composition for image generation

    Ming Zhong, Yelong Shen, Shuohang Wang, Yadong Lu, Yizhu Jiao, Siru Ouyang, Donghan Yu, Jiawei Han, and Weizhu Chen. Multi-lora composition for image generation. ArXiv, abs/2402.16843, 2024. 2 Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation Supple...

  59. [67]

    Providing the hyperparameters used in experiments for vision and vision-language models

  60. [68]

    Presenting the numerical results corresponding to the figures included in the main paper. B. Hyperparameters For experiments with DINO, the model was trained for 5000 steps, with evaluations every 100 steps. The best check- point, determined by validation set performance, was ...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.