REVIEW 4 major objections 4 minor 68 references
Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation
T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Fine-tuning with a small subset of LoRA blocks active can match full-LoRA target accuracy while preserving zero-shot and out-of-distribution knowledge.
desk verdict A broad, mostly consistent empirical study of gated LoRA for CLIP/DINO, but the headline CLIP claim needs a random-selection control and the unreported threshold tau makes the efficiency numbers unverifiable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a per-block gate built from a learnable scalar score $s_\ell$ and a straight-through estimator (STE) indicator $I_\tau(s_\ell)$ that multiplies each low-rank residual $A_\ell B_\ell$, trained with an $\ell^1$ sparsity penalty on the scores. The gate makes the set of active blocks a trainable, sparse choice: the scores decide which low-rank blocks stay on for the target task, and the $\ell^1$ term is what produces the tiny active-block percentages (1.39–6.25% on CLIP, 2.7–44% on DINO depending on $\lambda$, rank, and dataset). The straight-through estimator lets gradients flow through the binary decision so the selection itself is learned, not set by hand.
What would settle it
Repeat the rank-128 CLIP fine-tuning with the paper's equations but sweep the threshold $\tau$ across a wide range, say 0, 0.01, 0.1, and 0.5, while keeping $\lambda=1$; the claim fails if any run either leaves nearly all blocks active, drops ImageNet accuracy below LoRA's 81.77%, or loses the zero-shot retention advantage, since that would show the result depends on the unspecified threshold rather than on the gating mechanism itself.
Extended reading notes
Core claim
The central discovery is that the low-rank updates learned during fine-tuning are highly redundant: a small, task-dependent subset of blocks carries almost all of the adaptation signal, and the rest can be left off without losing target-task accuracy while preserving the pretrained feature space. Formally, the paper updates each block as $W_\ell = W_{0,\ell} + I_\tau(s_\ell)A_\ell B_\ell$, where $I_\tau(s_\ell)=1$ if $s_\ell\ge\tau$ and $0$ otherwise, and regularizes the gate scores with $\lambda\sum_\ell |s_\ell|$. With this update, LoRA at rank 128 on CLIP reaches 81.84% ImageNet-1K accuracy (versus 81.77% for full LoRA) using 6.25% of the blocks; zero-shot classification average rises to 61.48 versus 51.44 for LoRA, zero-shot retrieval loss falls to at most 5.73% versus about 28% for FLYP, and unmerged inference becomes up to 2.9x faster for LoRA and 5x faster for DoRA at rank 256. On DINO-ViT, activating as few as 2.7% of blocks keeps target accuracy comparable to LoRA while source-domain accuracy on ImageNet-100 is retained substantially better than under full fine-tuning.
Load-bearing premise
The load-bearing premise is that the sparsity penalty and an unspecified threshold $\tau$ together yield the reported small active-block counts; if $\tau$ is set differently or depends on the scale of the learned scores, the headline efficiency and forgetting numbers would not reproduce.
Editorial extensions
If this is right
- Fine-tuning a foundation model on a new dataset can be done with only a small fraction of its low-rank blocks active, cutting inference FLOPs and memory proportionally; the paper measures up to 2.9x (LoRA) and 5x (DoRA) faster unmerged inference at rank 256.
- Because the gate applies to any LoRA-variant method (the paper demonstrates DoRA), it is a plug-in that can reduce catastrophic forgetting across PEFT techniques.
- Higher ranks, which normally cause stronger forgetting, become usable: at rank 256 the gated LoRA reaches 82.31% ImageNet accuracy with 5.56% of blocks active, while plain LoRA's out-of-distribution mean drops to 53.81.
- The retained zero-shot abilities mean a fine-tuned CLIP can still serve as a general image-text retriever and zero-shot classifier, not just as a specialist on the target classes.
Reading between the lines
- Inference: the layer-activation maps, which concentrate in feedforward MLP blocks and in later CLIP vision layers, suggest that a fixed or pretrained subset of blocks could be enough, which would remove the need to learn gate scores at all.
- Inference: the reported preservation suggests low-rank updates are redundant across transformer residuals generally, implying selective LoRA could slot into continual-learning pipelines without rehearsal buffers.
- Inference: because $\tau$ is unspecified, a normalized gate such as a hard sigmoid over scores divided by the maximum score would make the method reproducible and remove dependence on the score scale.
- Inference: if sparsity is what preserves knowledge, then the method's benefit should be testable against an ablation that selects the same number of blocks by gradient magnitude or by the paper's block-activation heatmaps rather than training scores.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a PEFT method that extends Task Adaptive Parameter Sharing (TAPS) by adding an indicator function I_tau(s_i) that selectively activates LoRA blocks, with an L1 penalty on the block scores s_i. The central claim is that this gating mechanism allows fine-tuned models to match full-LoRA in-distribution accuracy while using only a small fraction of active blocks, thereby retaining zero-shot and OOD knowledge and reducing inference cost. Experiments are reported on DINO ViT-S/B-16 fine-tuned on six target datasets and on CLIP fine-tuned on ImageNet-1K, evaluated on OOD, zero-shot classification, and zero-shot retrieval benchmarks. The method is also applied to DoRA. The headline result is that CLIP at rank 128 with 6.25% active blocks reaches 81.84% ImageNet accuracy versus 81.77% for full LoRA, while zero-shot classification improves from 51.44 to 61.48 and retrieval retention is substantially better than full LoRA.
Significance. If the selective-activation mechanism is genuinely responsible for the observed preservation of pre-trained knowledge, the method is a simple, practical extension of LoRA-style PEFT with broad applicability. The paper’s strengths are its extensive empirical coverage: multiple backbones (DINO ViT-S/B, CLIP), multiple PEFT variants (LoRA and DoRA), many transfer datasets, several ranks and regularization strengths, and separate evaluations of OOD robustness, zero-shot classification, and zero-shot retrieval. The inference-FLOPs analysis is a useful practical contribution. However, the central scientific claim—that learned selection, rather than merely reducing the number of adapted parameters, yields the retention gains—is not fully supported for the main CLIP experiment, and the reproducibility of the reported active-block percentages is compromised by the unspecified threshold tau. The paper also overstates its contribution in the title and abstract relative to the explicit limitation stated in Section 5.
major comments (4)
- [Sec. 3.2, Eq. (3)] This is a load-bearing issue: the claimed efficiency gains and the claimed trade-off curves depend on where the threshold sits.
- [Sec. 4.2, Fig. 6 and Table 10] The central claim 'selective activation preserves knowledge' is not distinguishable from 'fewer adapted parameters preserve knowledge' without this control.
- [Sec. 5, Conclusion] The current framing overclaims the contribution relative to what is measured in the experiments.
- [Tables 10-15] This concern is particularly relevant to the DINO random-selection comparison in Table 1, where the claimed advantage is about 2%.
minor comments (4)
- [Supplementary Table 4] The Linear row in Table 4 lists CIFAR-100=72.07, IN-100=87.20, Mean=88.32, but the mean of 72.07 and 87.20 is 79.64. This appears to be a copy-paste error from Table 3 (where the mean 88.45 is consistent with the CIFAR-10 numbers). Please correct.
- [Sec. 2, Related Works] Given that AdaLoRA is cited as an adaptive rank-allocation method, a direct empirical comparison with AdaLoRA (or a brief explanation of why it is not compared) would help position the contribution. Currently the method is compared only with full LoRA/DoRA and with full fine-tuning/linear probing.
- [Sec. 4.3, Fig. 9] The FLOPs comparison is presented for the unmerged setting where inactive blocks are skipped. It would be helpful to state explicitly whether the reported speedups assume that active LoRA blocks are not merged into the base weights, since merging would eliminate the inference benefit; the current text implies but does not state this.
- [Sec. 4.4, Fig. 10] The figure caption says 'across 8 runs' for CLIP, but the text describes 7 runs at one lambda plus 6 runs at other lambdas. Please reconcile the count and clarify whether the shown activations are pooled over all runs or averaged per configuration.
Circularity Check
No significant circularity: the central LoRA-plus-indicator result is measured against held-out OOD and zero-shot benchmarks, and the self-citations only supply setup and inspiration, not the reported numbers.
full rationale
The paper's derivation chain is empirical rather than definitional. Equation (2) defines the adapted weight as W_l = W_0,l + I_tau(s_i) A_l B_l, and Equation (3) defines the indicator via an unspecified threshold tau, but the headline claims (e.g., 6.25% active blocks matching LoRA on ImageNet, zero-shot retention within 5.73%) are measured outcomes on held-out benchmarks, not quantities derived from the definition. There is no fitted parameter that is later renamed as a prediction: the reported active-block percentages are post-training counts, and the ID/OOD/zero-shot accuracies come from external evaluation datasets, so the central comparisons are not forced by construction. The paper does cite the authors' own prior work: reference [3] (same first and last authors) is used for the problem setup, namely 'Following [3]', and reference [53] (TAPS, with overlapping author Avinash Ravichandran) is the stated inspiration for the indicator function. These self-citations provide framing and a borrowed mechanism, but they do not by themselves imply the reported accuracy or retention numbers; the method is additionally supported by a random-selection control (Table 1) on DINO/CIFAR-100, even though no such control is given for the headline CLIP experiment. The conclusion's admission that the method 'does not directly address catastrophic forgetting' is an internal inconsistency with the title and abstract, but inconsistency is not circularity. The missing threshold tau is a reproducibility concern, not a circularity concern: it affects whether the exact 6.25% figure can be audited, but it does not make any prediction equal to an input by definition. Overall, the central claim retains independent empirical content, so the circularity score is low.
Assumptions & free parameters
free parameters (4)
- threshold tau =
not reported
- regularization lambda =
0.1, 0.5, 1.0, 1.5, chosen per dataset/rank
- LoRA rank =
4, 8, 16, 32, 64, 128, 256
- K-NN K =
20
assumptions (4)
- standard math Straight-through estimator provides usable gradients for the discrete indicator
- domain assumption Pretrained CLIP and DINO features are the knowledge to preserve; adapting a few LoRA blocks retains them
- domain assumption K-NN on ImageNet-100 (DINO) and zero-shot retrieval on COCO/Flickr (CLIP) measure catastrophic forgetting
- domain assumption Inference efficiency is assessed with LoRA/DoRA adapters left unmerged
Cite this review
Pith. "Pith review of Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation." pith.science (2026). https://pith.science/paper/4UKPXFMQ
@misc{pith2026250115377,
author = {Pith},
title = {Pith review of: Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation},
year = {2026},
howpublished = {\url{https://pith.science/paper/4UKPXFMQ}},
note = {Machine review of arXiv:2501.15377}
}
read the original abstract
Adapting deep learning models to new domains often requires computationally intensive retraining and risks catastrophic forgetting. While fine-tuning enables domain-specific adaptation, it can reduce robustness to distribution shifts, impacting out-of-distribution (OOD) performance. Pre-trained zero-shot models like CLIP offer strong generalization but may suffer degraded robustness after fine-tuning. Building on Task Adaptive Parameter Sharing (TAPS), we propose a simple yet effective extension as a parameter-efficient fine-tuning (PEFT) method, using an indicator function to selectively activate Low-Rank Adaptation (LoRA) blocks. Our approach minimizes knowledge loss, retains its generalization strengths under domain shifts, and significantly reduces computational costs compared to traditional fine-tuning. We demonstrate that effective fine-tuning can be achieved with as few as 5\% of active blocks, substantially improving efficiency. Evaluations on pre-trained models such as CLIP and DINO-ViT demonstrate our method's broad applicability and effectiveness in maintaining performance and knowledge retention.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Agiza, Marina Neseem, and Sherief Reda
Ahmed A. Agiza, Marina Neseem, and Sherief Reda. Mt- lora: A low-rank adaptation approach for efficient multi-task learning. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16196–16205, 2024. 2
work page 2024
-
[2]
data2vec: A general framework for self-supervised learning in speech, vision and language
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. data2vec: A general framework for self-supervised learning in speech, vision and language. ArXiv, abs/2202.03555, 2022. 2
arXiv 2022
-
[3]
Parameter efficient fine-tuning of self-supervised vits without catastrophic forgetting
Reza Akbarian Bafghi, Nidhin Harilal, Claire Monteleoni, and Maziar Raissi. Parameter efficient fine-tuning of self-supervised vits without catastrophic forgetting. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 3679–3684, 2024. 2, 3, 4
work page 2024
-
[4]
Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Joshua B. Tenenbaum, and Boris Katz. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In Neural Information Processing Systems, 2019. 6
work page 2019
-
[5]
Dan Biderman, Jose Gonzalez Ortiz, Jacob Portes, Mansheej Paul, Philip Greengard, Connor Jennings, Daniel King, Sam Havens, Vitaliy Chiley, Jonathan Frankle, Cody Blakeney, and John P. Cunningham. Lora learns less and forgets less. ArXiv, abs/2405.09673, 2024. 2
arXiv 2024
-
[6]
Food-101 - mining discriminative components with random forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101 - mining discriminative components with random forests. In European Conference on Computer Vision, 2014. 4, 6
work page 2014
-
[7]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 2, 4
work page 2021
-
[8]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Ge- offrey E. Hinton. A simple framework for contrastive learn- ing of visual representations. ArXiv, abs/2002.05709, 2020. 2
arXiv 2002
Show all 68 references
-
[9]
Describing textures in the wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. 2014 IEEE Conference on Computer Vision and Pat- tern Recognition, pages 3606–3613, 2013. 4, 6
2014
-
[10]
Ng, and Honglak Lee
Adam Coates, A. Ng, and Honglak Lee. An analysis of single-layer networks in unsupervised feature learning. InIn- ternational Conference on Artificial Intelligence and Statis- tics, 2011. 6
2011
-
[11]
Compacter: Efficient low-rank hypercomplex adapter layers
Joe Davison. Compacter: Efficient low-rank hypercomplex adapter layers. In Neural Information Processing Systems ,
-
[12]
Li, and Li Fei-Fei
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, K. Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009. 5
2009
-
[13]
The mnist database of handwritten digit images for machine learning research [best of the web]
Li Deng. The mnist database of handwritten digit images for machine learning research [best of the web]. IEEE Signal Processing Magazine, 29:141–142, 2012. 6
2012
-
[14]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...
2010 arXiv
-
[15]
Data determines distributional robustness in contrastive language image pre-training (clip)
Alex Fang, Gabriel Ilharco, Mitchell Wortsman, Yu Wan, Vaishaal Shankar, Achal Dave, and Ludwig Schmidt. Data determines distributional robustness in contrastive language image pre-training (clip). In International Conference on Machine Learning, 2022. 2
2022
-
[16]
Mixture-of-loras: An efficient multitask tuning method for large language models
Wenfeng Feng, Chuzhan Hao, Yuewei Zhang, Yu Han, and Hao Wang. Mixture-of-loras: An efficient multitask tuning method for large language models. ArXiv, abs/2403.03432,
-
[17]
Finetune like you pretrain: Im- proved finetuning of zero-shot vision models
Sachin Goyal, Ananya Kumar, Sankalp Garg, Zico Kolter, and Aditi Raghunathan. Finetune like you pretrain: Im- proved finetuning of zero-shot vision models. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19338–19347, 2022. 2, 3
2023
-
[18]
Richemond, Elena Buchatskaya, Carl Do- ersch, Bernardo ´Avila Pires, Zhaohan Daniel Guo, Mo- hammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, R´emi Munos, and Michal Valko
Jean-Bastien Grill, Florian Strub, Florent Altch’e, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Do- ersch, Bernardo ´Avila Pires, Zhaohan Daniel Guo, Mo- hammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, R´emi Munos, and Michal Valko. Bootstrap your own ...
2006 arXiv
-
[19]
Anchor- based robust finetuning of vision-language models
Jinwei Han, Zhiwen Lin, Zhongyi Sun, Yingguo Gao, Ke Yan, Shouhong Ding, Yuan Gao, and Gui-Song Xia. Anchor- based robust finetuning of vision-language models. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26909–26918, 2024. 1, 3
2024
-
[20]
Towards a unified view of parameter-efficient transfer learning
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg- Kirkpatrick, and Graham Neubig. Towards a unified view of parameter-efficient transfer learning. ArXiv, abs/2110.04366, 2021. 2
2021 arXiv
-
[21]
Zhang, Shaoqing Ren, and Jian Sun
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. 2016 IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) , pages 770–778, 2015. 1
2016
-
[22]
Natural adversarial exam- ples
Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Stein- hardt, and Dawn Xiaodong Song. Natural adversarial exam- ples. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15257–15266, 2019. 5
2021
-
[23]
The many faces of robust- ness: A critical analysis of out-of-distribution generalization
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kada- vath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Lixuan Zhu, Samyak Parajuli, Mike Guo, Dawn Xiaodong Song, Ja- cob Steinhardt, and Justin Gilmer. The many faces of robust- ness: A critical analysis of out-of-distribution...
2021
-
[24]
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for nlp. ArXiv, abs/1902.00751, 2019. 2
1902 arXiv
-
[25]
Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen
J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. ArXiv, abs/2106.09685, 2021. 1, 2, 3
2021 arXiv
-
[26]
Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V . Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Con- ference on Machine Learning, 2021. 2
2021
-
[27]
Belongie, Bharath Hariharan, and Ser Nam Lim
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge J. Belongie, Bharath Hariharan, and Ser Nam Lim. Vi- sual prompt tuning. ArXiv, abs/2203.12119, 2022. 2
2022 arXiv
-
[28]
Mora: High- rank updating for parameter-efficient fine-tuning
Ting Jiang, Shaohan Huang, Shengyue Luo, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang, Deqing Wang, and Fuzhen Zhuang. Mora: High- rank updating for parameter-efficient fine-tuning. ArXiv, abs/2405.12130, 2024. 2
2024 arXiv
-
[29]
Vera: Vector-based random matrix adaptation.ArXiv, abs/2310.11454, 2023
Dawid Jan Kopiczko, Tijmen Blankevoort, and Yuki Markus Asano. Vera: Vector-based random matrix adaptation.ArXiv, abs/2310.11454, 2023. 2
2023 arXiv
-
[30]
Learning multiple layers of features from tiny images
Alex Krizhevsky. Learning multiple layers of features from tiny images. 2009. 4, 6
2009
-
[31]
Fine-tuning can distort pretrained features and underperform out-of-distribution
Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. Fine-tuning can distort pretrained features and underperform out-of-distribution. ArXiv, abs/2202.10054, 2022. 1, 3
2022 arXiv
-
[32]
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In Confer- ence on Empirical Methods in Natural Language Processing,
-
[33]
Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, 2014. 6
2014
-
[34]
Grounding dino: Marrying dino with grounded pre-training for open-set object detec- tion
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chun yue Li, Jianwei Yang, Hang Su, Jun-Juan Zhu, and Lei Zhang. Grounding dino: Marrying dino with grounded pre-training for open-set object detec- tion. ArXiv, abs/2303.05499, 2023. 3
2023 arXiv
-
[35]
Blaschko, and Andrea Vedaldi
Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew B. Blaschko, and Andrea Vedaldi. Fine-grained visual classi- fication of aircraft. ArXiv, abs/1306.5151, 2013. 6
2013 arXiv
-
[36]
Saft: Towards out-of-distribution generalization in fine-tuning
Bac Nguyen, Stefan Uhlich, Fabien Cardinaux, Lukas Mauch, Marzieh Edraki, and Aaron Courville. Saft: Towards out-of-distribution generalization in fine-tuning. ArXiv, abs/2407.03036, 2024. 3
2024 arXiv
-
[37]
Continual vision-language representation learning with off-diagonal information
Zixuan Ni, Longhui Wei, Siliang Tang, Yueting Zhuang, and Qi Tian. Continual vision-language representation learning with off-diagonal information. In International Conference on Machine Learning, 2023. 3
2023
-
[38]
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing, pages 722–729, 2008. 4, 6
2008
-
[39]
Parkhi, Andrea Vedaldi, Andrew Zisserman, and C
Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V . Jawahar. Cats and dogs. 2012 IEEE Conference on Computer Vision and Pattern Recognition , pages 3498– 3505, 2012. 4, 6
2012
-
[40]
Girshick, Piotr Doll ´ar, Trevor Dar- rell, and Bharath Hariharan
Deepak Pathak, Ross B. Girshick, Piotr Doll ´ar, Trevor Dar- rell, and Bharath Hariharan. Learning features by watching objects move. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 6024–6033, 2016. 3
2017
-
[41]
Plummer, Liwei Wang, Christopher M
Bryan A. Plummer, Liwei Wang, Christopher M. Cervantes, Juan C. Caicedo, J. Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspon- dences for richer image-to-sentence models. International Journal of Computer Vision, 123:74 – 93, 2015. 6
2015
-
[42]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In International Con...
2021
-
[43]
Residual prompt tuning: Improving prompt tuning with residual reparameterization
Anastasia Razdaibiedina, Yuning Mao, Rui Hou, Madian Khabsa, Mike Lewis, Jimmy Ba, and Amjad Almahairi. Residual prompt tuning: Improving prompt tuning with residual reparameterization. In Annual Meeting of the As- sociation for Computational Linguistics, 2023. 2
2023
-
[44]
Do imagenet classifiers generalize to im- agenet? In International Conference on Machine Learning,
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to im- agenet? In International Conference on Machine Learning,
-
[45]
Tied-lora: Enhancing parameter efficiency of lora with weight tying
Adithya Renduchintala, Tugrul Konuk, and Oleksii Kuchaiev. Tied-lora: Enhancing parameter efficiency of lora with weight tying. In North American Chapter of the Association for Computational Linguistics, 2023. 2
2023
-
[46]
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. ArXiv, abs/2111.02114, 2021. 2
2021 arXiv
-
[47]
Ziplora: Any subject in any style by effectively merging loras.ArXiv, abs/2311.13600, 2023
Viraj Shah, Nataniel Ruiz, Forrester Cole, Erika Lu, Svet- lana Lazebnik, Yuanzhen Li, and Varun Jampani. Ziplora: Any subject in any style by effectively merging loras.ArXiv, abs/2311.13600, 2023. 2
2023
-
[48]
Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning. In An- nual Meeting of the Association for Computational Linguis- tics, 2018. 2
2018
-
[49]
Improved deep metric learning with multi- class n-pair loss objective
Kihyuk Sohn. Improved deep metric learning with multi- class n-pair loss objective. InNeural Information Processing Systems, 2016. 2
2016
-
[50]
A closer look at the robustness of contrastive language-image pre-training (clip)
Weijie Tu, Weijian Deng, and Tom Gedeon. A closer look at the robustness of contrastive language-image pre-training (clip). ArXiv, abs/2402.07410, 2024. 1
2024 arXiv
-
[51]
van de Ven, Nicholas Soures, and Dhireesha Ku- dithipudi
Gido M. van de Ven, Nicholas Soures, and Dhireesha Ku- dithipudi. Continual learning and catastrophic forgetting. ArXiv, abs/2403.05175, 2024. 1
2024 arXiv
-
[52]
Repre- sentation learning with contrastive predictive coding
A ¨aron van den Oord, Yazhe Li, and Oriol Vinyals. Repre- sentation learning with contrastive predictive coding. ArXiv, abs/1807.03748, 2018. 2
2018 arXiv
-
[53]
Fowlkes, Rahul Bhotika, and Ste- fan 0 Soatto
Matthew Wallingford, Hao Li, Alessandro Achille, Avinash Ravichandran, Charless C. Fowlkes, Rahul Bhotika, and Ste- fan 0 Soatto. Task adaptive parameter sharing for multi-task learning. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 7551–75...
2022
-
[54]
Xing, and Zachary Chase Lipton
Haohan Wang, Songwei Ge, Eric P. Xing, and Zachary Chase Lipton. Learning robust global representations by penalizing local predictive power. In Neural Information Processing Systems, 2019. 6
2019
-
[55]
Do clips always generalize better than imagenet models? ArXiv, abs/2403.11497, 2024
Qizhou Wang, Yong Lin, Yongqiang Chen, Ludwig Schmidt, Bo Han, and Tong Zhang. Do clips always generalize better than imagenet models? ArXiv, abs/2403.11497, 2024. 3
2024 arXiv
-
[56]
Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Mike Li, Jong Wook Kim, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, and Ludwig Schmidt. Robust fine-tuning of zero-shot models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 7949–7961, 2021. 1, 2, 3
2022
-
[57]
A practi- cal contrastive learning framework for single-image super- resolution
Gang Wu, Junjun Jiang, and Xianming Liu. A practi- cal contrastive learning framework for single-image super- resolution. IEEE Transactions on Neural Networks and Learning Systems, 2023. 3
2023
-
[58]
Ehinger, Aude Oliva, and Antonio Torralba
Jianxiong Xiao, James Hays, Krista A. Ehinger, Aude Oliva, and Antonio Torralba. Sun database: Large-scale scene recognition from abbey to zoo. 2010 IEEE Computer Soci- ety Conference on Computer Vision and Pattern Recognition, pages 3485–3492, 2010. 6
2010
-
[59]
Detco: Unsuper- vised contrastive learning for object detection
Enze Xie, Jian Ding, Wenhai Wang, Xiaohang Zhan, Hang Xu, Peize Sun, Zhenguo Li, and Ping Luo. Detco: Unsuper- vised contrastive learning for object detection. In Proceed- ings of the IEEE/CVF international conference on computer vision, pages 8392–8401, 2021. 3
2021
-
[60]
In- stance localization for self-supervised detection pretraining
Ceyuan Yang, Zhirong Wu, Bolei Zhou, and Stephen Lin. In- stance localization for self-supervised detection pretraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3987–3996, 2021. 3
2021
-
[61]
Dora: Weight-decomposed low-rank adaptation
Shih yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, and Min-Hung Chen. Dora: Weight-decomposed low-rank adaptation. ArXiv, abs/2402.09353, 2024. 1, 2
2024 arXiv
-
[62]
Lyu, Shuai Zhang, S
Penghang Yin, J. Lyu, Shuai Zhang, S. Osher, Yingyong Qi, and Jack Xin. Understanding straight-through esti- mator in training activation quantized neural nets. ArXiv, abs/1903.05662, 2019. 2
1903 arXiv
-
[63]
Low-rank few-shot adaptation of vision-language models.2024 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition Work- shops (CVPRW), pages 1593–1603, 2024
Maxime Zanella and Ismail Ben Ayed. Low-rank few-shot adaptation of vision-language models.2024 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition Work- shops (CVPRW), pages 1593–1603, 2024. 1
2024
-
[64]
Adalora: Adaptive budget allocation for parameter-efficient fine-tuning
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Nikos Karampatziakis, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. Adalora: Adaptive budget allocation for parameter-efficient fine-tuning. 2023. 2
2023
-
[65]
Preventing zero-shot transfer degradation in continual learning of vision-language mod- els
Zangwei Zheng, Mingyu Ma, Kai Wang, Ziheng Qin, Xi- angyu Yue, and Yang You. Preventing zero-shot transfer degradation in continual learning of vision-language mod- els. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 19068–19079, 2023. 1
2023
-
[66]
Multi-lora composition for image generation
Ming Zhong, Yelong Shen, Shuohang Wang, Yadong Lu, Yizhu Jiao, Siru Ouyang, Donghan Yu, Jiawei Han, and Weizhu Chen. Multi-lora composition for image generation. ArXiv, abs/2402.16843, 2024. 2 Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation Supple...
2024 arXiv
-
[67]
Providing the hyperparameters used in experiments for vision and vision-language models
-
[68]
Presenting the numerical results corresponding to the figures included in the main paper. B. Hyperparameters For experiments with DINO, the model was trained for 5000 steps, with evaluations every 100 steps. The best check- point, determined by validation set performance, was ...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.