REVIEW 4 major objections 5 minor 1 cited by
Revisiting LoRA through the Lens of Parameter Redundancy: Spectral Encoding Helps
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read LoRA tolerates 60% parameter masking at fixed rank; reparameterizing the survivors through inverse Fourier or wavelet transforms yields SeLoRA, which beats LoRA, DoRA, and HiRA on reasoning and code with fewer parameters.
desk verdict Useful PEFT method with honest limitations, but single-run results and an untested random mask leave the headline gains underdetermined. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the reparameterization map $\mathcal{T}$ applied to a sparsely masked spectral matrix: $\tilde{A} = \mathcal{T}(F_A)$ and $\tilde{B} = \mathcal{T}(F_B)$, where only the entries of $F_A, F_B$ at a fixed random index set $\Omega$ are trainable and all others are frozen at zero, with $|\Omega| = \lfloor(1-\eta)rd\rfloor$ controlled by the sparse ratio $\eta$. The map converts a sparse set of spectral coefficients into a dense spatial-domain matrix, so the adapter keeps LoRA's low-rank update structure while drawing its parameters from a spectral subspace. Initialization scales $F_A$ so that $\text{Var}(\mathcal{T}(F_A))$ matches a Xavier- or Kaiming-initialized auxiliary matrix, while $\tilde{B}$ starts at zero, preserving LoRA's training stability. The machinery does two jobs at once: the masking experiments establish the sparsity property, and the spectral basis supplies the inductive bias that lets a sparse coefficient set reconstruct a high-quality dense adapter.
What would settle it
Run the Figure 1 experiment as a sweep: fine-tune LLaMA-2-7B and LLaMA-3-8B at rank 32 with random masks of 20, 40, 60, and 80 percent across at least five seeds, and additionally redraw the shared index set $\Omega$ several times at a fixed sparsity level. If dense-to-masked accuracy differences at 60 percent exceed seed-level noise, or if accuracy varies substantially across draws of $\Omega$, the sparsity property and SeLoRA's motivation fail; stable parity would confirm the paper's central claim.
Extended reading notes
Core claim
The central claim is that a low-rank adapter can be learned inside a sparse spectral subspace without losing expressiveness: the update is written as $W' = W_0 + \tilde{B}\tilde{A}$ with $\tilde{A} = \mathcal{T}(F_A)$ and $\tilde{B} = \mathcal{T}(F_B)$, where $F_A$ and $F_B$ are spectral matrices whose learnable entries sit only at a randomly chosen index set $\Omega$ shared across all modules, and $\mathcal{T}$ is either the real part of the inverse 2D discrete Fourier transform or the inverse 2D wavelet transform (Haar by default). At rank 32, both SeLoRA variants match or exceed LoRA's accuracy on eight commonsense benchmarks while using about 0.28-0.5 percent of model parameters versus LoRA's 0.70-0.83 percent, and the wavelet variant raises average accuracy on math and code benchmarks by roughly 2 points. The paper further shows the plug-in property: wrapping DoRA and HiRA in the same spectral reparameterization (SeDoRA, SeHiRA) improves those baselines as well. A subspace analysis reports that SeLoRA's updates amplify task-relevant directions of $W$ less strongly than LoRA's, cutting the reverse amplification factor from 0.17 to 0.04, which the authors read as evidence that spectral encoding changes what LoRA learns, not merely how it is stored.
Load-bearing premise
The argument rests on the claim that randomly masking up to 60 percent of LoRA's parameters at rank 32 genuinely costs nothing in accuracy; if that parity is a fluke of one seed, one rank, or the particular sparsity levels chosen, then learning through a fixed random spectral mask has no reason to work, and the paper reports no sensitivity analysis for the random index set $\Omega$ that every module shares.
Editorial extensions
If this is right
- Under the sparsity property, density rather than rank is the right lever for shrinking LoRA: at rank 32, SeLoRA holds LoRA-grade accuracy with roughly 40 percent fewer trainable parameters (0.50 percent vs 0.83 percent on LLaMA-2-7B).
- SeLoRA acts as a generic plug-in: the same spectral wrapper improves DoRA and HiRA as well, with SeDoRAW and SeHiRAW adding up to 1.8 and 2.1 points over their bases.
- Spectral encoding buys the expressiveness of much larger ranks: SeLoRA at $r=32$ matches plain LoRA at $r=256$ on LLaMA-3-8B commonsense reasoning.
- SeLoRA converts data into accuracy more efficiently: with 25 percent of the training data it already beats LoRA trained on the full set, and the gap widens as data grows.
- Wavelet bases are the more robust instantiation: SeLoRAW outperforms SeLoRAF on most benchmarks, while the choice among Haar, Daubechies-4, Biorthogonal, and Coiflets changes results only marginally.
Reading between the lines
- If the sparsity property generalizes, it suggests a spectral lottery-ticket picture: a fixed random mask in the spectral domain acts like a static sparse connectivity pattern, with the basis rather than the mask carrying the expressiveness; a testable extension is whether data-dependent masks chosen by gradient sensitivity beat the random $\Omega$ at extreme sparsity.
- The fourfold drop in reverse amplification factor (0.17 to 0.04) hints that spectral encoding may reduce interference with pretrained knowledge, which would predict that SeLoRA-style adapters forget less under sequential fine-tuning, a test the paper does not run.
- Since SeLoRA with $\eta=0$ differs from plain LoRA only in the basis, comparing them at matched parameter counts isolates what the spectral prior itself buys; the paper's gains at a fixed budget already suggest the prior, not just the sparsity, is doing the work.
- The shared $\Omega$ across modules means the model learns one common sparse spectral subspace for all adapted weights; a natural stress test is whether $\Omega$ can be frozen at initialization and reused across tasks or backbones without retuning.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies parameter redundancy in LoRA fine-tuning, reports empirical evidence that random density masking of LoRA parameters preserves accuracy (the 'sparsity property'), and proposes SeLoRA, a reparameterization that learns only a sparse set of spectral coefficients and maps them back to the spatial domain via inverse Fourier or inverse wavelet transforms. The method is evaluated on LLaMA-2-7B and LLaMA-3-8B across commonsense reasoning, mathematical reasoning, and code generation, and is also plugged into DoRA and HiRA. The reported results show consistent accuracy gains of roughly 0.6 to 2.4 points over LoRA, DoRA, and HiRA at reduced parameter counts, together with ablations on sparsity ratio, rank, training data scale, module placement, wavelet basis, and training time.
Significance. If the central quantitative claim is robust, SeLoRA is a practically valuable contribution: it is simple, model-agnostic, compatible with existing LoRA variants, and shows consistent improvements at lower parameter budgets across multiple tasks and backbones. The paper is also strong in breadth: it includes careful comparisons with multiple baselines, analyses of wavelet bases, module sensitivity, rank scaling, data scaling, and efficiency, plus an honest limitations section. However, the headline improvements rest entirely on single-run point estimates with no reported variance and no released code, the shared random spectral mask is fixed for all experiments, and the Fourier variant has an unanalyzed parameter-redundancy issue. These concerns currently underdetermine the central claim, although the wavelet variant and the overall framework are plausible and worth revising rather than discarding.
major comments (4)
- [§3.1, Tables 1–2, Figure 1] All reported accuracies are single-run point estimates without standard deviations, seed counts, or confidence intervals, and no code is provided. The headline margins are typically +0.6 to +2.4 points, which is within the range of run-to-run variation commonly observed in instruction tuning of 7B–8B models. This is load-bearing because the paper's central claim is that SeLoRA consistently outperforms LoRA, DoRA, and HiRA. The motivating 'sparsity property' in Figure 1 is likewise presented as point curves with no error bars, so the claim that masking up to 60% of parameters preserves accuracy is not statistically supported. The authors should report multiple seeds with mean and variance, and ideally release code and seeds.
- [§2.2, Eq. (3), Section 3] The index set Ω is randomly initialized, fixed, and shared across all modules, and every SeLoRA result is obtained from a single random draw of Ω. The paper provides no sensitivity analysis over Ω, so it is unknown whether the observed gains reflect a robust property of spectral encoding or a favorable mask draw. Since Ω directly controls which spectral locations are learnable and the method's performance is comparable to baseline margins, the authors should either (a) show results across multiple independent Ω samples and report variance, or (b) provide evidence of mask invariance, such as a sweep of differently seeded masks.
- [§2.2, Eq. (4)] For real-valued spectral matrices FA, the real part of the inverse DFT is invariant under the simultaneous index exchange (u,v) ↔ (r−u, d−v), because the corresponding complex exponentials are complex conjugates and contribute identically to Re[F^{-1}(FA)]. Consequently, the linear map from spectral parameters to spatial matrices is non-injective and the reported parameter count for SeLoRAF overstates the number of independent degrees of freedom, potentially by a factor close to two (along the axes the pairing is slightly different). This matters for the parameter-efficiency claims made for the Fourier variant in Tables 1 and 2. The wavelet variant in Eq. (5)–(6) does not have this collapse, but the Fourier variant needs an explicit discussion, a corrected effective-parameter accounting, or a reformulation in terms of Hermitian-symmetric coefficients.
- [§3.1 and §4] The sparse ratio η is set separately for each task, model, and variant (e.g., 0.4 and 0.6 for commonsense, 0.2 and 0.4 for math and code), but the paper does not describe a validation protocol or held-out model selection for η. If η is chosen by evaluating on the same test benchmarks, the reported improvements may include selection bias. The authors should specify how η was selected, report the validation data used, or show that the qualitative conclusions are insensitive to η over a reasonable interval.
minor comments (5)
- [§3.1] Typo: 'alphaca-chat prompt template' should be 'Alpaca-chat prompt template'.
- [Table 3 caption] Typo: 'Peformance variations' should be 'Performance variations'.
- [Figure 1 caption] The caption reads 'Sparse Ratioη'; there should be a space between 'Ratio' and 'η'.
- [§3.1, Baselines] HiRA is implemented by the authors because the official code is not public. The description says this is 'based on the optimal configurations reported in their original papers,' but the paper should report the exact hyperparameters used for the HiRA implementation and note any adjustments, since this baseline is critical to the comparison and was not independently reproduced.
- [Tables 8–9] The full rank and data-scale results are useful, but they are presented without any uncertainty estimates; adding even a small number of repeated runs would strengthen the rank-scalability and data-scalability conclusions.
Circularity Check
No significant circularity: SeLoRA is an independent reparameterization evaluated against external benchmarks, and the only same-group references are non-load-bearing baselines.
full rationale
The claimed derivation chain is not circular. SeLoRA is introduced in Section 2.2 by an explicit reparameterization, W' = W0 + \tilde B \tilde A with \tilde A = T(F_A) and \tilde B = T(F_B), where F_A and F_B are zero outside a randomly initialized set Omega and T is an inverse Fourier or wavelet transform (Eqs. 2-6). This is a construction; no result is predicted from a quantity that was fitted to the same result. The motivating 'sparsity property' is an independent empirical observation about LoRA itself: Figure 1 compares LoRA at reduced rank with LoRA at fixed rank and randomized masking, and the claim that masking up to 60% preserves accuracy is about LoRA, not about SeLoRA. The headline improvements are then measured on standard external benchmarks (Commonsense170K, MetaMathQA, GSM8K/MATH, HumanEval/MBPP) with officially reproduced baselines. Table 1 and Table 2 report point estimates, and the Limitations section concedes that gains disappear at high rank (Table 9, rank 512: LoRA 86.2, SeLoRAW 86.6), which is an honest falsifiable check rather than a forced outcome. The same-group FourierFT baseline (Gao et al., 2024) and the circular-convolution related-work citation (Chen et al., 2024a) are not load-bearing for the central comparison against LoRA, DoRA, and HiRA, so they do not make the argument circular. The fixed random Omega and the absence of seed or mask sensitivity are reproducibility risks, but they are not equation-level circularity.
Assumptions & free parameters
free parameters (1)
- sparse ratio eta =
0.4/0.6 for commonsense on LLaMA-2/LLaMA-3; 0.2/0.4 for math and code
assumptions (3)
- standard math The inverse Fourier and wavelet transforms in Eqs. 4 and 6 are valid and preserve LoRA's update schema W' = W0 + B A.
- domain assumption A sparse subset of spectral coefficients at randomly chosen shared locations Omega spans a subspace of adaptation matrices expressive enough for fine-tuning.
- domain assumption LoRA's density redundancy, characterized as masking up to 60 percent of parameters at fixed rank preserving accuracy, generalizes across tasks and model families.
Cite this review
Pith. "Pith review of Revisiting LoRA through the Lens of Parameter Redundancy: Spectral Encoding Helps." pith.science (2026). https://pith.science/paper/GPE24OLR
@misc{pith2026250616787,
author = {Pith},
title = {Pith review of: Revisiting LoRA through the Lens of Parameter Redundancy: Spectral Encoding Helps},
year = {2026},
howpublished = {\url{https://pith.science/paper/GPE24OLR}},
note = {Machine review of arXiv:2506.16787}
}
read the original abstract
Low-Rank Adaptation (LoRA) has emerged as a prominent technique for fine-tuning large foundation models. Despite its successes, the substantial parameter redundancy, which limits the capacity and efficiency of LoRA, has been recognized as a bottleneck. In this work, we systematically investigate the impact of redundancy in fine-tuning LoRA and reveal that reducing density redundancy does not degrade expressiveness. Based on this insight, we introduce \underline{S}pectral-\underline{e}ncoding \underline{L}ow-\underline{R}ank \underline{A}daptation (SeLoRA), which harnesses the robust expressiveness of spectral bases to re-parameterize LoRA from a sparse spectral subspace. Designed with simplicity, SeLoRA enables seamless integration with various LoRA variants for performance boosting, serving as a scalable plug-and-play framework. Extensive experiments substantiate that SeLoRA achieves greater efficiency with fewer parameters, delivering superior performance enhancements over strong baselines on various downstream tasks, including commonsense reasoning, math reasoning, and code generation.
Figures
Forward citations
Cited by 1 Pith paper
-
ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics
Standard Conditional Flow Matching loss is a misleading early plateau; physics-informed metrics keep improving, so ScatterPrism and multi-metric diagnostics are needed for kinematic fidelity.
Reference graph
Works this paper leans on
-
[1]
Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta. 2021. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)
2021
-
[2]
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732
arXiv 2021
-
[3]
Bobby Azad, Reza Azad, Sania Eskandari, Afshin Bozorgpour, Amirhossein Kazerouni, Islem Rekik, and Dorit Merhof. 2023. Foundational models in medical imaging: A comprehensive survey and future vision. arXiv preprint arXiv:2310.18689
arXiv 2023
-
[4]
Klaudia Ba azy, Mohammadreza Banaei, Karl Aberer, and Jacek Tabor. 2024. Lora-xs: Low-rank adaptation with extremely small number of parameters. arXiv preprint arXiv:2405.17604
arXiv 2024
-
[5]
Nadav Benedek and Lior Wolf. 2024. Prilora: Pruned and rank-increasing low-rank adaptation. arXiv preprint arXiv:2401.11316
arXiv 2024
-
[6]
Srinadh Bhojanapalli, Ayan Chakrabarti, Andreas Veit, Michal Lukasik, Himanshu Jain, Frederick Liu, Yin-Wen Chang, and Sanjiv Kumar. 2021. Leveraging redundancy in attention with reuse transformers. arXiv preprint arXiv:2110.06821
arXiv 2021
-
[7]
Dan Biderman, Jose Gonzalez Ortiz, Jacob Portes, Mansheej Paul, Philip Greengard, Connor Jennings, Daniel King, Sam Havens, Vitaliy Chiley, Jonathan Frankle, et al. 2024. Lora learns less and forgets less. arXiv preprint arXiv:2405.09673
arXiv 2024
-
[8]
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence
2020
Show all 93 references
-
[9]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877--1901
2020
-
[10]
Eric L Buehler and Markus J Buehler. 2024. X-lora: Mixture of low-rank adapter experts, a flexible framework for large language models with applications in protein mechanics and molecular design. APL Machine Learning, 2(2)
2024
-
[11]
Aochuan Chen, Jiashun Cheng, Zijing Liu, Ziqi Gao, Fugee Tsung, Yu Li, and Jia Li. 2024 a . Parameter-efficient fine-tuning via circular convolution. arXiv preprint arXiv:2407.19342
2024 arXiv
-
[12]
Aochuan Chen, Yuguang Yao, Pin-Yu Chen, Yihua Zhang, and Sijia Liu. 2023 a . Understanding and improving visual prompting: A label-mapping perspective. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19133--19143
2023
-
[13]
Aochuan Chen, Yimeng Zhang, Jinghan Jia, James Diffenderfer, Jiancheng Liu, Konstantinos Parasyris, Yihua Zhang, Zheng Zhang, Bhavya Kailkhura, and Sijia Liu. 2023 b . Deepzero: Scaling up zeroth-order optimization for deep model training. arXiv preprint arXiv:2310.02025
2023 arXiv
-
[14]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374
2021 arXiv
-
[15]
Nuo Chen, Yuhan Li, Jianheng Tang, and Jia Li. 2024 b . Graphwiz: An instruction-following language model for graph computational problems. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 353--364
2024
-
[16]
Nuo Chen, Yan Wang, Haiyun Jiang, Deng Cai, Yuhan Li, Ziyang Chen, Longyue Wang, and Jia Li. 2023 c . Large language models meet harry potter: A dataset for aligning dialogue agents with characters. In Findings of the Association for Computational Linguistics: EMNLP 2023, page...
2023
-
[17]
Nuo Chen, Ning Wu, Jianhui Chang, and Jia Li. 2024 c . Controlmath: Controllable data generation promotes math generalist models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 12201--12217
2024
-
[18]
Jiashun Cheng, Man Li, Jia Li, and Fugee Tsung. 2023. Wiener graph deconvolutional network improves graph self-supervised learning. In Proceedings of the AAAI conference on artificial intelligence, volume 37, pages 7131--7139
2023
-
[19]
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044
2019 arXiv
-
[20]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457
2018 arXiv
-
[21]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168
2021 arXiv
-
[22]
Fahim Dalvi, Hassan Sajjad, Nadir Durrani, and Yonatan Belinkov. 2020. Analyzing redundancy in pretrained transformer models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics
2020
-
[23]
Ning Ding, Xingtai Lv, Qiaosen Wang, Yulin Chen, Bowen Zhou, Zhiyuan Liu, and Maosong Sun. 2023. https://openreview.net/forum?id=jxgz7FEqWq Sparse low-rank adaptation of pre-trained language models . In The 2023 Conference on Empirical Methods in Natural Language Processing
2023
-
[24]
Marco F Duarte and Richard G Baraniuk. 2013. Spectral compressive sensing. Applied and Computational Harmonic Analysis, 35(1):111--129
2013
-
[25]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783
2024 arXiv
-
[26]
Jonathan Frankle and Michael Carbin. 2018. The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635
2018 arXiv
-
[27]
Elias Frantar and Dan Alistarh. 2023. Sparsegpt: Massive language models can be accurately pruned in one-shot. In International Conference on Machine Learning, pages 10323--10337. PMLR
2023
-
[28]
Ziqi Gao, Qichao Wang, Aochuan Chen, Zijing Liu, Bingzhe Wu, Liang Chen, and Jia Li. 2024. Parameter-efficient fine-tuning with discrete fourier transform. arXiv preprint arXiv:2405.03003
2024 arXiv
-
[29]
Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pages 249--256. JMLR Workshop and Conference Proceedings
2010
-
[30]
Naibin Gu, Peng Fu, Xiyu Liu, Bowen Shen, Zheng Lin, and Weiping Wang. 2024. Light-peft: Lightening parameter-efficient fine-tuning via early pruning. In Findings of the Association for Computational Linguistics
2024
-
[31]
Song Han, Huizi Mao, and William J Dally. 2015 a . Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149
2015 arXiv
-
[32]
Song Han, Jeff Pool, John Tran, and William Dally. 2015 b . Learning both weights and connections for efficient neural network. Advances in neural information processing systems, 28
2015
-
[33]
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2021. Towards a unified view of parameter-efficient transfer learning. arXiv preprint arXiv:2110.04366
2021 arXiv
-
[34]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, pages 1026--1034
2015
-
[35]
Shwai He, Liang Ding, Daize Dong, Miao Zhang, and Dacheng Tao. 2022. Sparseadapter: An easy approach for improving the parameter-efficiency of adapters. arXiv preprint arXiv:2210.04284
2022 arXiv
-
[36]
Shwai He, Guoheng Sun, Zheyu Shen, and Ang Li. 2024. What matters in transformers? not all attention is needed. arXiv preprint arXiv:2406.15786
2024 arXiv
-
[37]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300
2020 arXiv
-
[38]
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for nlp. In International conference on machine learning, pages 2790--2799. PMLR
2019
-
[39]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685
2021 arXiv
-
[40]
Zhiqiang Hu, Lei Wang, Yihuai Lan, Wanyu Xu, Ee-Peng Lim, Lidong Bing, Xing Xu, Soujanya Poria, and Roy Ka-Wei Lee. 2023. Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models. arXiv preprint arXiv:2304.01933
2023 arXiv
-
[41]
Qiushi Huang, Tom Ko, Zhan Zhuang, Lilian Tang, and Yu Zhang. 2025. https://openreview.net/forum?id=TwJrTz9cRS Hi RA : Parameter-efficient hadamard high-rank adaptation for large language models . In The Thirteenth International Conference on Learning Representations
2025
-
[42]
Kazuki Irie and J \"u rgen Schmidhuber. 2021. Training and generating neural networks in compressed weight space. arXiv preprint arXiv:2112.15545
2021 arXiv
-
[43]
Shuyang Jiang, Yusheng Liao, Yanfeng Wang, Ya Zhang, and Yu Wang. 2025. https://openreview.net/forum?id=ZV7CLf0RHK Fine-tuning with reserved majority for noise reduction . In The Thirteenth International Conference on Learning Representations
2025
-
[44]
Shuyang Jiang, Yusheng Liao, Ya Zhang, Yanfeng Wang, and Yu Wang. 2024 a . https://openreview.net/forum?id=XxSME6GE1G TAIA : Large language models are out-of-distribution data learners . In The Thirty-eighth Annual Conference on Neural Information Processing Systems
2024
-
[45]
Ting Jiang, Shaohan Huang, Shengyue Luo, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang, Deqing Wang, et al. 2024 b . Mora: High-rank updating for parameter-efficient fine-tuning. arXiv preprint arXiv:2405.12130
2024 arXiv
-
[46]
Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. 2021. Compacter: Efficient low-rank hypercomplex adapter layers. Advances in Neural Information Processing Systems, 34:1022--1035
2021
-
[47]
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. 2023. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4015--4026
2023
-
[48]
Dawid Jan Kopiczko, Tijmen Blankevoort, and Yuki Markus Asano. 2023. Vera: Vector-based random matrix adaptation. arXiv preprint arXiv:2310.11454
2023 arXiv
-
[49]
Jan Koutnik, Faustino Gomez, and J \"u rgen Schmidhuber. 2010. Evolving neural networks in compressed weight space. In Proceedings of the 12th annual conference on Genetic and evolutionary computation, pages 619--626
2010
-
[50]
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip HS Torr. 2018. Snip: Single-shot network pruning based on connection sensitivity. arXiv preprint arXiv:1810.02340
2018 arXiv
-
[51]
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. https://arxiv.org/abs/2104.08691 The power of scale for parameter-efficient prompt tuning . Preprint, arXiv:2104.08691
2021 arXiv
-
[52]
Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190
2021 arXiv
-
[53]
Yang Li, Shaobo Han, and Shihao Ji. 2024 a . https://openreview.net/forum?id=kuCY0mW4Q3 VB -lo RA : Extreme parameter efficient fine-tuning with vector banks . In The Thirty-eighth Annual Conference on Neural Information Processing Systems
2024
-
[54]
Yang Li, Shaobo Han, and Shihao Ji. 2024 b . Vb-lora: Extreme parameter efficient fine-tuning with vector banks. arXiv preprint arXiv:2405.15179
2024 arXiv
-
[55]
Yuhan Li, Peisong Wang, Xiao Zhu, Aochuan Chen, Haiyun Jiang, Deng Cai, Victor W Chan, and Jia Li. 2024 c . Glbench: A comprehensive benchmark for graph with large language models. In Advances in Neural Information Processing Systems, volume 37, pages 42349--42368
2024
-
[56]
Yuhan Li, Xinni Zhang, Linhao Luo, Heng Chang, Yuxiang Ren, Irwin King, and Jia Li. 2025. G-refer: Graph retrieval-augmented large language model for explainable recommendation. In Proceedings of the ACM on Web Conference 2025, pages 240--251
2025
-
[57]
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2024 a . Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems, 36
2024
-
[58]
Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, and Min-Hung Chen. 2024 b . Dora: Weight-decomposed low-rank adaptation. arXiv preprint arXiv:2402.09353
2024 arXiv
-
[59]
Shiwei Liu, Tianlong Chen, Xiaohan Chen, Li Shen, Decebal Constantin Mocanu, Zhangyang Wang, and Mykola Pechenizkiy. 2022. The unreasonable effectiveness of random pruning: Return of the most naive baseline for sparse training. arXiv preprint arXiv:2202.02643
2022 arXiv
-
[60]
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2021. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint arXiv:2110.07602
2021 arXiv
-
[61]
Ilya Loshchilov and Frank Hutter. 2019. https://openreview.net/forum?id=Bkg6RiCqY7 Decoupled weight decay regularization . In International Conference on Learning Representations
2019
-
[62]
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2024. https://openreview.net/forum?id=UnUwSIgK5W Wizardcoder: Empowering code large language models with evol-instruct . In The Twelfth International Confe...
2024
-
[63]
Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan. 2022. Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft
2022
-
[64]
Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, and Weipeng Chen. 2024. Shortgpt: Layers in large language models are more redundant than you expect. arXiv preprint arXiv:2403.03853
2024 arXiv
-
[65]
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789
2018 arXiv
-
[66]
Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H Nguyen, Madeleine Gibescu, and Antonio Liotta. 2018. Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science. Nature communications, 9(1):2383
2018
-
[67]
Mahdi Nikdan, Soroush Tabesh, and Dan Alistarh. 2024. Rosa: Accurate parameter-efficient fine-tuning via robust adaptation. arXiv preprint arXiv:2401.04679
2024 arXiv
-
[68]
Henri J Nussbaumer and Henri J Nussbaumer. 1982. The fast Fourier transform. Springer
1982
-
[69]
Jonas Pfeiffer, Aishwarya Kamath, Andreas R \"u ckl \'e , Kyunghyun Cho, and Iryna Gurevych. 2020. Adapterfusion: Non-destructive task composition for transfer learning. arXiv preprint arXiv:2005.00247
2020 arXiv
-
[70]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[71]
Adithya Renduchintala, Tugrul Konuk, and Oleksii Kuchaiev. 2023. Tied-lora: Enhacing parameter efficiency of lora with weight tying. arXiv preprint arXiv:2311.09578
2023 arXiv
-
[72]
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99--106
2021
-
[73]
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. 2019. Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728
2019 arXiv
-
[74]
Arijit Sehanobish, Avinava Dubey, Krzysztof Choromanski, Somnath Basu Roy Chowdhury, Deepali Jain, Vikas Sindhwani, and Snigdha Chaturvedi. 2024. Structured unrestricted-rank matrices for parameter efficient fine-tuning. arXiv preprint arXiv:2406.17740
2024 arXiv
-
[75]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 a . Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971
2023 arXiv
-
[76]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 b . Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288
2023 arXiv
-
[77]
Sjoerd Van Steenkiste, Jan Koutn \' k, Kurt Driessens, and J \"u rgen Schmidhuber. 2016. A wavelet-based encoding for neuroevolution. In Proceedings of the Genetic and Evolutionary Computation Conference 2016, pages 517--524
2016
-
[78]
Marinus T Vlaardingerbroek and Jacques A Boer. 2013. Magnetic resonance imaging: theory and practice. Springer Science & Business Media
2013
-
[79]
Chaoqi Wang, Guodong Zhang, and Roger Grosse. 2020. Picking winning tickets before training by preserving gradient flow. arXiv preprint arXiv:2002.07376
2020 arXiv
-
[80]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837
2022
-
[81]
Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang. 2024. Magicoder: Empowering code generation with oss-instruct. In Forty-first International Conference on Machine Learning
2024
-
[82]
Moritz Wolter, Felix Blanke, Jochen Garcke, and Charles Tapley Hoyt. 2024. ptwt-the pytorch wavelet toolbox. Journal of Machine Learning Research, 25(80):1--7
2024
-
[83]
Moritz Wolter, Shaohui Lin, and Angela Yao. 2020. Neural network compression via learnable wavelet transforms. In Artificial Neural Networks and Machine Learning--ICANN 2020: 29th International Conference on Artificial Neural Networks, Bratislava, Slovakia, September 15--18, 2...
2020
-
[84]
Yifan Yang, Jiajun Zhou, Ngai Wong, and Zheng Zhang. 2024. Loretta: Low-rank economic tensor-train adaptation for ultra-low-parameter fine-tuning of large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational ...
2024
-
[85]
Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li. 2024. Language models are super mario: Absorbing abilities from homologous models as a free lunch. In Forty-first International Conference on Machine Learning
2024
-
[86]
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2023. Metamath: Bootstrap your own mathematical questions for large language models. arXiv preprint arXiv:2309.12284
2023 arXiv
-
[87]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830
2019 arXiv
-
[88]
Longteng Zhang, Lin Zhang, Shaohuai Shi, Xiaowen Chu, and Bo Li. 2023 a . Lora-fa: Memory-efficient low-rank adaptation for large language models fine-tuning. arXiv preprint arXiv:2308.03303
2023 arXiv
-
[89]
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Nikos Karampatziakis, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. 2023 b . Adalora: Adaptive budget allocation for parameter-efficient fine-tuning. arXiv preprint arXiv:2303.10512
2023 arXiv
-
[90]
Yihua Zhang, Yuguang Yao, Parikshit Ram, Pu Zhao, Tianlong Chen, Mingyi Hong, Yanzhi Wang, and Sijia Liu. 2022. Advancing model pruning via bi-level optimization. Advances in Neural Information Processing Systems, 35:18309--18326
2022
-
[91]
Bowen Zhao, Hannaneh Hajishirzi, and Qingqing Cao. 2024. https://openreview.net/forum?id=sb81Xl50JG APT : Adaptive pruning and tuning pretrained language models for efficient training and inference . In Forty-first International Conference on Machine Learning
2024
-
[92]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[93]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.