REVIEW 3 major objections 7 minor 44 references
Train More Parameters But Mind Their Placement: Insights into Language Adaptation with PEFT
T0 review · 3 major / 7 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Increasing trainable parameters improves language adaptation, but only when placed in feed-forward layers; prefix tuning and (IA)3 hurt.
desk verdict A careful, honest ablation study of PEFT placement for Icelandic adaptation of a 1B model; the single-run evaluation is the main soft spot, but the findings are plausible and worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by the contrast between four parameter-efficient fine-tuning mechanisms placed inside a fixed Transformer: LoRA, which adds low-rank decomposition matrices to chosen weights; bottleneck adapters, which insert down- and up-projections between layers; (IA)3, which multiplies activations by learned vectors; and prefix tuning, which prepends learnable prefix vectors to the input. The paper varies the rank or reduction factor of each method and, for LoRA, the module (query/value attention versus feed-forward) and layer range, then reads off BERTScore and ROUGE-L on the RUV Radio News summarisation task. The key contrast is that parameter count alone does not explain performance: feed-forward LoRA and bottleneck adapters convert parameters into scores efficiently, attention LoRA does not, and prefix tuning interferes with generation despite its parameter count.
What would settle it
Rerun the full ablation with five random seeds and a human preference study on the generated summaries; a reversal in the ranking, say attention LoRA matching or beating feed-forward LoRA, would falsify the paper's central claim.
Extended reading notes
Core claim
The central discovery is a placement-and-capacity ranking for PEFT-based language adaptation. With 250,000 Icelandic text chunks of up to 1,024 tokens, the paper finds that adaptation quality rises with the number of trainable parameters, but only when the parameters are put in the right modules. Feed-forward LoRA with rank 256 reaches BERTScore 65.60 / ROUGE-L 09.72 in 0-shot summarisation, beating attention-placed LoRA of the same rank, and combining both modules is not better than feed-forward alone. Bottleneck adapters with reduction factor 4 are competitive, while attention LoRA needs far more parameters (rank 1024) to approach feed-forward results, and prefix tuning and (IA)3 actively hurt the model, with prefix tuning collapsing in 1- and 5-shot settings. The paper argues this shows feed-forward modules are the most promising target and that sufficient learning capacity is necessary.
Load-bearing premise
The ranking is built on single runs of BERTScore and ROUGE-L on one news-summarisation dataset, so if those automatic scores are noisy or do not track generation quality, the recommended setups could change.
Editorial extensions
If this is right
- Practitioners adapting instruction-tuned models with unstructured text should prefer feed-forward LoRA at high rank, or bottleneck adapters with a small reduction factor, over attention LoRA.
- More trainable parameters help, but placement can dominate: attention LoRA at rank 1024 still trails feed-forward LoRA at rank 256, so parameter budgets should go to feed-forward modules first.
- Prefix tuning and (IA)3 should be avoided for text-only language adaptation of instruction-tuned models.
- Context degradation from limited adaptation context can be reduced by adapting only the final layers, at a slight cost in 0-shot performance.
- In-context learning with target-language demonstrations is a viable alternative: the 1-shot baseline already matches many adapted setups.
Reading between the lines
- Beyond the paper's explicit comparisons, its placement ranking suggests that for a fixed parameter budget, PEFT design should target where language-specific knowledge is stored, not just how many parameters are trained.
- A testable extension the paper leaves implicit: train the same adapters with a 4,096-token adaptation context; if the 5-shot degradation from attention LoRA disappears, the short-context hypothesis is confirmed.
- The result that curated data gave no benefit is about extractive summarisation; on a task requiring language-specific world knowledge, the ordering might shift.
- Human evaluation of the top setups could reveal quality differences BERTScore and ROUGE-L miss, since those metrics reward lexical overlap.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports an empirical study of parameter-efficient fine-tuning (PEFT) methods for adapting Llama-3.2-1B-Instruct to Icelandic using unstructured text. The author compares LoRA in attention and feed-forward modules across several ranks, bottleneck adapters, (IA)3, and prefix tuning, and evaluates 0-shot, 1-shot, and 5-shot abstractive summarization on the RÚV Radio News dataset with BERTScore and ROUGE-L. The main findings are that more trainable parameters help within a method family, feed-forward LoRA is the best configuration, bottleneck adapters are second, attention-placed LoRA is weaker for its parameter count, prefix tuning and (IA)3 are unsuitable, and restricting adapters to the final layers mitigates degradation at longer contexts. The code, prompt generator, and trained adapters are released.
Significance. If the results hold, the paper provides practical, actionable guidance for language adaptation of small instruction-tuned LLMs: place LoRA in feed-forward layers, provide sufficient learning capacity, avoid prefix tuning for unstructured-text adaptation, and consider final-layer-only adaptation to preserve long-context abilities. The study is carefully structured, with parameter-count comparisons, three targeted ablations (module placement, layer selection, and corpus choice), and a consistent evaluation protocol. The release of code and adapters is a concrete strength, as is the use of both a web-crawled and a curated adaptation corpus. The main caveats are that all conclusions rest on single-run point estimates and on automatic summarization metrics for one language, one model, and one task, and that potential overlap between the adaptation corpus and the evaluation data is not examined.
major comments (3)
- [§2.2, §2.4] The adaptation sample in §2.2 consists of 250,000 random chunks from the Icelandic CC100 corpus after CCNet filtering, and the evaluation in §2.4 is abstractive summarization on RÚV Radio News (RRN). CC100 is built from Common Crawl and RRN is Icelandic broadcaster news content, so RRN articles may be present in the adaptation sample. If a training chunk contains an RRN main body followed by the reference introduction, the causal-language-modeling objective can memorize the reference summary, and high-capacity configurations such as LoRA-ff-256 and Bottlen.-4 would be rewarded for memorization rather than for language adaptation; the 30-token prefix-tuned model could not memorize at the same scale, which would artificially strengthen the 'prefix tuning is not suitable' conclusion. The same concern applies to the CCNet-versus-IGC comparison in Table 4 if CCNet contains the news domain and IGC does not. The manuscript reports no decontamination or n-gram overlap analysis, so this missing check is load-bearing for the headline ranking and for the abstract's general guidance.
- [Tables 1–4] Every result in the four tables is a single point estimate from one training run, with no standard deviations, seeds, or significance tests. Several comparisons that support the 'more trainable parameters is better' claim differ by less than one BERTScore point or less than one ROUGE-L point, e.g., LoRA-ff-128 versus LoRA-ff-256 in 1-shot (69.10/13.86 versus 69.06/13.89) and Bottlen.-16 versus Bottlen.-4 in 0-shot (63.33/8.38 versus 63.78/8.15). Without variance estimates or paired significance testing, the ranking of nearby configurations is not established, and the unconditional wording of the abstract ('are not suitable') is stronger than the evidence supports. Multiple seeds for at least the main configurations, or another variance-aware analysis, would make the ranking defensible.
- [Limitations] The Limitations section explicitly states that automatic summarization metrics are 'questioned' and that human evaluation is needed, yet the abstract and Section 3.1 present the ranking as conclusive for language adaptation on the basis of BERTScore and ROUGE-L on a single task. The central claim is prescriptive for practitioners, so the paper should either add convergent evidence (for example, a human evaluation or a second task on a subset of the main configurations) or qualify the conclusions as preliminary. This is not a demand for full evaluation of every ablation, but the strength of the headline claims should match the strength of the evidence.
minor comments (7)
- [§2.3] There is a typo: 'we use use α = 2r' should be 'we use α = 2r'.
- [§3.2] The sentence 'Moreover, it is slightly better than LoRA in both the attention and the feed-forward modules' is ambiguous; it should say that feed-forward LoRA slightly outperforms LoRA applied to both modules.
- [Figure 1] The x-axis labels are dense and the excluded methods (prefix tuning and (IA)3) are not shown; add a note or a supplementary table with their parameter counts so the reader can see why they are excluded.
- [Table 3] The note 'Self-attention (qv) LoRA rank 32' should appear in the table caption rather than below the table body, since it is needed to interpret the numbers.
- [§3.4] The sentence 'we do not test on any task where high-quality generation is important but on text summarisation' is potentially confusing: summarization is itself a generation task; the intended contrast is that the evaluation may reward copying rather than language quality. Rephrase for clarity.
- [Footnote 3] 'the no adapters model' should be 'the no-adapter model'.
- [§2.3] Prefix tuning is tested with a single prefix length of 30; a short prefix-length sweep would strengthen the 'not suitable' conclusion, since prefix length directly controls the method's capacity.
Circularity Check
No circularity: the paper reports direct empirical PEFT ablations against a no-adapter baseline on external Icelandic data, with no fitted constants, self-derived equations, or self-citation chains load-bearing for the conclusions.
full rationale
The manuscript's claims are empirical comparisons, not derivations. Section 2.2 fixes the adaptation corpora (CC100/CCNet and Icelandic Gigaword), Section 2.3 fixes the PEFT methods and hyperparameters (learning rate 5e-5, batch size 4, alpha = 2r, prefix length 30), and Section 2.4 fixes the external evaluation task (RÚV Radio News main-to-intro summarization with BERTScore and ROUGE-L). Tables 1-4 then report measured differences against a no-adapter baseline, and Figure 1 plots parameter count against BERTScore. None of these quantities is defined in terms of the conclusion, and no parameter is fitted to the evaluation set and then renamed as a prediction. The Limitations section explicitly flags that automatic summarization metrics are questioned and that human evaluation is missing; that is a validity and generalizability concern, not a circularity concern. The same applies to the paper's own remark in Section 3.4 that summarization 'can rely on copying chunks of text' and to the absence of an overlap/decontamination analysis between CC100 and RRN: possible train/evaluation contamination is an empirical risk, not a logical reduction of the result to its inputs. The related-work discussion cites other authors' findings as independent context, and the paper does not invoke any uniqueness theorem or prior self-citation to force its architectural conclusions. The central claim (more trainable parameters generally help, feed-forward LoRA and bottleneck adapters outperform attention-placed LoRA, prefix tuning and IA3 are unsuitable) rests entirely on the executed experiments, so there is no circular step to exhibit.
Assumptions & free parameters
free parameters (4)
- LoRA rank r =
8, 32, 64, 128, 256, or 1024 (ablation grid)
- Bottleneck adapter reduction factor =
4, 16, or 64
- Prefix length =
30 tokens
- Learning rate and scheduler =
5e-5 with linear scheduler
assumptions (4)
- domain assumption Automatic summarization metrics (BERTScore, ROUGE-L) on the RRN dataset are a valid proxy for language adaptation quality.
- domain assumption Llama-3.2-1B-Instruct is representative of smaller instruction-tuned LLMs for adaptation conclusions.
- domain assumption The adaptation data (CC100 Icelandic) was present in pretraining, so the effect measured is priming toward Icelandic rather than learning new factual knowledge.
- domain assumption Single-run results are stable enough for the reported rankings.
Cite this review
Pith. "Pith review of Train More Parameters But Mind Their Placement: Insights into Language Adaptation with PEFT." pith.science (2026). https://pith.science/paper/HTLZIAPN
@misc{pith2026241212674,
author = {Pith},
title = {Pith review of: Train More Parameters But Mind Their Placement: Insights into Language Adaptation with PEFT},
year = {2026},
howpublished = {\url{https://pith.science/paper/HTLZIAPN}},
note = {Machine review of arXiv:2412.12674}
}
read the original abstract
Smaller LLMs still face significant challenges even in medium-resourced languages, particularly when it comes to language-specific knowledge -- a problem not easily resolved with machine-translated data. In this case study on Icelandic, we aim to enhance the generation performance of an LLM by specialising it using unstructured text corpora. A key focus is on preventing interference with the models' capabilities of handling longer context during this adaptation. Through ablation studies using various parameter-efficient fine-tuning (PEFT) methods and setups, we find that increasing the number of trainable parameters leads to better and more robust language adaptation. LoRAs placed in the feed-forward layers and bottleneck adapters show promising results with sufficient parameters, while prefix tuning and (IA)3 are not suitable. Although improvements are consistent in 0-shot summarisation, some adapted models struggle with longer context lengths, an issue that can be mitigated by adapting only the final layers.
Figures
Reference graph
Works this paper leans on
-
[1]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Mikel Artetxe, Itziar Aldabe, Rodrigo Agerri, Olatz Perez-de Vi \ n aspre, and Aitor Soroa. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.499 Does corpus quality really matter for low-resource languages? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7383--7390, Abu Dhabi, United Arab Emirates. Associa...
-
[4]
Starka ur Barkarson, Stein \'o r Steingr \' msson, and Hildur Hafsteinsd \'o ttir. 2022. https://aclanthology.org/2022.lrec-1.254 Evolving large text corpora: Four versions of the I celandic G igaword corpus . In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 2371--2381, Marseille, France. European Language Resources Association
work page 2022
-
[5]
Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajanthan, Puneet K. Dokania, Philip H. S. Torr, and Marc'Aurelio Ranzato. 2019. http://arxiv.org/abs/1902.10486 On tiny episodic memories in continual learning
arXiv 2019
-
[6]
Pinzhen Chen, Shaoxiong Ji, Nikolay Bogoychev, Andrey Kutuzov, Barry Haddow, and Kenneth Heafield. 2024 a . https://aclanthology.org/2024.findings-eacl.90 Monolingual or multilingual instruction tuning: Which makes a better alpaca . In Findings of the Association for Computational Linguistics: EACL 2024, pages 1347--1356, St. Julian ' s, Malta. Associatio...
work page 2024
-
[7]
Pinzhen Chen, Simon Yu, Zhicheng Guo, and Barry Haddow. 2024 b . http://arxiv.org/abs/2406.12822 Is it good data for multilingual instruction tuning or just bad multilingual evaluation for large language models?
work page Pith review arXiv 2024
-
[8]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/v1/2020.acl-main.747 Unsupervised cross-lingual representation learning at scale . In Proceedings of the 58th Annual Meeting of the Association for Comp...
Show all 44 references
-
[9]
Fahim Faisal and Antonios Anastasopoulos. 2022. https://aclanthology.org/2022.aacl-main.34 Phylogeny-inspired adaptation of multilingual models to new languages . In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics ...
2022
-
[10]
Vlad Fomenko, Han Yu, Jongho Lee, Stanley Hsieh, and Weizhu Chen. 2024. http://arxiv.org/abs/2404.05086 A note on lora
2024 arXiv
-
[11]
Ruidan He, Linlin Liu, Hai Ye, Qingyu Tan, Bosheng Ding, Liying Cheng, Jiawei Low, Lidong Bing, and Luo Si. 2021. https://doi.org/10.18653/v1/2021.acl-long.172 On the effectiveness of adapter-based tuning for pretrained language model adaptation . In Proceedings of the 59th An...
2021 doi
-
[12]
Oskar Holmstr \"o m and Ehsan Doostmohammadi. 2023. https://aclanthology.org/2023.nodalida-1.62 Making instruction finetuning accessible to non- E nglish languages: A case study on S wedish models . In Proceedings of the 24th Nordic Conference on Computational Linguistics (NoD...
2023
-
[13]
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. https://proceedings.mlr.press/v97/houlsby19a.html Parameter-efficient transfer learning for NLP . In Proceedings of the 36th In...
2019
-
[14]
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. https://openreview.net/forum?id=nZeVKeeFYf9 Lo RA : Low-rank adaptation of large language models . In International Conference on Learning Representations
2022
-
[15]
Adam Ibrahim, Benjamin Th \'e rien, Kshitij Gupta, Mats Leon Richter, Quentin Gregory Anthony, Eugene Belilovsky, Timoth \'e e Lesort, and Irina Rish. 2024. https://openreview.net/forum?id=DimPeeCxKO Simple and scalable strategies to continually pre-train large language models...
2024
-
[16]
Zhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodriguez, Chunting Zhou, Graham Neubig, Xi Lin, Wen-tau Yih, and Srini Iyer. 2024. https://doi.org/10.18653/v1/2024.acl-long.296 Instruction-tuned language models are better knowledge learners . In Proceedings of the 62nd Annual ...
2024 doi
-
[17]
Tannon Kew, Florian Schottmann, and Rico Sennrich. 2023. http://arxiv.org/abs/2312.12683 Turning english-centric llms into polyglots: How much multilinguality is needed?
2023 arXiv
-
[18]
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitri...
2022
-
[19]
Xiang Lisa Li and Percy Liang. 2021. https://doi.org/10.18653/v1/2021.acl-long.353 Prefix-tuning: Optimizing continuous prompts for generation . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conferen...
2021 doi
-
[20]
Chin-Yew Lin. 2004. https://aclanthology.org/W04-1013 ROUGE : A package for automatic evaluation of summaries . In Text Summarization Branches Out, pages 74--81, Barcelona, Spain. Association for Computational Linguistics
2004
-
[21]
Haokun Liu, Derek Tam, Muqeeth Mohammed, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. 2022. https://openreview.net/forum?id=rBCvMG-JsPd Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning . In Advances in Neural Information Proc...
2022
-
[22]
LlamaTeam. 2024. http://arxiv.org/abs/2407.21783 The llama 3 herd of models
2024 arXiv
-
[23]
Michael Mccloskey and Neil J. Cohen. 1989. Catastrophic interference in connectionist networks: T he sequential learning problem. The Psychology of Learning and Motivation, 24:104--169
1989
-
[24]
Amil Merchant, Elahe Rahimtoroghi, Ellie Pavlick, and Ian Tenney. 2020. https://doi.org/10.18653/v1/2020.blackboxnlp-1.4 What happens to BERT embeddings during fine-tuning? In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, ...
2020 doi
-
[25]
Hedderich, and Dietrich Klakow
Marius Mosbach, Anna Khokhlova, Michael A. Hedderich, and Dietrich Klakow. 2020. https://doi.org/10.18653/v1/2020.blackboxnlp-1.7 On the interplay between fine-tuning and sentence-level probing for linguistic knowledge in pre-trained transformers . In Proceedings of the Third ...
2020 doi
-
[26]
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward ...
2023 doi
-
[27]
Dan Saattrup Nielsen, Kenneth Enevoldsen, and Peter Schneider-Kamp. 2024. http://arxiv.org/abs/2406.13469 Encoder vs decoder: Comparative analysis of encoder and decoder language models on multilingual nlu tasks
2024 arXiv
-
[28]
Rik van Noord, Taja Kuzman, Peter Rupnik, Nikola Ljube s i \'c , Miquel Espl \`a -Gomis, Gema Ram \' rez-S \'a nchez, and Antonio Toral. 2024. https://aclanthology.org/2024.lrec-main.465 Do language models care about text quality? evaluating web-crawled corpora across 11 langu...
2024
-
[29]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2024
-
[30]
Jupinder Parmar, Sanjev Satheesh, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. 2024. http://arxiv.org/abs/2407.07263 Reuse, don't retrain: A recipe for continued pretraining of language models
2024 arXiv
-
[31]
Jonas Pfeiffer, Ivan Vuli \'c , Iryna Gurevych, and Sebastian Ruder. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.617 MAD-X : A n A dapter- B ased F ramework for M ulti- T ask C ross- L ingual T ransfer . In Proceedings of the 2020 Conference on Empirical Methods in Natur...
2020 doi
-
[32]
Clifton Poth, Hannah Sterz, Indraneil Paul, Sukannya Purkayastha, Leon Engl \"a nder, Timo Imhof, Ivan Vuli \'c , Sebastian Ruder, Iryna Gurevych, and Jonas Pfeiffer. 2023. https://aclanthology.org/2023.emnlp-demo.13 Adapters: A unified library for parameter-efficient and modu...
2023
-
[33]
Evgeniia Razumovskaia, Ivan Vulić, and Anna Korhonen. 2024. http://arxiv.org/abs/2403.01929 Analyzing and adapting large language models for few-shot multilingual nlu: Are we there yet?
2024 arXiv
-
[34]
Stein \'o r Steingr \' msson, Sigr \'u n Helgad \'o ttir, Eir \' kur R \"o gnvaldsson, Starka ur Barkarson, and J \'o n Gu nason. 2018. https://aclanthology.org/L18-1690 R isam \'a lheild: A very large I celandic text corpus . In Proceedings of the Eleventh International Confe...
2018
-
[35]
\'o r Sverrisson and Hafsteinn Einarsson. 2023. https://aclanthology.org/2023.nodalida-1.3 Abstractive text summarization for I celandic . In Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa), pages 17--31, T \'o rshavn, Faroe Islands. Universit...
2023
-
[36]
Dai, and Quoc V Le
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022. https://openreview.net/forum?id=gEZrGCozdqR Finetuned language models are zero-shot learners . In International Conference on Learning Representations
2022
-
[37]
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzm \'a n, Armand Joulin, and Edouard Grave. 2020. https://aclanthology.org/2020.lrec-1.494 CCN et: Extracting high quality monolingual datasets from web crawl data . In Proceedings of the Twel...
2020
-
[38]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020 doi
-
[39]
Kai Yao, Penlei Gao, Lichun Li, Yuan Zhao, Xiaofeng Wang, Wei Wang, and Jianke Zhu. 2024. http://arxiv.org/abs/2410.11772 Layer-wise importance matters: Less memory for better performance in parameter-efficient fine-tuning of large language models
2024 arXiv
-
[40]
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. 2023. https://openreview.net/forum?id=lq62uWRJjiY Adaptive budget allocation for parameter-efficient fine-tuning . In The Eleventh International Conference on Learning Representations
2023
-
[41]
Renrui Zhang, Jiaming Han, Chris Liu, Aojun Zhou, Pan Lu, Yu Qiao, Hongsheng Li, and Peng Gao. 2024 a . https://openreview.net/forum?id=d4UiXAHN2W LL a MA -adapter: Efficient fine-tuning of large language models with zero-initialized attention . In The Twelfth International Co...
2024
-
[42]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://openreview.net/forum?id=SkeHuCVFDr Bertscore: Evaluating text generation with BERT . In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April...
2020
-
[43]
Hashimoto
Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B. Hashimoto. 2024 b . https://doi.org/10.1162/tacl_a_00632 Benchmarking large language models for news summarization . Transactions of the Association for Computational Linguistics, 12:39--57
2024 doi
-
[44]
Yichu Zhou and Vivek Srikumar. 2022. https://doi.org/10.18653/v1/2022.acl-long.75 A closer look at how fine-tuning changes BERT . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1046--1061, Dublin, Irela...
2022 doi
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.