REVIEW 4 major objections 5 minor 1 cited by
LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation
T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read This paper claims that short-text degradation after context extension stems from hidden-state drift and forgetting, and that distilling the original model's hidden states on short text restores the lost performance.
desk verdict Useful incremental method for preserving short-text skills during context extension, but the headline 'comparable or even better' long-text claim is contradicted by the paper's own 128K Llama-3-8B RULER numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a pair of distillation losses anchored to the original, pre-extension model. Short-text distillation ($\mathcal{L}_{\mathrm{short}}$) feeds the same short sequence with original positional indices to both models and maximizes cosine similarity between hidden states at $M$ layers chosen by attention-KL divergence plus the final layer. Short-to-long distillation ($\mathcal{L}_{\mathrm{s2l}}$) feeds the teacher a short sequence with normal indices and the student the same tokens with skipped positional indices that mimic long-range positions, then aligns only last-layer hidden states. These are added to the long-text language-modeling loss ($\mathcal{L}_{\mathrm{long}}$) with weights $\alpha_1$ and $\alpha_2$; the paper also defines hidden-state similarity and attention-KL divergence as diagnostic metrics for distribution drift.
What would settle it
Train a long-context model with the same data budget and the same three-way data mixture, but replace the hidden-state distillation loss with a parameter-space or output-space regularizer that keeps the extended model close to the original without optimizing cosine similarity of intermediate hidden states. If MMLU preservation matches LongReD while hidden-state similarity does not improve, the claimed mechanism is not necessary. Conversely, take a long-context model whose hidden-state similarity to the original is naturally high (for example, a lightly extended model) and check whether MMLU is preserved without any distillation; if not, similarity alone is insufficient.
Extended reading notes
Core claim
The central claim is that short-text degradation after context-window extension is driven by distribution drift and catastrophic forgetting, and that both can be counteracted by restoration distillation. LongReD combines ordinary long-text language modeling with two distillation objectives: short-text distillation, which maximizes cosine similarity between the extended model and the frozen original model on short sequences at a few selected layers, and short-to-long distillation, which uses skipped positional indices so that short text seen by the teacher aligns with long positions seen by the student. The paper shows this preserves short-text performance on 17 benchmarks while keeping RULER scores competitive with (and for Mistral-7B better than) continual pre-training baselines. It also shows that layer choice and the skipped-index sampling method (CREAM vs uniform) materially change the trade-off.
Load-bearing premise
The load-bearing premise is that making the extended model's hidden states resemble the original model's on short text is what actually preserves short-text task performance; if that correlation is a byproduct of extra training data or regularization instead of a cause, the method's motivation weakens, even though it might still work.
Editorial extensions
If this is right
- Short-text capabilities can be retained through context extension without a long-text penalty, so deployment of 32K-128K models on mixed-length workloads becomes safer.
- Distilling a small set of high-drift layers is better than distilling all layers, so efficient restoration is possible with modest extra compute (roughly 10% over the mixed-data baseline).
- The choice of skipped positional indices matters: CREAM wins at 4x extension while uniform sampling wins at 128K, so the sampling strategy should be tuned to the target context length.
- The diagnostic correlation between hidden-state similarity and MMLU preservation gives a cheap monitoring signal that could be used for early stopping during continual pretraining.
Reading between the lines
- The paper's correlation evidence does not prove causation; a plausible alternative is that distillation acts mainly as a regularizer that keeps the student close to the teacher, and a controlled test could compare LongReD against a weight-distance penalty at matched compute.
- If the distribution-drift story is right, LongReD's benefit should grow with extension ratio and RoPE base, since those create larger drift; this is testable by sweeping the base from 5e6 to 5e20.
- The same restoration idea could transfer to other post-training steps that modify representations, such as RLHF or domain adaptation, though the paper does not explore that direction.
- Skipped-position short-to-long distillation could serve as a data-efficient way to teach length generalization without generating long documents, which is a testable extension the paper's setup already enables.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the short-text performance degradation that follows RoPE-based context extension of LLMs. It first presents an empirical analysis attributing the degradation to two factors: distribution drift in hidden states and attention scores, and catastrophic forgetting during continual pre-training. It then proposes LongReD, a training framework that combines long-text language modeling with two distillation objectives: short-text distillation, which aligns hidden states of selected layers of the extended model with those of the original model on short texts, and short-to-long distillation, which aligns the final-layer output distribution on short texts under skipped positional indices. The method is evaluated by extending Llama-3-8B and Mistral-7B-v0.3 to 32K and 128K contexts using ABF or PI, measuring short-text performance on 17 benchmarks and long-text performance on RULER. The authors report that LongReD preserves short-text performance better than continual pre-training baselines while maintaining comparable long-text ability.
Significance. If the results are taken at face value, LongReD is a practical recipe for reducing short-text degradation during context extension, and the paper contributes a useful diagnostic framework based on hidden-state similarity and attention KL divergence. The ablations over distillation layers, distillation length, and the two loss weights are informative, and the release of code is a concrete asset for reproducibility. The formal bound in Appendix D.1 relating RoPE base to attention-score drift is a useful analysis contribution. However, the headline long-text claim is not supported by the paper's own Table 3 for the 128K Llama-3-8B ABF configuration, and the comparison against the Mix+CPT baseline is weakened by that baseline's collapse on reading-comprehension benchmarks. These issues do not invalidate the short-text preservation results, but they do require a substantial revision of the claims and the evidence.
major comments (4)
- [§5.2, Table 3] The abstract's claim of 'comparable or even better capacity to handle long texts' is directly contradicted by the 128K Llama-3-8B ABF row: RULER average is 69.70 for Long CPT, 69.64 for Mix CPT, but only 64.93 for LongReD-C and 68.41 for LongReD-U. Since RULER is the only long-text benchmark used, this is a direct counterexample within the reported results. The paper should either qualify the long-text claim to configurations where it holds (e.g., 32K and Mistral-128K) or provide multiple-seed evidence or an explanation for the 128K Llama drop.
- [§5.1, Table 16] The Mix+CPT baseline collapses on reading comprehension: for 32K and 128K ABF it reports SquadV2 scores of 0.89 and 1.08 and DROP scores of 0.17 and 0.21, while LongReD preserves scores around 70–74 on SquadV2 and around 49 on DROP. This collapse makes the 'consistently outperforms most baselines' claim partly an artifact of a broken baseline. The authors should report why Mix+CPT fails on these tasks, compare against Long CPT as the primary baseline, or otherwise show that the short-text improvement is not driven by this baseline failure.
- [Appendix B.1, Table 8] The D2 data mixture proportions sum to 1.856 (0.476 + 0.722 + 0.230 + 0.095 + 0.095 + 0.096 + 0.095 + 0.047), not to 1, so the exact training data cannot be reconstructed from the paper. The authors should provide normalized proportions or clearly state whether these are unnormalized weights; without this, the experimental setup is not reproducible.
- [§3.1 and §4.2, Eq. (8)] The diagnostic relationship in Figure 2 uses hidden-state similarity, which is exactly the quantity optimized by L_short in Eq. (8). The correlation between this metric and MMLU preservation is therefore partly self-confirming and does not by itself establish the causal claim that reducing hidden-state drift restores short-text performance. A control experiment that optimizes a different regularizer (e.g., output-distribution KL on short texts) or that evaluates the same distillation objective on held-out short-text tasks would strengthen the causal interpretation.
minor comments (5)
- [Table 15] The Mistral-7B-v0.3 128K ABF Mix rows list 'LongReD-C' twice; the second row (70.84, 65.6, 63.87, 55.48, 78.5) appears to be LongReD-U and should be relabeled.
- [Appendix A.1] The text says the positional indices are split into 'phead, ptail, and ptail'; this should read 'phead, pmid, and ptail'.
- [Eq. (18)] The notation 'L_f inal' contains a spurious space and should be typeset as L_final or L_final.
- [Limitations] The Limitations section contains a sentence fragment: 'However, a notable degradation in short-text performance after lightweight continual pretraining on only several billion tokens.' This should be completed and clarified.
- [Abstract and Conclusion] The abstract says 'comparable or even better' long-text performance, while the Conclusion says only 'comparable long-text modeling abilities'; these two claims should be aligned with the evidence actually reported.
Circularity Check
No significant circularity: the similarity metric doubles as the objective, but the central short-text claims are validated on independent benchmarks.
full rationale
The paper's closest approach to circularity is the overlap between the diagnostic metric and the training objective: Eq. 5 defines hidden-state cosine similarity Sim(H_l, H_hat_l), and Eq. 8 defines L_short = - sum Sim(...), so optimizing L_short directly maximizes the same quantity that Figure 2 correlates with MMLU preservation. This overlap is real but does not make the central claim circular, because the paper's central claim is about short-text benchmark performance, not about similarity. The similarity objective is a proxy; the paper demonstrates on held-out benchmarks (MMLU, HumanEval, PIQA, TriviaQA, etc., Table 3 and Appendix F) that training with this proxy improves those distinct metrics relative to Long CPT and Mix CPT baselines, and ablations (Table 4) show that removing L_short degrades short-text performance. The Figure 2 correlation is used as motivation, not as a fitted prediction of the outcome. Self-citations appear (Dong et al. 2024a for positional vectors; Men et al. 2024 and Dong et al. 2025 for layer selection) but they support auxiliary techniques that are also tested by ablations in this paper (Tables 5-7), and no load-bearing uniqueness theorem or ansatz is imported from the authors' prior work. The RULER gap for Llama-3-8B 128K LongReD-C (64.93 vs 69.70 Long CPT) is a support and correctness concern about the abstract's long-text claim, not a circularity of the derivation. Overall, the derivation chain is self-contained: the method is defined by explicit losses, and its claims are checked on independent benchmarks.
Assumptions & free parameters
free parameters (5)
- alpha_1 =
5 (32K), 2 (128K)
- alpha_2 =
10 (32K), 15 (128K)
- M (number of distillation layers) =
3 for Llama-3-8B 128K; 6 otherwise
- Ts (short-text distillation sequence length) =
1024
- D1:D2:D3 token ratio =
4:3:1
assumptions (6)
- domain assumption RoPE base scaling causes distribution drift, and larger bases cause larger drift with monotonic B(Theta) bounds.
- domain assumption Hidden-state cosine similarity and attention KL divergence are valid proxies for task-relevant distribution drift.
- domain assumption Reducing distribution drift on short texts via hidden-state distillation restores short-text task performance.
- domain assumption Catastrophic forgetting is a major cause of short-text degradation and can be mitigated by replay or distillation.
- domain assumption Skipped positional indices on short texts faithfully simulate long-position processing.
- standard math Standard mathematical tools: triangle inequality and Abel summation.
Cite this review
Pith. "Pith review of LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation." pith.science (2026). https://pith.science/paper/GJWP2GLQ
@misc{pith2026250207365,
author = {Pith},
title = {Pith review of: LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation},
year = {2026},
howpublished = {\url{https://pith.science/paper/GJWP2GLQ}},
note = {Machine review of arXiv:2502.07365}
}
read the original abstract
Large language models (LLMs) have gained extended context windows through scaling positional encodings and lightweight continual pre-training. However, this often leads to degraded performance on short-text tasks, while the reasons for this degradation remain insufficiently explored. In this work, we identify two primary factors contributing to this issue: distribution drift in hidden states and attention scores, and catastrophic forgetting during continual pre-training. To address these challenges, we propose Long Context Pre-training with Restoration Distillation (LongReD), a novel approach designed to mitigate short-text performance degradation through minimizing the distribution discrepancy between the extended and original models. Besides training on long texts, LongReD distills the hidden state of selected layers from the original model on short texts. Additionally, LongReD also introduces a short-to-long distillation, aligning the output distribution on short texts with that on long texts by leveraging skipped positional indices. Experiments on common text benchmarks demonstrate that LongReD effectively preserves the model's short-text performance while maintaining comparable or even better capacity to handle long texts than baselines. Our code is available at https://github.com/RUCAIBox/LongReD.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
LinearARD restores RoPE-scaled LLMs by exact linear-memory KL distillation of Q/Q, K/K, and V/V self-relations, reaching ~95% short-text recovery with 4.25M tokens.
Reference graph
Works this paper leans on
-
[1]
Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. 2024. On-policy distillation of language models: Learning from self-generated mistakes. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net
2024
-
[2]
Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J
Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. 2021. Program synthesis with large language models. CoRR, abs/2108.07732
arXiv 2021
-
[3]
Jiang, Jia Deng, Stella Biderman, and Sean Welleck
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen Marcus McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, and Sean Welleck. 2024. Llemma: An open language model for mathematics. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net
work page 2024
-
[4]
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. PIQA: reasoning about physical commonsense in natural language. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advan...
work page 2020
-
[5]
bloc97. 2023. NTK-Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation
work page 2023
-
[6]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert - Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litw...
2020
-
[7]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pond \' e de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bava...
arXiv 2021
-
[8]
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023. Extending context window of large language models via positional interpolation. CoRR, abs/2306.15595
arXiv 2023
Show all 81 references
-
[9]
Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia. 2024 a . Longlora: Efficient fine-tuning of long-context large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 20...
2024
-
[10]
Zhipeng Chen, Yingqian Min, Beichen Zhang, Jie Chen, Jinhao Jiang, Daixuan Cheng, Wayne Xin Zhao, Zheng Liu, Xu Miao, Yang Lu, Lei Fang, Zhongyuan Wang, and Ji - Rong Wen. 2025. An empirical study on eliciting and improving r1-like reasoning models. CoRR, abs/2503.04548
2025 arXiv
-
[11]
Zhipeng Chen, Liang Song, Kun Zhou, Wayne Xin Zhao, Bingning Wang, Weipeng Chen, and Ji - Rong Wen. 2024 b . Extracting and transferring abilities for building multi-lingual ability-enhanced large language models. CoRR, abs/2410.07825
2024 arXiv
-
[12]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. https://lmsys.org/blog/2023-03-30-vicuna/ Vicuna: An open-source chatbot impressing gpt-4 with 90\
2023
-
[13]
Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen - tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018. Quac: Question answering in context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - No...
2018
-
[14]
Christopher Clark, Kenton Lee, Ming - Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for C...
2019
-
[15]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the AI2 reasoning challenge. CoRR, abs/1803.05457
2018 arXiv
-
[16]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. CoRR, abs/2110.14168
2021 arXiv
-
[17]
OpenCompass Contributors. 2023. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass
2023
-
[18]
Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang. 2024. Longrope: Extending LLM context window beyond 2 million tokens. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27...
2024
-
[19]
Zican Dong, Junyi Li, Xin Men, Wayne Xin Zhao, Bingbing Wang, Zhen Tian, Weipeng Chen, and Ji - Rong Wen. 2024 a . Exploring context window of large language models via decomposed positional vectors. CoRR, abs/2405.18009
2024 arXiv
-
[20]
Zican Dong, Han Peng, Peiyu Liu, Wayne Xin Zhao, Dong Wu, Feng Xiao, and Zhifeng Wang. 2025. Domain-specific pruning of large mixture-of-experts models with few-shot demonstrations. CoRR, abs/2504.06792
2025 arXiv
-
[21]
Zican Dong, Tianyi Tang, Junyi Li, and Wayne Xin Zhao. 2023. A survey on long text modeling with transformers. CoRR, abs/2302.14502
2023 arXiv
-
[22]
Zican Dong, Tianyi Tang, Junyi Li, Wayne Xin Zhao, and Ji - Rong Wen. 2024 b . BAMBOO: A comprehensive benchmark for evaluating long text modeling capacities of large language models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Langu...
2024
-
[23]
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for C...
2019
-
[24]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al - Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Z...
2024 arXiv
-
[25]
emozilla. 2023. Dynamically Scaled RoPE further increases performance of long context LLaMA with zero fine-tuning
2023
-
[26]
Yao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue, Hannaneh Hajishirzi, Yoon Kim, and Hao Peng. 2024. Data engineering for scaling language models to 128k context. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net
2024
-
[27]
Tianyu Gao, Alexander Wettig, Howard Yen, and Danqi Chen. 2024. How to train long-context language models (effectively). CoRR, abs/2410.02660
2024
-
[28]
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2024. Minillm: Knowledge distillation of large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net
2024
-
[29]
Chi Han, Qifan Wang, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang. 2023. Lm-infinite: Simple on-the-fly length generalization for large language models. CoRR, abs/2308.16137
2023 arXiv
-
[30]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 a . Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . O...
2021
-
[31]
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021 b . Measuring mathematical problem solving with the MATH dataset. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchma...
2021
-
[32]
Hinton, Oriol Vinyals, and Jeffrey Dean
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015. Distilling the knowledge in a neural network. CoRR, abs/1503.02531
2015 arXiv
-
[33]
Namgyu Ho, Laura Schmid, and Se - Young Yun. 2023. Large language models are reasoning teachers. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 , pages 14852--14882....
2023
-
[34]
Cheng - Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. 2024. RULER: what's the real context size of your long-context language models? CoRR, abs/2404.06654
2024 arXiv
-
[35]
Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, Xinrong Zhang, Zhen Leng Thai, Kai Zhang, Chongyi Wang, Yuan Yao, Chenyang Zhao, Jie Zhou, Jie Cai, Zhongwu Zhai, Ning Ding, Chao Jia, Guoyang Zeng, Dahai Li, Z...
2024 arXiv
-
[36]
Zhiyuan Hu, Yuliang Liu, Jinman Zhao, Suyuchen Wang, Yan Wang, Wei Shen, Qing Gu, Anh Tuan Luu, See-Kiong Ng, Zhiwei Jiang, et al. 2024 b . Longrecipe: Recipe for efficient long context generalization in large language models. CoRR, abs/2409.00509
2024 arXiv
-
[37]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L \' e lio Renard Lavaud, Marie - Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, ...
2023 arXiv
-
[38]
Jinhao Jiang, Junyi Li, Wayne Xin Zhao, Yang Song, Tao Zhang, and Ji - Rong Wen. 2024. Mix-cpt: A domain adaptation framework via decoupling knowledge learning and format alignment. CoRR, abs/2407.10804
2024 arXiv
-
[39]
Weld, and Luke Zettlemoyer
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Can...
2017
-
[40]
Chen Liang, Simiao Zuo, Qingru Zhang, Pengcheng He, Weizhu Chen, and Tuo Zhao. 2023. Less is more: Task-aware layer-wise distillation for language model compression. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA , volume 202...
2023
-
[41]
Hao Liu, Matei Zaharia, and Pieter Abbeel. 2023. Ring attention with blockwise transformers for near-infinite context. CoRR, abs/2310.01889
2023 arXiv
-
[42]
Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net
2019
-
[43]
McAuley, Han Hu, Torsten Scholak, S \' e bastien Paquet, Jennifer Robinson, Carolyn Jane Anderson, Nicolas Chapados, and et al
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy - Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul,...
2024 arXiv
-
[44]
Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, and Weipeng Chen. 2024. Shortgpt: Layers in large language models are more redundant than you expect. CoRR, abs/2403.03853
2024 arXiv
-
[45]
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? A new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October ...
2018
-
[46]
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah. 2023. Orca: Progressive learning from complex explanation traces of GPT-4 . CoRR, abs/2306.02707
2023 arXiv
-
[47]
Saurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jan Kautz, and Pavlo Molchanov. 2024. Compact language models via pruning and knowledge distillation. CoRR, abs/2407.14679
2024 arXiv
-
[48]
OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774
2023 arXiv
-
[49]
Denis Paperno, Germ \' a n Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fern \' a ndez. 2016. The LAMBADA dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th Ann...
2016
-
[50]
Leonid Pekelis, Michael Feil, Forrest Moret, Mark Huang, and Tiffany Peng. 2024. https://doi.org/10.57967/hf/3372 Llama 3 gradient: A series of long context models
2024 doi
-
[51]
Guilherme Penedo, Hynek Kydl \' cek, Loubna Ben Allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro von Werra, and Thomas Wolf. 2024. The fineweb datasets: Decanting the web for the finest text data at scale. CoRR, abs/2406.17557
2024 arXiv
-
[52]
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023 a . Instruction tuning with GPT-4 . CoRR, abs/2304.03277
2023 arXiv
-
[53]
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. 2023 b . Yarn: Efficient context window extension of large language models. CoRR, abs/2309.00071
2023 arXiv
-
[54]
Han Peng, Jinhao Jiang, Zican Dong, Wayne Xin Zhao, and Lei Fang. 2025. Cafe: Retrieval head-based coarse-to-fine information seeking to enhance multi-document qa capability. arXiv preprint arXiv:2505.10063
2025 arXiv
-
[55]
Smith, and Mike Lewis
Ofir Press, Noah A. Smith, and Mike Lewis. 2022. Train short, test long: Attention with linear biases enables input length extrapolation. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net
2022
-
[56]
Rae, Anna Potapenko, Siddhant M
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. 2020. Compressive transformers for long-range sequence modelling. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . Ope...
2020
-
[57]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100, 000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016 ...
2016
-
[58]
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019. Social iqa: Commonsense reasoning about social interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Na...
2019
-
[59]
Daria Soboleva, Faisal Al-Khateeb, Robert Myers, Jacob R Steeves, Joel Hestness, and Nolan Dey. 2023. SlimPajama: A 627B token cleaned and deduplicated version of RedPajama
2023
-
[60]
Jianlin Su, Murtadha H. M. Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063
2024
-
[61]
Le, Ed H
Mirac Suzgun, Nathan Scales, Nathanael Sch \" a rli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. 2023. Challenging big-bench tasks and whether chain-of-thought can solve them. In Findings of the Associatio...
2023
-
[62]
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...
2019
-
[63]
Tianyi Tang, Yiwen Hu, Bingqian Li, Wenyang Luo, Zijing Qin, Haoxiang Sun, Jiapeng Wang, Shiyi Xu, Xiaoxue Cheng, Geyang Guo, Han Peng, Bowen Zheng, Yiru Tang, Yingqian Min, Yushuo Chen, Jie Chen, Yuanqian Zhao, Luran Ding, Yuhao Wang, Zican Dong, Chunxuan Xia, Junyi Li, Kun Z...
2024 arXiv
-
[64]
Xinyu Tang, Xiaolei Wang, Zhihao Lv, Yingqian Min, Wayne Xin Zhao, Binbin Hu, Ziqi Liu, and Zhiqiang Zhang. 2025 a . Unlocking general long chain-of-thought reasoning capabilities of large language models via representation engineering. CoRR, abs/2503.11314
2025 arXiv
-
[65]
Xinyu Tang, Xiaolei Wang, Wayne Xin Zhao, Siyuan Lu, Yaliang Li, and Ji - Rong Wen. 2025 b . Unleashing the potential of large language models as prompt optimizers: Analogical analysis with gradient-based model optimizers. In AAAI-25, Sponsored by the Association for the Advan...
2025
-
[66]
Xinyu Tang, Xiaolei Wang, Xin Zhao, and Ji - Rong Wen. 2025 c . DAWN-ICL: strategic planning of problem-solving trajectories for zero-shot in-context learning. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Lin...
2025
-
[67]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton - Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu,...
2023 arXiv
-
[68]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30
2017
-
[69]
Yuhao Wang, Ruiyang Ren, Junyi Li, Xin Zhao, Jing Liu, and Ji - Rong Wen. 2024. REAR: A relevance-aware retrieval-augmented framework for open-domain question answering. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miam...
2024
-
[70]
Yuhao Wang, Ruiyang Ren, Yucheng Wang, Wayne Xin Zhao, Jing Liu, Hua Wu, and Haifeng Wang. 2025. Unveiling knowledge utilization mechanisms in llm-based retrieval-augmented generation. CoRR, abs/2505.11995
2025 arXiv
-
[71]
Tong Wu, Yanpeng Zhao, and Zilong Zheng. 2024 a . An efficient recipe for long context extension via middle-focused positional encoding. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
2024
-
[72]
Tongtong Wu, Linhao Luo, Yuan - Fang Li, Shirui Pan, Thuy - Trang Vu, and Gholamreza Haffari. 2024 b . Continual learning for large language models: A survey. CoRR, abs/2402.01364
2024 arXiv
-
[73]
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2023. Efficient streaming language models with attention sinks. CoRR, abs/2309.17453
2023 arXiv
-
[74]
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, S...
2024
-
[75]
Xiaohan Xu, Ming Li, Chongyang Tao, Tao Shen, Reynold Cheng, Jinyang Li, Can Xu, Dacheng Tao, and Tianyi Zhou. 2024. A survey on knowledge distillation of large language models. CoRR, abs/2402.13116
2024 arXiv
-
[76]
Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong, Mingdao Liu, Minlie Huang, Peng...
2024 arXiv
-
[77]
Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu. 2024. Long context transfer from language to vision. CoRR, abs/2406.16852
2024 arXiv
-
[78]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian - Yun Nie, and Ji - Ro...
2023 arXiv
-
[79]
Dawei Zhu, Nan Yang, Liang Wang, Yifan Song, Wenhao Wu, Furu Wei, and Sujian Li. 2024. Pose: Efficient context window extension of llms via positional skip-wise training. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11,...
2024
-
[80]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[81]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.