REVIEW 4 major objections 4 minor 1 cited by
Memory Limitations of Prompt Tuning in Transformers
T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Prompt tuning in transformers can store at most linearly as much information as the prompt length, and beyond that the fraction of reachable outputs decays exponentially.
desk verdict New linear-scaling bound for prompt tuning is solid; the exponential-decay theorem has a real packing-scale error that inflates the rate, but the qualitative point likely survives. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key mechanism is a volume-versus-Lipschitz counting argument. A transformer, or its mean-field generalization as a map between Wasserstein spaces of probability measures, is L-Lipschitz on bounded inputs, so any ε/L-ball of prompts can only produce outputs within ε of one another. The number of distinguishable output sequences or distributions is therefore bounded by how many such balls fit in the output space, measured through covering and packing numbers. For the mean-field theorem, the crucial lower bound is that the set of discrete output distributions has Wasserstein covering number at least exp(1/ε^d), giving the exponential decay in k.
What would settle it
Compute or lower-bound the Wasserstein covering number of the set of empirical measures with n atoms in the d-dimensional ball for fixed n; if it grows only polynomially in 1/ε rather than exponentially, Theorem 4.10's exponential decay cannot hold for outputs with n atoms. Alternately, exhibit a fixed transformer whose prompt of length mp reliably reproduces k ≫ mp/m prescribed input/output pairs within tolerance ε—Theorem 4.7 says such a transformer can succeed only on a vanishingly small fraction of output sequences.
Extended reading notes
Core claim
The central claim is that prompt tuning has an inherent memory ceiling. Theorem 4.7 bounds, in volume terms, the fraction of output sequences that a fixed transformer can approximately produce by varying only the prepended prompt: once the number k of stored input/output pairs exceeds a constant multiple of mp/m, that fraction decays exponentially in k. Theorem 4.10 removes dependence on prompt length by passing to the mean-field transformer acting on probability measures; it states that the proportion of output distributions that are ε-accessible through any prompt is at most O(exp(-k 3^d / ε^d)) once k is large enough relative to the Lipschitz constant, embedding radius, dimension, and tol
Load-bearing premise
The load-bearing premise is that the space of possible prompt-produced output distributions is as large as the Wasserstein covering lower bound N(G, W_q, ε) ≥ exp(1/ε^d)/C; if finite-atom empirical measures actually have smaller metric entropy, the exponential unaccessibility of Theorem 4.10 collapses to a weaker polynomial decay.
Editorial extensions
If this is right
- For in-context learning, a pre-prompt listing k example pairs can reliably memorize at most k ≈ C·mp/m pairs; adding more examples cannot make all of them reliably recallable through the prompt alone.
- Long-context performance degradation is not merely a training or optimization artifact: for a fixed trained transformer, the fraction of output distributions reachable by any prompt shrinks exponentially in the number of stored pairs, independent of context size.
- The linear scaling k ∈ O(mp/m) is optimal for encoding information as input/output pairs in the prompt, which implies soft-prompt optimization can gain at most linearly over discrete prompt engineering in this memorization setting.
- The main results also hold for masked (causal) self-attention, so decoder-only transformer language models are covered by the limitation.
- For single-layer transformers, the reachable output set is essentially a low-dimensional subspace, so even approximate memorization of two input/output pairs sharing a token fails for generic transformers.
Reading between the lines
- A testable corollary the authors do not draw: measuring how many random key-value pairs a fixed LLM can retrieve through prompt search as mp grows should show a sharp threshold with slope about 1/m, set by the effective Lipschitz constant and embedding radius of the model.
- The exponential factor in embedding dimension d suggests that high-dimensional token embeddings amplify this memorization bottleneck: they make the total output space huge while the set reachable by any prompt stays relatively small.
- If the bound is tight, retrieval-augmented generation or weight updates are not merely conveniences but structural necessities: external memory changes the input distribution or the map itself, bypassing the covering-number constraint that limits prompt-only memorization.
- The mean-field theorem counts prompts as empirical distributions, so its exponential decay rate should hold regardless of prompt token count once the prompt distribution is rich enough; a natural extension would be to make the pre-exponential dependence on prompt length explicit.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies memorization through prompt tuning in transformers. It introduces a notion of ε-accessibility for output sequences and output distributions and proves two main results. Theorem 4.7 states that the number k of input/output pairs of length m that a transformer can memorize via a prompt of length m_p is at most O(m_p/m): the proportion of ε-accessible output sequences decays exponentially with k. Theorem 4.10, in a mean-field/Wasserstein formulation, claims that the proportion of output distributions accessible via arbitrary-length prompts decays as O(exp(-k 3^d/ε^d)), independent of prompt length, for a transformer with Lipschitz constant L, embedding radius r, and dimension d. Section 5 adds statements that single-layer transformers have very limited prompt-tuning expressivity. Theorems 4.7 and 4.10 are also claimed to hold for masked self-attention.
Significance. If correct, the paper would be a valuable theoretical complement to empirical observations of long-context degradation: it would show that prompt-based memorization is bounded by prompt length and that the fraction of target distributions accessible by prompt tuning shrinks exponentially in the number of stored items. The use of mean-field transformers and Wasserstein metric entropy is appropriate, and the main derivations are explicit in L, r, d, and ε, with no fitted constants. However, the central quantitative claim of Theorem 4.10 is currently not established as stated because of a packing-scale error in Appendix D, and Section 5 contains statements that are internally inconsistent. The qualitative conclusion may survive after repair, but the paper requires substantial revision before the advertised results can be accepted.
major comments (4)
- [Appendix D / Theorem 4.10] The packing step inverts the covering scale. To count 3ε-separated target distributions one needs M(G,3ε) ≥ N(G,3ε) ≥ C^{-1} exp(1/(3ε)^d) = C^{-1} exp(3^{-d}/ε^d). The proof instead uses the lower bound exp(3^d/ε^d), which corresponds to scale ε/3. An ε/3-separated family is not a 3ε-packing: one accessible output can lie within ε of many such targets, so the Cin/Cout counting in the proof is invalid. Consequently Theorem 4.10's rate O(exp(-k 3^d/ε^d)) and the displayed threshold are not established. The same argument, after repair, yields at best O(exp(-k/(3ε)^d)) (up to constants), still exponential in k but with different ε-dependence.
- [Definitions 4.4–4.5 and Theorem 4.10] The object counted in Theorem 4.10 is ambiguous. Definition 4.5 defines accessibility for a single output distribution μY, but the proof in Appendix D counts k-tuples of independent output spaces, since Cout is the k-th power of the single-space packing number. If the intended statement is about k-tuples (μY_1,...,μY_k), the definition and theorem should say so. Moreover, 'proportion' over P_c has no canonical volume; it must be defined explicitly as a metric-entropy/packing ratio. Without this clarification the theorem is not a well-defined quantitative claim.
- [Theorem 5.7 / Appendix E] The dimension formula is inconsistent. For h=1 and d=5, the formula gives (d-2)!/(d-4)! = 6, but R^5 has no 6-dimensional subspace. The formula appears to count ordered orthogonal frames rather than the dimension of a vector space. In addition, the statement quantifies (y1,...,y_{h+1}) but the conclusion only refers to i∈{1,2}. The proof of Lemma E.1 is a chain of unexplained identities and does not rigorously establish the advertised accessibility statement. Theorem 5.7 needs to be restated and reproved.
- [Theorem 5.8 / Assumption 5.5] Theorem 5.8 uses invertibility of the MLP via Behrmann et al. (the condition ∥W1∥2·∥W2∥2<1) in its proof, but this assumption is absent from the theorem statement. The conclusion 'τ([P,xi,x0])^{-1} ≥ (1−∥W1∥2·∥W2∥2)r/2' compares a vector to a scalar; presumably a norm is intended. The hypotheses and conclusion of the theorem must be stated precisely.
minor comments (4)
- [Throughout] Typos and wording issues include 'refered' in the Introduction, 'independant' in §4.3, and 'Lipshitz' in the footnote; a careful proofreading pass is needed.
- [Section 5 notation] The notation τ([P,xi,x0])^{-1} uses -1 as a last-column index per Section 1.3, but in Section 5 it appears without recalling this convention and is easy to misread as an inverse; define it explicitly at first use.
- [Theorem 4.7] The phrase 'proportion (in terms of volume)' should be formalized: the proof counts packing balls inside B^{dm}(0,r) and uses Cin/Cout as a volume-ratio surrogate. The reference measure and possible boundary effects should be stated.
- [Proposition 3.6 / Appendix B] The proof imports Kloeckner's critical-exponent result without stating the theorem or verifying its hypotheses for the set G of empirical measures. Since this lower bound is load-bearing for Theorem 4.10, please state the external theorem precisely and justify that G has the required critical exponent.
Circularity Check
No significant circularity: the main theorems are derived from external metric-entropy and Lipschitz estimates, and no fitted parameter is disguised as a prediction.
full rationale
The paper's central claims (Theorem 4.7 and Theorem 4.10) are derived, not assumed: they combine (i) external Lipschitz constants for (mean-field) self-attention (Castin et al. 2024, Geshkovski et al. 2023), (ii) external covering/packing estimates for Wasserstein space over discrete measures (Nguyen 2013, Kloeckner 2012/2014), and (iii) a discretization/volume argument from Vershynin. There are no self-citations by the current authors, and no prior result is cited that itself assumes the conclusion of this paper. No parameter is fitted to data and then renamed as a prediction: the constants L, r, d, eps appear explicitly as assumptions, and the exponential-decay bound is a consequence of the covering-number lower bound, not an input to it. The theorem about linear-in-prompt-length memorization (Theorem 4.7) uses the same covering/packing logic and does not presuppose the limitation it proves. The Section 5 results generalize Wang et al. by relaxing assumptions and using dimension-counting; they do not import a uniqueness or forced-choice conclusion from the authors' own prior work. A possible scale error in Appendix D (using epsilon/3-separated outputs as if they were 3epsilon-separated) would be a correctness defect in the packing argument, not a circularity: the proof would still be an attempted derivation from external estimates rather than a reduction of the theorem to its own assumptions. Therefore the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (5)
- domain assumption For inputs in a bounded ball, the transformer and its mean-field generalization are L-Lipschitz with constants given in Propositions 2.11, 2.12, 2.16 (from Castin et al. 2024, Geshkovski et al. 2023).
- standard math The covering numbers of the space of probability measures G satisfy N(G, W_q, epsilon) <= exp(O(epsilon^{-d})) and N(G, W_q, epsilon) >= (1/C) exp(1/epsilon^d) (Propositions 3.5 and 3.6, citing Nguyen 2013 and Kloeckner 2012/2014).
- domain assumption Input tokens, including pre-prompt tokens, have norm bounded by r (the embedding radius).
- domain assumption The MLP is invertible when ||W1||_2 * ||W2||_2 < 1, used in the proof of Theorem 5.8 (from Behrmann et al. 2019).
- domain assumption Mean-field transformer layers generalize finite transformer layers on empirical measures: T(M(X)) = M(tau(X)).
Cite this review
Pith. "Pith review of Memory Limitations of Prompt Tuning in Transformers." pith.science (2026). https://pith.science/paper/SPI2MQQC
@misc{pith2026250900421,
author = {Pith},
title = {Pith review of: Memory Limitations of Prompt Tuning in Transformers},
year = {2026},
howpublished = {\url{https://pith.science/paper/SPI2MQQC}},
note = {Machine review of arXiv:2509.00421}
}
read the original abstract
Despite the empirical success of prompt tuning in adapting pretrained language models to new tasks, theoretical analyses of its capabilities remain limited. Existing theoretical work primarily addresses universal approximation properties, demonstrating results comparable to standard weight tuning. In this paper, we explore a different aspect of the theory of transformers: the memorization capability of prompt tuning. We provide two principal theoretical contributions. First, we prove that the amount of information memorized by a transformer cannot scale faster than linearly with the prompt length. Second, and more importantly, we present the first formal proof of a phenomenon empirically observed in large language models: performance degradation in transformers with extended contexts. We rigorously demonstrate that transformers inherently have limited memory, constraining the amount of information they can retain, regardless of the context size. This finding offers a fundamental understanding of the intrinsic limitations of transformer architectures, particularly their ability to handle long sequences.
Forward citations
Cited by 1 Pith paper
-
Training-Free Universal Approximation by Prompting Random Transformers
Frozen random-weight attention transformers can emulate kernel regression and approximate Hölder functions at minimax-optimal rates, with soft prompts constructed by solving linear systems.
Reference graph
Works this paper leans on
-
[1]
Jens Behrmann, Will Grathwohl, Ricky T. Q. Chen, David Duvenaud, and Joern-Henrik Jacobsen. Invertible residual networks. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 573--582. PMLR, 09--15 Jun 2019. URL https://pr...
work page 2019
-
[2]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
1901
-
[3]
How smooth is attention? In ICML, 2024
Valérie Castin, Pierre Ablin, and Gabriel Peyré. How smooth is attention? In ICML, 2024. URL https://arxiv.org/abs/2312.14820
arXiv 2024
-
[4]
PLOT : Prompt learning with optimal transport for vision-language models
Guangyi Chen, Weiran Yao, Xiangchen Song, Xinyue Li, Yongming Rao, and Kun Zhang. PLOT : Prompt learning with optimal transport for vision-language models. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=zqwryBoXYnh
work page 2023
-
[5]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sashank Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bra...
2023
-
[6]
BERT : Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, V...
2019
-
[7]
Longrope: extending llm context window beyond 2 million tokens
Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang. Longrope: extending llm context window beyond 2 million tokens. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org, 2024
work page 2024
-
[8]
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, Xu Sun, Lei Li, and Zhifang Sui. A survey on in-context learning. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1107--1128, Miami, ...
Show all 55 references
-
[9]
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. Attention is not all you need: Pure attention loses rank doubly exponentially with depth. In International conference on machine learning, pages 2793--2803. PMLR, 2021
2021
-
[10]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...
2021
-
[11]
Nemesis: Normalizing the soft-prompt vectors of vision-language models
Shuai Fu, Xiequn Wang, Qiushi Huang, and Yu Zhang. Nemesis: Normalizing the soft-prompt vectors of vision-language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=zmJDzPh1Dm
2024
-
[12]
Protein multimer structure prediction via prompt learning
Ziqi Gao, Xiangguo Sun, Zijing Liu, Yu Li, Hong Cheng, and Jia Li. Protein multimer structure prediction via prompt learning. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=OHpvivXrQr
2024
-
[13]
The emergence of clusters in self-attention dynamics
Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. The emergence of clusters in self-attention dynamics. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages...
2023
-
[14]
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. Universal language model fine-tuning for text classification. In Iryna Gurevych and Yusuke Miyao, editors, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 328--339, Melbou...
2018 doi
-
[15]
RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling, 2024
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, and Boris Ginsburg. RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=kIoBbc76Sy
2024
-
[16]
Fundamental limits of prompt tuning transformers: Universality, capacity and efficiency
Jerry Yao-Chieh Hu, Wei-Po Wang, Ammar Gilani, Chenyang Li, Zhao Song, and Han Liu. Fundamental limits of prompt tuning transformers: Universality, capacity and efficiency. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net...
2025
-
[17]
Long-context LLM s meet RAG : Overcoming challenges for long inputs in RAG
Bowen Jin, Jinsung Yoon, Jiawei Han, and Sercan O Arik. Long-context LLM s meet RAG : Overcoming challenges for long inputs in RAG . In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=oU3tpaR8fm
2025
-
[18]
Are transformers with one layer self-attention using low-rank weight matrices universal approximators? In The Twelfth International Conference on Learning Representations, 2024
Tokio Kajitsuka and Issei Sato. Are transformers with one layer self-attention using low-rank weight matrices universal approximators? In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=nJnky5K944
2024
-
[19]
On the optimal memorization capacity of transformers
Tokio Kajitsuka and Issei Sato. On the optimal memorization capacity of transformers. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=UGVYezlLcZ
2025
-
[20]
Maple: Multi-modal prompt learning
Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. Maple: Multi-modal prompt learning. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19113--19122, 2023. doi:10.1109/CVPR52729.2023.01832
2023
-
[21]
The lipschitz constant of self-attention
Hyunjik Kim, George Papamakarios, and Andriy Mnih. The lipschitz constant of self-attention. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 5562--5571....
2021
-
[22]
Provable memorization capacity of transformers
Junghwan Kim, Michelle Kim, and Barzan Mozafari. Provable memorization capacity of transformers. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=8JCg5xJCTPR
2023
-
[23]
A generalization of hausdorff dimension applied to hilbert cubes and wasserstein spaces
Benoit Kloeckner. A generalization of hausdorff dimension applied to hilbert cubes and wasserstein spaces. Journal of Topology and Analysis, 04 0 (02): 0 203--235, 2012. doi:10.1142/S1793525312500094. URL https://doi.org/10.1142/S1793525312500094
2012 doi
-
[24]
Kloeckner
Benoît R. Kloeckner. A geometric study of wasserstein spaces: Ultrametrics. Mathematika, 61 0 (1): 0 162–178, May 2014. ISSN 2041-7942. doi:10.1112/s0025579314000059. URL http://dx.doi.org/10.1112/S0025579314000059
2014 doi
-
[25]
Attention is not only a weight: Analyzing transformers with vector norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. Attention is not only a weight: Analyzing transformers with vector norms. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Langua...
2020 doi
-
[26]
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS '22, Red Hook, NY, USA, 2022. Curran Associates ...
2022
-
[27]
Summary of a haystack: A challenge to long-context LLM s and RAG systems
Philippe Laban, Alexander Fabbri, Caiming Xiong, and Chien-Sheng Wu. Summary of a haystack: A challenge to long-context LLM s and RAG systems. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Lang...
2024 doi
-
[28]
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processin...
2021 doi
-
[29]
Same task, more tokens: the impact of input length on the reasoning performance of large language models
Mosh Levy, Alon Jacoby, and Yoav Goldberg. Same task, more tokens: the impact of input length on the reasoning performance of large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computa...
2024 doi
-
[30]
Jiaqi Li, Mengmeng Wang, Zilong Zheng, and Muhan Zhang. L oo GLE : Can long-context language models understand long contexts? In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vol...
2024 doi
-
[31]
Extending context window in large language models with segmented base adjustment for rotary position embeddings
Rongsheng Li, Jin Xu, Zhixiong Cao, Hai-Tao Zheng, and Hong-Gee Kim. Extending context window in large language models with segmented base adjustment for rotary position embeddings. Applied Sciences, 14 0 (7): 0 3076, 2024 b
2024
-
[32]
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Pap...
2021 doi
-
[33]
Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12: 0 157--173, 2024. doi:10.1162/tacl_a_00638...
2024 doi
-
[34]
Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. Generating wikipedia by summarizing long sequences. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=Hyg0vbWC-
2018
-
[35]
P -tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. P -tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting ...
2022 doi
-
[36]
Memorization capacity of multi-head attention in transformers
Sadegh Mahdavi, Renjie Liao, and Christos Thrampoulidis. Memorization capacity of multi-head attention in transformers. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=MrR3rMxqqv
2024
-
[37]
A theoretical framework for prompt engineering: Approximating smooth functions with transformer prompts, 2025
Ryumei Nakada, Wenlong Ji, Tianxi Cai, James Zou, and Linjun Zhang. A theoretical framework for prompt engineering: Approximating smooth functions with transformer prompts, 2025. URL https://arxiv.org/abs/2503.20561
2025 arXiv
-
[38]
Convergence of latent mixing measures in finite and infinite mixture models
XuanLong Nguyen. Convergence of latent mixing measures in finite and infinite mixture models. The Annals of Statistics, 41 0 (1): 0 370--400, 2013. ISSN 00905364, 21688966. URL http://www.jstor.org/stable/41806611
2013
-
[39]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
-
[40]
On the role of attention in prompt-tuning
Samet Oymak, Ankit Singh Rawat, Mahdi Soltanolkotabi, and Christos Thrampoulidis. On the role of attention in prompt-tuning. In Proceedings of the 40th International Conference on Machine Learning, ICML'23. JMLR.org, 2023
2023
-
[41]
Prompting a pretrained transformer can be a universal approximator
Aleksandar Petrov, Adel Bibi, and Philip Torr. Prompting a pretrained transformer can be a universal approximator. In ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2024 a . URL https://openreview.net/forum?id=z7LOXziWxH
2024
-
[42]
When do prompting and prefix-tuning work? a theory of capabilities and limitations
Aleksandar Petrov, Philip Torr, and Adel Bibi. When do prompting and prefix-tuning work? a theory of capabilities and limitations. In The Twelfth International Conference on Learning Representations, 2024 b . URL https://openreview.net/forum?id=JewzobRhay
2024
-
[43]
Sander, Pierre Ablin, Mathieu Blondel, and Gabriel Peyr\'e
Michael E. Sander, Pierre Ablin, Mathieu Blondel, and Gabriel Peyr\'e. Sinkformers: Transformers with doubly stochastic attention. In Gustau Camps-Valls, Francisco J. R. Ruiz, and Isabel Valera, editors, Proceedings of The 25th International Conference on Artificial Intelligen...
2022
-
[44]
Optimal transport for applied mathematicians
Filippo Santambrogio. Optimal transport for applied mathematicians. Springer, 2015
2015
-
[45]
De PT : Decomposed prompt tuning for parameter-efficient fine-tuning
Zhengxiang Shi and Aldo Lipani. De PT : Decomposed prompt tuning for parameter-efficient fine-tuning. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=KjegfPGRde
2024
-
[46]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural In...
2017
-
[47]
High-Dimensional Probability: An Introduction with Applications in Data Science
Roman Vershynin. High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 2nd edition, 2025
2025
-
[48]
Optimal Transport: Old and New, volume 338 of Grundlehren der mathematischen Wissenschaften
Cédric Villani. Optimal Transport: Old and New, volume 338 of Grundlehren der mathematischen Wissenschaften. Springer, Berlin, 2008
2008
-
[49]
Universality and limitations of prompt tuning
Yihan Wang, Jatin Chauhan, Wei Wang, and Cho-Jui Hsieh. Universality and limitations of prompt tuning. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 75623--75643. Curran Asso...
2023
-
[50]
Multitask prompt tuning enables parameter-efficient transfer learning
Zhen Wang, Rameswar Panda, Leonid Karlinsky, Rogerio Feris, Huan Sun, and Yoon Kim. Multitask prompt tuning enables parameter-efficient transfer learning. In The Eleventh International Conference on Learning Representations, 2023 b . URL https://openreview.net/forum?id=Nk2pDtuhTq
2023
-
[51]
Chi, Quoc V
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the 36th International Conference on Neural Information Processing Sys...
2022
-
[52]
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, and Tom Goldstein. Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://ope...
2023
-
[53]
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[54]
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12104--12113, 2022
2022
-
[55]
Point transformer
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun. Point transformer. In Proceedings of the IEEE/CVF international conference on computer vision, pages 16259--16268, 2021
2021
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.