REVIEW 3 major objections 5 minor 62 references
Length Representations in Large Language Models
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The second attention layer of an LLM carries a length dial: probing finds individual hidden units, and scaling them lengthens or shortens output without losing meaning.
desk verdict Useful empirical mapping of length control to early attention units, but the position confound and in-sample selection make the disentanglement claim weaker than the abstract suggests. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing instruments are (1) a two-layer neural-network regressor that predicts the generation time step from hidden states and per-unit R2 values naming the length units, and (2) a causal scaling intervention that multiplies those top-k units in the second layer's attention output by positive or negative constants during decoding. The second-layer attention output is the object: it is where length predictions concentrate, and the scaling experiment is what turns correlational probes into evidence of control. The Priming prompt, which states source length and target kept-token count, is the condition that shifts activation onto a distinct set of units and makes single-unit scaling effective.
What would settle it
Train the per-unit probes only on hidden states taken at the first generated token, a fixed position, across summaries that end at different lengths; if the top-k 'length units' no longer predict final output length with R2 far above zero, the probes were reading absolute position rather than length, and the scaling result would not prove length control.
Extended reading notes
Core claim
Stated on the paper's own terms: large language models encode output sequence length as a locally concentrated, unit-level feature in the attention output of the second transformer layer, and this feature is partially separable from the semantic content being generated. A two-layer regression probe trained to predict the generation time step from hidden states gives the highest R2 for that layer's attention output across model families and precisions; per-unit probes then single out a handful of hidden units whose activity tracks length, and the top-1 unit alone can shift compression ratio when scaled. Multiplying these units by negative constants makes the model generate longer summaries, positive constants shorter ones, and human evaluation under the length-priming prompt shows informativeness is preserved. The same units recur after fine-tuning, and units found on summarization also steer translation and story generation, which the authors read as evidence of a learned length mechanism rather than a prompt artifact.
Load-bearing premise
The identification of 'length units' assumes that high probe accuracy for the generation time step means a hidden unit encodes output length, rather than the position of the token being generated or the identity of that token.
Editorial extensions
If this is right
- Prompt-free length control is possible for standard multi-head-attention models by rescaling the top-k units in the second layer's attention output.
- Length information survives 4-bit and 8-bit quantization, so compressed models retain the internal length code.
- The same top-3 length units are active under Priming prompts in both zero-shot and fine-tuned models, supporting the view that in-context learning recruits the same circuitry that fine-tuning strengthens.
- Length units identified on sentence compression transfer to machine translation and story generation, indicating a shared length axis across tasks.
- Scaling the least-active units does not change the generated text, which is the control showing the top-k effect is not a generic perturbation.
Reading between the lines
- If length is a genuinely separable early-layer feature, then production systems could steer length by a direct unit-level gain instead of prompt phrasing, and the same identification-plus-scaling recipe could be tried for other discrete generation attributes such as formality, stance, or domain.
- The paper's own negative result for grouped-query attention in zero-shot settings suggests that shared key/value projections move or scatter the length code; a testable extension is to probe the shared projections and per-head output rather than the concatenated hidden units.
- Because the probe target is the current generation step, the identified units may encode how far the sequence has gone rather than when it should stop; an experiment that conditions on fixed total lengths or predicts remaining steps would separate these two readings.
- The transfer of summarization-identified units to translation and story generation implies a task-general length axis; if confirmed, length steering should work by the same unit gains across decoding strategies with only calibration of the gain value.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether large language models encode output sequence length as an internal, partially disentangled feature. Using Google sentence summarization data with three prompt formats (No-constraint, Length, Priming), the authors train a two-layer neural network to predict the generation time step from hidden states extracted at different transformer components and layers, reporting R2 scores across Llama-2, Llama-3, Phi-3, and Qwen-2.5 models under different precision settings. They find that the layer-2 attention output yields the highest R2, identify individual hidden units with high per-unit R2, and show that scaling the top-k units with positive or negative factors changes summary length while leaving the smallest-k units ineffective. They also report human evaluation results, cross-task transfer to machine translation and story generation, and shared units between zero-shot and fine-tuned settings. The paper concludes that length information is partially disentangled from semantic information and that LLMs possess internal, unit-level length control mechanisms.
Significance. If the central claim holds, the paper makes a useful contribution to mechanistic interpretability of LLMs by identifying unit-level length-controlling directions that are partly independent of content, with potential applications to controllable generation. The strengths of the paper are its breadth of models, sizes, precision settings, and prompting conditions; the inclusion of an actual intervention (scaling units) with a smallest-k control; human evaluation; cross-task transfer experiments; and released code. However, the interpretation of the probe as measuring length rather than position is the load-bearing step. Because the probe target is the generation time step, which coincides with absolute position during generation and is confounded with positional encoding and token order, the current evidence does not yet separate length representations from position representations. The modest per-unit R2 values and the fact that unit selection and scaling evaluation share the same corpus further weaken the identification step. These issues are fixable, so the result is potentially valuable but not yet established.
major comments (3)
- [Section 3.1 and Table 4] The probe target is the generation time step n, which is identical to the absolute position of the token being generated. Because Llama, Phi, and Qwen hidden states carry positional information through RoPE and through token order, the high R2 of the layer-2 attention output may simply reflect how well position is decodable, not how much planned output length is represented. The per-unit R2 values in Table 4 are modest (0.11–0.42), so the top-k units selected by this target could be position/order units rather than length units; in that case, the scaling intervention in Figure 2 would show that perturbing position-correlated directions changes stopping behavior, which is a real but different phenomenon. This does not establish the paper's central disentanglement claim. I ask for a position-controlled probe, for example, train the regression on hidden states collected at a fixed generation step across outputs of different total lengths, or predict the remaining number of steps while conditioning on n, or remove/randomize the positional contribution, and show that the same units remain predictive of length beyond position. The authors' own remark in Section 4 that the final-layer increase may 'reinforce positional context' indicates the confound is present and needs to be addressed explicitly.
- [Section 5.2, Figure 2, Table 4] The length units are selected by per-unit R2 on the same Google summarization corpus that is later used for the scaling evaluation and human evaluation; no cross-validation or held-out unit selection is reported. If the selection step overfits to corpus-specific generation patterns, the scaling result in Figure 2 is partly a confirmation of the selection criterion. I request that units be selected on a training/validation split and evaluated on a held-out test split, and that the cross-task results in Appendix D be reported both with units selected independently per task and with units transferred from summarization. The current design does not separate discovery from evaluation.
- [Section 5.2, Figure 3, Limitations] The claim that LLMs have robust internal mechanisms for length control is directly qualified by the authors' own finding that scaling top-k units in grouped-query attention models (Qwen-2.5, Llama-3, Phi-3) does not control length in zero-shot settings, with success only after fine-tuning. Since the abstract and conclusion claim a general LLM mechanism, this limitation should be promoted into the main claims, or the claim should be explicitly restricted to standard multi-head attention models. In its current form, the paper demonstrates a mechanism in Llama-2 family models; the broader conclusion exceeds the evidence.
minor comments (5)
- [Throughout] There are several typos and reference formatting errors, including 'Googlesentence' in Section 3, 'Juseon-Do Juseon-Do' in the bibliography, a doubled comma in the Goh et al. entry, and 'Guangyi, Zhang' in the Llama 3 reference; these should be corrected.
- [Table 4] The column 'Avg 30' is not defined; please state whether it is the mean R2 over the top-30 units, and report standard errors or a statistical test for the differences between top-unit R2 values across prompting conditions.
- [Section 5.1] The sentence describing the dagger symbol, 'the improvement for scales between 10 and -10 is significant', is unclear; specify the exact paired comparison and report confidence intervals for the human evaluation scores in Table 5.
- [Figures 2 and 3] The horizontal axis is labeled only as 'scale'; please label the axis as the multiplier applied to the selected units, state the range, and clarify whether error bars are standard errors across the five runs or across the test instances.
- [Appendix D] Please specify which encoder/decoder model was used for the machine translation experiment, whether the transferred length units were selected on the same summarization split as in the main experiments, and why BLEU-1/2 rather than a length-sensitive metric was selected for the translation evaluation.
Circularity Check
The probe target is the generation time step, so the identified 'length units' are selected and validated on the same position/time-step quantity; the core length-disentanglement claim is partially definitional.
-
self definitional
[Section 3.1 (Models and Methods); Section 5 (Table 4, Figure 2)]
"During token generation, we saved each output with its corresponding numeric time step value, excluding the input token prompts (Kaplan et al., 2024). For instance, we saved n with its corresponding output when the model generated the n-th token. ... By investigating how well the model can predict the generation time step, we can gain insights into how length representations are encoded within the LLM’s hidden states."
The probe target Y is the generation time step n, which is exactly the absolute position of a generated token. Hidden states at step n carry positional/order information, so high R2 can reflect position decoding rather than a planned output-length attribute. The paper then selects 'length-related' units by per-unit R2 for predicting n (Table 4) and validates them by measuring ΔCR, the change in the number of generated summary tokens—i.e., the number of generated time steps. Selection and validation are thus tied to the same variable n, making the identification of 'length units' partly definitional: units are chosen because they predict n, and the demonstration that scaling them changes output 'length' is again a change in n.
full rationale
The central causal intervention has genuine independent content: scaling top-k units changes output length directionally, while scaling smallest-k units does not, and effects are consistent across prompts and tasks. If the result were purely a renamed fit, the smallest-k control and cross-task transfer would not be expected. However, the load-bearing identification step is not self-contained. The probe target in Section 3.1 is the generation time step n, which is identical to token position during generation; therefore the very high R2 values in Tables 2 and 3 are plausibly explained by positional or token-order information. Table 4 selects units by per-unit R2 for predicting n, and Figure 2 measures success via ΔCR, a function of the number of generated time steps. The same variable is thus used to select and to validate, so the label 'length unit' is partly a restatement of the selection criterion. The paper's own Section 4 remark that final-layer length information may 'reinforce positional context' acknowledges the confound. The self-citation of InstructCMP (Juseon-Do et al., 2024) for prompt templates is not load-bearing for the mechanism claim, since the causal intervention is evaluated on the authors' own models and data. Overall, this is partial circularity in the identification step, not complete circularity; a position-controlled probe would be required to convert the interesting causal effect into evidence for a disentangled length representation.
Assumptions & free parameters
assumptions (4)
- domain assumption Probing classifiers can identify functionally relevant directions and units in LLM hidden states despite known limitations.
- domain assumption Generation time step can serve as a proxy for output sequence length and is not confounded by positional encoding or token identity.
- domain assumption Scaling individual hidden units is a causally valid intervention for testing their role in generation.
- domain assumption Length-specific prompts from InstructCMP (Juseon-Do et al., 2024) are representative of external length control.
Cite this review
Pith. "Pith review of Length Representations in Large Language Models." pith.science (2026). https://pith.science/paper/OHEN4AQG
@misc{pith2026250720398,
author = {Pith},
title = {Pith review of: Length Representations in Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/OHEN4AQG}},
note = {Machine review of arXiv:2507.20398}
}
read the original abstract
Large language models (LLMs) have shown remarkable capabilities across various tasks, that are learned from massive amounts of text-based data. Although LLMs can control output sequence length, particularly in instruction-based settings, the internal mechanisms behind this control have been unexplored yet. In this study, we provide empirical evidence on how output sequence length information is encoded within the internal representations in LLMs. In particular, our findings show that multi-head attention mechanisms are critical in determining output sequence length, which can be adjusted in a disentangled manner. By scaling specific hidden units within the model, we can control the output sequence length without losing the informativeness of the generated text, thereby indicating that length information is partially disentangled from semantic information. Moreover, some hidden units become increasingly active as prompts become more length-specific, thus reflecting the model's internal awareness of this attribute. Our findings suggest that LLMs have learned robust and adaptable internal mechanisms for controlling output length without any external control.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, Weizhu Chen, Yen-Chun Chen, Yi-Ling Chen, Hao Cheng, Parul Chopra, Xiyang Dai, Matt...
arXiv 2024
-
[4]
Gasper Beguš, Maksymilian Dąbkowski, and Ryan Rhodes. 2023. https://arxiv.org/abs/2305.00948 Large linguistic models: Analyzing theoretical linguistic abilities of llms . In arXiv
arXiv 2023
-
[5]
Yonatan Belinkov. 2022. https://doi.org/10.1162/coli_a_00422 Probing classifiers: Promises, shortcomings, and advances . Computational Linguistics, 48(1):207--219
-
[6]
Ond r ej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aur \'e lie N \'e v \'e ol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016. ...
-
[7]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
2020
-
[8]
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023. https://arxiv.org/abs/2303.12712 Sparks of artificial general intelligence: Early experiments with gpt-4 . In arXiv
arXiv 2023
Show all 62 references
-
[9]
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei. 2023. https://doi.org/10.18653/v1/2023.findings-acl.247 Why can GPT learn in-context? language models secretly perform gradient descent as meta-optimizers . In Findings of the Association for Compu...
2023 doi
-
[10]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized llms. In arXiv preprint arXiv:2305.14314
2023 arXiv
-
[11]
Besnik Fetahu, Zhiyu Chen, Oleg Rokhlenko, and Shervin Malmasi. 2023. https://doi.org/10.18653/v1/2023.emnlp-industry.63 I nstruct PTS : Instruction-tuning LLM s for product title summarization . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Pr...
2023 doi
-
[12]
Katja Filippova and Yasemin Altun. 2013. https://aclanthology.org/D13-1155 Overcoming the lack of parallel data in sentence compression . In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1481--1491, Seattle, Washington, USA. Asso...
2013
-
[13]
Demian Ghalandari, Chris Hokamp, and Georgiana Ifrim. 2022. https://doi.org/10.18653/v1/2022.acl-long.90 Efficient unsupervised sentence compression by fine-tuning transformers with reinforcement learning . In Proceedings of the 60th Annual Meeting of the Association for Compu...
2022 doi
-
[14]
Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, , and Chris Olah. 2021. Multimodal neurons in artificial neural networks. Distill, (6(3):e30)
2021
-
[15]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[16]
Prakhar Gupta, Jeffrey Bigham, Yulia Tsvetkov, and Amy Pavel. 2021. https://doi.org/10.18653/v1/2021.naacl-main.240 Controlling dialogue generation with semantic exemplars . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computationa...
2021 doi
-
[17]
Wes Gurnee and Max Tegmark. 2024. https://openreview.net/forum?id=jE8xbmvFin Language models represent space and time . In The Twelfth International Conference on Learning Representations
2024
-
[18]
Michael Hanna, Ollie Liu, and Alexandre Variengien. 2023. https://arxiv.org/abs/2305.00586 How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model . In arXiv
2023 arXiv
-
[19]
Junxian He, Wojciech Kryscinski, Bryan McCann, Nazneen Rajani, and Caiming Xiong. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.396 CTRL sum: Towards generic controllable text summarization . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Pr...
2022 doi
-
[20]
Benjamin Heinzerling and Kentaro Inui. 2024. https://doi.org/10.18653/v1/2024.acl-short.18 Monotonic representation of numeric attributes in language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), p...
2024 doi
-
[21]
Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Xing Wang, Shuming Shi, and Zhaopeng Tu. 2023. https://arxiv.org/abs/2301.08745 Is chatgpt a good translator? yes with gpt-4 as the engine . In arXiv
2023 arXiv
-
[22]
Renlong Jie, Xiaojun Meng, Lifeng Shang, Xin Jiang, and Qun Liu. 2024. https://doi.org/10.18653/v1/2024.findings-acl.63 Prompt-based length controlled generation with multiple control types . In Findings of the Association for Computational Linguistics: ACL 2024, pages 1067--1...
2024 doi
-
[23]
Juseon-Do Juseon-Do, Hidetaka Kamigaito, Manabu Okumura, and Jingun Kwon. 2024. https://aclanthology.org/2024.findings-acl.532 I nstruct CMP : Length control in sentence compression through instruction-based large language models . In Findings of the Association for Computatio...
2024
-
[24]
Guy Kaplan, Matanel Oren, Yuval Reif, and Roy Schwartz. 2024. https://arxiv.org/abs/2410.05864 From tokens to words: On the inner lexicon of llms . Preprint, arXiv:2410.05864
2024 arXiv
-
[25]
Yuta Kikuchi, Graham Neubig, Ryohei Sasano, Hiroya Takamura, and Manabu Okumura. 2016. https://doi.org/10.18653/v1/D16-1140 Controlling output length in neural encoder-decoders . In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1...
2016 doi
-
[26]
Philipp Koehn. 2004. Statistical significance tests for machine translation evaluation. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics
2004
-
[27]
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2023. https://arxiv.org/abs/2205.11916 Large language models are zero-shot reasoners . In arXiv
2023 arXiv
-
[28]
Jingun Kwon, Hidetaka Kamigaito, and Manabu Okumura. 2023. https://doi.org/10.18653/v1/2023.findings-eacl.45 Abstractive document summarization with summary-length prediction . In Findings of the Association for Computational Linguistics: EACL 2023, pages 618--624, Dubrovnik, ...
2023 doi
-
[29]
Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, and Jimmy Huang. 2023. https://doi.org/10.18653/v1/2023.findings-acl.29 A systematic study and comprehensive evaluation of C hat GPT on benchmark datasets . In Findings of the Associa...
2023 doi
-
[30]
Haoran Li, Junnan Zhu, Jiajun Zhang, Chengqing Zong, and Xiaodong He. 2020. https://doi.org/10.1609/aaai.v34i05.6333 Keywords-guided abstractive sentence summarization . Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):8196--8203
2020 doi
-
[31]
Chin-Yew Lin. 2004. https://aclanthology.org/W04-1013 ROUGE : A package for automatic evaluation of summaries . In Text Summarization Branches Out, pages 74--81, Barcelona, Spain. Association for Computational Linguistics
2004
-
[32]
Bang Liu, Haojie Wei, Di Niu, Haolan Chen, and Yancheng He. 2020. https://doi.org/10.1145/3366423.3380270 Asking questions the human way: Scalable question-answer generation from text corpus . In Proceedings of The Web Conference 2020, WWW '20, page 2032–2043, New York, NY, US...
2020
-
[33]
Yizhu Liu, Qi Jia, and Kenny Zhu. 2022. https://doi.org/10.18653/v1/2022.acl-long.474 Length control in abstractive summarization by pretraining information selection . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...
2022 doi
-
[34]
Yizhu Liu, Zhiyi Luo, and Kenny Zhu. 2018. https://doi.org/10.18653/v1/D18-1444 Controlling length in abstractive summarization using a convolutional neural network . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4110--4119, B...
2018 doi
-
[35]
Takuya Makino, Tomoya Iwakura, Hiroya Takamura, and Manabu Okumura. 2019. https://doi.org/10.18653/v1/P19-1099 Global optimization under length constraint for neural text summarization . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics...
2019 doi
-
[36]
Meta. 2024. https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
2024
-
[37]
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016. https://doi.org/10.18653/v1/N16-1098 A corpus and cloze evaluation for deeper understanding of commonsense stories . In Proceedings of the 2...
2016 doi
-
[38]
Kenton Murray and David Chiang. 2018. https://doi.org/10.18653/v1/W18-6322 Correcting length bias in neural machine translation . In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 212--223, Brussels, Belgium. Association for Computational Li...
2018 doi
-
[39]
Jingcheng Niu, Wenjie Lu, and Gerald Penn. 2022. https://aclanthology.org/2022.coling-1.278 Does BERT rediscover a classical NLP pipeline? In Proceedings of the 29th International Conference on Computational Linguistics, pages 3143--3153, Gyeongju, Republic of Korea. Internati...
2022
-
[40]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2022 arXiv
-
[41]
Keqin Peng, Liang Ding, Qihuang Zhong, Li Shen, Xuebo Liu, Min Zhang, Yuanxin Ouyang, and Dacheng Tao. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.373 Towards making the most of C hat GPT for machine translation . In Findings of the Association for Computational Ling...
2023 doi
-
[42]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, ...
2025 arXiv
-
[43]
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. page 9. OpenAI
2019
-
[44]
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020. https://doi.org/10.1162/tacl_a_00349 A primer in BERT ology: What we know about how BERT works . Transactions of the Association for Computational Linguistics, 8:842--866
2020 doi
-
[45]
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell. 2023. https://arxiv.org/abs/2207.13243 Toward transparent ai: A survey on interpreting the inner structures of deep neural networks . In arXiv
2023 arXiv
-
[46]
Hassan Sajjad, Nadir Durrani, and Fahim Dalvi. 2022. https://doi.org/10.1162/tacl_a_00519 Neuron-level interpretation of deep nlp models: A survey . Transactions of the Association for Computational Linguistics, 10:1285--1303
2022 doi
-
[47]
Raphael Schumann, Lili Mou, Yao Lu, Olga Vechtomova, and Katja Markert. 2020. https://doi.org/10.18653/v1/2020.acl-main.452 Discrete optimization for unsupervised sentence summarization with word-level extraction . In Proceedings of the 58th Annual Meeting of the Association f...
2020 doi
-
[48]
Wei Shen, Rui Zheng, Wenyu Zhan, Jun Zhao, Shihan Dou, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.188 Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback . In Findings of the Association ...
2023 doi
-
[49]
Xing Shi, Kevin Knight, and Deniz Yuret. 2016. https://doi.org/10.18653/v1/D16-1248 Why neural translations are the right length . In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2278--2282, Austin, Texas. Association for Comput...
2016 doi
-
[50]
Sho Takase and Naoaki Okazaki. 2019. https://doi.org/10.18653/v1/N19-1401 Positional encoding to control output sequence length . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies,...
2019 doi
-
[51]
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. https://doi.org/10.18653/v1/P19-1452 BERT rediscovers the classical NLP pipeline . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4593--4601, Florence, Italy. Association for ...
2019 doi
-
[52]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...
2023 arXiv
-
[53]
Noah Wang, Feiyu Duan, Yibo Zhang, Wangchunshu Zhou, Ke Xu, Wenhao Huang, and Jie Fu. 2024. https://aclanthology.org/2024.findings-emnlp.983 P osition ID : LLM s can control lengths, copy and paste with explicit positional awareness . In Findings of the Association for Computa...
2024
-
[54]
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2022. Finetuned language models are zero-shot learners. In International Conference on Learning Representations
2022
-
[55]
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith St...
2016 arXiv
-
[56]
Tingyu Xie, Qi Li, Jian Zhang, Yan Zhang, Zuozhu Liu, and Hongwei Wang. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.493 Empirical study of zero-shot NER with C hat GPT . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7935...
2023 doi
-
[57]
Tingyu Xie, Qi Li, Yan Zhang, Zuozhu Liu, and Hongwei Wang. 2024. https://doi.org/10.18653/v1/2024.naacl-short.49 Self-improving for zero-shot named entity recognition with large language models . In Proceedings of the 2024 Conference of the North American Chapter of the Assoc...
2024 doi
-
[58]
Xi Ye, Srinivasan Iyer, Asli Celikyilmaz, Veselin Stoyanov, Greg Durrett, and Ramakanth Pasunuru. 2023. https://doi.org/10.18653/v1/2023.findings-acl.273 Complementary explanations for effective in-context learning . In Findings of the Association for Computational Linguistics...
2023 doi
-
[59]
Weizhe Yuan, Ilia Kulikov, Ping Yu, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston, and Jing Xu. 2024. https://arxiv.org/abs/2406.17744 Following length constraints in instructions . In arXiv
2024 arXiv
-
[60]
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, and Ed Chi. 2023. https://arxiv.org/abs/2205.10625 Least-to-most prompting enables complex reasoning in large language models . In arXiv
2023 arXiv
-
[61]
Hattie Zhou, Azade Nova, Hugo Larochelle, Aaron Courville, Behnam Neyshabur, and Hanie Sedghi. 2022. https://arxiv.org/abs/2211.09066 Teaching algorithmic reasoning via in-context learning . In arXiv
2022 arXiv
-
[62]
Zhang Zhuocheng, Shuhao Gu, Min Zhang, and Yang Feng. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.773 Addressing the length bias challenge in document-level neural machine translation . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 1...
2023 doi
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.