REVIEW 5 major objections 7 minor 1 cited by
AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling
T0 review · 5 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read AbbIE-D, a Transformer that reuses its middle block with an extra residual connection, achieves better perplexity than a standard Transformer at the same token budget and keeps improving on zero-shot reasoning tasks when given more…
desk verdict A plausible but over-claimed recipe for test-time compute scaling; the residual-injection variant is worth a look, but the headline numbers and the Depth baseline need fixing before the central comparison holds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the recursive Body recurrence. For AbbIE-D it is $h_{k+1} = B(h_k) + h_k$, where $B$ is the stack of Transformer blocks in the Body and $h_0$ is the output of the Head; adding $h_k$ back on each pass strengthens the contribution of the original input relative to the cumulative attention and feed-forward updates. This extra residual makes the iterates converge toward a fixed point in concept space, the latent region between token embedding and unembedding, which is what the paper links to generalization beyond the training iteration count. A standard Transformer is the special case $r=1$, so AbbIE is a direct architectural generalization rather than a separate family.
What would settle it
Train AbbIE-D at 200M and 350M with two iterations, five seeds each, then evaluate at $r=2$, 4, 8, and 32 on the four benchmarks plus held-out tasks; the upward-generalization claim is falsified if mean accuracy at $r=8$ does not exceed mean accuracy at $r=2$, or if inter-iteration distance stops decreasing at the larger iteration counts.
Extended reading notes
Core claim
The discovery is that upward generalization in iteration count follows from a simple structural change: adding a residual connection around the entire reused Body stack. With training iterations fixed at r=2, the 350M AbbIE-D model's perplexity keeps decreasing up to r=4 and its accuracy on HellaSwag, LAMBADA, ARC-Easy, and CommonsenseQA keeps improving up to r=8, four times the training count; the 200M variant shows weaker effects, suggesting a size threshold. The same training budget that produces this behavior also yields roughly 5% better perplexity than a standard Transformer, up to 12% higher zero-shot accuracy on HellaSwag, and the only above-random CommonsenseQA scores among the compared models. AbbIE-C, which relies only on the Transformer's internal residual stream and omits the extra inter-iteration residual, does not converge and is excluded.
Load-bearing premise
The paper assumes that fixed-point convergence measured on 16 random samples per benchmark and the results of a single 350M training run are enough to establish upward generalization as a stable property of the architecture.
Editorial extensions
If this is right
- At the same token budget, AbbIE-D beats a standard Transformer on perplexity by about 5% while remaining a drop-in replacement, since running it once matches standard perplexity.
- Because performance continues to improve at $r=8$ after training at $r=2$, iteration count becomes a test-time compute dial that can be raised for harder inputs without retraining or specialized data.
- The FLOP cost of AbbIE relative to a standard Transformer falls as training extends beyond the compute-optimal point, so in the long-training regime common in large-scale runs the overhead approaches that of a standard model.
- Zero-shot in-context learning gains persist even after perplexity begins to worsen at higher iteration counts, indicating that the iteration procedure changes task-relevant representations and not just reduces loss.
- Recurrent transformers that require many training iterations or random input injection are not necessary for upward generalization; two iterations and the inter-iteration residual suffice at this scale.
Reading between the lines
- If the single-seed 350M result is representative, iteration count could become a per-input adaptive budget, with the model stopping once it reaches its fixed point; this would make inference cost proportional to input difficulty.
- The decoupling between perplexity and in-context learning at high iteration counts suggests latent-space iteration may act as a form of implicit reasoning, and one could test this by probing whether intermediate iterates show increasingly structured semantic representations.
- The same fixed-point principle could transfer to other autoregressive architectures, such as encoder-decoder or multimodal models, by placing the extra residual around the reused component; this is an untested extension.
- Multi-seed and multi-size studies are needed to confirm that the upward-generalization threshold sits between 200M and 350M parameters rather than being a random artifact of one run.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces AbbIE, an encoder-only Transformer variant partitioned into Head, Body, and Tail groups, in which the Body stack is applied iteratively in latent space. Two variants are defined: AbbIE-C (direct recursion over the Body) and AbbIE-D (the Body output is added to its input via an extra inter-iteration residual). The authors train with two Body iterations and evaluate at test-time iteration counts of 1, 2, 4, 8, and 32. They claim AbbIE-D matches or beats a standard Transformer at equal token budget, achieves lower perplexity, and “upward generalizes”: performance improves at test-time iteration counts beyond the training count, up to 4× the training iteration count. The experiments compare AbbIE with a standard Transformer baseline (Std) and a recurrent-depth baseline (Depth) at 200M and 350M parameter scales on HellaSwag, LAMBADA, ARC-Easy, and CommonsenseQA.
Significance. The central idea is attractive: if a recurrent Transformer can upward-generalize from only two training iterations without specialized data or projections, it offers a practical, drop-in mechanism for test-time compute scaling that is complementary to parameter scaling. The paper has several strengths: the architecture is simple and clearly described, the training protocol is standard, the authors fix the token budget across comparisons, and they disclose the small number of large-model runs. However, the headline claims are only partially supported by the reported numbers. The abstract's 12% ICL improvement is not recoverable from Table 2, the 5% perplexity improvement is not tabulated, the 350M upward-generalization result rests on a single training run, and the Depth baseline appears to be trained under a protocol that may not match the original method. If the evidence were strengthened through corrected baselines, error bars, and a clearer match between claims and tables, the paper would make a useful contribution.
major comments (5)
- [§3.1, Fig. 6b, Table 2] The Depth baseline training protocol is not specified. Section 5 states that the original recurrent-depth method of Geiping et al. (2025) is trained over a range of randomly sampled iteration counts, but the experimental section never states whether the Depth implementation in this paper received that protocol. The collapse of Depth at every iteration count other than r=2, visible in Fig. 6b and Table 2, is exactly what one expects from a model trained at fixed r=2. Because Depth is the principal iterative baseline, the claim that AbbIE “far outperforms alternative iterative methods” is not established unless Depth was trained with the same sampled-iteration protocol as the original method. Please specify the protocol; if Depth was trained at fixed r=2, re-run it with sampled iterations or clearly state the deviation and justify why the comparison is fair.
- [Abstract, Table 2, Discussion] The “up to 12% improvement” claim is not recoverable from the reported numbers. In the 350M rows of Table 2, the relative improvements of AbbIE-D at r=8 over Std are approximately 7.1% on HellaSwag (34.8/32.5), 7.7% on LAMBADA (28.0/26.0), and 6.4% on ARC-Easy (51.3/48.2). Comparing against Depth(r=2) gives similar magnitudes. If the 12% figure comes from a different calculation (for example, against the random baseline or against Depth at a collapsed setting), the derivation must be reported. The same applies to the 5% perplexity improvement, which is mentioned in Fig. 5b and the Discussion but never appears as a tabulated value. The abstract and Discussion should state only the gains that are directly readable from the tables.
- [§3.2, §4.3, Table 2] The upward-generalization result rests on very limited statistical evidence. The 350M model, which is the only configuration showing clear gains at r=4 and r=8, was trained with a single seed (Section 3.2). Table 2 reports the “best result” without error bars or confidence intervals, and Fig. 4 uses only 16 random samples with one initial state per sample. Given that the central claim of the paper is that two training iterations suffice for upward generalization at larger scale, the authors should either provide multiple 350M seeds with error bars, or at minimum clearly label the single-run status in every figure and table that supports the scaling claim, and temper the strength of the conclusion accordingly.
- [§4.1, Fig. 4] The fixed-point comparison between AbbIE-C and AbbIE-D is based on two different quantities. For AbbIE-C, the inter-iteration distance is ‖B(h_k) − h_k‖, because h_{k+1} = B(h_k). For AbbIE-D, the distance is ‖B(h_k)‖, because h_{k+1} = B(h_k) + h_k. These are different functionals with different fixed-point conditions (B(h*) = h* versus B(h*) = 0), so the observation that “AbbIE-C diverges while AbbIE-D converges” does not directly follow from comparing the two curves. Please plot a common distance measure, such as ‖h_{k+1} − h_k‖ or a normalized state difference, for both variants, and state explicitly which quantity is shown.
- [§4.2, Fig. 5b] The perplexity improvement claim is not quantified in the text. Figure 5b shows curves that appear separated, and the caption says “roughly 5% better perplexity,” but no numerical Perplexity values or standard deviations are given. Since 200M models were trained with five seeds, the authors should report mean and standard error for Std, Depth, and AbbIE-D at the compute-optimal token budget, and state the actual relative improvement. This is necessary to verify both the direction and the magnitude of the claimed perplexity gain.
minor comments (7)
- [Fig. 1 caption] The caption says “increasing training iterations (1, 2, 4, and 8)”, but the models are trained with two iterations and evaluated at those test-time counts; please rephrase to avoid implying varying training iteration counts.
- [§2.1] The term “detokenization” is used to mean the transition from token space to concept space inside the model, which is not the standard meaning of detokenization in language modeling; consider a different term or explicitly define the usage.
- [§2.2, Eq. (1)-(2)] The notation around path independence is loose: Eq. (1) uses f^∞(x, z0), while Eq. (2) writes f^{(n)}(x_n) and relates it to f^{(n+1)}(x_{n+1}); the role of the initial state z0 and the statement that “we omit z0” should be clarified, since path independence is defined with respect to arbitrary z0.
- [§3.2] The learning-rate schedule is described as “Warmup-Stable-Decay”, but the text only describes linear warmup and cosine decay; please specify what the “stable” phase is, or rename the schedule.
- [Table 2] There are small typographical issues: “Commensense Question Answering” should be “Commonsense Question Answering”, and the note in the table body is informal; consider moving protocol caveats to the caption.
- [§4.3] The sentence “we see a performance significantly higher than the random baseline” uses “significantly” without a statistical test; please replace with “numerically higher” unless a test is performed.
- [§4.3] The phrase “which no other general recursive transformer has shown” is absolute; please soften to “to our knowledge” or cite a negative result.
Circularity Check
No significant circularity: the central claims are direct empirical measurements, and the path-independence discussion is external motivation rather than a derived premise.
full rationale
This paper is an empirical architecture-comparison study rather than a derivation from fitted quantities. The central claims (lower perplexity at equal token budget, upward generalization at r=4/8 after training at r=2, and ICL gains) are obtained by training models and measuring standard metrics. No parameter is fitted to the evaluation benchmarks and then reported as a prediction: the perplexity and ICL numbers in Table 2 and Figures 1, 5, and 6 are direct measurements, not outputs of a fitted model. The path-independence discussion in Sections 2.2 and 5 is used only to motivate the input-reinjection design of AbbIE-C versus AbbIE-D; the paper does not derive the upward-generalization result from path independence, and it tests convergence empirically in Section 4.1 and Figure 4. The only overlapping-author citation is Shen et al. (2025), used in a list with Fan et al. (2019) and Elhoushi et al. (2024) to support the peripheral observation that Transformer blocks can be reused; this citation is not load-bearing, and no uniqueness theorem is imported from the authors' own prior work. The weaker points identified in the manuscript (single 350M run, 16-sample convergence analysis, and the Depth baseline training protocol) are statistical and experimental-fairness concerns about evidence strength, not circular reductions. Therefore no step in the derivation reduces to its own input by construction.
Assumptions & free parameters
free parameters (1)
- Number of training iterations =
2
assumptions (4)
- domain assumption Residual dependence on the original input is sufficient for path independence and upward generalization.
- domain assumption Inter-iteration distance measured on 16 samples per benchmark is representative of the full distribution.
- domain assumption Perplexity improvements at 200M-350M scale transfer to larger models.
- domain assumption Chinchilla compute-optimal token budgets apply to recursive training.
Cite this review
Pith. "Pith review of AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling." pith.science (2026). https://pith.science/paper/XUDA3LUR
@misc{pith2026250708567,
author = {Pith},
title = {Pith review of: AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling},
year = {2026},
howpublished = {\url{https://pith.science/paper/XUDA3LUR}},
note = {Machine review of arXiv:2507.08567}
}
read the original abstract
We introduce the Autoregressive Block-Based Iterative Encoder (AbbIE), a novel recursive generalization of the encoder-only Transformer architecture, which achieves better perplexity than a standard Transformer and allows for the dynamic scaling of compute resources at test time. This simple, recursive approach is a complement to scaling large language model (LLM) performance through parameter and token counts. AbbIE performs its iterations in latent space, but unlike latent reasoning models, does not require a specialized dataset or training protocol. We show that AbbIE upward generalizes (ability to generalize to arbitrary iteration lengths) at test time by only using 2 iterations during train time, far outperforming alternative iterative methods. AbbIE's ability to scale its computational expenditure based on the complexity of the task gives it an up to \textbf{12\%} improvement in zero-shot in-context learning tasks versus other iterative and standard methods and up to 5\% improvement in language perplexity. The results from this study open a new avenue to Transformer performance scaling. We perform all of our evaluations on model sizes up to 350M parameters.
Forward citations
Cited by 1 Pith paper
-
LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning
LatentChem reasons in continuous latent space for chemistry, achieving a 59.88% non-tie win rate over explicit CoT on ChemCoTBench with a 10.84x average reduction in reasoning overhead.
Reference graph
Works this paper leans on
-
[1]
J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebron, and S. Sanghai. GQA : Training generalized multi-query transformer models from multi-head checkpoints. In H. Bouamor, J. Pino, and K. Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4895--4901, Singapore, Dec. 2023. Association for...
-
[2]
L. B. Allal, A. Lozhkov, E. Bakouch, G. M. Bl \' a zquez, G. Penedo, L. Tunstall, A. Marafioti, H. Kydl \' cek, A. P. Lajar \' n, V. Srivastav, J. Lochner, C. Fahlgren, X. Nguyen, C. Fourrier, B. Burtenshaw, H. Larcher, H. Zhao, C. Zakka, M. Morlon, C. Raffel, L. von Werra, and T. Wolf. Smollm2: When smol goes big - data-centric training of a small langua...
-
[3]
L. B. Allal, A. Lozhkov, E. Bakouch, G. M. Blázquez, G. Penedo, L. Tunstall, A. Marafioti, H. Kydlíček, A. P. Lajarín, V. Srivastav, J. Lochner, C. Fahlgren, X.-S. Nguyen, C. Fourrier, B. Burtenshaw, H. Larcher, H. Zhao, C. Zakka, M. Morlon, C. Raffel, L. von Werra, and T. Wolf. Smollm2: When smol goes big -- data-centric training of a small language mode...
arXiv 2025
-
[4]
Amd-llama-135m: A 135m parameter language model trained on amd instinct mi250 accelerators
AMD . Amd-llama-135m: A 135m parameter language model trained on amd instinct mi250 accelerators. https://huggingface.co/amd/AMD-Llama-135m, 2023. Accessed: 2025-05-15
work page 2023
-
[5]
C. Anil, A. Pokle, K. Liang, J. Treutlein, Y. Wu, S. Bai, J. Z. Kolter, and R. B. Grosse. Path independent equilibrium models can better exploit test-time computation. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Syste...
work page 2022
-
[6]
S. Bae, A. Fisch, H. Harutyunyan, Z. Ji, S. Kim, and T. Schuster. Relaxed recursive transformers: Effective parameter sharing with layer-wise lora. arXiv preprint arXiv:2410.20672, 2024
arXiv 2024
-
[7]
L. Ben Allal, A. Lozhkov, G. Penedo, T. Wolf, and L. von Werra. Cosmopedia, February 2024. URL https://huggingface.co/datasets/HuggingFaceTB/cosmopedia
work page 2024
- [8]
Show all 52 references
-
[9]
Clark, I
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1, 2018
2018 arXiv
- [10]
-
[11]
Dehghani, S
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser. Universal transformers. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. URL https://openreview.net/forum?id=HyzdRiR9Y7
2019
-
[12]
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. Bosma, Z. Zhou, T. Wang, Y. E. Wang, K. Webster, M. Pellat, K. Robinson, K. Meier-Hellstern, T. Duke, L. Dixon, K. Zhang, Q. V. Le, Y. Wu, Z. Chen, and C. Cui...
2022 arXiv
-
[13]
Elfwing, E
S. Elfwing, E. Uchibe, and K. Doya. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017. URL https://arxiv.org/abs/1702.03118
2017 arXiv
-
[14]
Elhage, T
N. Elhage, T. Hume, C. Olsson, N. Nanda, T. Henighan, S. Johnston, S. ElShowk, N. Joseph, N. DasSarma, B. Mann, D. Hernandez, A. Askell, K. Ndousse, A. Jones, D. Drain, A. Chen, Y. Bai, D. Ganguli, L. Lovitt, Z. Hatfield-Dodds, J. Kernion, T. Conerly, S. Kravec, S. Fort, S. Ka...
2022
-
[15]
Elhoushi, A
M. Elhoushi, A. Shrivastava, D. Liskovich, B. Hosmer, B. Wasti, L. Lai, A. Mahmoud, B. Acun, S. Agarwal, A. Roman, A. A. Aly, B. Chen, and C. Wu. Layerskip: Enabling early exit inference and self-speculative decoding. In L. Ku, A. Martins, and V. Srikumar, editors, Proceedings...
2024
-
[16]
A. Fan, E. Grave, and A. Joulin. Reducing transformer depth on demand with structured dropout. arXiv preprint arXiv:1909.11556, 2019
1909 arXiv
-
[17]
Fedus, B
W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23 0 (120): 0 1--39, 2022
2022
-
[18]
Geiping, S
J. Geiping, S. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson, B. Kailkhura, A. Bhatele, and T. Goldstein. Scaling up test-time compute with latent reasoning: A recurrent depth approach, 2025. URL https://arxiv.org/abs/2502.05171
2025 arXiv
-
[19]
Giannou, S
A. Giannou, S. Rajput, J. Sohn, K. Lee, J. D. Lee, and D. Papailiopoulos. Looped transformers as programmable computers. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, editors, International Conference on Machine Learning, ICML 2023, 23-29 July ...
2023
-
[20]
Grattafiori, A
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru,...
2024 arXiv
-
[21]
Gu and T
A. Gu and T. Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023
2023 arXiv
-
[22]
Gurnee, N
W. Gurnee, N. Nanda, M. Pauly, K. Harvey, D. Troitskii, and D. Bertsimas. Finding neurons in a haystack: Case studies with sparse probing. Trans. Mach. Learn. Res., 2023, 2023. URL https://openreview.net/forum?id=JYs1R9IMJr
2023
- [23]
-
[24]
Hoffmann, S
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifr...
-
[25]
Hsieh, C.-L
C.-Y. Hsieh, C.-L. Li, C.-K. Yeh, H. Nakhost, Y. Fujii, A. Ratner, R. Krishna, C.-Y. Lee, and T. Pfister. Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes. arXiv preprint arXiv:2305.02301, 2023
2023 arXiv
-
[26]
Hägele, E
A. Hägele, E. Bakouch, A. Kosson, L. B. Allal, L. V. Werra, and M. Jaggi. Scaling laws and compute-optimal training beyond fixed training durations, 2024. URL https://arxiv.org/abs/2405.18392
2024 arXiv
-
[27]
Jumper, R
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera-Paredes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Petersen, D. Reiman, E. C...
2021
- [28]
-
[29]
Kaplan, S
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models. CoRR, abs/2001.08361, 2020. URL https://arxiv.org/abs/2001.08361
2001 arXiv
-
[30]
Katz and Y
S. Katz and Y. Belinkov. VISIT: visualizing and interpreting the semantic information flow of transformers. In H. Bouamor, J. Pino, and K. Bali, editors, Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023 , pages 14094--14113....
2023 doi
-
[31]
Kocetkov, R
D. Kocetkov, R. Li, L. Ben Allal, J. Li, C. Mou, C. Muñoz Ferrandis, Y. Jernite, M. Mitchell, S. Hughes, T. Wolf, D. Bahdanau, L. von Werra, and H. de Vries. The stack: 3 tb of permissively licensed source code. Preprint, 2022
2022
-
[32]
H. Lu, Y. Zhou, S. Liu, Z. Wang, M. W. Mahoney, and Y. Yang. Alphapruning: Using heavy-tailed self regularization theory for improved layer-wise pruning of large language models. Advances in Neural Information Processing Systems, 37: 0 9117--9152, 2024
2024
-
[33]
McCandlish, J
S. McCandlish, J. Kaplan, D. Amodei, and O. D. Team. An empirical model of large-batch training, 2018. URL https://arxiv.org/abs/1812.06162
2018 arXiv
-
[34]
M. I. Nye, A. J. Andreassen, G. Gur - Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, C. Sutton, and A. Odena. Show your work: Scratchpads for intermediate computation with language models. CoRR, abs/2112.00114, 2021. URL https://arxiv.org...
2021 arXiv
-
[35]
Paperno, G
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández. The lambada dataset: Word prediction requiring a broad discourse context, 2016. URL https://arxiv.org/abs/1606.06031
2016 arXiv
-
[36]
Penedo, H
G. Penedo, H. Kydl \' cek, L. B. Allal, A. Lozhkov, M. Mitchell, C. A. Raffel, L. von Werra, and T. Wolf. The fineweb datasets: Decanting the web for the finest text data at scale. In A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang, editor...
2024
-
[37]
Penedo, H
G. Penedo, H. Kydlíček, L. B. allal, A. Lozhkov, M. Mitchell, C. Raffel, L. V. Werra, and T. Wolf. The fineweb datasets: Decanting the web for the finest text data at scale, 2024 b . URL https://arxiv.org/abs/2406.17557
2024 arXiv
-
[38]
B. Peng, E. Alcaide, Q. G. Anthony, A. Albalak, S. Arcadinho, S. Biderman, H. Cao, X. Cheng, M. N. Chung, L. Derczynski, et al. Rwkv: Reinventing rnns for the transformer era. In The 2023 Conference on Empirical Methods in Natural Language Processing
2023
-
[39]
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, A. Levskaya, J. Heek, K. Xiao, S. Agrawal, and J. Dean. Efficiently scaling transformer inference, 2022. URL https://arxiv.org/abs/2211.05102
2022 arXiv
-
[40]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. OpenAI, 2019. URL https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf. Accessed: 2024-11-15
2019
-
[41]
Sardana, J
N. Sardana, J. Portes, S. Doubov, and J. Frankle. Beyond chinchilla-optimal: Accounting for inference in language model scaling laws, 2025. URL https://arxiv.org/abs/2401.00448
2025 arXiv
-
[42]
W. F. Shen, X. Qiu, M. Kurmanji, A. Iacob, L. Sani, Y. Chen, N. Cancedda, and N. D. Lane. Lunar: Llm unlearning via neural activation redirection. arXiv preprint arXiv:2502.07218, 2025
2025
-
[43]
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomput., 568 0 (C), Feb. 2024. ISSN 0925-2312. doi:10.1016/j.neucom.2023.127063. URL https://doi.org/10.1016/j.neucom.2023.127063
2024
-
[44]
Talmor, J
A. Talmor, J. Herzig, N. Lourie, and J. Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge, 2019. URL https://arxiv.org/abs/1811.00937
2019 arXiv
-
[45]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. In I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Pro...
2017
-
[46]
von Platen and T
P. von Platen and T. Wolf. Smollm: Efficient language models for everyone. https://huggingface.co/blog/smollm, dec 2023. Accessed: 2025-05-13
2023
-
[47]
S. Wang, P. Zhou, J. Li, and H. Huang. 4-bit shampoo for memory-efficient network training. Advances in Neural Information Processing Systems, 37: 0 126997--127029, 2024
2024
-
[48]
Wolters, X
C. Wolters, X. Yang, U. Schlichtmann, and T. Suzumura. Memory is all you need: An overview of compute-in-memory architectures for accelerating large language model inference, 2024. URL https://arxiv.org/abs/2406.08413
2024 arXiv
-
[49]
Xiong, Y
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T. Liu. On layer normalization in the transformer architecture. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event , volume 119...
2020
-
[50]
L. Yang, K. Lee, R. D. Nowak, and D. Papailiopoulos. Looped transformers are better at learning learning algorithms. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL https://openreview.n...
2024
-
[51]
Zellers, A
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. Hellaswag: Can a machine really finish your sentence?, 2019. URL https://arxiv.org/abs/1905.07830
2019 arXiv
-
[52]
Zhang and R
B. Zhang and R. Sennrich. Root mean square layer normalization, 2019. URL https://arxiv.org/abs/1910.07467
2019 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.