Pith. sign in

REVIEW 5 major objections 7 minor 1 cited by

AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling

T0 review · 5 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read AbbIE-D, a Transformer that reuses its middle block with an extra residual connection, achieves better perplexity than a standard Transformer at the same token budget and keeps improving on zero-shot reasoning tasks when given more…

desk verdict A plausible but over-claimed recipe for test-time compute scaling; the residual-injection variant is worth a look, but the headline numbers and the Depth baseline need fixing before the central comparison holds. read the letter →

arxiv 2507.08567 v2 pith:XUDA3LUR submitted 2025-07-11 cs.LG

classification cs.LG
keywords autoregressiveblock-basediterativeencoderrecurrenttransformertest-timecomputescalingupwardgeneralizationfixed-pointconvergencelatentspacereasoningzero-shotin-contextlearningperplexity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces AbbIE, a recursive variant of the decoder-only Transformer that reuses a middle group of layers, called the Body, over several iterations in latent space. Its central claim is that the AbbIE-D variant, which adds one extra residual connection around the Body and trains with exactly two iterations, reaches lower perplexity than a standard Transformer at the same token budget and remains stable enough that more iterations at test time keep improving zero-shot in-context learning accuracy. This would matter because it gives model developers a way to spend more compute on hard inputs at inference without training on specialized data or changing the architecture's single-pass behavior. The authors frame AbbIE as a complement to scaling parameters and tokens, since it matches a standard Transformer when run once but can be dialed up when extra computation is available.

What carries the argument

The load-bearing object is the recursive Body recurrence. For AbbIE-D it is $h_{k+1} = B(h_k) + h_k$, where $B$ is the stack of Transformer blocks in the Body and $h_0$ is the output of the Head; adding $h_k$ back on each pass strengthens the contribution of the original input relative to the cumulative attention and feed-forward updates. This extra residual makes the iterates converge toward a fixed point in concept space, the latent region between token embedding and unembedding, which is what the paper links to generalization beyond the training iteration count. A standard Transformer is the special case $r=1$, so AbbIE is a direct architectural generalization rather than a separate family.

What would settle it

Train AbbIE-D at 200M and 350M with two iterations, five seeds each, then evaluate at $r=2$, 4, 8, and 32 on the four benchmarks plus held-out tasks; the upward-generalization claim is falsified if mean accuracy at $r=8$ does not exceed mean accuracy at $r=2$, or if inter-iteration distance stops decreasing at the larger iteration counts.

Watch

Extended reading notes

Core claim

The discovery is that upward generalization in iteration count follows from a simple structural change: adding a residual connection around the entire reused Body stack. With training iterations fixed at r=2, the 350M AbbIE-D model's perplexity keeps decreasing up to r=4 and its accuracy on HellaSwag, LAMBADA, ARC-Easy, and CommonsenseQA keeps improving up to r=8, four times the training count; the 200M variant shows weaker effects, suggesting a size threshold. The same training budget that produces this behavior also yields roughly 5% better perplexity than a standard Transformer, up to 12% higher zero-shot accuracy on HellaSwag, and the only above-random CommonsenseQA scores among the compared models. AbbIE-C, which relies only on the Transformer's internal residual stream and omits the extra inter-iteration residual, does not converge and is excluded.

Load-bearing premise

The paper assumes that fixed-point convergence measured on 16 random samples per benchmark and the results of a single 350M training run are enough to establish upward generalization as a stable property of the architecture.

Editorial extensions

If this is right

  • At the same token budget, AbbIE-D beats a standard Transformer on perplexity by about 5% while remaining a drop-in replacement, since running it once matches standard perplexity.
  • Because performance continues to improve at $r=8$ after training at $r=2$, iteration count becomes a test-time compute dial that can be raised for harder inputs without retraining or specialized data.
  • The FLOP cost of AbbIE relative to a standard Transformer falls as training extends beyond the compute-optimal point, so in the long-training regime common in large-scale runs the overhead approaches that of a standard model.
  • Zero-shot in-context learning gains persist even after perplexity begins to worsen at higher iteration counts, indicating that the iteration procedure changes task-relevant representations and not just reduces loss.
  • Recurrent transformers that require many training iterations or random input injection are not necessary for upward generalization; two iterations and the inter-iteration residual suffice at this scale.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the single-seed 350M result is representative, iteration count could become a per-input adaptive budget, with the model stopping once it reaches its fixed point; this would make inference cost proportional to input difficulty.
  • The decoupling between perplexity and in-context learning at high iteration counts suggests latent-space iteration may act as a form of implicit reasoning, and one could test this by probing whether intermediate iterates show increasingly structured semantic representations.
  • The same fixed-point principle could transfer to other autoregressive architectures, such as encoder-decoder or multimodal models, by placing the extra residual around the reused component; this is an untested extension.
  • Multi-seed and multi-size studies are needed to confirm that the upward-generalization threshold sits between 200M and 350M parameters rather than being a random artifact of one run.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. The paper introduces AbbIE, an encoder-only Transformer variant partitioned into Head, Body, and Tail groups, in which the Body stack is applied iteratively in latent space. Two variants are defined: AbbIE-C (direct recursion over the Body) and AbbIE-D (the Body output is added to its input via an extra inter-iteration residual). The authors train with two Body iterations and evaluate at test-time iteration counts of 1, 2, 4, 8, and 32. They claim AbbIE-D matches or beats a standard Transformer at equal token budget, achieves lower perplexity, and “upward generalizes”: performance improves at test-time iteration counts beyond the training count, up to 4× the training iteration count. The experiments compare AbbIE with a standard Transformer baseline (Std) and a recurrent-depth baseline (Depth) at 200M and 350M parameter scales on HellaSwag, LAMBADA, ARC-Easy, and CommonsenseQA.

Significance. The central idea is attractive: if a recurrent Transformer can upward-generalize from only two training iterations without specialized data or projections, it offers a practical, drop-in mechanism for test-time compute scaling that is complementary to parameter scaling. The paper has several strengths: the architecture is simple and clearly described, the training protocol is standard, the authors fix the token budget across comparisons, and they disclose the small number of large-model runs. However, the headline claims are only partially supported by the reported numbers. The abstract's 12% ICL improvement is not recoverable from Table 2, the 5% perplexity improvement is not tabulated, the 350M upward-generalization result rests on a single training run, and the Depth baseline appears to be trained under a protocol that may not match the original method. If the evidence were strengthened through corrected baselines, error bars, and a clearer match between claims and tables, the paper would make a useful contribution.

major comments (5)
  1. [§3.1, Fig. 6b, Table 2] The Depth baseline training protocol is not specified. Section 5 states that the original recurrent-depth method of Geiping et al. (2025) is trained over a range of randomly sampled iteration counts, but the experimental section never states whether the Depth implementation in this paper received that protocol. The collapse of Depth at every iteration count other than r=2, visible in Fig. 6b and Table 2, is exactly what one expects from a model trained at fixed r=2. Because Depth is the principal iterative baseline, the claim that AbbIE “far outperforms alternative iterative methods” is not established unless Depth was trained with the same sampled-iteration protocol as the original method. Please specify the protocol; if Depth was trained at fixed r=2, re-run it with sampled iterations or clearly state the deviation and justify why the comparison is fair.
  2. [Abstract, Table 2, Discussion] The “up to 12% improvement” claim is not recoverable from the reported numbers. In the 350M rows of Table 2, the relative improvements of AbbIE-D at r=8 over Std are approximately 7.1% on HellaSwag (34.8/32.5), 7.7% on LAMBADA (28.0/26.0), and 6.4% on ARC-Easy (51.3/48.2). Comparing against Depth(r=2) gives similar magnitudes. If the 12% figure comes from a different calculation (for example, against the random baseline or against Depth at a collapsed setting), the derivation must be reported. The same applies to the 5% perplexity improvement, which is mentioned in Fig. 5b and the Discussion but never appears as a tabulated value. The abstract and Discussion should state only the gains that are directly readable from the tables.
  3. [§3.2, §4.3, Table 2] The upward-generalization result rests on very limited statistical evidence. The 350M model, which is the only configuration showing clear gains at r=4 and r=8, was trained with a single seed (Section 3.2). Table 2 reports the “best result” without error bars or confidence intervals, and Fig. 4 uses only 16 random samples with one initial state per sample. Given that the central claim of the paper is that two training iterations suffice for upward generalization at larger scale, the authors should either provide multiple 350M seeds with error bars, or at minimum clearly label the single-run status in every figure and table that supports the scaling claim, and temper the strength of the conclusion accordingly.
  4. [§4.1, Fig. 4] The fixed-point comparison between AbbIE-C and AbbIE-D is based on two different quantities. For AbbIE-C, the inter-iteration distance is ‖B(h_k) − h_k‖, because h_{k+1} = B(h_k). For AbbIE-D, the distance is ‖B(h_k)‖, because h_{k+1} = B(h_k) + h_k. These are different functionals with different fixed-point conditions (B(h*) = h* versus B(h*) = 0), so the observation that “AbbIE-C diverges while AbbIE-D converges” does not directly follow from comparing the two curves. Please plot a common distance measure, such as ‖h_{k+1} − h_k‖ or a normalized state difference, for both variants, and state explicitly which quantity is shown.
  5. [§4.2, Fig. 5b] The perplexity improvement claim is not quantified in the text. Figure 5b shows curves that appear separated, and the caption says “roughly 5% better perplexity,” but no numerical Perplexity values or standard deviations are given. Since 200M models were trained with five seeds, the authors should report mean and standard error for Std, Depth, and AbbIE-D at the compute-optimal token budget, and state the actual relative improvement. This is necessary to verify both the direction and the magnitude of the claimed perplexity gain.
minor comments (7)
  1. [Fig. 1 caption] The caption says “increasing training iterations (1, 2, 4, and 8)”, but the models are trained with two iterations and evaluated at those test-time counts; please rephrase to avoid implying varying training iteration counts.
  2. [§2.1] The term “detokenization” is used to mean the transition from token space to concept space inside the model, which is not the standard meaning of detokenization in language modeling; consider a different term or explicitly define the usage.
  3. [§2.2, Eq. (1)-(2)] The notation around path independence is loose: Eq. (1) uses f^∞(x, z0), while Eq. (2) writes f^{(n)}(x_n) and relates it to f^{(n+1)}(x_{n+1}); the role of the initial state z0 and the statement that “we omit z0” should be clarified, since path independence is defined with respect to arbitrary z0.
  4. [§3.2] The learning-rate schedule is described as “Warmup-Stable-Decay”, but the text only describes linear warmup and cosine decay; please specify what the “stable” phase is, or rename the schedule.
  5. [Table 2] There are small typographical issues: “Commensense Question Answering” should be “Commonsense Question Answering”, and the note in the table body is informal; consider moving protocol caveats to the caption.
  6. [§4.3] The sentence “we see a performance significantly higher than the random baseline” uses “significantly” without a statistical test; please replace with “numerically higher” unless a test is performed.
  7. [§4.3] The phrase “which no other general recursive transformer has shown” is absolute; please soften to “to our knowledge” or cite a negative result.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the central claims are direct empirical measurements, and the path-independence discussion is external motivation rather than a derived premise.

full rationale

This paper is an empirical architecture-comparison study rather than a derivation from fitted quantities. The central claims (lower perplexity at equal token budget, upward generalization at r=4/8 after training at r=2, and ICL gains) are obtained by training models and measuring standard metrics. No parameter is fitted to the evaluation benchmarks and then reported as a prediction: the perplexity and ICL numbers in Table 2 and Figures 1, 5, and 6 are direct measurements, not outputs of a fitted model. The path-independence discussion in Sections 2.2 and 5 is used only to motivate the input-reinjection design of AbbIE-C versus AbbIE-D; the paper does not derive the upward-generalization result from path independence, and it tests convergence empirically in Section 4.1 and Figure 4. The only overlapping-author citation is Shen et al. (2025), used in a list with Fan et al. (2019) and Elhoushi et al. (2024) to support the peripheral observation that Transformer blocks can be reused; this citation is not load-bearing, and no uniqueness theorem is imported from the authors' own prior work. The weaker points identified in the manuscript (single 350M run, 16-sample convergence analysis, and the Depth baseline training protocol) are statistical and experimental-fairness concerns about evidence strength, not circular reductions. Therefore no step in the derivation reduces to its own input by construction.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The method introduces no new fitted constants or invented entities. The load-bearing assumptions are about the sufficiency of residual input injection, the representativeness of small-sample convergence evidence, and the transferability of small-scale scaling trends.

free parameters (1)
  • Number of training iterations = 2
    The central claim is that two iterations during training suffice for upward generalization. This number is chosen by hand, not swept or fitted, and the method's reported advantage depends on it.
assumptions (4)
  • domain assumption Residual dependence on the original input is sufficient for path independence and upward generalization.
    Section 2.2 omits the noise-injection component that Anil et al. (2022) identified as necessary for path independence; the paper assumes the residual stream alone suffices, with no proof.
  • domain assumption Inter-iteration distance measured on 16 samples per benchmark is representative of the full distribution.
    Section 4.1 and Fig. 4 draw the fixed-point conclusion from a small sample; the paper assumes this extends to all inputs.
  • domain assumption Perplexity improvements at 200M-350M scale transfer to larger models.
    Section 4.2 extrapolates FLOP-efficiency trends and scaling behavior beyond the tested scale.
  • domain assumption Chinchilla compute-optimal token budgets apply to recursive training.
    Section 3.2 sets COT as 20 tokens per parameter following Hoffmann et al.; the paper assumes this definition is valid when layers are reused.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling." pith.science (2026). https://pith.science/paper/XUDA3LUR

@misc{pith2026250708567,
  author       = {Pith},
  title        = {Pith review of: AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XUDA3LUR}},
  note         = {Machine review of arXiv:2507.08567}
}
read the original abstract

We introduce the Autoregressive Block-Based Iterative Encoder (AbbIE), a novel recursive generalization of the encoder-only Transformer architecture, which achieves better perplexity than a standard Transformer and allows for the dynamic scaling of compute resources at test time. This simple, recursive approach is a complement to scaling large language model (LLM) performance through parameter and token counts. AbbIE performs its iterations in latent space, but unlike latent reasoning models, does not require a specialized dataset or training protocol. We show that AbbIE upward generalizes (ability to generalize to arbitrary iteration lengths) at test time by only using 2 iterations during train time, far outperforming alternative iterative methods. AbbIE's ability to scale its computational expenditure based on the complexity of the task gives it an up to \textbf{12\%} improvement in zero-shot in-context learning tasks versus other iterative and standard methods and up to 5\% improvement in language perplexity. The results from this study open a new avenue to Transformer performance scaling. We perform all of our evaluations on model sizes up to 350M parameters.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning

    physics.chem-ph 2026-02 conditional novelty 6.0 of 10

    LatentChem reasons in continuous latent space for chemistry, achieving a 59.88% non-tie win rate over explicit CoT on ChemCoTBench with a 10.84x average reduction in reasoning overhead.

Reference graph

Works this paper leans on

52 extracted references · 17 canonical work pages · cited by 1 Pith paper

  1. [1]

    Ainslie, J

    J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebron, and S. Sanghai. GQA : Training generalized multi-query transformer models from multi-head checkpoints. In H. Bouamor, J. Pino, and K. Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4895--4901, Singapore, Dec. 2023. Association for...

  2. [2]

    L. B. Allal, A. Lozhkov, E. Bakouch, G. M. Bl \' a zquez, G. Penedo, L. Tunstall, A. Marafioti, H. Kydl \' cek, A. P. Lajar \' n, V. Srivastav, J. Lochner, C. Fahlgren, X. Nguyen, C. Fourrier, B. Burtenshaw, H. Larcher, H. Zhao, C. Zakka, M. Morlon, C. Raffel, L. von Werra, and T. Wolf. Smollm2: When smol goes big - data-centric training of a small langua...

  3. [3]

    L. B. Allal, A. Lozhkov, E. Bakouch, G. M. Blázquez, G. Penedo, L. Tunstall, A. Marafioti, H. Kydlíček, A. P. Lajarín, V. Srivastav, J. Lochner, C. Fahlgren, X.-S. Nguyen, C. Fourrier, B. Burtenshaw, H. Larcher, H. Zhao, C. Zakka, M. Morlon, C. Raffel, L. von Werra, and T. Wolf. Smollm2: When smol goes big -- data-centric training of a small language mode...

  4. [4]

    Amd-llama-135m: A 135m parameter language model trained on amd instinct mi250 accelerators

    AMD . Amd-llama-135m: A 135m parameter language model trained on amd instinct mi250 accelerators. https://huggingface.co/amd/AMD-Llama-135m, 2023. Accessed: 2025-05-15

  5. [5]

    C. Anil, A. Pokle, K. Liang, J. Treutlein, Y. Wu, S. Bai, J. Z. Kolter, and R. B. Grosse. Path independent equilibrium models can better exploit test-time computation. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Syste...

  6. [6]

    S. Bae, A. Fisch, H. Harutyunyan, Z. Ji, S. Kim, and T. Schuster. Relaxed recursive transformers: Effective parameter sharing with layer-wise lora. arXiv preprint arXiv:2410.20672, 2024

  7. [7]

    Ben Allal, A

    L. Ben Allal, A. Lozhkov, G. Penedo, T. Wolf, and L. von Werra. Cosmopedia, February 2024. URL https://huggingface.co/datasets/HuggingFaceTB/cosmopedia

  8. [8]

    Biran, D

    E. Biran, D. Gottesman, S. Yang, M. Geva, and A. Globerson. Hopping too late: Exploring the limitations of large language models on multi-hop queries, 2024. URL https://arxiv.org/abs/2406.12775

Show all 52 references
  1. [9]

    Clark, I

    P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1, 2018

  2. [10]

    DeepSeek - AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu...

  3. [11]

    Dehghani, S

    M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser. Universal transformers. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. URL https://openreview.net/forum?id=HyzdRiR9Y7

  4. [12]

    N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. Bosma, Z. Zhou, T. Wang, Y. E. Wang, K. Webster, M. Pellat, K. Robinson, K. Meier-Hellstern, T. Duke, L. Dixon, K. Zhang, Q. V. Le, Y. Wu, Z. Chen, and C. Cui...

  5. [13]

    Elfwing, E

    S. Elfwing, E. Uchibe, and K. Doya. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017. URL https://arxiv.org/abs/1702.03118

  6. [14]

    Elhage, T

    N. Elhage, T. Hume, C. Olsson, N. Nanda, T. Henighan, S. Johnston, S. ElShowk, N. Joseph, N. DasSarma, B. Mann, D. Hernandez, A. Askell, K. Ndousse, A. Jones, D. Drain, A. Chen, Y. Bai, D. Ganguli, L. Lovitt, Z. Hatfield-Dodds, J. Kernion, T. Conerly, S. Kravec, S. Fort, S. Ka...

  7. [15]

    Elhoushi, A

    M. Elhoushi, A. Shrivastava, D. Liskovich, B. Hosmer, B. Wasti, L. Lai, A. Mahmoud, B. Acun, S. Agarwal, A. Roman, A. A. Aly, B. Chen, and C. Wu. Layerskip: Enabling early exit inference and self-speculative decoding. In L. Ku, A. Martins, and V. Srikumar, editors, Proceedings...

  8. [16]

    A. Fan, E. Grave, and A. Joulin. Reducing transformer depth on demand with structured dropout. arXiv preprint arXiv:1909.11556, 2019

  9. [17]

    Fedus, B

    W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23 0 (120): 0 1--39, 2022

  10. [18]

    Geiping, S

    J. Geiping, S. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson, B. Kailkhura, A. Bhatele, and T. Goldstein. Scaling up test-time compute with latent reasoning: A recurrent depth approach, 2025. URL https://arxiv.org/abs/2502.05171

  11. [19]

    Giannou, S

    A. Giannou, S. Rajput, J. Sohn, K. Lee, J. D. Lee, and D. Papailiopoulos. Looped transformers as programmable computers. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, editors, International Conference on Machine Learning, ICML 2023, 23-29 July ...

  12. [20]

    Grattafiori, A

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru,...

  13. [21]

    Gu and T

    A. Gu and T. Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023

  14. [22]

    Gurnee, N

    W. Gurnee, N. Nanda, M. Pauly, K. Harvey, D. Troitskii, and D. Bertsimas. Finding neurons in a haystack: Case studies with sparse probing. Trans. Mach. Learn. Res., 2023, 2023. URL https://openreview.net/forum?id=JYs1R9IMJr

  15. [23]

    S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y. Tian. Training large language models to reason in a continuous latent space. CoRR, abs/2412.06769, 2024. doi:10.48550/ARXIV.2412.06769. URL https://doi.org/10.48550/arXiv.2412.06769

  16. [24]

    Hoffmann, S

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifr...

  17. [25]

    Hsieh, C.-L

    C.-Y. Hsieh, C.-L. Li, C.-K. Yeh, H. Nakhost, Y. Fujii, A. Ratner, R. Krishna, C.-Y. Lee, and T. Pfister. Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes. arXiv preprint arXiv:2305.02301, 2023

  18. [26]

    Hägele, E

    A. Hägele, E. Bakouch, A. Kosson, L. B. Allal, L. V. Werra, and M. Jaggi. Scaling laws and compute-optimal training beyond fixed training durations, 2024. URL https://arxiv.org/abs/2405.18392

  19. [27]

    Jumper, R

    J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera-Paredes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Petersen, D. Reiman, E. C...

  20. [28]

    Kaplan, M

    G. Kaplan, M. Oren, Y. Reif, and R. Schwartz. From tokens to words: On the inner lexicon of llms. CoRR, abs/2410.05864, 2024. doi:10.48550/ARXIV.2410.05864. URL https://doi.org/10.48550/arXiv.2410.05864

  21. [29]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models. CoRR, abs/2001.08361, 2020. URL https://arxiv.org/abs/2001.08361

  22. [30]

    Katz and Y

    S. Katz and Y. Belinkov. VISIT: visualizing and interpreting the semantic information flow of transformers. In H. Bouamor, J. Pino, and K. Bali, editors, Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023 , pages 14094--14113....

  23. [31]

    Kocetkov, R

    D. Kocetkov, R. Li, L. Ben Allal, J. Li, C. Mou, C. Muñoz Ferrandis, Y. Jernite, M. Mitchell, S. Hughes, T. Wolf, D. Bahdanau, L. von Werra, and H. de Vries. The stack: 3 tb of permissively licensed source code. Preprint, 2022

  24. [32]

    H. Lu, Y. Zhou, S. Liu, Z. Wang, M. W. Mahoney, and Y. Yang. Alphapruning: Using heavy-tailed self regularization theory for improved layer-wise pruning of large language models. Advances in Neural Information Processing Systems, 37: 0 9117--9152, 2024

  25. [33]

    McCandlish, J

    S. McCandlish, J. Kaplan, D. Amodei, and O. D. Team. An empirical model of large-batch training, 2018. URL https://arxiv.org/abs/1812.06162

  26. [34]

    M. I. Nye, A. J. Andreassen, G. Gur - Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, C. Sutton, and A. Odena. Show your work: Scratchpads for intermediate computation with language models. CoRR, abs/2112.00114, 2021. URL https://arxiv.org...

  27. [35]

    Paperno, G

    D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández. The lambada dataset: Word prediction requiring a broad discourse context, 2016. URL https://arxiv.org/abs/1606.06031

  28. [36]

    Penedo, H

    G. Penedo, H. Kydl \' cek, L. B. Allal, A. Lozhkov, M. Mitchell, C. A. Raffel, L. von Werra, and T. Wolf. The fineweb datasets: Decanting the web for the finest text data at scale. In A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang, editor...

  29. [37]

    Penedo, H

    G. Penedo, H. Kydlíček, L. B. allal, A. Lozhkov, M. Mitchell, C. Raffel, L. V. Werra, and T. Wolf. The fineweb datasets: Decanting the web for the finest text data at scale, 2024 b . URL https://arxiv.org/abs/2406.17557

  30. [38]

    B. Peng, E. Alcaide, Q. G. Anthony, A. Albalak, S. Arcadinho, S. Biderman, H. Cao, X. Cheng, M. N. Chung, L. Derczynski, et al. Rwkv: Reinventing rnns for the transformer era. In The 2023 Conference on Empirical Methods in Natural Language Processing

  31. [39]

    R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, A. Levskaya, J. Heek, K. Xiao, S. Agrawal, and J. Dean. Efficiently scaling transformer inference, 2022. URL https://arxiv.org/abs/2211.05102

  32. [40]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. OpenAI, 2019. URL https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf. Accessed: 2024-11-15

  33. [41]

    Sardana, J

    N. Sardana, J. Portes, S. Doubov, and J. Frankle. Beyond chinchilla-optimal: Accounting for inference in language model scaling laws, 2025. URL https://arxiv.org/abs/2401.00448

  34. [42]

    W. F. Shen, X. Qiu, M. Kurmanji, A. Iacob, L. Sani, Y. Chen, N. Cancedda, and N. D. Lane. Lunar: Llm unlearning via neural activation redirection. arXiv preprint arXiv:2502.07218, 2025

  35. [43]

    J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomput., 568 0 (C), Feb. 2024. ISSN 0925-2312. doi:10.1016/j.neucom.2023.127063. URL https://doi.org/10.1016/j.neucom.2023.127063

  36. [44]

    Talmor, J

    A. Talmor, J. Herzig, N. Lourie, and J. Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge, 2019. URL https://arxiv.org/abs/1811.00937

  37. [45]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. In I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Pro...

  38. [46]

    von Platen and T

    P. von Platen and T. Wolf. Smollm: Efficient language models for everyone. https://huggingface.co/blog/smollm, dec 2023. Accessed: 2025-05-13

  39. [47]

    S. Wang, P. Zhou, J. Li, and H. Huang. 4-bit shampoo for memory-efficient network training. Advances in Neural Information Processing Systems, 37: 0 126997--127029, 2024

  40. [48]

    Wolters, X

    C. Wolters, X. Yang, U. Schlichtmann, and T. Suzumura. Memory is all you need: An overview of compute-in-memory architectures for accelerating large language model inference, 2024. URL https://arxiv.org/abs/2406.08413

  41. [49]

    Xiong, Y

    R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T. Liu. On layer normalization in the transformer architecture. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event , volume 119...

  42. [50]

    L. Yang, K. Lee, R. D. Nowak, and D. Papailiopoulos. Looped transformers are better at learning learning algorithms. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL https://openreview.n...

  43. [51]

    Zellers, A

    R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. Hellaswag: Can a machine really finish your sentence?, 2019. URL https://arxiv.org/abs/1905.07830

  44. [52]

    Zhang and R

    B. Zhang and R. Sennrich. Root mean square layer normalization, 2019. URL https://arxiv.org/abs/1910.07467

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.