Pith. sign in

REVIEW 7 cited by

Looped Transformers for Length Generalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.15647 v5 pith:3KQZCGFO submitted 2024-09-24 cs.LG

classification cs.LG
keywords transformerslengthgeneralizationloopedtasksinputslength-generalizableoperation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work has shown that Transformers trained from scratch can successfully solve various arithmetic and algorithmic tasks, such as adding numbers and computing parity. While these Transformers generalize well on unseen inputs of the same length, they struggle with length generalization, i.e., handling inputs of unseen lengths. In this work, we demonstrate that looped Transformers with an adaptive number of steps significantly improve length generalization. We focus on tasks with a known iterative solution, involving multiple iterations of a RASP-L operation - a length-generalizable operation that can be expressed by a finite-sized Transformer. We train looped Transformers using our proposed learning algorithm and observe that they learn highly length-generalizable solutions for various tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Weight-tied looped transformers on group prefix products implement a linear computation frontier whose speed matches the training loop budget, and a new convergence-time instrument reveals it.

  2. Enhancing Auto-regressive Chain-of-Thought through Loop-Aligned Reasoning

    cs.CL 2025-02 conditional novelty 7.0 of 10

    Loop-aligned supervision lets a looped Transformer generate CoT chains beyond training length, and those chains improve an auto-regressive CoT model's length generalization.

  3. ELT: Elastic Looped Transformers for Visual Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Weight-shared looped transformers trained with intra-loop self-distillation match MaskGIT-class FID/FVD at roughly 4x fewer parameters and support any-time inference across loop counts.

  4. Channel-Wise MLPs Improve the Generalization of Recurrent Convolutional Networks

    cs.LG 2025-08 conditional novelty 5.0 of 10

    Adding a gated channel-wise MLP to a recurrent convolutional network raises median exact-match accuracy on 185 Re-ARC tasks from 78.75% to 92.19% in-distribution and from 2.34% to 14.58% on harder out-of-distribution tasks.

  5. Extrapolation by Association: Length Generalization Transfer in Transformers

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Length generalization on a short-trained main task can be inherited from a longer-trained related auxiliary task trained jointly with it.

  6. Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs

    cs.LG 2025-07 reject novelty 4.0 of 10

    Pretrained LLM layers can be skipped/repeated per input to build custom paths, but the search uses ground-truth answers, so the accuracy gains are fitted, not predicted.

  7. SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought

    cs.AI 2025-05 conditional novelty 4.0 of 10

    SCOUT combines progressive distillation with a cross-attention module to make recursive latent reasoning work through fine-tuning, yielding up to 1.8% accuracy gains over standard fine-tuning.

Pith tools