REVIEW 6 cited by
Scaling Laws for Linear Complexity Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The interest in linear complexity models for large language models is on the rise, although their scaling capacity remains uncertain. In this study, we present the scaling laws for linear complexity language models to establish a foundation for their scalability. Specifically, we examine the scaling behaviors of three efficient linear architectures. These include TNL, a linear attention model with data-independent decay; HGRN2, a linear RNN with data-dependent decay; and cosFormer2, a linear attention model without decay. We also include LLaMA as a baseline architecture for softmax attention for comparison. These models were trained with six variants, ranging from 70M to 7B parameters on a 300B-token corpus, and evaluated with a total of 1,376 intermediate checkpoints on various downstream tasks. These tasks include validation loss, commonsense reasoning, and information retrieval and generation. The study reveals that existing linear complexity language models exhibit similar scaling capabilities as conventional transformer-based models while also demonstrating superior linguistic proficiency and knowledge retention.
Forward citations
Cited by 6 Pith papers
-
Small transformer architectures for task switching
On a new task-switching benchmark, a cisformer with expressive attention reaches about 95% accuracy, while standard transformers, LSTMs, and MLPs all stay near or below 60%.
-
A Systematic Analysis of Hybrid Linear Attention
A systematic 72-model study shows standalone linear-attention quality does not predict hybrid performance, and a 3:1 to 6:1 linear-to-full ratio best balances recall against KV-cache cost.
-
Fast and Simplex: 2-Simplicial Attention in Triton
2-simplicial attention, implemented in Triton with a sliding window, is claimed to yield a steeper loss-versus-parameters scaling exponent than dot-product attention on math and reasoning benchmarks.
-
Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective
A new linear attention with spatial-aware decay at row boundaries lowers FID for autoregressive image generation on ImageNet relative to the softmax LlamaGen baseline, but the description of the core mask is internall...
-
Essential-Web v1.0: 24T tokens of organized web data
A 24T-token web corpus with automatic document-level taxonomy labels enables competitive domain-specific datasets via simple filters.
-
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
LASP-2 trains linear-attention transformers with long sequences by exchanging GPU memory states in one all-gather step, improving throughput over prior sequence-parallel methods.
Discussion (0). Continue with ORCID to comment.