Pith. sign in

REVIEW 7 cited by

AI and Memory Wall

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14123 v1 pith:P3FJJATJ submitted 2024-03-21 cs.LG cs.ARcs.DC

classification cs.LGcs.ARcs.DC
keywords memorybandwidthbottlenecktrainingcomputedecodermodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The availability of unprecedented unsupervised training data, along with neural scaling laws, has resulted in an unprecedented surge in model size and compute requirements for serving/training LLMs. However, the main performance bottleneck is increasingly shifting to memory bandwidth. Over the past 20 years, peak server hardware FLOPS has been scaling at 3.0x/2yrs, outpacing the growth of DRAM and interconnect bandwidth, which have only scaled at 1.6 and 1.4 times every 2 years, respectively. This disparity has made memory, rather than compute, the primary bottleneck in AI applications, particularly in serving. Here, we analyze encoder and decoder Transformer models and show how memory bandwidth can become the dominant bottleneck for decoder models. We argue for a redesign in model architecture, training, and deployment strategies to overcome this memory limitation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Structured Recurrent Mixers for Massively Parallelized Sequence Generation

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    Structured Recurrent Mixers provide a dual parallel-recurrent representation for sequence models, claiming superior training efficiency, information capacity, and inference throughput over linear complexity alternatives.

  2. Masked Gated Linear Unit

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A masked, single-weight-matrix gating scheme, MGLU, matches Gated Linear Unit accuracy on tested language tasks while reducing per-token memory reads by up to 47 percent.

  3. A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A skill-graph random-walk model gives closed-form accuracy-versus-compute formulas for four reasoning strategies and connects them to training scaling.

  4. Hardware-Efficient Attention for Fast Decoding

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Grouped-Tied Attention and Grouped Latent Attention reduce KV-cache memory and speed up LLM decoding by up to 2x while matching the quality of GQA and MLA at up to 1.47B parameters.

  5. Matrix-Free Methods for Finite-Strain Elasticity: Automatic Code Generation with No Performance Overhead

    math.NA 2025-05 conditional novelty 6.0 of 10

    Code generated by automatic differentiation runs as fast or faster than hand-written quadrature kernels in matrix-free finite-strain elasticity, with lower development effort.

  6. Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Ternary language models trained on 1.2 trillion tokens continue to improve, and a new GPU kernel speeds up their inference up to 5x end-to-end.

  7. Toward a Global Regime for Compute Governance: Building the Pause Button

    cs.CY 2025-06 conditional novelty 4.0 of 10

    The paper argues that a global, enforceable compute pause is achievable through a layered framework of hardware controls, supply chain tracking, and regulation.

Pith tools