Pith. sign in

REVIEW 5 cited by

Softmax is not Enough (for Sharp Size Generalisation)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.01104 v3 pith:7OB6TWNU submitted 2024-10-01 cs.LG cs.AIcs.ITmath.IT

classification cs.LGcs.AIcs.ITmath.IT
keywords softmaxsharpcircuitsfunctioninputsperformsizesystems
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A key property of reasoning systems is the ability to make sharp decisions on their input data. For contemporary AI systems, a key carrier of sharp behaviour is the softmax function, with its capability to perform differentiable query-key lookups. It is a common belief that the predictive power of networks leveraging softmax arises from "circuits" which sharply perform certain kinds of computations consistently across many diverse inputs. However, for these circuits to be robust, they would need to generalise well to arbitrary valid inputs. In this paper, we dispel this myth: even for tasks as simple as finding the maximum key, any learned circuitry must disperse as the number of items grows at test time. We attribute this to a fundamental limitation of the softmax function to robustly approximate sharp functions with increasing problem size, prove this phenomenon theoretically, and propose adaptive temperature as an ad-hoc technique for improving the sharpness of softmax at inference time.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Learned replacement non-linearities show transformers are rarely optimal for algorithmic tasks, with benefits that are task-specific, while language/code gains are smaller and more transferable.

  2. What makes a good feedforward computational graph?

    cs.LG 2025-02 conditional novelty 7.0 of 10

    The authors define mixing time and minimax fidelity for feedforward graphs, use them to design a recursive sparse graph (FS) with polylogarithmic mixing time, and show it matches dense attention on parity and retrieval tasks.

  3. Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Softmax temperature controls the rank of learned representations: high temperature induces rank-deficit bias, compresses features, hurts OOD generalization, and improves OOD detection.

  4. Pippo: High-Resolution Multi-View Humans from a Single Image

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A single-image multi-view diffusion transformer generates 1K-resolution turnaround views of humans, with attention biasing for many views and a new reprojection-error metric.

  5. SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SEMA combines window attention with global token averaging, motivated by a dispersion theorem for generalized attention, and reports 0.2 to 0.7 percent top-1 accuracy gains over comparable vision Mamba and MILA models.

Pith tools