Pith. sign in

REVIEW 5 minor 7 references

Training a residual-free compressed ReLU network under L4 loss elicits computation in superposition via sparse binary codes over neurons.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 13:18 UTC pith:QLFFZH5O

load-bearing objection L4 (or any p>2) turns the residual-free compressed-computation toy into a clean, reverse-engineered sparse-binary-code + pseudoinverse instance of CiS; residual gaps are quantified and do not erase the mechanism.

arxiv 2607.04800 v1 pith:QLFFZH5O submitted 2026-07-06 cs.LG

Compressed Computation under L⁴ Loss is likely Computation in Superposition

classification cs.LG
keywords computation in superpositioncompressed computationL4 losssparse binary codesReLU networkstoy modelsmechanistic interpretability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Neural networks are known to pack more concepts than dimensions through representation in superposition, but whether they also pack more nonlinear functions than neurons—computation in superposition—has lacked trained toy models that can be fully reverse-engineered. This paper shows that a single-hidden-layer ReLU network with 50 neurons, asked to compute the elementwise ReLU of 100 sparse input features, learns such a solution when trained under L4 loss rather than the usual L2. The network assigns each feature a short sparse binary codeword over the neurons and decodes with a near-pseudoinverse of the encoder. Given those codewords, three scalars recover most of the trained performance, and hand-designed biregular codes of the same family match it. The result supplies a concrete, reverse-engineerable trained instance of computation in superposition that earlier L2 versions of the same setup did not produce.

Core claim

Under L4 (or any Lp with p>2), the residual-free compressed-computation model—50 ReLU neurons computing the ReLU of each of 100 sparse features—learns a solution that appears to compute all features in superposition: each feature is given a sparse binary codeword over neurons and is read out by a scaled pseudoinverse of the encoder. A three-scalar ansatz on those codewords recovers about 1.13 times the trained L4 loss, and designed biregular codes reach about 1.12 times, far below non-superposition baselines.

What carries the argument

Sparse binary codewords over neurons decoded by a scaled pseudoinverse of the encoder. Encoder entries split into large on-code positives and small off-code negatives, defining a binary matrix M; the three-scalar family Win = a M^T + b(1−M^T), Wout = c (Win)+ then recovers most performance, and overlap-minimized biregular codes of length K≈5 work equally well.

Load-bearing premise

Even, low per-feature error together with a binary-code and near-pseudoinverse description is enough to count as genuine computation in superposition rather than some other distributed approximation.

What would settle it

A non-superposition construction that still yields the same narrow, even per-feature MSE distribution under L4, or a codeword-swap test in which transplanting one feature’s hidden magnitudes onto another feature’s codeword neurons fails to decode to that feature.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Losses with exponent above 2 can be used as a practical training knob to favor even, all-features solutions over winner-take-all feature selection.
  • The trained network is an empirical instance of sparse binary coding for computation in superposition, arising by gradient descent rather than hand design.
  • The reverse-engineered solution can serve as ground truth for testing parameter-decomposition and other interpretability methods on a known computation-in-superposition mechanism.
  • The same residual-free, higher-order-loss recipe can be applied to larger or multi-layer models to check whether the sparse-code-plus-pseudoinverse motif persists.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Higher-order losses may systematically favor interference-minimizing codes whenever the number of sparse tasks exceeds hidden width.
  • The remaining gap from non-uniform on-code values and residual decoder structure may implement a soft load-balancing or error-correction step predicted by theoretical constructions of computation in superposition.
  • Applying the same L4 recipe to multi-layer or attention models would test whether computation in superposition scales beyond a single ReLU layer.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper studies a residual-free compressed-computation toy model: a single-hidden-layer ReLU network with 50 neurons that must compute the elementwise ReLU of 100 sparse continuous features (p=0.02). Training under L4 (or any Lp with p>2) rather than L2 elicits a solution that spreads error evenly across all features, far below naive and emulate-bias baselines. The authors reverse-engineer the solution as a sparse binary code over neurons (near-biregular, K~5.5) decoded by a near-pseudoinverse of the encoder (Frobenius cosine 0.996). A three-scalar ansatz on the extracted codewords recovers ~1.13 imes the trained L4 loss; hand-designed overlap-minimized biregular codes reach ~1.12 imes. Residual gaps (non-uniform on-code values, decoder residual) are quantified in Table 1. The same mechanism appears with and without a random embedding.

Significance. Computation in superposition has been largely theoretical or hand-designed; this is one of the first trained, reverse-engineered toy models of it. The multi-pronged evidence (per-feature MSE distributions, encoder bimodality, code regularity, 100% swap-test success, decoder–pinv agreement, three-scalar and hand-designed reconstructions) is quantitative and reproducible, with code released. The residual-free setup cleanly isolates superposition from the mixing-matrix artifact identified by Bhagat et al. The result supplies a concrete ground-truth solution for testing parameter-decomposition methods (APD/SPD) and empirically instantiates the sparse-binary-code constructions of Hänni et al. Even if the remaining encoder/decoder residuals are not fully explained, the paper substantially advances the empirical study of CiS.

minor comments (5)
  1. [Section 4.1 / Figure 2] Section 4.1 and Figure 2: the coefficient-of-variation statistic is well motivated, but a short explicit statement of how many seeds and what batch size underlie the violin plots would improve reproducibility of the L2-vs-Lp contrast.
  2. [Section 4.2 / Figure 3] Section 4.2 / Figure 3: the on/off threshold of 0.05 is stated but not justified; a brief sensitivity check (or note that the two clusters are well separated) would strengthen the binary-code extraction.
  3. [Table 1] Table 1: the row labels are dense; a short caption sentence clarifying the difference between “trained support, free per-entry, pinv” and “full trained encoder, pinv” would help readers parse the ablation.
  4. [Appendix C] Appendix C: the edge-swap objective (sum of squared off-diagonal overlaps) is clear, but the acceptance criterion and number of iterations could be moved into the main text or a short methods paragraph for readers who do not open the appendix.
  5. [Section 2] Related Work: Gibson (2026) is concurrent and carefully distinguished; a one-sentence note on whether the L4-vs-L2 phenomenon appears in that identification task would be a useful cross-check if the authors have tried it.

Circularity Check

1 steps flagged

No significant circularity: reverse-engineering extracts a binary code from the trained encoder and fits a 3-scalar ansatz, but independent hand-designed codes and L4-vs-L2 baselines keep the claim non-tautological.

specific steps
  1. fitted input called prediction [Sec. 4.3 / Table 1 (3-scalar ansatz)]
    "Fitting these three scalars to the trained network’s codewords gives an L4 loss ∼1.13× that of the trained network... Trained support, 3 scalars pinv 1.13×"

    The binary support M is read off the trained encoder by thresholding, then a,b,c are fitted to the same network; the resulting loss ratio is therefore a description of the network rather than an independent prediction. The circularity is mild because the identical three-parameter family is also evaluated on independently generated biregular and random codes (Fig. 7) that never saw the trained weights, and those codes reach comparable loss.

full rationale

The paper is an empirical reverse-engineering study of a residual-free compressed-computation network trained under L4 (or Lp, p>2). The central claims rest on (i) the L4-trained network achieving uniformly low per-feature MSE far below the naive and emulate-bias baselines (Fig. 2, CV ~0.03), (ii) encoder bimodality yielding near-biregular sparse binary codewords, a 100% codeword-swap test, and decoder–pseudoinverse Frobenius cosine 0.996 (Sec. 4.2), and (iii) a three-scalar ansatz on those codewords recovering ~1.13× loss, with independently generated overlap-minimized biregular codes reaching ~1.12× (Fig. 7, Table 1). Extracting M from Win and then fitting a,b,c is ordinary description of a trained network, not a first-principles derivation that reduces a claimed prediction to its own inputs by construction. The same ansatz family is validated on synthetic codes never seen by the trained network, and the L2 control produces the qualitatively different naïve solution. Self-citations (Braun et al., Bhagat et al.) supply the toy-model setup and the residual-removal motivation; they are reconfirmed under L2 and are not load-bearing uniqueness theorems that force the L4 CiS claim. No equation equates a fitted quantity to a quantity later presented as an independent prediction. Residual unexplained structure (non-uniform on-code values, decoder residual) is quantified rather than papered over. Hence the derivation chain is self-contained against external baselines; circularity is at most the mild self-reference inherent to any reverse-engineering description.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 1 invented entities

The central claim rests on a standard sparse-input ReLU regression task, the modeling choice to remove the residual connection (so that beating the naive baseline requires superposition), and the empirical observation that L4 loss produces a binary-code solution. Free parameters are the three ansatz scalars and the codeword length K; no new physical entities are postulated. Domain assumptions are the usual ones for toy models of superposition (sparsity, ReLU, single hidden layer).

free parameters (5)
  • on-code value a
    Scalar fitted in the three-parameter ansatz Win = a M^T + b(1-M^T); recovers most of trained performance when combined with b and c.
  • off-code value b
    Second free scalar of the three-parameter ansatz; small negative value observed empirically (~-0.02).
  • decoder scale c
    Third free scalar; multiplies the pseudoinverse of the encoder in the ansatz.
  • codeword length K
    Chosen or observed (trained mean ~5.5; designed codes sweep K); controls capacity and interference.
  • sparsity p
    Input feature activation probability fixed at 0.02; controls expected number of simultaneous active features and therefore interference regime.
axioms (4)
  • domain assumption Removing the residual connection sets the target exactly to y=ReLU(x), so any substantial improvement over the naive half-features baseline must come from superposition.
    Stated in Section 3, following Bhagat et al. (2025); necessary for the claim that the observed solution is CiS rather than residual mixing.
  • domain assumption Uniform low per-feature error (low coefficient of variation of MSE) plus neuron sharing is evidence of computation in superposition.
    Used throughout Section 4.1 as the operational definition of CiS for this toy model.
  • domain assumption ReLU networks with sparse continuous inputs can be reverse-engineered by inspecting encoder entry distributions and decoder–pseudoinverse agreement.
    Methodological premise of Sections 4.2–4.3; standard in mechanistic interpretability of toy models.
  • standard math Standard facts of linear algebra (pseudoinverse, Frobenius cosine, edge-swap graph rewiring).
    Used for decoder analysis and code design (Appendix C).
invented entities (1)
  • sparse binary codeword over neurons (M_j) independent evidence
    purpose: Internal identifier that the trained network assigns to each feature; support of large positive encoder entries.
    Extracted post-hoc from the trained encoder; not postulated a priori but discovered. Independent evidence is partial: hand-designed codes with the same structure work nearly as well.

pith-pipeline@v1.1.0-grok45 · 14578 in / 3042 out tokens · 21011 ms · 2026-07-11T13:18:32.567081+00:00 · methodology

0 comments
read the original abstract

Neural networks are thought to represent concepts as directions in their activation space, and superposition lets them encode more concepts than they have dimensions. It is natural to ask whether they can also compute more functions than they have neurons, i.e., perform computation in superposition. In this regime many functions of sparse inputs are evaluated by a layer with fewer neurons than there are functions to compute. Representation in superposition is by now fairly well understood, but computation in superposition is not, and there are few toy models of it arising through training rather than being hand designed. As a toy model of computation in superposition we study the compressed-computation setup: a single-hidden-layer ReLU network with 50 neurons that must compute the ReLU of each of 100 sparse input features. We show that training it under an $L^4$ loss (the mean fourth power of the error), rather than the usual $L^2$, elicits a solution that appears to compute all features in superposition. We then reverse-engineer this solution. We find that the network assigns each feature a sparse binary codeword over neurons and decodes it with a pseudoinverse of the encoder. Given these codewords, a description with only three scalars recovers most of the network's performance, and we validate it by building equivalent networks from hand-designed codes.

Figures

Figures reproduced from arXiv: 2607.04800 by Francisco Ferreira da Silva, Stefan Heimersheim.

Figure 1
Figure 1. Figure 1: Network architecture: a linear encoder Win, 50 ReLU neurons without bias, and a linear decoder Wout. The input and output are 100-dimensional; the task is to compute the elementwise ReLU of the input. superposition, is what lets the model beat the naive baseline of ignoring half the features. Removing the residual connec￾tion sets M = 0, so the target is exactly y = ReLU(x) and any substantial improvement … view at source ↗
Figure 2
Figure 2. Figure 2: Per-feature MSE on x > 0 samples: the strongest non￾superposition baseline and the trained networks at several loss exponents. Each violin pools the per-feature MSE (over the 100 features, and over 10 seeds for the trained networks); black dots are medians. The emulate-bias baseline and L 2 (reds) are bimodal — half the features handled, half left near the do-nothing baseline of 1/3 (dashed), which the bia… view at source ↗
Figure 3
Figure 3. Figure 3: shows the distribution of the entries of the encoder Win, pooled over all 100 feature columns. The entries fall into two well-separated groups (values reported as mean ± standard deviation): a large mass of small negative values (−0.02 ± 0.01) and a smaller mass of large positive values (0.34 ± 0.06). We call the support of the large entries the codeword of feature j, written Mj ∈ {0, 1} N — the subset of … view at source ↗
Figure 4
Figure 4. Figure 4: The binary code is regular: number of large encoder entries (those exceeding 0.05, i.e. on-code entries) per feature column (blue, left axis) and per neuron row (red, right axis). Every feature uses 5-7 neurons (mean 5.5) and every neuron is shared by 10-12 codewords (mean 10.9); the tight bimodal distribution shows the code is close to biregular. If the codewords serve as the network’s internal identifier… view at source ↗
Figure 6
Figure 6. Figure 6: Mechanism for a single active feature: the encoder pro￾duces large positives on the feature’s codeword neurons and small negatives elsewhere; ReLU zeros the negatives; the pseudoinverse decoder reads the codeword and outputs a peak at the same index. Fitting these three scalars to the trained network’s code￾words gives an L 4 loss ∼1.13× that of the trained network, far below the non-superposition baseline… view at source ↗
Figure 7
Figure 7. Figure 7: L 4 loss of 3-parameter ansatz networks: Loss of networks built from synthetic binary codes (biregular and random) as a function of codeword length K, normalized by the trained model’s loss. Solid lines apply overlap-minimizing edge swaps (five seeds each; error bars are standard deviations, scatter points are individual seeds); dashed lines are the same codes without edge swaps. Biregular K = 5 with edge … view at source ↗
Figure 8
Figure 8. Figure 8: shows this embedded architecture, with the fixed embedding WE and unembedding W⊤ E restored relative to the axis-aligned model of [PITH_FULL_IMAGE:figures/full_fig_p006_8.png] view at source ↗
Figure 10
Figure 10. Figure 10: Embedded model, effective-encoder entries (cf. Fig￾ure 3). Pooled histogram of the entries of the effective encoder Wfin = WinW⊤ E ; the same bimodal split into a small off-code cluster and a larger on-code cluster appears [PITH_FULL_IMAGE:figures/full_fig_p007_10.png] view at source ↗
Figure 9
Figure 9. Figure 9: Embedded model, per-feature MSE (cf [PITH_FULL_IMAGE:figures/full_fig_p007_9.png] view at source ↗
Figure 12
Figure 12. Figure 12: Single-feature input–output response. For each fea￾ture j in turn, feature j alone is driven with value v and we read the output at index j; the blue curve is the mean over all 100 fea￾tures. The trained network traces a ReLU attenuated to mean slope ≈ 0.81, below the true ReLU (dashed). The 5–95% band across features is too narrow to be visible: every feature is attenuated almost identically, as expected… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

7 extracted references · 3 linked inside Pith

  1. [1]

    and Shavit, N

    Adler, M. and Shavit, N. On the complexity of neu- ral computation in superposition.arXiv preprint arXiv:2409.15318,

  2. [2]

    Compressed computation is (probably) not com- putation in superposition

    Bhagat, J., Molas-Medina, S., Giglemiani, G., and Heimer- sheim, S. Compressed computation is (probably) not com- putation in superposition. InMechanistic Interpretability Workshop at NeurIPS 2025,

  3. [4]

    Bushnaq, L., Braun, D., and Sharkey, L

    URL https://ar xiv.org/abs/2501.14926. Bushnaq, L., Braun, D., and Sharkey, L. Stochastic param- eter decomposition.arXiv preprint arXiv:2506.20790,

  4. [5]

    Toy models of superposition.arXiv preprint arXiv:2209.10652,

    Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., et al. Toy models of superposition.arXiv preprint arXiv:2209.10652,

  5. [7]

    URL https://ar xiv.org/abs/2408.05451. A. The embedded model The main text studies a model in which the 50 ReLU neu- rons act directly on the 100 input features. Braun et al. (2025)’s original compressed-computation model instead embeds the 100-dimensional input into R1000 through a fixed random matrix WE with unit-norm rows: the 50 neu- rons read and wri...

  6. [8]

    and SPD (Bushnaq et al., 2025), for which the non-axis-aligned features of the embedded model are the more challenging and realistic setting. For the narrower claim that the trained network computes in superposition, the axis-aligned model in the main text is cleaner: with axis-aligned input features there is no random embedding that could, even in princi...

  7. [9]

    for the embedded model. 6 Compressed Computation underL 4 Loss is likely Computation in Superposition Table 2.The solution is unchanged by adding the random embedding.Headline quantities from Sections 4.1 and 5 for the axis-aligned model (main text) and an otherwise identical model trained with Braun et al. (2025)’s random embedding. CV is the coefficient...