Pith. sign in

REVIEW 2 cited by

Group Crosscoders for Mechanistic Analysis of Symmetry

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.24184 v2 pith:XGAHIJEH submitted 2024-10-31 cs.LG

classification cs.LG
keywords groupcrosscodersanalysisfeaturesnetworksneuralsymmetryfeature
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We introduce group crosscoders, an extension of crosscoders that systematically discover and analyse symmetrical features in neural networks. While neural networks often develop equivariant representations without explicit architectural constraints, understanding these emergent symmetries has traditionally relied on manual analysis. Group crosscoders automate this process by performing dictionary learning across transformed versions of inputs under a symmetry group. Applied to InceptionV1's mixed3b layer using the dihedral group $\mathrm{D}_{32}$, our method reveals several key insights: First, it naturally clusters features into interpretable families that correspond to previously hypothesised feature types, providing more precise separation than standard sparse autoencoders. Second, our transform block analysis enables the automatic characterisation of feature symmetries, revealing how different geometric features (such as curves versus lines) exhibit distinct patterns of invariance and equivariance. These results demonstrate that group crosscoders can provide systematic insights into how neural networks represent symmetry, offering a promising new tool for mechanistic interpretability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding sparse autoencoder scaling in the presence of feature manifolds

    cs.LG 2025-09 conditional novelty 5.0 of 10

    When loss improvement from tiling a common feature manifold decays slower than feature frequency decays, sparse autoencoders allocate most latents to that manifold and discover sublinearly many features.

  2. Naturally Computed Scale Invariance in the Residual Stream of ResNet18

    cs.CV 2025-04 conditional novelty 5.0 of 10

    ResNet18's residual stream can build scale invariance by summing a smaller-scale feature from the block input with a larger-scale feature from the pre-sum output, and ablating these channels mildly impairs scale-robus...

Pith tools