Pith. sign in

REVIEW 1 cited by

Neural Characteristic Activation Analysis and Geometric Parameterization for ReLU Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.15912 v5 pith:ZNRZ4SOM submitted 2023-05-25 cs.LG stat.ML

classification cs.LGstat.ML
keywords neuralparameterizationreluactivationanalysischaracteristicconvergencegeneralization
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce a novel approach for analyzing the training dynamics of ReLU networks by examining the characteristic activation boundaries of individual ReLU neurons. Our proposed analysis reveals a critical instability in common neural network parameterizations and normalizations during stochastic optimization, which impedes fast convergence and hurts generalization performance. Addressing this, we propose Geometric Parameterization (GmP), a novel neural network parameterization technique that effectively separates the radial and angular components of weights in the hyperspherical coordinate system. We show theoretically that GmP resolves the aforementioned instability issue. We report empirical results on various models and benchmarks to verify GmP's advantages of optimization stability, convergence speed and generalization performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Capacity Matters: a Proof-of-Concept for Transformer Memorization on Real-World Data

    cs.CL 2025-06 conditional novelty 5.0 of 10

    On structured SNOMED-derived memorization tasks, small transformers memorize most when embedding size is large and depth is kept low.

Pith tools