REVIEW 3 major objections 5 minor 14 references
KAN-PCA recovers classical PCA exactly when its edge functions are forced to be affine, and otherwise captures slightly more in-sample variance on stock returns.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 16:25 UTC pith:GL6E4L6A
load-bearing objection Sound special-case theorem that PCA is the affine limit of a KAN-encoder/linear-decoder autoencoder; honest leakage correction; tiny empirical edge that vanishes OOS and never isolates the crisis regime that motivates the work. the 3 major comments →
Nonlinear Factor Decomposition via Kolmogorov-Arnold Networks: A Spectral Approach to Asset Return Analysis
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
When every spline activation in the KAN encoder is forced to have vanishing second derivative (i.e., to be affine), the optimal KAN-PCA encoder and linear decoder recover exactly the classical PCA projection onto the top-k eigenvectors of the sample covariance; with free splines the same architecture captures modestly more training reconstruction variance than train-fit PCA at equal factor count.
What carries the argument
Theorem 4 (KAN-PCA reduces to PCA in the linear limit): affine edge functions turn each KAN layer into ordinary matrix multiplication, so the reconstruction objective becomes the classical linear autoencoder whose unique optimum is the PCA subspace.
Load-bearing premise
That a modest gain in reconstruction R-squared on a fixed twenty-stock, three-factor sample (with the main crisis inside the training window) is enough to show that spline-mediated nonlinear factors improve risk decomposition relative to linear PCA.
What would settle it
Re-run the identical three-factor comparison on a larger universe (order of 200 stocks) with the main stress period held strictly out of sample; if free-spline KAN-PCA no longer exceeds train-fit PCA even in-sample, or falls materially behind out of sample, the claimed practical advantage of the larger function class disappears.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes KAN-PCA, an autoencoder with a Kolmogorov–Arnold Network encoder and a linear decoder for extracting nonlinear factors from asset returns. The central theoretical claim is that if all edge activations are constrained to be affine, the model recovers classical PCA exactly (Theorem 4 / Corollary 5), via the standard Baldi–Hornik linear-autoencoder optimum. Empirically, on 20 S&P 500 stocks with k=3 factors and a corrected 70/10/20 protocol, KAN-PCA slightly exceeds train-fit PCA in-sample (65.56% vs 64.71% explained variance) but is marginally behind out-of-sample (56.73% vs 57.32%). The authors document and correct a data-leakage issue from earlier training, and present edge-function and factor-path visualizations as qualitative evidence of nonlinearity.
Significance. The special-case recovery of PCA is a clean and useful characterization: it places KAN-PCA in a strictly larger function class with a known linear limit, which is stronger than most generic nonlinear autoencoder papers in asset pricing. The methodological honesty about leakage and the switch to a training-only benchmark are also strengths. If the larger function class delivered a robust out-of-sample or crisis-regime advantage on a meaningful risk metric, the contribution would be of clear interest to empirical asset pricing and risk management. As written, the theory is solid but the empirical support for the crisis-motivation is thin, so the paper’s significance currently rests more on the formal inclusion of PCA than on demonstrated economic improvement.
major comments (3)
- Abstract vs body inconsistency on the headline numbers. The abstract still reports full-sample reconstruction R² of 66.57% (KAN-PCA) vs 62.99% (PCA) and claims that KAN-PCA “matches PCA out-of-sample after correcting for data leakage.” The body (Table 1, §5.3) uses the corrected train-only protocol and reports train 65.56% vs 64.71% and OOS 56.73% vs 57.32% (PCA ahead). The abstract must be rewritten to match the leakage-corrected benchmark that the paper itself treats as the valid comparison; otherwise the central empirical claim is misstated at the point of first contact.
- The evaluation design does not test the paper’s stated motivation. Introduction and Abstract argue that PCA “misrepresents the dependence structure during market stress” when correlations change sharply. In §5.1 the COVID crash lies inside the training window (2015–2020), and the test set (2022–2024) is not a comparable crisis regime. Table 1 reports only aggregate reconstruction R²; there is no crisis-window residual-covariance comparison, no regime-split R², and no economic risk metric (e.g., factor-hedged portfolio volatility, VaR, or correlation-structure error). Figs. 1–2 show a March-2020 spike and curved edges, but these are not controlled tests of the claimed failure mode. Without an evaluation that isolates stress periods or residual dependence, the empirical half does not support the motivation that justifies the larger function class.
- §5.3–5.4: the reported OOS gap is small and favors PCA (0.59 pp), while the in-sample gain is also small (~0.85 pp). With free parameters including architecture [20–10–3], progressive grid 3→5→10, spline/entropy penalties, and a 20-name universe, it is unclear whether the in-sample edge is genuine nonlinear structure or residual flexibility. The paper should either (i) provide statistical uncertainty (bootstrap or block-bootstrap CIs on the R² differences), (ii) report results for multiple k and larger universes as promised in the Conclusion, or (iii) reframe the contribution as primarily theoretical with a modest empirical illustration, rather than as evidence that KAN-PCA improves factor decomposition under the conditions that motivate the method.
minor comments (5)
- Notation in §3 is hard to parse in places (e.g., the covariance and spectral displays with many subscripts and missing symbols in the arXiv text). A cleaned notation table for Σ, U_k, φ_{ij}, and the reconstruction loss would help.
- Lemma 1 sets the intercept b=0 for zero-mean inputs; a one-sentence remark that this is without loss of generality for the recovered subspace (already implied) would make the reduction fully self-contained.
- Figure 1 caption claims contrast with “the linear response of classical PCA” but does not plot PCA factors alongside KAN factors; adding that panel would make the qualitative claim checkable.
- Related Work could briefly note how KAN-PCA differs from kernel PCA in that the nonlinearity is learned per edge rather than fixed by a kernel; the distinction is present but could be sharper for readers coming from the KPCA literature.
- Several OCR/encoding artifacts remain in the manuscript text (e.g., “� � = 65�56%”, “� ��”). These should be cleaned before resubmission.
Circularity Check
No circularity: the PCA-recovery theorem is a legitimate special-case reduction to an external classical result, and the R^{2} numbers are direct reconstruction scores.
full rationale
The load-bearing theoretical claim (Theorem 4 / Corollary 5) is that constraining all spline activations so that their second derivatives vanish reduces the KAN encoder to a linear map, after which the reconstruction objective becomes the classical linear autoencoder problem whose optimum is the top-k PCA subspace (Lemmas 1–3, citing Baldi–Hornik 1989 and Ky Fan). That is a proper special-case inclusion, not a self-definitional loop: the target (PCA) is an external, independently established optimum, not something defined from the KAN-PCA construction. The empirical half reports train and OOS reconstruction explained variance for free-spline KAN-PCA versus train-fit PCA on the same 20-name panel (Table 1); these are direct reconstruction scores under a fixed protocol, not fitted parameters renamed as predictions. There is no self-citation chain, no uniqueness theorem imported from the same authors, and no ansatz smuggled in via prior work by the author. Architecture and grid schedule are chosen on validation data, which is ordinary model selection rather than circular derivation. Score 0 with empty steps is therefore the correct finding.
Axiom & Free-Parameter Ledger
free parameters (5)
- number of factors k
- encoder width / depth [20-10-3]
- B-spline grid schedule (3→5→10) and degree
- spline and entropy regularization strengths
- train/val/test split fractions and stock universe (20 names)
axioms (6)
- standard math Empirical covariance of standardized returns admits a spectral decomposition with orthonormal eigenvectors ordered by eigenvalues (Classical PCA section).
- standard math A linear autoencoder with squared reconstruction loss has global optimum at the top-k PCA subspace (Lemma 3, citing Baldi & Hornik 1989).
- standard math Univariate functions with vanishing second derivative are affine; for zero-mean inputs intercepts can be dropped without changing the recovered subspace (Lemma 1).
- domain assumption Market crises induce dependence structures that fixed linear factors estimated over full samples misrepresent (Introduction).
- domain assumption Reconstruction R² / explained variance on standardized log-returns is the appropriate success metric for nonlinear factor quality (Experiments).
- ad hoc to paper KAN edge activations of the form w_b SiLU + w_s spline are an adequate function class for asset-to-factor maps (Definition 2–3).
invented entities (1)
-
KAN-PCA (KAN encoder + linear decoder autoencoder for return factors)
no independent evidence
read the original abstract
KAN-PCA is an autoencoder that uses a KAN as encoder and a linear map as decoder. It generalizes classical PCA by replacing linear projections with learned B-spline functions on each edge. The motivation is to capture more variance than classical PCA, which becomes inefficient during market crises when the linear assumption breaks down and correlations between assets change dramatically. We prove that if the spline activations are forced to be linear, KAN-PCA yields exactly the same results as classical PCA, establishing PCA as a special case. Experiments on 20 S&P 500 stocks (2015-2024) show that KAN-PCA achieves a reconstruction R^2 of 66.57%, compared to 62.99% for classical PCA with the same 3 factors, while matching PCA out-of-sample after correcting for data leakage in the training procedure.
Reference graph
Works this paper leans on
-
[1]
and Hornik, K
Baldi, P. and Hornik, K. (1989). Neural networks and principal component analysis: Learning from examples without local minima.Neural Networks, 2(1):53–58
1989
-
[2]
(2001).A Practical Guide to Splines(Revised Edition)
de Boor, C. (2001).A Practical Guide to Splines(Revised Edition). Applied Mathe- matical Sciences, vol. 27. Springer-Verlag, New York
2001
-
[3]
Elfwing, S., Uchibe, E., and Doya, K. (2018). Sigmoid-weighted linear units for neural network function approximation in reinforcement learning.Neural Networks, 107:3–11
2018
-
[4]
Fama, E. F. and French, K. R. (1993). Common risk factors in the returns on stocks and bonds.Journal of Financial Economics, 33(1):3–56
1993
-
[5]
Fan, K. (1949). On a theorem of Weyl concerning eigenvalues of linear transforma- tions. I.Proceedings of the National Academy of Sciences USA, 35(11):652–655
1949
-
[6]
Gu, S., Kelly, B., and Xiu, D. (2021). Autoencoder asset pricing models.Journal of Econometrics, 222(1):429–450
2021
-
[7]
Hinton, G. E. and Salakhutdinov, R. R. (2006). Reducing the dimensionality of data with neural networks.Science, 313(5786):504–507
2006
-
[8]
Kolmogorov, A. N. (1957). On the representation of continuous functions of many 14 variables by superposition of continuous functions of one variable and addition.Doklady Akademii Nauk SSSR, 114(5):953–956. (In Russian.)
1957
-
[9]
Liu, Z., Wang, Y., Vaidya, S., Ruehle, F., Halverson, J., Soljaˇ ci´ c, M., Hou, T. Y., and Tegmark, M. (2024a). KAN: Kolmogorov-Arnold Networks.arXiv:2404.19756
-
[10]
Liu, Z., Ma, P., Wang, Y., Matusik, W., and Tegmark, M. (2024b). KAN 2.0: Kolmogorov-Arnold Networks Meet Science.arXiv:2408.10205
-
[11]
Moradi, M., Panahi, S., Bollt, E., and Lai, Y.-C. (2024). Kolmogorov-Arnold Network Autoencoders.arXiv:2410.02077
Pith/arXiv arXiv 2024
-
[12]
Sch¨ olkopf, B., Smola, A., and M¨ uller, K.-R. (1998). Nonlinear component analysis as a kernel eigenvalue problem.Neural Computation, 10(5):1299–1319
1998
-
[13]
Wang, T. and Singh, S. (2024). KAN based Autoencoders for Factor Models. arXiv:2408.02694
Pith/arXiv arXiv 2024
-
[14]
Yu, F., Hu, R., Lin, Y., Ma, Y., Huang, Z., and Li, W. (2024). KAE: Kolmogorov- Arnold Auto-Encoder for Representation Learning.arXiv:2501.00420. 15
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.