Pith. sign in

REVIEW 3 major objections 5 minor 14 references

KAN-PCA recovers classical PCA exactly when its edge functions are forced to be affine, and otherwise captures slightly more in-sample variance on stock returns.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 16:25 UTC pith:GL6E4L6A

load-bearing objection Sound special-case theorem that PCA is the affine limit of a KAN-encoder/linear-decoder autoencoder; honest leakage correction; tiny empirical edge that vanishes OOS and never isolates the crisis regime that motivates the work. the 3 major comments →

arxiv 2603.28257 v2 pith:GL6E4L6A submitted 2026-03-30 q-fin.ST cs.LG

Nonlinear Factor Decomposition via Kolmogorov-Arnold Networks: A Spectral Approach to Asset Return Analysis

classification q-fin.ST cs.LG
keywords KAN-PCAKolmogorov-Arnold Networksprincipal component analysisnonlinear factor modelsasset returnsB-splinesautoencodersspectral decomposition
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Classical PCA extracts a few common risk factors from stock returns by linear projections, but the paper argues that this linear picture can misrepresent dependence when markets go into stress and correlations jump. The authors replace those linear projections with a Kolmogorov-Arnold Network encoder whose every edge is a learned B-spline, while keeping a linear decoder so the factors stay interpretable. They prove that if the splines are constrained to be affine, the model collapses exactly to ordinary PCA, so PCA is a strict special case inside a larger function class. On twenty large-cap S&P 500 names, free splines raise training explained variance a little (65.56 percent versus 64.71 percent with three factors) and produce visibly nonlinear, often asymmetric edge curves; after a corrected train-only protocol the two methods are essentially tied out of sample. The work therefore supplies both a clean theoretical nesting of PCA inside a nonlinear model and direct visual evidence that some nonlinear factor structure is present in equity returns.

Core claim

When every spline activation in the KAN encoder is forced to have vanishing second derivative (i.e., to be affine), the optimal KAN-PCA encoder and linear decoder recover exactly the classical PCA projection onto the top-k eigenvectors of the sample covariance; with free splines the same architecture captures modestly more training reconstruction variance than train-fit PCA at equal factor count.

What carries the argument

Theorem 4 (KAN-PCA reduces to PCA in the linear limit): affine edge functions turn each KAN layer into ordinary matrix multiplication, so the reconstruction objective becomes the classical linear autoencoder whose unique optimum is the PCA subspace.

Load-bearing premise

That a modest gain in reconstruction R-squared on a fixed twenty-stock, three-factor sample (with the main crisis inside the training window) is enough to show that spline-mediated nonlinear factors improve risk decomposition relative to linear PCA.

What would settle it

Re-run the identical three-factor comparison on a larger universe (order of 200 stocks) with the main stress period held strictly out of sample; if free-spline KAN-PCA no longer exceeds train-fit PCA even in-sample, or falls materially behind out of sample, the claimed practical advantage of the larger function class disappears.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes KAN-PCA, an autoencoder with a Kolmogorov–Arnold Network encoder and a linear decoder for extracting nonlinear factors from asset returns. The central theoretical claim is that if all edge activations are constrained to be affine, the model recovers classical PCA exactly (Theorem 4 / Corollary 5), via the standard Baldi–Hornik linear-autoencoder optimum. Empirically, on 20 S&P 500 stocks with k=3 factors and a corrected 70/10/20 protocol, KAN-PCA slightly exceeds train-fit PCA in-sample (65.56% vs 64.71% explained variance) but is marginally behind out-of-sample (56.73% vs 57.32%). The authors document and correct a data-leakage issue from earlier training, and present edge-function and factor-path visualizations as qualitative evidence of nonlinearity.

Significance. The special-case recovery of PCA is a clean and useful characterization: it places KAN-PCA in a strictly larger function class with a known linear limit, which is stronger than most generic nonlinear autoencoder papers in asset pricing. The methodological honesty about leakage and the switch to a training-only benchmark are also strengths. If the larger function class delivered a robust out-of-sample or crisis-regime advantage on a meaningful risk metric, the contribution would be of clear interest to empirical asset pricing and risk management. As written, the theory is solid but the empirical support for the crisis-motivation is thin, so the paper’s significance currently rests more on the formal inclusion of PCA than on demonstrated economic improvement.

major comments (3)
  1. Abstract vs body inconsistency on the headline numbers. The abstract still reports full-sample reconstruction R² of 66.57% (KAN-PCA) vs 62.99% (PCA) and claims that KAN-PCA “matches PCA out-of-sample after correcting for data leakage.” The body (Table 1, §5.3) uses the corrected train-only protocol and reports train 65.56% vs 64.71% and OOS 56.73% vs 57.32% (PCA ahead). The abstract must be rewritten to match the leakage-corrected benchmark that the paper itself treats as the valid comparison; otherwise the central empirical claim is misstated at the point of first contact.
  2. The evaluation design does not test the paper’s stated motivation. Introduction and Abstract argue that PCA “misrepresents the dependence structure during market stress” when correlations change sharply. In §5.1 the COVID crash lies inside the training window (2015–2020), and the test set (2022–2024) is not a comparable crisis regime. Table 1 reports only aggregate reconstruction R²; there is no crisis-window residual-covariance comparison, no regime-split R², and no economic risk metric (e.g., factor-hedged portfolio volatility, VaR, or correlation-structure error). Figs. 1–2 show a March-2020 spike and curved edges, but these are not controlled tests of the claimed failure mode. Without an evaluation that isolates stress periods or residual dependence, the empirical half does not support the motivation that justifies the larger function class.
  3. §5.3–5.4: the reported OOS gap is small and favors PCA (0.59 pp), while the in-sample gain is also small (~0.85 pp). With free parameters including architecture [20–10–3], progressive grid 3→5→10, spline/entropy penalties, and a 20-name universe, it is unclear whether the in-sample edge is genuine nonlinear structure or residual flexibility. The paper should either (i) provide statistical uncertainty (bootstrap or block-bootstrap CIs on the R² differences), (ii) report results for multiple k and larger universes as promised in the Conclusion, or (iii) reframe the contribution as primarily theoretical with a modest empirical illustration, rather than as evidence that KAN-PCA improves factor decomposition under the conditions that motivate the method.
minor comments (5)
  1. Notation in §3 is hard to parse in places (e.g., the covariance and spectral displays with many subscripts and missing symbols in the arXiv text). A cleaned notation table for Σ, U_k, φ_{ij}, and the reconstruction loss would help.
  2. Lemma 1 sets the intercept b=0 for zero-mean inputs; a one-sentence remark that this is without loss of generality for the recovered subspace (already implied) would make the reduction fully self-contained.
  3. Figure 1 caption claims contrast with “the linear response of classical PCA” but does not plot PCA factors alongside KAN factors; adding that panel would make the qualitative claim checkable.
  4. Related Work could briefly note how KAN-PCA differs from kernel PCA in that the nonlinearity is learned per edge rather than fixed by a kernel; the distinction is present but could be sharper for readers coming from the KPCA literature.
  5. Several OCR/encoding artifacts remain in the manuscript text (e.g., “� � = 65�56%”, “� ��”). These should be cleaned before resubmission.

Circularity Check

0 steps flagged

No circularity: the PCA-recovery theorem is a legitimate special-case reduction to an external classical result, and the R^{2} numbers are direct reconstruction scores.

full rationale

The load-bearing theoretical claim (Theorem 4 / Corollary 5) is that constraining all spline activations so that their second derivatives vanish reduces the KAN encoder to a linear map, after which the reconstruction objective becomes the classical linear autoencoder problem whose optimum is the top-k PCA subspace (Lemmas 1–3, citing Baldi–Hornik 1989 and Ky Fan). That is a proper special-case inclusion, not a self-definitional loop: the target (PCA) is an external, independently established optimum, not something defined from the KAN-PCA construction. The empirical half reports train and OOS reconstruction explained variance for free-spline KAN-PCA versus train-fit PCA on the same 20-name panel (Table 1); these are direct reconstruction scores under a fixed protocol, not fitted parameters renamed as predictions. There is no self-citation chain, no uniqueness theorem imported from the same authors, and no ansatz smuggled in via prior work by the author. Architecture and grid schedule are chosen on validation data, which is ordinary model selection rather than circular derivation. Score 0 with empty steps is therefore the correct finding.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 1 invented entities

The theory rests on standard linear algebra and the classical linear-autoencoder–PCA equivalence; the empirical claim rests on modeling choices (k, architecture, grids, penalties) and the assumption that reconstruction R² on 20 standardized names is the right yardstick. No new physical entities are postulated—only a named model class.

free parameters (5)
  • number of factors k
    Fixed at 3 for all main comparisons; drives both PCA and KAN-PCA capacity and the reported R² gaps.
  • encoder width / depth [20-10-3]
    Chosen after iteration; intermediate width 10 is not derived and affects expressivity vs PCA’s single linear map.
  • B-spline grid schedule (3→5→10) and degree
    Progressive grid extension and spline degree control nonlinearity; selected via training procedure rather than theory.
  • spline and entropy regularization strengths
    λ decreasing from 1e-2 to 1e-4 and entropy penalty 0.1 are hand-set; they directly trade off fit vs smoothness and thus in-sample R².
  • train/val/test split fractions and stock universe (20 names)
    70/10/20 and the specific ticker list are design choices that define the empirical claim’s scope.
axioms (6)
  • standard math Empirical covariance of standardized returns admits a spectral decomposition with orthonormal eigenvectors ordered by eigenvalues (Classical PCA section).
    Spectral theorem for symmetric PSD matrices; used to define PCA baseline and Lemma 3.
  • standard math A linear autoencoder with squared reconstruction loss has global optimum at the top-k PCA subspace (Lemma 3, citing Baldi & Hornik 1989).
    Load-bearing external theorem that turns the affine-KAN reduction into ‘equals PCA’.
  • standard math Univariate functions with vanishing second derivative are affine; for zero-mean inputs intercepts can be dropped without changing the recovered subspace (Lemma 1).
    Elementary calculus/linear algebra step linking constrained splines to matrix multiplies.
  • domain assumption Market crises induce dependence structures that fixed linear factors estimated over full samples misrepresent (Introduction).
    Motivates KAN-PCA; not proved in the paper and only weakly probed by the OOS window.
  • domain assumption Reconstruction R² / explained variance on standardized log-returns is the appropriate success metric for nonlinear factor quality (Experiments).
    No pricing, Sharpe, or risk-forecast metric is used; the central empirical comparison stands or falls on this choice.
  • ad hoc to paper KAN edge activations of the form w_b SiLU + w_s spline are an adequate function class for asset-to-factor maps (Definition 2–3).
    Architecture choice inherited from Liu et al. KAN; not derived from asset-pricing theory.
invented entities (1)
  • KAN-PCA (KAN encoder + linear decoder autoencoder for return factors) no independent evidence
    purpose: Name and define the model class that strictly contains PCA and allows inspectable nonlinear edge loadings.
    Methodological construct rather than a new physical object; independent evidence is only the paper’s own fits and plots, not an external measurement.

pith-pipeline@v1.1.0-grok45 · 12574 in / 3722 out tokens · 38571 ms · 2026-07-13T16:25:22.305321+00:00 · methodology

0 comments
read the original abstract

KAN-PCA is an autoencoder that uses a KAN as encoder and a linear map as decoder. It generalizes classical PCA by replacing linear projections with learned B-spline functions on each edge. The motivation is to capture more variance than classical PCA, which becomes inefficient during market crises when the linear assumption breaks down and correlations between assets change dramatically. We prove that if the spline activations are forced to be linear, KAN-PCA yields exactly the same results as classical PCA, establishing PCA as a special case. Experiments on 20 S&P 500 stocks (2015-2024) show that KAN-PCA achieves a reconstruction R^2 of 66.57%, compared to 62.99% for classical PCA with the same 3 factors, while matching PCA out-of-sample after correcting for data leakage in the training procedure.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

14 extracted references · 5 linked inside Pith

  1. [1]

    and Hornik, K

    Baldi, P. and Hornik, K. (1989). Neural networks and principal component analysis: Learning from examples without local minima.Neural Networks, 2(1):53–58

  2. [2]

    (2001).A Practical Guide to Splines(Revised Edition)

    de Boor, C. (2001).A Practical Guide to Splines(Revised Edition). Applied Mathe- matical Sciences, vol. 27. Springer-Verlag, New York

  3. [3]

    Elfwing, S., Uchibe, E., and Doya, K. (2018). Sigmoid-weighted linear units for neural network function approximation in reinforcement learning.Neural Networks, 107:3–11

  4. [4]

    Fama, E. F. and French, K. R. (1993). Common risk factors in the returns on stocks and bonds.Journal of Financial Economics, 33(1):3–56

  5. [5]

    Fan, K. (1949). On a theorem of Weyl concerning eigenvalues of linear transforma- tions. I.Proceedings of the National Academy of Sciences USA, 35(11):652–655

  6. [6]

    Gu, S., Kelly, B., and Xiu, D. (2021). Autoencoder asset pricing models.Journal of Econometrics, 222(1):429–450

  7. [7]

    Hinton, G. E. and Salakhutdinov, R. R. (2006). Reducing the dimensionality of data with neural networks.Science, 313(5786):504–507

  8. [8]

    Kolmogorov, A. N. (1957). On the representation of continuous functions of many 14 variables by superposition of continuous functions of one variable and addition.Doklady Akademii Nauk SSSR, 114(5):953–956. (In Russian.)

  9. [9]

    Y., and Tegmark, M

    Liu, Z., Wang, Y., Vaidya, S., Ruehle, F., Halverson, J., Soljaˇ ci´ c, M., Hou, T. Y., and Tegmark, M. (2024a). KAN: Kolmogorov-Arnold Networks.arXiv:2404.19756

  10. [10]

    Liu, Z., Ma, P., Wang, Y., Matusik, W., and Tegmark, M. (2024b). KAN 2.0: Kolmogorov-Arnold Networks Meet Science.arXiv:2408.10205

  11. [11]

    Moradi, M., Panahi, S., Bollt, E., and Lai, Y.-C. (2024). Kolmogorov-Arnold Network Autoencoders.arXiv:2410.02077

  12. [12]

    Sch¨ olkopf, B., Smola, A., and M¨ uller, K.-R. (1998). Nonlinear component analysis as a kernel eigenvalue problem.Neural Computation, 10(5):1299–1319

  13. [13]

    and Singh, S

    Wang, T. and Singh, S. (2024). KAN based Autoencoders for Factor Models. arXiv:2408.02694

  14. [14]

    Yu, F., Hu, R., Lin, Y., Ma, Y., Huang, Z., and Li, W. (2024). KAE: Kolmogorov- Arnold Auto-Encoder for Representation Learning.arXiv:2501.00420. 15