Pith. sign in

REVIEW 4 minor 50 references

Foundations of Independent Component Analysis

T0 review · 0 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper proves that linear ICA is identifiable up to permutation, scale, translation and sign whenever the sources are mutually independent and Gaussian-free, even under additive Gaussian noise with arbitrary covariance.

desk verdict A careful, self-contained re-proof of the classical ICA identifiability theory with a clean Gaussian-free formulation; not much new, but the rigor and clarity earn it a serious referee. read the letter →

arxiv 2608.13229 v1 pith:5YWPMUHR submitted 2026-08-13 math.ST cs.LGmath.PRstat.MLstat.TH

classification math.STcs.LGmath.PRstat.MLstat.TH MSC 62H2560E10
keywords independentcomponentanalysisidentifiabilityGaussian-freesourcescharacteristicfunctionscumulantgeneratingKagan–Linnik–RaotheoremequivariantgradientdescentLiNGAM
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's central claim is that linear independent component analysis is identifiable in the strongest sense—sources recoverable up to translation, permutation, scale, and sign—provided the sources are mutually independent and Gaussian-free, meaning that no non-degenerate Gaussian distribution can be split off from any source as a convolution factor, and then even when the observed mixture carries additive Gaussian noise with arbitrary, possibly degenerate covariance. The authors prove this by building the full theory of characteristic functions from scratch, then proving a two-sided identifiability theorem (Theorem 6.12) whose engine is a decomposition result: every real-valued random variable splits essentially uniquely into a Gaussian-free part plus independent Gaussian noise, and the Gaussian-free parts are what the mixture pins down. Along the way they prove the classical Kagan–Linnik–Rao dichotomy with a self-contained finite-difference argument, and show that merely non-Gaussian sources leave a residual additive Gaussian ambiguity, while Gaussian sources are not identifiable at all. The paper also derives the online equivariant gradient descent algorithm and shows that its separating fixed point is locally stable exactly when the identifiability theory declares the model identifiable. A sympathetic reader should care because these results delineate precisely when blind source separation is solvable in principle and what ambiguity must remain.

What carries the argument

The argument is carried by characteristic functions and their distinguished logarithms (cumulant generating functions), combined with three classical rigidity results: Marcinkiewicz' theorem, which says the exponential of a polynomial is a characteristic function only in the Gaussian case; Cramér's decomposition theorem, which says a Gaussian sum can only have Gaussian independent summands; and the Kagan–Linnik–Rao theorem (Theorem 4.2), which this paper proves by a finite-difference argument over ridge functions, such that any column of one mixing matrix that is not proportional to a column of the other forces its source to be Gaussian. Around that core the paper wraps Lemma 5.2, which trades Gaussian noise vectors for extra columns of the mixing matrix and back, and Theorem 6.7, which splits every source into a Gaussian-free part plus independent Gaussian noise. The estimation half then shows that the natural-gradient (relative gradient) update for the mixing matrix induces a dynamics on the global system matrix whose stability condition—in terms of the quantity ζj = −βjσj²—matches the identifiability condition exactly when the model score equals the true score.

What would settle it

Try to find two full-rank representations of the same observed law, A¹Z¹+E¹ = A²Z²+E², where every source in both representations is Gaussian-free in the sense that no non-degenerate Gaussian convolution factor can be split off, and check whether the sources are related only by permutation, scale, and translation. Exhibiting such a pair without this relation—or computing for a concrete law that the set of splittable Gaussian scales is not a closed interval, contradicting Lemma 6.6(iii)—would refute the paper's central claim.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is Theorem 6.12: under mutual independence, Gaussian-free sources, and full column rank of the mixing matrix, two representations of the same observed law must be related by a permutation, a diagonal scaling, and a translation, with both the source laws and the noise covariance then forced to agree. The key structural insight is Theorem 6.7, which shows that every real-valued random variable decomposes, essentially uniquely, as a Gaussian-free random variable plus independent Gaussian noise, and that the maximal Gaussian scale is always attained. This makes 'Gaussian-free' the exact hypothesis that removes the additive Gaussian ambiguity that remains when sources are merely non-Gaussian (Theorem 5.5), and it explains why Gaussian sources are hopeless: they can be rotated by any orthogonal matrix without changing the observed law. In the complete noiseless case, the theory specialises to the classical statement that a square invertible mixture is identifiable if and only if at most one source is Gaussian (Corollary 7.4), and the same machinery proves LiNGAM's causal order is identified.

Load-bearing premise

The proof of the strongest theorem (Theorem 6.12) collapses if a source can be written as a Gaussian plus something independent—i.e., if it fails to be Gaussian-free—because then the residual Gaussian ambiguity of the merely non-Gaussian case survives.

Editorial extensions

If this is right

  • If both candidate source vectors are Gaussian-free, the whole model—mixing matrix, source laws, and noise covariance—is identified up to permutation, scale, and shift, even when the additive Gaussian noise is degenerate and arbitrarily correlated across coordinates.
  • Merely non-Gaussian sources are not enough: the sources are then determined only up to an additive componentwise Gaussian noise, so the practical lesson is that blindly applying ICA to non-Gaussian but Gaussian-contaminated sources overstates what can be recovered.
  • In the complete noiseless square case, identifiability holds if and only if at most one source is Gaussian, and the recovered sources are determined up to permutation and sign once centred and scaled.
  • The online equivariant gradient descent algorithm attains the same boundary: with correctly specified source scores, the separating solution is locally asymptotically stable exactly when the model is identifiable, so identifiability and algorithm stability coincide.
  • LiNGAM removes the permutation ambiguity for acyclic causal models, so the causal order and coefficients are fully identified from the law of the observations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The Gaussian-free hypothesis is strictly stronger than non-Gaussianity—mixtures such as ½N(−1,1)+½N(1,1) are non-Gaussian but not Gaussian-free—so methods that check σmax(Z)=0 (e.g., via characteristic-function decay or support) could be used to certify in advance whether the stronger identifiability conclusion applies to a given dataset.
  • The splitting theorem suggests a natural quantitative measure of residual ambiguity: the maximal Gaussian scale σmax(Z) acts as a 'Gaussian content' of a source, and one could design partial identifiability statements for misspecified models that bound how much of the source still can be attributed to noise.
  • The stability quantity ζj for a misspecified score is a computable functional of the source law; one could adaptively choose the nonlinearity per source, tuning it so that ζj>1 holds, which would make the algorithm's convergence guarantee data-dependent rather than assumed.
  • The finite-difference proof of the column dichotomy is self-contained and may carry over to identifiability questions beyond linear ICA, such as nonlinear or time-varying mixtures, where the same 'every unwanted direction is annihilated by a difference operator' strategy could be replayed.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The manuscript is a self-contained mathematical treatment of linear independent component analysis (ICA). It develops the theory of characteristic functions of probability measures on R^d, including analyticity, cumulants, and the Gaussian characterisation theorems, and then proves a sequence of identifiability results for the linear ICA model under successively stronger source assumptions: non-constant, non-Gaussian, and Gaussian-free. The central result is Theorem 6.12, which identifies Gaussian-free independent sources up to permutation, scale, and translation even in the presence of additive Gaussian noise with an arbitrary covariance matrix. The second half of the paper studies the complete noiseless square ICA model, derives the relative-gradient (equivariant) online algorithm, proves local stability of separating solutions under explicit conditions on model scores, and derives LiNGAM identifiability as a corollary. Full proofs are provided in two appendices, with the Kagan–Linnik–Rao theorem proved by a self-contained finite-difference argument.

Significance. If the results are correct, this is a valuable rigorous reference for the mathematical foundations of ICA. The paper's main strengths are its completeness: the identifiability statements are proved from first principles, the source assumptions and remaining ambiguities are stated precisely, and the proof of the central Theorem 6.12 is internally coherent, with the Gaussian-free hypothesis used exactly where it is needed. The treatment of splittable Gaussian scales and the maximal Gaussian decomposition (Theorem 6.7) is a useful clarification of a notion that is often left informal. The paper also gives a careful account of the relationship between the relative gradient and the natural metric, and it states explicitly which results are quoted rather than proved. The contribution is more expository and foundational than revolutionary, but it meets a genuine need for a precise, self-contained presentation of the classical identifiability theory and its modern refinements.

minor comments (4)
  1. [Section 4, Proposition 4.4] The proof asserts that for a continuous map T one has supp L(T(Z)) = T(supp L(Z)) and refers to T(supp L(Z)) as a closed set. This is not true in general: a linear map need not be a closed map, since a continuous image of a closed set need not be closed (for instance, a projection of the closed hyperbola xy=1 has non-closed image). The conclusion of the proposition is nevertheless correct; the proof should argue directly that aff(supp L(X)) = T(aff(supp L(Z))) using the fact that affine hulls commute with affine maps and are unchanged by taking closures.
  2. [Section 6.2, Theorem 6.12 proof] The sentence 'the G_j were constructed componentwise, so we may and do take G to have independent components' deserves an explicit justification. The componentwise identities in Eq. (136) fix only the marginal laws of the components of Z(2); to pass to the vector identity Eq. (143), one should state that the G_j can be recoupled independently on an enlarged probability space, with Z(2) then defined by Eq. (143), and that the resulting vector has independent components with the required marginals. As written, the step is correct but requires the reader to reconstruct the coupling argument.
  3. [Section 7.4, after Corollary 7.23] The claimed stable non-separating equilibrium for k=2 with the specific value R* = 0.80993 is stated without derivation. Likewise, the entries in Table 2 are said to be obtained by numerical quadrature but no code or computational details are supplied. A short explanation of how these quantities were computed, or a reference to reproducible code, would strengthen the presentation.
  4. [Section 7.4, Remark 7.27] The statement that the online stochastic approximation algorithm converges almost surely to locally stable equilibria of the ODE under 'the usual regularity and boundedness conditions' is informal and is not proved. Since Theorem 7.20 establishes local stability only for the mean dynamics Eq. (202), the remark should either state precise hypotheses sufficient for the stochastic approximation result or explicitly label that convergence statement as heuristic.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the ICA identifiability results are proved from characteristic-function first principles, and self-citations appear only in a non-load-bearing survey.

full rationale

The paper's central derivation chain is acyclic. The engine theorem, Theorem 4.2, is stated with attribution to Kagan–Linnik–Rao but is proved in full in Appendix B via a self-contained finite-difference argument using only the distinguished logarithm, Fréchet's functional equation, and Marcinkiewicz' theorem, all of which are also proved in the paper or quoted as standard classical results independent of ICA. The later identifiability theorems (Theorem 5.5, Theorem 6.12, Corollary 7.4) are derived from Theorem 4.2 with explicit bookkeeping of Gaussian noise and Gaussian-free hypotheses; they do not assume their own conclusions. The Gaussian-free definition is an explicit modelling assumption, not a restatement of identifiability, and the paper acknowledges that it is strictly stronger than non-Gaussianity and that it fails for mixtures of Gaussians. The estimation section derives the relative-gradient algorithm as an exact gradient for a right-invariant metric and analyzes its stability with standard dynamical-systems tools; no fitted parameter is renamed as a prediction. The only self-citations (Pandeva and Forré 2023a,b; Pandeva et al. 2025) occur in Section 7.7's survey of generalizations and are not load-bearing for any theorem in the paper. The only unproved external inputs are classical results such as Bochner's theorem and the unstable-manifold theorem, which are independent of ICA and are explicitly identified as such. Thus the paper does not exhibit self-definitional, fitted-input, self-citation-load-bearing, or ansatz-smuggling circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new entities, particles, or forces. All assumptions are explicitly stated as definitions or theorem hypotheses, and the proof machinery relies on standard mathematical theorems. The only potentially debatable modeling choice is the Gaussian-free condition on sources, which is a hypothesis of the identifiability theorems rather than a mathematical axiom.

assumptions (3)
  • standard math Bochner's theorem: a continuous positive definite function with value 1 at 0 is a characteristic function.
    Quoted in Theorem 3.3 and used in Remark 3.8 and Theorem 6.7 to certify that certain functions are characteristic functions. It is the only classical result used without proof.
  • standard math Unstable-manifold theorem for maps.
    Used in the proof of Theorem 7.20 to upgrade linear instability to nonlinear instability of the fixed point.
  • standard math Standard measure-theoretic background: dominated convergence, Fubini, Prokhorov's theorem, complex analysis up to Morera and identity theorem.
    Listed in Appendix A as background used without proof. These are universally accepted tools in probability theory.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Foundations of Independent Component Analysis." pith.science (2026). https://pith.science/paper/5YWPMUHR

@misc{pith2026260813229,
  author       = {Pith},
  title        = {Pith review of: Foundations of Independent Component Analysis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5YWPMUHR}},
  note         = {Machine review of arXiv:2608.13229}
}
abstract

We present the mathematical foundations of linear independent component analysis (ICA) models based on standard literature in a self-contained note. It is aimed at readers with a background in measure-theoretic probability theory. We first develop the theory of the characteristic functions of probability measures on $\mathbb{R}^d$, including their analyticity and the way in which they determine and characterise the distributions. We then focus on several identifiability results of ICA models with successively strengthened assumptions on the sources: from merely non-constant, to non-Gaussian, to Gaussian-free independent sources. Under the strictest assumptions, we show that the independent sources are identifiable up to translation, permutation, scales and signs, and this even in the presence of additive Gaussian noise. Furthermore, we present the online equivariant gradient descent ICA algorithm for recovering the independent sources from data, in the standard complete noiseless non-Gaussian ICA setting.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

50 extracted references · 31 canonical work pages

  1. [1]

    Amari, A

    S. Amari, A. Cichocki, and H. H. Yang. A new learning algorithm for blind signal separation. In D. S. Touretzky, M. C. Mozer, and M. E. Hasselmo, editors, Advances in Neural Information Processing Systems 8 (NIPS 1995), pages 757--763. MIT Press, Cambridge, MA, 1996. https://papers.nips.cc/paper/1115-a-new-learning-algorithm-for-blind-signal-separation

  2. [2]

    Amari, T.-P

    S. Amari, T.-P. Chen, and A. Cichocki. Stability analysis of learning algorithms for blind source separation. Neural Networks, 10(8):1345--1351, 1997. doi:10.1016/S0893-6080(97)00039-7

  3. [3]

    Amari and J.-F

    S. Amari and J.-F. Cardoso. Blind source separation---semiparametric statistical approach. IEEE Transactions on Signal Processing, 45(11):2692--2700, 1997. doi:10.1109/78.650095

  4. [4]

    S. Amari. Natural gradient works efficiently in learning. Neural Computation, 10(2):251--276, 1998. doi:10.1162/089976698300017746

  5. [5]

    K. P. Balanda and H. L. MacGillivray. Kurtosis: a critical review. The American Statistician, 42(2):111--119, 1988. doi:10.1080/00031305.1988.10475539

  6. [6]

    A. J. Bell and T. J. Sejnowski. An information-maximization approach to blind separation and blind deconvolution. Neural Computation, 7(6):1129--1159, 1995. doi:10.1162/neco.1995.7.6.1129

  7. [7]

    Cardoso and B

    J.-F. Cardoso and B. H. Laheld. Equivariant adaptive source separation. IEEE Transactions on Signal Processing, 44(12):3017--3030, 1996. doi:10.1109/78.553476

  8. [8]

    J.-F. Cardoso. Infomax and maximum likelihood for blind source separation. IEEE Signal Processing Letters, 4(4):112--114, 1997. doi:10.1109/97.566704

Show all 50 references
  1. [9]

    J.-F. Cardoso. Multidimensional independent component analysis. In Proceedings of the 1998 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP '98), volume 4, pages 1941--1944, Seattle, WA, 1998a. IEEE. doi:10.1109/ICASSP.1998.681443

  2. [10]

    J.-F. Cardoso. Blind signal separation: statistical principles. Proceedings of the IEEE, 86(10):2009--2025, 1998b. doi:10.1109/5.720250

  3. [11]

    P. Comon. Independent component analysis, a new concept? Signal Processing, 36(3):287--314, 1994. doi:10.1016/0165-1684(94)90029-9

  4. [12]

    Comon and C

    P. Comon and C. Jutten, editors. Handbook of Blind Source Separation: Independent Component Analysis and Applications. Academic Press (Elsevier), Oxford, 1st edition, 2010. ISBN 978-0-12-374726-6 (print), 978-0-08-088494-3 (e-book)

  5. [13]

    Cram\'er

    H. Cram\'er. \"Uber eine Eigenschaft der normalen Verteilungsfunktion. Mathematische Zeitschrift, 41(1):405--414, 1936. doi:10.1007/BF01180430

  6. [14]

    R. B. Darlington. Is kurtosis really ``peakedness?'' The American Statistician, 24(2):19--22, 1970. doi:10.1080/00031305.1970.10478885

  7. [15]

    G. Darmois. Analyse g\'en\'erale des liaisons stochastiques: \'etude particuli\`ere de l'analyse factorielle lin\'eaire. Revue de l'Institut International de Statistique / Review of the International Statistical Institute, 21(1/2):2--8, 1953. doi:10.2307/1401511

  8. [16]

    Eriksson and V

    J. Eriksson and V. Koivunen. Identifiability, separability, and uniqueness of linear ICA models. IEEE Signal Processing Letters, 11(7):601--604, 2004. doi:10.1109/LSP.2004.830118

  9. [17]

    Eriksson and V

    J. Eriksson and V. Koivunen. Complex random vectors and ICA models: identifiability, uniqueness, and separability. IEEE Transactions on Information Theory, 52(3):1017--1029, 2006. doi:10.1109/TIT.2005.864440. Preprint: https://arxiv.org/abs/cs/0512063

  10. [18]

    W. Feller. An Introduction to Probability Theory and Its Applications, volume II. Wiley Series in Probability and Mathematical Statistics. John Wiley & Sons, New York, 2nd edition, 1971. 669 pp. ISBN 978-0-471-25709-7

  11. [19]

    Fr\'echet

    M. Fr\'echet. Une d\'efinition fonctionnelle des polynomes. Nouvelles annales de math\'ematiques, 4e s\'erie, 9:145--162, 1909. https://www.numdam.org/item/NAM_1909_4_9__145_0/

  12. [20]

    R. Ger. On extensions of polynomial functions. Results in Mathematics, 26(3--4):281--289, 1994. doi:10.1007/BF03323050

  13. [21]

    S. G. Ghurye and I. Olkin. A characterization of the multivariate normal distribution. The Annals of Mathematical Statistics, 33(2):533--541, 1962. doi:10.1214/aoms/1177704579

  14. [22]

    ugelgen, V. Stimper, B. Sch\

    L. Gresele, J. von K\"ugelgen, V. Stimper, B. Sch\"olkopf, and M. Besserve. Independent mechanism analysis, a new concept? In Advances in Neural Information Processing Systems 34 (NeurIPS 2021), pages 28233--28248. Curran Associates, 2021. https://proceedings.neurips.cc/paper/...

  15. [23]

    P. O. Hoyer, D. Janzing, J. M. Mooij, J. Peters, and B. Sch\"olkopf. Nonlinear causal discovery with additive noise models. In Advances in Neural Information Processing Systems 21 (NIPS 2008), pages 689--696, 2008. https://proceedings.neurips.cc/paper/2008/hash/f7664060cc52bc6...

  16. [24]

    Hyv\"arinen and H

    A. Hyv\"arinen and H. Morioka. Unsupervised feature extraction by time-contrastive learning and nonlinear ICA. In Advances in Neural Information Processing Systems 29 (NIPS 2016), pages 3765--3773. Curran Associates, 2016. https://proceedings.neurips.cc/paper/2016/hash/d305281...

  17. [25]

    Hyv\"arinen and H

    A. Hyv\"arinen and H. Morioka. Nonlinear ICA of temporally dependent stationary sources. In A. Singh and J. Zhu, editors, Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS), volume 54 of Proceedings of Machine Learning Research...

  18. [26]

    Hyv\"arinen and E

    A. Hyv\"arinen and E. Oja. Independent component analysis: algorithms and applications. Neural Networks, 13(4--5):411--430, 2000. doi:10.1016/S0893-6080(00)00026-5

  19. [27]

    Hyv\"arinen and P

    A. Hyv\"arinen and P. Pajunen. Nonlinear independent component analysis: existence and uniqueness results. Neural Networks, 12(3):429--439, 1999. doi:10.1016/S0893-6080(98)00140-3

  20. [28]

    Hyv\"arinen, J

    A. Hyv\"arinen, J. Karhunen, and E. Oja. Independent Component Analysis. Wiley Series on Adaptive and Learning Systems for Signal Processing, Communications, and Control. John Wiley & Sons, New York, 2001. 504 pp. ISBN 978-0-471-40540-5. doi:10.1002/0471221317

  21. [29]

    A. M. Kagan, Yu. V. Linnik, and C. R. Rao. Characterization Problems in Mathematical Statistics. Wiley Series in Probability and Mathematical Statistics. John Wiley & Sons, New York, 1973. xii+499 pp. ISBN 978-0-471-45421-2. Translated from the Russian by B. Ramachandran; Russ...

  22. [30]

    Kallenberg

    O. Kallenberg. Foundations of Modern Probability, volume 99 of Probability Theory and Stochastic Modelling. Springer, Cham, 3rd edition, 2021. xii+946 pp. ISBN 978-3-030-61870-4. doi:10.1007/978-3-030-61871-1

  23. [31]

    Kaplansky

    I. Kaplansky. A common error concerning kurtosis. Journal of the American Statistical Association, 40(230):259, 1945. doi:10.1080/01621459.1945.10501856

  24. [32]

    Khemakhem, D

    I. Khemakhem, D. P. Kingma, R. P. Monti, and A. Hyv\"arinen. Variational autoencoders and nonlinear ICA: a unifying framework. In S. Chiappa and R. Calandra, editors, Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics (AISTATS), volume 1...

  25. [33]

    T. Kim, T. Eltoft, and T.-W. Lee. Independent vector analysis: an extension of ICA to multivariate components. In Independent Component Analysis and Blind Signal Separation (ICA 2006), volume 3889 of Lecture Notes in Computer Science, pages 165--172. Springer, Berlin, Heidelbe...

  26. [34]

    A. Klenke. Probability Theory: A Comprehensive Course. Universitext. Springer, Cham, 3rd edition, 2020. ISBN 978-3-030-56401-8. doi:10.1007/978-3-030-56402-5. L\'evy's continuity theorem is Section 15.3

  27. [35]

    T.-W. Lee, M. Girolami, and T. J. Sejnowski. Independent component analysis using an extended infomax algorithm for mixed subgaussian and supergaussian sources. Neural Computation, 11(2):417--441, 1999. doi:10.1162/089976699300016719

  28. [36]

    Yu. V. Linnik and I. V. Ostrovskii. Decomposition of Random Variables and Vectors, volume 48 of Translations of Mathematical Monographs. American Mathematical Society, Providence, RI, 1977. ix+380 pp. ISBN 978-0-8218-1598-4. Translated from the Russian; translation edited by J...

  29. [37]

    E. Lukacs. Characteristic Functions. Charles Griffin & Company Limited, London, 2nd, revised and enlarged edition, 1970. x+350 pp. ISBN 0-85264-170-2. LCCN 70-513840

  30. [38]

    D. J. C. MacKay. Maximum likelihood and covariant algorithms for independent component analysis. Unpublished report, Cavendish Laboratory, University of Cambridge, 1996. https://www.inference.org.uk/mackay/ica.pdf . Version 3.8, 8 January 1999, with minor corrections of 8 October 2002

  31. [39]

    D. J. C. MacKay. Information Theory, Inference, and Learning Algorithms. Cambridge University Press, Cambridge, 2003. xii+628 pp. ISBN 978-0-521-64298-9. Chapter 34, ``Independent Component Analysis and Latent Variable Modelling'', pp. 437--444. https://www.inference.org.uk/ma...

  32. [40]

    Marcinkiewicz

    J. Marcinkiewicz. Sur une propri\'et\'e de la loi de Gauss. Mathematische Zeitschrift, 44(1):612--618, 1939. doi:10.1007/BF01210677

  33. [41]

    J. J. A. Moors. The meaning of kurtosis: Darlington reexamined. The American Statistician, 40(4):283--284, 1986. doi:10.1080/00031305.1986.10475415

  34. [42]

    Pandeva and P

    T. Pandeva and P. Forr\'e. Multi-view independent component analysis with shared and individual sources. In R. J. Evans and I. Shpitser, editors, Proceedings of the 39th Conference on Uncertainty in Artificial Intelligence (UAI), volume 216 of Proceedings of Machine Learning R...

  35. [43]

    Pandeva and P

    T. Pandeva and P. Forr\'e. Multi-view independent component analysis for omics data integration. ICLR 2023 Workshop on Machine Learning and Global Health, 2023b. https://openreview.net/forum?id=r5KL-AfXt75

  36. [44]

    Pandeva, M

    T. Pandeva, M. J. Jonker, L. Hamoen, J. Mooij, and P. Forr\'e. Robust multi-view co-expression network inference. In B. Huang and M. Drton, editors, Proceedings of the 4th Conference on Causal Learning and Reasoning (CLeaR), volume 275 of Proceedings of Machine Learning Resear...

  37. [45]

    G. P\'olya. Remarks on characteristic functions. In J. Neyman, editor, Proceedings of the Berkeley Symposium on Mathematical Statistics and Probability, pages 115--123. University of California Press, Berkeley and Los Angeles, 1949. https://digitalassets.lib.berkeley.edu/math/...

  38. [46]

    Robbins and S

    H. Robbins and S. Monro. A stochastic approximation method. The Annals of Mathematical Statistics, 22(3):400--407, 1951. doi:10.1214/aoms/1177729586

  39. [47]

    Shimizu, P

    S. Shimizu, P. O. Hoyer, A. Hyv\"arinen, and A. Kerminen. A linear non-Gaussian acyclic model for causal discovery. Journal of Machine Learning Research, 7:2003--2030, 2006. https://jmlr.org/papers/v7/shimizu06a.html

  40. [48]

    Shimizu, T

    S. Shimizu, T. Inazumi, Y. Sogawa, A. Hyv\"arinen, Y. Kawahara, T. Washio, P. O. Hoyer, and K. Bollen. DirectLiNGAM: a direct method for learning a linear non-Gaussian structural equation model. Journal of Machine Learning Research, 12:1225--1248, 2011. https://jmlr.org/papers...

  41. [49]

    V. P. Skitovich. Linear forms of independent random variables and the normal distribution law. Izvestiya Akademii Nauk SSSR, Seriya Matematicheskaya, 18(2):185--200, 1954. In Russian. https://www.mathnet.ru/eng/im3497 . Announced in Doklady Akademii Nauk SSSR (N.S.), 89:217--219, 1953

  42. [50]

    P. H. Westfall. Kurtosis as peakedness, 1905--2014. R.I.P. The American Statistician, 68(3):191--195, 2014. doi:10.1080/00031305.2014.917055

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.