Pith. sign in

REVIEW 2 major objections 4 minor 107 references

A hybrid neural network reconstructs the supernova distance ladder without a cosmological model and pins H0 near 69.6 km/s/Mpc.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-10 20:05 UTC pith:U6WOII63

load-bearing objection Competent hybrid-network extension of LADDER that delivers a usable μ(z) reconstruction and consistent H0 inside flat ΛCDM, but the superiority claim rests on non-significant MSE differences and unlinked stability. the 2 major comments →

arxiv 2607.06959 v1 pith:U6WOII63 submitted 2026-07-08 astro-ph.CO

KAN-LSTM-Transformer Neural Networks, MFV and Cosmological Parameters

classification astro-ph.CO
keywords cosmic distance ladderType Ia supernovaeKolmogorov-Arnold networksLSTMTransformerHubble constantMost Frequent Valueflat ΛCDM
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that a hybrid network called KLT-Net, which stacks Kolmogorov–Arnold layers, LSTM memory cells and a Transformer encoder, can reconstruct the distance modulus of Type Ia supernovae directly from the Pantheon catalogue without assuming any cosmological model. Once that reconstruction is in hand, the authors combine it with a Most-Frequent-Value estimate of the absolute magnitude and standard Bayesian fitting inside flat ΛCDM to obtain a tightly constrained Hubble constant. They claim the hybrid architecture is the most accurate and the most stable of the machine-learning and classical regression models they tested, and that the three independent statistical routes (χ², MCMC and Hessian) all agree. A sympathetic reader cares because a stable, model-independent distance ladder is a practical tool for the next generation of large supernova surveys and for any later exploration of dark-energy or modified-gravity models.

Core claim

KLT-Net produces the lowest mean-squared error (0.015727) and the lowest cross-seed coefficient of variation among all architectures examined, and the resulting non-parametric distance modulus, together with an MFV absolute magnitude of –19.377, yields H0 = 69.576^{+0.483}_{-0.482} km s^{-1} Mpc^{-1} and Ωm0 = 0.301^{+0.039}_{-0.036} inside flat ΛCDM, with χ², MCMC and Hessian analyses mutually consistent.

What carries the argument

KLT-Net: a three-stage hybrid that first lets LSTM capture local redshift-sequence correlations (including covariance), then lets KAN layers replace linear weights by learnable splines for compact non-linear feature transformation, and finally lets a Transformer self-attention block extract global evolutionary patterns across the full redshift range.

Load-bearing premise

That the hybrid network’s superior training stability across random seeds is enough reason to prefer it for cosmological inference even though formal statistical tests find no significant accuracy difference versus simpler alternatives.

What would settle it

Retrain the identical architecture and the LADDER baseline on the same Pantheon+ sample with a larger set of random seeds; if the resulting posterior distributions for H0 and Ωm differ by more than the quoted 1σ uncertainties, or if a paired test now finds a significant MSE gap, the claim that KLT-Net is the uniquely preferred reconstructor collapses.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A model-independent μ(z) ladder can be fed into any subsequent dark-energy or modified-gravity analysis without re-fitting the supernova photometry.
  • The same KLT architecture can be retrained on forthcoming LSST or Euclid supernova catalogues to test high-redshift extrapolation.
  • MFV plus bootstrap supplies a robust absolute-magnitude prior that can be reused in any other ladder-based H0 determination.
  • Hessian-matrix validation becomes a quick consistency check for any Bayesian cosmological pipeline that uses neural-network predictions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the reconstruction itself never assumes flat ΛCDM, the same μ(z) could later be used to test whether a late-time transition in absolute magnitude is required, an idea the paper only mentions in the introduction.
  • The stability argument may matter more for next-generation surveys that will have far denser high-redshift sampling, where small training fluctuations could otherwise bias dark-energy equation-of-state constraints.
  • If the Transformer’s global attention is the main source of stability, simpler attention-augmented LSTM models might achieve the same cosmological performance with fewer parameters.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes KLT-Net, a hybrid of Kolmogorov-Arnold networks, LSTM and Transformer layers, as a data-driven non-parametric reconstructor of the distance modulus μ(z) from Pantheon SN Ia apparent magnitudes (including covariance). After comparing against classical regressors (DTR, RF, SVM, KNN, MLP) and performing multi-seed ablations against LADDER and partial hybrids, the authors select KLT-Net on the basis of lowest MSE (0.015727) and lowest cross-seed coefficient of variation. Absolute magnitude MB is estimated via the Most Frequent Value (MFV) procedure applied to a literature compilation; the reconstructed μ(z) is then inserted into a flat-ΛCDM likelihood whose parameters H0 and Ωm0 are constrained by χ² minimization, emcee MCMC and Hessian-matrix error propagation, yielding mutually consistent values H0 ≈ 69.58 km s⁻¹ Mpc⁻¹ and Ωm0 ≈ 0.30.

Significance. If the hybrid architecture demonstrably improves the reliability of the reconstructed distance ladder for downstream cosmological inference, the work supplies a practical, covariance-aware alternative to Gaussian processes and pure LSTM (LADDER). Strengths that should be credited include the systematic ten-seed ablation study with both parametric and non-parametric significance tests, the explicit cross-validation of three independent inference engines (χ², MCMC, Hessian), and the transparent application of MFV plus bootstrap to the absolute-magnitude prior. These elements make the pipeline more reproducible than many black-box ML cosmology papers. The scientific advance remains incremental relative to Shah et al. (2024) and does not resolve the Hubble tension, but a well-validated hybrid reconstructor would still be a useful community tool for future large SN samples.

major comments (2)
  1. §3 (Ablation Study), Tables 2–3: every pairwise MSE comparison of KLT versus LADDER, LT, KT and KL returns p > 0.05 by both paired t-test and Wilcoxon signed-rank test. The accuracy advantage is therefore not statistically established. The manuscript elevates residual training stability (lowest CV = 3.45 %) as sufficient justification for feeding only the KLT reconstruction into the §4 cosmological pipeline. Without reporting the identical MFV-MB + flat-ΛCDM analysis driven by the LADDER (or ablated) reconstructions, it remains unproven that the lower CV tightens or de-biases the headline posteriors H0 = 69.576^{+0.483}_{-0.482}, Ωm0 = 0.301^{+0.039}_{-0.036}. This comparison is load-bearing for the claim that KLT-Net is the preferred architecture for cosmological parameter estimation.
  2. §4.2 and Eqs. (28)–(29): the likelihood is written as a simple diagonal χ² on μSN(zi) − μth(zi). The text repeatedly asserts that “full covariance information” of Pantheon+ is incorporated, yet the published Csys matrix never appears in the likelihood or in the network loss. Clarification is required whether the network was trained with a multivariate Gaussian loss that includes Csys, and whether the same matrix is used (or marginalized) when the reconstructed μ(z) is passed to the cosmological sampler. If the covariance is ignored at either stage, the reported 1σ uncertainties on H0 and Ωm0 are under-estimated.
minor comments (4)
  1. Figure 1 caption and §2.2.4: the conceptual diagram is described but the precise layer dimensions, spline order, number of Transformer heads and training hyper-parameters are never tabulated; a short architecture table would aid reproducibility.
  2. §4.1: the literature sample of MB values used for the MFV calculation is not listed; a supplementary table of the adopted references and their quoted MB would allow independent verification of the MFV and bootstrap intervals.
  3. Throughout: several typographical inconsistencies appear (e.g., “ΛCDM” vs “Λ CDM”, “Pantheon” vs “Pantheon+”, missing spaces before units). A careful copy-edit pass is needed.
  4. Eq. (2) and surrounding text: the conversion from m(z) to μ(z) assumes a constant MB; the later discussion of possible redshift evolution of MB is not propagated into the reconstruction uncertainty.

Circularity Check

1 steps flagged

No load-bearing circularity: KLT-Net reconstruction of μ(z) is data-driven and model-independent; subsequent flat-ΛCDM inference of H0/Ωm and MFV MB are ordinary sequential steps, not reductions by construction.

specific steps
  1. self citation load bearing [§4.1 (MFV for MB) and references [77,80,81,83]]
    "Using these previous results and the bootstrap method, the holistic estimate of MB for all data is -19.3772 and the dihesion ε=0.0291. ... Similar outcomes were observed in the literature cited, as noted by Mukherjee et al. [7]. This implies that error distributions are not Gaussian, thereby bolstering the credibility and rationality of the MFV estimate."

    The MFV procedure itself is Steiner’s (external); the authors cite their own prior applications of MFV to other astrophysical quantities. This is ordinary self-citation of method reuse, not a load-bearing uniqueness theorem that forces the present MB or H0 values. It does not make the cosmological-parameter result circular.

full rationale

The derivation chain is: (i) train KLT-Net (and ablations) on Pantheon apparent magnitudes + covariance to reconstruct μ(z) non-parametrically (lowest MSE 0.015727, lowest CV 3.45 % across seeds); (ii) obtain MB via MFV + bootstrap on a literature compilation; (iii) feed the reconstructed μ(z) into χ²/MCMC/Hessian under the flat-ΛCDM likelihood to obtain H0 ≈ 69.58 and Ωm0 ≈ 0.30. Step (i) does not embed ΛCDM or H0 (explicitly stated in §5); the network is validated against external baselines (LADDER, DTR, RF, etc.) and ablations with paired tests (Table 3, all p > 0.05). Step (ii) applies Steiner’s MFV (with self-citations only for prior applications of the same estimator, not for a uniqueness theorem that forces the result). Step (iii) is standard parametric inference after a non-parametric reconstruction; the paper itself flags the residual model dependence and does not claim a first-principles or model-free derivation of H0. No equation reduces a claimed prediction to a fitted input by construction, no self-citation supplies a load-bearing uniqueness claim, and no ansatz is smuggled. The preference for KLT on stability alone (despite non-significant MSE differences) is a methodological weakness, not circularity. Score 1 only for the minor, non-load-bearing self-citations of the authors’ earlier MFV applications.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 1 invented entities

The central numerical claims rest on (i) the hybrid network’s ability to learn μ(z) from Pantheon under MSE, (ii) the MFV summary of heterogeneous literature M_B, and (iii) the flat-ΛCDM likelihood that converts μ(z) into H0 and Ωm. Free parameters include the fitted cosmological pair and the MFV M_B; axioms include FLRW geometry, Gaussian SN errors, and the Kolmogorov–Arnold representation used to motivate KAN layers. KLT-Net itself is the sole invented architectural entity.

free parameters (4)
  • H0 (flat ΛCDM) = 69.576^{+0.483}_{-0.482} km s⁻¹ Mpc⁻¹
    Fitted by χ² / MCMC / Hessian to the KLT-Net μ(z) curve; reported central value 69.576 km s⁻¹ Mpc⁻¹.
  • Ωm0 (flat ΛCDM) = 0.301^{+0.039}_{-0.036}
    Jointly fitted with H0 under the same likelihood; reported 0.301^{+0.039}_{-0.036}.
  • MB (MFV of literature) = -19.3772 (ε=0.0291; 68% CI [-19.389, -19.352])
    Most Frequent Value computed from published absolute-magnitude estimates; used as fixed calibration when converting m(z) to μ(z).
  • KLT-Net training hyperparameters
    Layer widths, spline order, learning rate, number of heads, dropout, etc., are chosen by the authors and control the reconstructed μ(z); exact values are not tabulated.
axioms (5)
  • domain assumption Spatially flat FLRW luminosity-distance integral (Eq. 1) relates H(z) to DL(z).
    Invoked throughout §2.1 and §4 to convert reconstructed μ into cosmological parameters.
  • domain assumption Flat ΛCDM expansion history H(z)=H0√[Ωm0(1+z)³+(1-Ωm0)] (Eq. 22).
    Assumed for all final H0/Ωm inference in §4.2; paper acknowledges this re-introduces model dependence.
  • domain assumption Gaussian likelihood for distance-modulus residuals (Eqs. 28–29).
    Standard SN Ia assumption used for χ², MCMC and Hessian.
  • standard math Kolmogorov–Arnold representation theorem justifies replacing MLP weights by univariate splines.
    Cited via Liu et al. to motivate the KAN module (§2.2.1).
  • ad hoc to paper Literature M_B estimates form a sample whose central tendency is best captured by MFV rather than mean/median.
    Justified by non-Gaussian residuals in Fig. 9; underpins the absolute-magnitude calibration.
invented entities (1)
  • KLT-Net (KAN-LSTM-Transformer hybrid) no independent evidence
    purpose: Non-parametric, covariance-aware reconstruction of μ(z) claimed to improve stability over pure LSTM (LADDER).
    Architectural invention of the paper; no independent external validation beyond the authors’ own ablations.

pith-pipeline@v1.1.0-grok45 · 33931 in / 3566 out tokens · 45199 ms · 2026-07-10T20:05:57.529632+00:00 · methodology

0 comments
read the original abstract

Reconstructing the cosmic distance ladder directly from observations is a crucial issue in cosmology. In this paper, we present a novel method for modeling the cosmic distance ladder and estimating cosmological parameters through the use of Kolmogorov-Arnold networks (KAN), Long Short-Term Memory (LSTM), and Transformer networks (collectively referred to as KLT-Net), based on the apparent magnitude data from the Pantheon SN Ia compilation. As a data-driven, non-parametric method for reconstructing the distance modulus $\mu(z)$, KLT-Net is shown to be highly effective in capturing the intricate, nonlinear measurement distributions. After validating against various statistical and machine learning models, we have identified it as the most effective choice among the considered alternatives and ablation experiments. Subsequently, the statistical inference of $H_0$ and $\Omega_{\rm m}$ adopts the flat $\Lambda$CDM framework. Moreover, we introduce the Most Frequent Value (MFV) approach to evaluate the absolute magnitude, $M_B$, from existing literature data. In addition, we employ the Hessian matrix to validate the Bayesian method, demonstrating that the Hubble constant can be precisely constrained from the KLT-Net predictions within the flat $\Lambda$CDM framework. The integration of KLT-Net, the MFV approach, and Bayesian statistics establishes a robust framework for inferring cosmological parameters. This methodology facilitates future cosmological research, particularly in the analysis of complex datasets and the exploration of high-dimensional parameter spaces.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

107 extracted references · 107 canonical work pages · 44 internal anchors

  1. [1]

    Cosmology with the Large Synoptic Survey Telescope: an Overview

    Zhan, H., Tyson, J.A.: Cosmology with the Large Synoptic Sur- vey Telescope: an overview. Reports on Progress in Physics81(6), 066901 (2018) https://arxiv.org/abs/1707.06948 [astro-ph.CO]. https: //doi.org/10.1088/1361-6633/aab1bd

  2. [2]

    Role of Future SNIa Data from Rubin LSST in Reinvestigating Cosmological Models

    Shah, R., Mitra, A., Mukherjee, P., Pal, B., Pal, S.: Role of future SNIa data from Rubin LSST in reinvestigating cosmological models. MNRAS530(3), 2627–2636 (2024) https://arxiv.org/abs/2305.08786 [astro-ph.CO]. https://doi.org/10.1093/mnras/stae1016

  3. [3]

    PRD111(4), 043503 (2025) https://arxiv.org/abs/2406

    Kumar, D., Mitra, A., Adil, S.A., Sen, A.A.: Exploring alternative cos- mologies with the LSST: Simulated forecasts and current observational constraints. PRD111(4), 043503 (2025) https://arxiv.org/abs/2406. 06757 [astro-ph.CO]. https://doi.org/10.1103/PhysRevD.111.043503

  4. [4]

    Reconstructing the Hubble parameter with future Gravitational Wave missions using Machine Learning

    Mukherjee, P., Shah, R., Bhaumik, A., Pal, S.: Reconstructing the Hub- ble Parameter with Future Gravitational-wave Missions Using Machine Learning. ApJ960(1), 61 (2024) https://arxiv.org/abs/2303.05169 [astro-ph.CO]. https://doi.org/10.3847/1538-4357/ad055f

  5. [6]

    A&A663, 112 (2022) https://arxiv.org/abs/2203

    Moya, A., L´ opez-Sastre, R.J.: Stellar mass and radius estimation using artificial intelligence. A&A663, 112 (2022) https://arxiv.org/abs/2203. 06027 [astro-ph.SR]. https://doi.org/10.1051/0004-6361/202142930

  6. [7]

    Mukherjee, P., Dialektopoulos, K.F., Said, J.L., Mifsud, J.: A possible late-time transition of M B inferred via neural networks. J. Cosmology Astropart. Phys.2024(9), 060 (2024) https://arxiv.org/abs/2402.10502 [astro-ph.CO]. https://doi.org/10.1088/1475-7516/2024/09/060

  7. [8]

    Neural network reconstruction of scalar-tensor cosmology

    Dialektopoulos, K.F., Mukherjee, P., Said, J.L., Mifsud, J.: Neural net- work reconstruction of scalar-tensor cosmology. Physics of the Dark Universe43, 101383 (2024) https://arxiv.org/abs/2305.15500 [gr-qc]. https://doi.org/10.1016/j.dark.2023.101383

  8. [9]

    Fortunato, J.A.S., Bacon, D.J., Hip´ olito-Ricaldi, W.S., Wands, D.: Springer Nature 2021 LATEX template Article Title31 Fast Radio Bursts and Artificial Neural Networks: a cosmological- model-independent estimation of the Hubble constant. J. Cosmology Astropart. Phys.2025(1), 018 (2025) https://arxiv.org/abs/2407.03532 [astro-ph.CO]. https://doi.org/10.1...

  9. [10]

    Di Valentino, E., Levi Said, J., Riess, A., Pollo, A., Poulin, V., G´ omez- Valent, A., Weltman, A., Palmese, A., Huang, C.D., van de Bruck, C., Shekhar Saraf, C., Kuo, C.-Y., Uhlemann, C., Grand´ on, D., Paz, D., Eckert, D., Teixeira, E.M., Saridakis, E.N., Colg´ ain, E. ´O., Beut- ler, F., Niedermann, F., Bajardi, F., Barenboim, G., Gubitosi, G., Musell...

  10. [11]

    ApJ977(1), 120 (2024) https://arxiv.org/abs/2408

    Riess, A.G., Scolnic, D., Anand, G.S., Breuval, L., Casertano, S., Macri, L.M., Li, S., Yuan, W., Huang, C.D., Jha, S., Murakami, Y.S., Beaton, R., Brout, D., Wu, T., Addison, G.E., Bennett, C., Anderson, R.I., Fil- ippenko, A.V., Carr, A.: JWST Validates HST Distance Measurements: Selection of Supernova Subsample Explains Differences in JWST Esti- mates ...

  11. [12]

    JWST Observations Reject Unrecognized Crowding of Cepheid Photometry as an Explanation for the Hubble Tension at 8 sigma Confidence

    Riess, A.G., Anand, G.S., Yuan, W., Casertano, S., Dolphin, A., Macri, L.M., Breuval, L., Scolnic, D., Perrin, M., Anderson, R.I.: JWST Obser- vations Reject Unrecognized Crowding of Cepheid Photometry as an Explanation for the Hubble Tension at 8σConfidence. ApJL962(1), 17 (2024) https://arxiv.org/abs/2401.04773 [astro-ph.CO]. https://doi. org/10.3847/20...

  12. [13]

    Systematics in the Cepheid and TRGB Distance Scales: Metallicity Sensitivity of the Wesenheit Leavitt Law

    Madore, B.F., Freedman, W.L.: Systematics in the Cepheid and TRGB Distance Scales: Metallicity Sensitivity of the Wesenheit Leavitt Law. ApJ961(2), 166 (2024) https://arxiv.org/abs/2309.10859 [astro-ph.SR]. https://doi.org/10.3847/1538-4357/acfaea

  13. [14]

    Calibration of the Tip of the Red Giant Branch (TRGB)

    Freedman, W.L., Madore, B.F., Hoyt, T., Jang, I.S., Beaton, R., Lee, M.G., Monson, A., Neeley, J., Rich, J.: Calibration of the Tip of the Red Giant Branch. ApJ891(1), 57 (2020) https://arxiv.org/abs/2002.01550 [astro-ph.GA]. https://doi.org/10.3847/1538-4357/ab7339

  14. [15]

    Freedman, W.L., Madore, B.F.: Progress in direct measurements of the Hubble constant. J. Cosmology Astropart. Phys.2023(11), 050 (2023) https://arxiv.org/abs/2309.05618 [astro-ph.CO]. https://doi. org/10.1088/1475-7516/2023/11/050

  15. [16]

    MNRAS495(3), 2630–2644 (2020) https://arxiv.org/abs/1910

    Camarena, D., Marra, V.: A new method to build the (inverse) distance ladder. MNRAS495(3), 2630–2644 (2020) https://arxiv.org/abs/1910. 14125 [astro-ph.CO]. https://doi.org/10.1093/mnras/staa770

  16. [17]

    European Physical Journal C84(8), 873 (2024)

    Gong, X., Liu, T., Wang, J.: Inverse distance ladder method for determining ¡inline-formula id=“IEq1”¿¡mml:math¿¡mml:msub¿¡mml:mi¿H¡/mml:mi¿¡mml:mn¿0¡/mml:mn¿¡/mml:msub¿¡/mml:math¿¡/inline- formula¿ from angular diameter distances of time-delay lenses and supernova observations. European Physical Journal C84(8), 873 (2024). https://doi.org/10.1140/epjc/s1...

  17. [18]

    Adame, A.G., Aguilar, J., Ahlen, S., Alam, S., Alexander, D.M., Alvarez, M., Alves, O., Anand, A., Andrade, U., Armengaud, E., Avila, S., Aviles, Springer Nature 2021 LATEX template Article Title33 A., Awan, H., Bahr-Kalus, B., Bailey, S., Baltay, C., Bault, A., Behera, J., BenZvi, S., Bera, A., Beutler, F., Bianchi, D., Blake, C., Blum, R., Brieden, S., ...

  18. [19]

    New Hubble Space Telescope Discoveries of Type Ia Supernovae at z > 1: Narrowing Constraints on the Early Behavior of Dark Energy

    Riess, A.G., Strolger, L.-G., Casertano, S., Ferguson, H.C., Mobasher, B., Gold, B., Challis, P.J., Filippenko, A.V., Jha, S., Li, W., Tonry, J., Foley, R., Kirshner, R.P., Dickinson, M., MacDonald, E., Eisenstein, D., Livio, M., Younger, J., Xu, C., Dahl´ en, T., Stern, D.: New Hubble Space Telescope Discoveries of Type Ia Supernovae at z ¿= 1: Narrowing...

  19. [20]

    The Complete Light-curve Sample of Spectroscopically Confirmed Type Ia Supernovae from Pan-STARRS1 and Cosmological Constraints from The Combined Pantheon Sample

    Scolnic, D.M., Jones, D.O., Rest, A., Pan, Y.C., Chornock, R., Foley, R.J., Huber, M.E., Kessler, R., Narayan, G., Riess, A.G., Rodney, S., Berger, E., Brout, D.J., Challis, P.J., Drout, M., Finkbeiner, D., Lun- nan, R., Kirshner, R.P., Sanders, N.E., Schlafly, E., Smartt, S., Stubbs, C.W., Tonry, J., Wood-Vasey, W.M., Foley, M., Hand, J., Johnson, E., Bu...

  20. [21]

    CoLFI: Cosmological Likelihood-free Inference with Neural Density Estimators

    Wang, G.-J., Cheng, C., Ma, Y.-Z., Xia, J.-Q., Abebe, A., Beesham, A.: CoLFI: Cosmological Likelihood-free Inference with Neural Density Estimators. ApJS268(1), 7 (2023) https://arxiv.org/abs/2306.11102 [astro-ph.CO]. https://doi.org/10.3847/1538-4365/ace113

  21. [22]

    Machine Learning for Observational Cosmology

    Moriwaki, K., Nishimichi, T., Yoshida, N.: Machine learning for obser- vational cosmology. Reports on Progress in Physics86(7), 076901 (2023) https://arxiv.org/abs/2303.15794 [astro-ph.IM]. https://doi.org/ 10.1088/1361-6633/acd2ea

  22. [23]

    KAN: Kolmogorov-Arnold Networks

    Liu, Z., Wang, Y., Vaidya, S., Ruehle, F., Halverson, J., Soljaˇ ci´ c, M., Hou, T.Y., Tegmark, M.: KAN: Kolmogorov-Arnold Networks. arXiv e-prints, 2404–19756 (2024) https://arxiv.org/abs/2404.19756 [cs.LG]. https://doi.org/10.48550/arXiv.2404.19756

  23. [24]

    Attention Is All You Need

    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention Is All You Need. arXiv e-prints, 1706–03762 (2017) https://arxiv.org/abs/1706.03762 [cs.CL]. https://doi.org/10.48550/arXiv.1706.03762

  24. [25]

    Zhong, F., Luo, R., Napolitano, N.R., Tortora, C., Li, R., Zhu, X., Busillo, V., Koopmans, L.V.E., Longo, G.: Galaxy–Galaxy Strong Lensing with U-Net (GGSL-UNet). I. Extracting Two-dimensional Information from Multiband Images in Ground and Space Observa- tions. ApJS277(1), 12 (2025) https://arxiv.org/abs/2410.02936 [astro- ph.GA]. https://doi.org/10.3847...

  25. [26]

    ApJ980(1), 66 (2025) https://arxiv.org/abs/2407

    R´ o˙ za´ nski, T., Ting, Y.-S., Jab lo´ nska, M.: TransformerPayne: Enhancing Spectral Emulation Accuracy and Data Efficiency by Capturing Long- range Correlations. ApJ980(1), 66 (2025) https://arxiv.org/abs/2407. Springer Nature 2021 LATEX template Article Title35 05751 [astro-ph.IM]. https://doi.org/10.3847/1538-4357/ad9b99

  26. [27]

    Advancing Cosmological Parameter Estimation and Hubble Parameter Reconstruction with Long Short-Term Memory and Efficient-Kolmogorov-Arnold Networks

    Cui, J., Biesiada, M., Liu, A., Wen, C., Liu, T., Wang, J.: Advancing Cos- mological Parameter Estimation and Hubble Parameter Reconstruction with Long Short-Term Memory and Efficient-Kolmogorov-Arnold Net- works. arXiv e-prints, 2504–00392 (2025) https://arxiv.org/abs/2504. 00392 [astro-ph.CO]. https://doi.org/10.48550/arXiv.2504.00392

  27. [28]

    The Pantheon+ Analysis: The Full Dataset and Light-Curve Release

    Scolnic, D., Brout, D., Carr, A., Riess, A.G., Davis, T.M., Dwomoh, A., Jones, D.O., Ali, N., Charvu, P., Chen, R., Peterson, E.R., Popovic, B., Rose, B.M., Wood, C.M., Brown, P.J., Chambers, K., Coulter, D.A., Dettman, K.G., Dimitriadis, G., Filippenko, A.V., Foley, R.J., Jha, S.W., Kilpatrick, C.D., Kirshner, R.P., Pan, Y.-C., Rest, A., Rojas- Bravo, C....

  28. [29]

    https://doi.org/ 10.1016/C2017-0-01943-2

    Dodelson, S., Schmidt, F.: Modern Cosmology, (2020). https://doi.org/ 10.1016/C2017-0-01943-2

  29. [30]

    On the use of the local prior on the absolute magnitude of Type Ia supernovae in cosmological inference

    Camarena, D., Marra, V.: On the use of the local prior on the absolute magnitude of Type Ia supernovae in cosmological inference. MNRAS504(4), 5164–5171 (2021) https://arxiv.org/abs/2101.08641 [astro-ph.CO]. https://doi.org/10.1093/mnras/stab1200

  30. [31]

    KAN 2.0: Kolmogorov-Arnold Networks Meet Science

    Liu, Z., Ma, P., Wang, Y., Matusik, W., Tegmark, M.: KAN 2.0: Kolmogorov-Arnold Networks Meet Science. arXiv e-prints, 2408–10205 (2024) https://arxiv.org/abs/2408.10205 [cs.LG]. https://doi.org/10. 48550/arXiv.2408.10205

  31. [32]

    Convolutional Kolmogorov-Arnold Networks

    Dylan Bodner, A., Santiago Tepsich, A., Natan Spolski, J., Pourteau, S.: Convolutional Kolmogorov-Arnold Networks. arXiv e-prints, 2406– 13155 (2024) https://arxiv.org/abs/2406.13155 [cs.CV]. https://doi.org/ 10.48550/arXiv.2406.13155

  32. [33]

    A Critical Review of Recurrent Neural Networks for Sequence Learning

    Lipton, Z.C., Berkowitz, J., Elkan, C.: A critical review of recurrent neural networks for sequence learning. arXiv preprint arXiv:1506.00019 (2015)

  33. [34]

    Machine learning for Brain disorders, 117–138 (2023)

    Das, S., Tariq, A., Santos, T., Kantareddy, S.S., Banerjee, I.: Recurrent neural networks (rnns): architectures, training tricks, and introduction to influential research. Machine learning for Brain disorders, 117–138 (2023)

  34. [35]

    Artificial Intelligence Review53(8), 5929–5955 Springer Nature 2021 LATEX template 36Article Title (2020)

    Van Houdt, G., Mosquera, C., N´ apoles, G.: A review on the long short- term memory model. Artificial Intelligence Review53(8), 5929–5955 Springer Nature 2021 LATEX template 36Article Title (2020)

  35. [36]

    Information15(9), 517 (2024)

    Mienye, I.D., Swart, T.G., Obaido, G.: Recurrent neural networks: A comprehensive review of architectures, variants, and applications. Information15(9), 517 (2024)

  36. [37]

    Neural com- putation9(8), 1735–1780 (1997)

    Hochreiter, S., Schmidhuber, J.: Long short-term memory. Neural com- putation9(8), 1735–1780 (1997)

  37. [38]

    Neural computation31(7), 1235– 1270 (2019)

    Yu, Y., Si, X., Hu, C., Zhang, J.: A review of recurrent neural networks: Lstm cells and network architectures. Neural computation31(7), 1235– 1270 (2019)

  38. [39]

    A Comprehensive Overview and Comparative Analysis on Deep Learning Models: CNN, RNN, LSTM, GRU

    Shiri, F.M., Perumal, T., Mustapha, N., Mohamed, R.: A comprehensive overview and comparative analysis on deep learning models: Cnn, rnn, lstm, gru. arXiv preprint arXiv:2305.17473 (2023)

  39. [40]

    Neural computation12(10), 2451–2471 (2000)

    Gers, F.A., Schmidhuber, J., Cummins, F.: Learning to forget: Continual prediction with lstm. Neural computation12(10), 2451–2471 (2000)

  40. [41]

    Generating Sequences With Recurrent Neural Networks

    Graves, A.: Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850 (2013)

  41. [42]

    Reformer: The Efficient Transformer

    Kitaev, N., Kaiser, L., Levskaya, A.: Reformer: The efficient transformer. arXiv preprint arXiv:2001.04451 (2020)

  42. [43]

    Devlin, J., Chang, M.-W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional transformers for language understanding. In: Pro- ceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolo- gies, Volume 1 (long and Short Papers), pp. 4171–4186 (2019)

  43. [44]

    Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)

    Clevert, D.-A., Unterthiner, T., Hochreiter, S.: Fast and accurate deep network learning by exponential linear units (elus). arXiv preprint arXiv:1511.07289 (2015)

  44. [45]

    Springer, ??? (1999)

    Vapnik, V.: The Nature of Statistical Learning Theory. Springer, ??? (1999)

  45. [46]

    Springer, ??? (2009)

    Hastie, T., Tibshirani, R., Friedman, J.H., Friedman, J.H.: The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer, ??? (2009)

  46. [47]

    Some Aspects of Measurement Error in Linear Regression of Astronomical Data

    Kelly, B.C.: Some Aspects of Measurement Error in Linear Regression of Astronomical Data. ApJ665, 1489–1506 (2007) https://arxiv.org/abs/ 0705.2774. https://doi.org/10.1086/519947

  47. [48]

    Springer Nature 2021 LATEX template Article Title37 Cambridge University Press, ??? (2012)

    Feigelson, E.D., Babu, G.J.: Modern Statistical Methods for Astronomy. Springer Nature 2021 LATEX template Article Title37 Cambridge University Press, ??? (2012)

  48. [49]

    Chapman and Hall/CRC, ??? (1984)

    Breiman, L., Friedman, J., Olshen, R.A., Stone, C.J.: Classification and Regression Trees. Chapman and Hall/CRC, ??? (1984)

  49. [50]

    Wiley interdisciplinary reviews: data mining and knowledge discovery1(1), 14–23 (2011)

    Loh, W.-Y.: Classification and regression trees. Wiley interdisciplinary reviews: data mining and knowledge discovery1(1), 14–23 (2011)

  50. [51]

    Interna- tional Statistical Review82(3), 329–348 (2014)

    Loh, W.-Y.: Fifty years of classification and regression trees. Interna- tional Statistical Review82(3), 329–348 (2014)

  51. [52]

    Machine learning45, 5–32 (2001)

    Breiman, L.: Random forests. Machine learning45, 5–32 (2001)

  52. [53]

    Springer, ??? (2020)

    Genuer, R., Poggi, J.-M., Genuer, R., Poggi, J.-M.: Random Forests. Springer, ??? (2020)

  53. [54]

    PhD thesis, Universite de Liege (Belgium) (2014)

    Louppe, G.: Understanding random forests: From theory to practice. PhD thesis, Universite de Liege (Belgium) (2014)

  54. [55]

    Machine learning20, 273–297 (1995)

    Cortes, C., Vapnik, V.: Support-vector networks. Machine learning20, 273–297 (1995)

  55. [56]

    Noble, W.S.: What is a support vector machine? Nature biotechnology 24(12), 1565–1567 (2006)

  56. [57]

    Advances in neural information processing systems9(1996)

    Drucker, H., Burges, C.J., Kaufman, L., Smola, A., Vapnik, V.: Support vector regression machines. Advances in neural information processing systems9(1996)

  57. [58]

    Statistics and computing14, 199–222 (2004)

    Smola, A.J., Sch¨ olkopf, B.: A tutorial on support vector regression. Statistics and computing14, 199–222 (2004)

  58. [59]

    Analyst135(2), 230–267 (2010)

    Brereton, R.G., Lloyd, G.R.: Support vector machines for classification and regression. Analyst135(2), 230–267 (2010)

  59. [60]

    Efficient learning machines: Theories, concepts, and applications for engineers and system designers, 67–80 (2015)

    Awad, M., Khanna, R., Awad, M., Khanna, R.: Support vector regres- sion. Efficient learning machines: Theories, concepts, and applications for engineers and system designers, 67–80 (2015)

  60. [61]

    Guo, G., Wang, H., Bell, D., Bi, Y., Greer, K.: Knn model-based approach in classification. In: On The Move to Meaningful Internet Systems 2003: CoopIS, DOA, and ODBASE: OTM Confederated Inter- national Conferences, CoopIS, DOA, and ODBASE 2003, Catania, Sicily, Italy, November 3-7, 2003. Proceedings, pp. 986–996 (2003). Springer

  61. [62]

    In: The Top Ten Algorithms in Data Mining, pp

    Steinbach, M., Tan, P.-N.: knn: k-nearest neighbors. In: The Top Ten Algorithms in Data Mining, pp. 165–176. Chapman and Hall/CRC, ??? (2009) Springer Nature 2021 LATEX template 38Article Title

  62. [63]

    Neurocomputing251, 26–34 (2017)

    Song, Y., Liang, J., Lu, J., Zhao, X.: An efficient instance selection algo- rithm for k nearest neighbor regression. Neurocomputing251, 26–34 (2017)

  63. [64]

    Acta numerica8, 143–195 (1999)

    Pinkus, A.: Approximation theory of the mlp model in neural networks. Acta numerica8, 143–195 (1999)

  64. [65]

    In: Geomatic Approaches for Modeling Land Change Scenarios, pp

    Taud, H., Mas, J.-F.: Multilayer perceptron (mlp). In: Geomatic Approaches for Modeling Land Change Scenarios, pp. 451–455. Springer, ??? (2017)

  65. [66]

    KAN or MLP: A Fairer Comparison

    Yu, R., Yu, W., Wang, X.: Kan or mlp: A fairer comparison. arXiv preprint arXiv:2407.16674 (2024)

  66. [67]

    Jour- nal of Hydrology649, 132430 (2025)

    Ren, D., Hu, Q., Zhang, T.: EKLT: Kolmogorov-Arnold attention-driven LSTM with Transformer model for river water level prediction. Jour- nal of Hydrology649, 132430 (2025). https://doi.org/10.1016/j.jhydrol. 2024.132430

  67. [68]

    A Comprehensive Measurement of the Local Value of the Hubble Constant with 1 km/s/Mpc Uncertainty from the Hubble Space Telescope and the SH0ES Team

    Riess, A.G., Yuan, W., Macri, L.M., Scolnic, D., Brout, D., Caser- tano, S., Jones, D.O., Murakami, Y., Anand, G.S., Breuval, L., Brink, T.G., Filippenko, A.V., Hoffmann, S., Jha, S.W., D’arcy Kenworthy, W., Mackenty, J., Stahl, B.E., Zheng, W.: A Comprehensive Measure- ment of the Local Value of the Hubble Constant with 1 km s −1 Mpc−1 Uncertainty from t...

  68. [69]

    Planck Collaboration, Aghanim, N., Akrami, Y., Ashdown, M., Aumont, J., Baccigalupi, C., Ballardini, M., Banday, A.J., Barreiro, R.B., Bar- tolo, N., Basak, S., Battye, R., Benabed, K., Bernard, J.-P., Bersanelli, M., Bielewicz, P., Bock, J.J., Bond, J.R., Borrill, J., Bouchet, F.R., Boulanger, F., Bucher, M., Burigana, C., Butler, R.C., Calabrese, E., Ca...

  69. [70]

    Nature Reviews Physics2(1), 10–12 (2020) https://arxiv.org/abs/2001

    Riess, A.G.: The expansion of the Universe is faster than expected. Nature Reviews Physics2(1), 10–12 (2020) https://arxiv.org/abs/2001. 03624 [astro-ph.CO]. https://doi.org/10.1038/s42254-019-0137-0

  70. [71]

    Liddle, A.R.: An Introduction to Modern Cosmology, Third Edition, (2015)

  71. [72]

    Cosmology-informed neural networks to solve the background dynamics of the Universe

    Chantada, A.T., Landau, S.J., Protopapas, P., Sc´ occola, C.G., Gar- raffo, C.: Cosmology-informed neural networks to solve the background dynamics of the Universe. PRD107(6), 063523 (2023) https://arxiv. org/abs/2205.02945 [astro-ph.CO]. https://doi.org/10.1103/PhysRevD. 107.063523

  72. [73]

    Steiner, F.: Most frequent value procedures. Geophys. Trans.34(2-3), 139–260 (1988)

  73. [74]

    (ed.): Optimum Methods in Statistics

    Steiner, F. (ed.): Optimum Methods in Statistics. Akademia Kiado, Budapest, Hungary, ??? (1997)

  74. [75]

    Acta Geod Geophys.49, 95–104 (2014)

    Szegedi, H., Dobroka, M.: On the use of steiner weights in inversion- based fourier transformation: robustification of a previously published algorithm. Acta Geod Geophys.49, 95–104 (2014)

  75. [76]

    Szab´ o, N.P., Balogh, G.P., Stickel, J.: Most frequent value-based factor analysis of direct-push logging data. Geophys. Prospect.66(3), 530–548 Springer Nature 2021 LATEX template 40Article Title (2018). https://doi.org/10.1111/1365-2478.12573

  76. [77]

    MFV approach to robust estimate of neutron lifetime

    Zhang, J., Zhang, S., Zhang, Z.-R., Zhang, P., Li, W.-B., Hong, Y.: MFV approach to robust estimate of neutron lifetime. European Phys- ical Journal C82(12), 1106 (2022) https://arxiv.org/abs/2212.05890 [physics.data-an]. https://doi.org/10.1140/epjc/s10052-022-11071-9

  77. [78]

    Application of the Most Frequent Value Method for $^{39}$Ar Half-Life Determination

    Golovko, V.V.: Application of the most frequent value method for 3 9Ar half-life determination. European Physical Journal C83(10), 930 (2023) https://arxiv.org/abs/2310.06867 [physics.geo-ph]. https://doi.org/10. 1140/epjc/s10052-023-12113-6

  78. [79]

    arXiv e-prints, 2410–19988 (2024) https://arxiv.org/abs/2410.19988 [physics.data-an]

    Golovko, V.V.: Estimation of Ru-97 Half-Life Using the Most Fre- quent Value Method and Bootstrapping Techniques. arXiv e-prints, 2410–19988 (2024) https://arxiv.org/abs/2410.19988 [physics.data-an]. https://doi.org/10.48550/arXiv.2410.19988

  79. [80]

    MNRAS468, 5014–5019 (2017)

    Zhang, J.: Most frequent value statistics and distribution of 7Li abun- dance observations. MNRAS468, 5014–5019 (2017). https://doi.org/10. 1093/mnras/stx627

  80. [81]

    PASP130(990), 084502 (2018)

    Zhang, J.: Most Frequent Value Statistics and the Hubble Constant. PASP130(990), 084502 (2018). https://doi.org/10.1088/1538-3873/ aac767

Showing first 80 references.