Pith. sign in

REVIEW 3 major objections 71 references

Hard-wired shape constraints on a neural kernel beat soft penalties for recovering memory and nonlocal kernels from sparse, noisy data.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 06:59 UTC pith:ULALAVOH

load-bearing objection Hard Bernstein MC-KAN constraints beat soft Cheb-KAN penalties for sparse/noisy 2D kernel recovery within a clearly scoped positive-monotone-convex class; solid synthetic evidence, no hidden flaw. the 3 major comments →

arxiv 2607.11110 v1 pith:ULALAVOH submitted 2026-07-13 cs.LG cs.NAmath.NA

Neural Discovery of Memory and Nonlocal Kernels in Integro-Differential Equations with Constrained Kolmogorov--Arnold Networks

classification cs.LG cs.NAmath.NA
keywords integro-differential equationskernel discoveryKolmogorov–Arnold networksdifferentiable solversmemory kernelsnonlocal modelsinverse problemssymbolic regression
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Memory and nonlocal kernels in integro-differential equations are often unknown, and recovering them from sparse or noisy observations is an ill-posed inverse problem. This paper embeds a trainable Kolmogorov–Arnold network inside a differentiable numerical solver and optimizes the network so that the simulated field matches the data. Two ways of imposing positivity, monotone decrease, and convexity are compared: a Bernstein-based Monotone–Convex KAN that hard-enforces those properties through coefficient constraints, and a Chebyshev KAN that only soft-penalizes violations. On one-dimensional Volterra and viscoelastic-wave benchmarks both recover the correct closed-form kernels; on a sparse, noisy two-dimensional nonlocal reaction–diffusion problem the hard-constrained network yields substantially lower kernel error. The practical claim is that building the physical shape into the architecture itself is more reliable than hoping a penalty term will keep the optimizer honest when data are scarce and multidimensional.

Core claim

Enforcing positivity, monotonic decrease, and convexity by construction inside a Bernstein-polynomial Monotone–Convex KAN produces more accurate and stable recovery of multidimensional memory and nonlocal kernels from sparse, noisy observations than encouraging the same properties with soft penalty terms in a Chebyshev KAN, while both architectures recover the correct functional form on simpler one-dimensional problems.

What carries the argument

The Monotone–Convex KAN (MC-KAN): a layered Kolmogorov–Arnold network whose univariate edge functions are Bernstein polynomials whose coefficients are reparameterized so that positivity, monotone decrease, and convexity hold by construction and propagate through every layer (Appendix A).

Load-bearing premise

The unknown kernel is assumed from the start to be positive, monotonically decreasing, and convex along each lag; if the true kernel is oscillatory, sign-changing, or non-monotone, the hard-constraint architecture does not apply.

What would settle it

Replace the true kernel in the two-dimensional nonlocal reaction–diffusion experiment with an oscillatory or sign-changing kernel of comparable magnitude, retrain both networks under the same sparse noisy observation regime, and check whether MC-KAN still produces lower kernel reconstruction error than Cheb-KAN (or whether either recovers a usable kernel at all).

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • For dissipative memory kernels of the positive-monotone-convex class, a hard-constrained Bernstein KAN can be dropped into existing differentiable IDE solvers without problem-specific adjoints or specialized observation setups.
  • Symbolic regression applied after training can recover closed-form exponential or stretched-exponential laws even when the raw network output is only a pointwise approximation.
  • On sparse multidimensional inverse problems the gap between hard and soft constraint enforcement widens with noise, so architectural constraints become the preferred regularizer when data are scarce.
  • The same coefficient-reparameterization idea can be reused for other univariate shape constraints once the Bernstein (or analogous) basis is fixed.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same hard-constraint pattern could be tried on other completely-monotone kernels (Mittag–Leffler, multi-exponential Prony series) that appear in fractional viscoelasticity and anomalous diffusion without changing the solver infrastructure.
  • When the true kernel is known a priori to violate the shape assumptions, the soft-penalty Cheb-KAN (or an unconstrained KAN) remains the only option inside this framework, suggesting a hybrid switch based on a cheap pre-test of monotonicity.
  • Extending the staged slice-plus-residual symbolic procedure to fully coupled multivariate expressions would close the largest practical gap left open by the two-dimensional experiment.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper proposes a discretize-then-optimize framework for recovering memory and nonlocal kernels in integro-differential equations from sparse/noisy spatiotemporal data. The unknown kernel is parameterized by a constrained KAN inside a differentiable IDE solver; two variants are compared: MC-KAN (Bernstein edges with coefficient constraints that enforce positivity, monotone decrease, and convexity by construction) and Cheb-KAN (same shape properties via soft penalties). After training, PySR is used to extract closed-form expressions. Three synthetic benchmarks are studied: a 1D Volterra IDE (exponential kernel, noise sweep), a 1D viscoelastic wave PIDE (KWW kernel, temporal sparsity), and a 2D nonlocal reaction–diffusion equation (anisotropic coupled stretched-exponential kernel, sparsity+noise). On 1D problems both methods recover the correct functional form with comparable solution accuracy; on the sparse/noisy 2D problem MC-KAN consistently yields lower kernel L2 error (e.g., 12.36% vs 21.95% at σ=0.15 with 7 snapshots on a 32×32 grid). Appendix A proves that the MC-KAN composition preserves the axis-wise shape constraints.

Significance. Kernel identification in IDEs is a classical ill-posed inverse problem; a general, differentiable-solver approach that does not require case-specific adjoints or specialized observations is of clear interest to SciML and nonlocal continuum modeling. The hard-vs-soft comparison is carefully controlled (matched parameter counts, multi-seed means±std, explicit λ calibration), and the composition proof for MC-KAN (Theorem A.3) is a genuine technical contribution. Within the stated class of positive, monotone-decreasing, convex kernels the empirical claim is well supported and the 2D robustness gap is substantial. Limitations (synthetic data only, architecture scoped to that shape class, incomplete recovery of the 2D coupling term) are acknowledged by the authors and do not erase the value of the controlled comparison.

major comments (3)
  1. The central robustness claim is scoped to kernels satisfying Eq. (4) (positivity, monotone decrease, convexity along each lag). Table 1 and Appendix A build the entire hard-constraint apparatus only for that class; §4 correctly notes that oscillatory or sign-changing kernels are out of scope. The abstract and introduction should state this premise as a modeling assumption rather than as a universal advantage of hard constraints, so that the 44% 2D gap is not over-read as applying to arbitrary nonlocal kernels.
  2. All three benchmarks use synthetic data generated from known analytical kernels (exponential, KWW, anisotropic stretched-exp with coupling). The inverse problem is therefore well-specified and free of model-form error. A short discussion or additional experiment on misspecification (e.g., true kernel outside the MC-KAN shape class, or mild model error in the local operators) would strengthen the claim that hard constraints remain advantageous under realistic conditions; without it the transferability argument rests only on the synthetic suite.
  3. In Experiment III the 2D coupling residual is recovered only approximately (δC_rec ≈ 0.156 τ1^0.481 τ2 vs true 0.289 τ1^0.5 τ2^0.5, relative L2 0.20). The staged slice-plus-residual PySR procedure is pragmatic but incomplete; the paper should either improve the multivariate symbolic step or clearly qualify that closed-form recovery of fully coupled multidimensional kernels remains open.

Circularity Check

0 steps flagged

No significant circularity: kernel recovery is optimized against held-out solution data under explicit shape priors; self-citations are background only.

full rationale

The paper's central claim is an empirical comparison (hard Bernstein MC-KAN vs soft-penalty Cheb-KAN) on synthetic IDE benchmarks, with kernels recovered by minimizing data mismatch of a differentiable forward solver and evaluated against held-out ground-truth kernels (E(K), Tables 2–6). Shape constraints (Eq. 4) are stated domain priors for fading-memory kernels, enforced by construction in MC-KAN (Table 1, Thm. A.3) or by soft penalties; they restrict the hypothesis class but do not tautologically force the recovered exponential/KWW forms or the 2D error gap. Symbolic regression is post-hoc. Self-citations (e.g., Faroughi-group KAN/SciML papers) appear only as related-work motivation and are not load-bearing for uniqueness, uniqueness theorems, or the reported robustness numbers. The derivation chain is therefore self-contained against external synthetic benchmarks; the only minor note is ordinary author-overlap citation of prior KAN tooling, which does not reduce any prediction to its inputs by construction.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 1 invented entities

The method assumes the IDE structure (operators T, N, F, B, source) is known and only K is unknown; imposes a specific shape class on K; and introduces MC-KAN as the hard-constraint vehicle. Free knobs (λ, N, M, α, architecture widths, optimizer schedules) are tuned per experiment. No new physical particles or forces—only an architectural entity and standard inverse-problem priors.

free parameters (5)
  • Cheb-KAN constraint weight λ = 0.1 / 0.3 / 0.01 by experiment
    Chosen per benchmark by short calibration so initial constraint penalty matches data-loss order (Exp. I λ=0.1; II λ=0.3; III λ=0.01). Affects soft-constraint baseline fairness.
  • Bernstein degree N and Chebyshev degree M = N=6, M=3
    N=6 and M=3 fixed empirically for capacity vs parameter count; not derived from theory.
  • Power-law normalization exponents α_i = 0.75 (fixed after noise-free fit)
    In Exp. III, α1=α2=0.75 fixed after noise-free optimization of learnable α; changes near-origin resolution and reported E(K).
  • Network widths/depths and Adam/L-BFGS schedules
    Architectures chosen to match parameter counts; learning rates and epoch budgets differ by experiment (e.g., Exp. III Adam LR 1e-1 MC-KAN vs 1e-3 Cheb-KAN). Hand-tuned training hyperparameters.
  • Kernel amplitude normalization K(0+)=1 (Exp. II) = K(0+)=1
    Instantaneous modulus fixed by dividing network output by value at τ_min=1e-3 to remove amplitude degeneracy; shape-only identification by design.
axioms (5)
  • domain assumption The governing IDE form (Eq. 1) including operators T, N, F, B and source S is known; only the kernel K is unknown.
    Stated in §2.1 problem setup; identification is kernel-only, not full equation discovery.
  • domain assumption Target kernels are positive, monotonically decreasing, and convex along each lag dimension (Eq. 4).
    Motivated by fading-memory viscoelasticity/anomalous diffusion; required for MC-KAN guarantees and soft penalties.
  • standard math Bernstein coefficient finite-difference conditions (Table 1) suffice for positivity, monotone decrease, and convexity of edge functions, and compose through the network (Appendix A).
    Classical Bernstein basis properties plus composition lemmas; proved axis-wise, not joint convexity.
  • domain assumption Discretize-then-optimize gradients through the chosen numerical solver equal the training signal for the continuous inverse problem.
    Standard SciML assumption; solver matches reference discretization per experiment.
  • ad hoc to paper Synthetic observations are generated from known analytical kernels (exponential, KWW, anisotropic stretched-exp with coupling) plus optional Gaussian noise.
    All three experiments use manufactured data; no experimental measurements validate the pipeline.
invented entities (1)
  • Monotone–Convex KAN (MC-KAN) no independent evidence
    purpose: Parameterize unknown memory/nonlocal kernels with Bernstein edges whose coefficients enforce positivity, monotone decrease, and convexity by construction.
    Architectural construct introduced in §2.2.2; independent evidence is only the synthetic recovery experiments and the composition proof, not external physical measurement of MC-KAN itself.

pith-pipeline@v1.1.0-grok45 · 34856 in / 3697 out tokens · 37108 ms · 2026-07-14T06:59:01.342350+00:00 · methodology

0 comments
read the original abstract

Discovering the memory or nonlocal kernel governing an integro-differential equation (IDE) from sparse and noisy observations is an ill-posed inverse problem. Existing identification methods often rely on problem-specific analytical derivations, specialized observation requirements, or restrictive assumptions about the kernel, limiting their applicability across different classes of IDEs. In this work, we propose a differentiable-solver-based framework for discovering memory and nonlocal kernels directly from spatiotemporal observations. Within the solver, the unknown kernel is represented using a constrained Kolmogorov--Arnold Network (KAN) parameterization, with the physical constraints imposed through two different approaches: a Bernstein-polynomial-based Monotone--Convex KAN (MC-KAN), whose coefficient constraints enforce positivity, monotonic decrease, and convexity by construction, and a Chebyshev-based KAN (Cheb-KAN), in which the same properties are encouraged through soft penalty terms. After training, symbolic regression is applied to the learned kernels to obtain interpretable closed-form representations. We evaluate both methods on benchmarks spanning a one-dimensional Volterra equation, a one-dimensional viscoelastic wave partial integro-differential equation, and a two-dimensional nonlocal reaction-diffusion equation with an anisotropic coupled kernel. For the 1D problems, both methods recover the correct kernel functional form and achieve comparable solution-reconstruction accuracy. In contrast, for the sparse and noisy 2D nonlocal problem, the hard-constrained MC-KAN consistently achieves lower kernel reconstruction errors than the soft-constrained Cheb-KAN. Our results demonstrate that enforcing physically motivated shape constraints by construction provides greater robustness than soft penalties for multidimensional kernel discovery from sparse and noisy observations.

Figures

Figures reproduced from arXiv: 2607.11110 by Aruzhan Tleubek, Salah A Faroughi.

Figure 1
Figure 1. Figure 1: Schematic of the neural-network-embedded kernel discovery framework. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: At all noise levels, both methods closely follow the reference exponential kernel. The hard-constrained [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Recovered kernels in Experiment I under increasing noise, showing the best-performing architecture (lowest [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Recovered kernels in Experiment II under temporal sparsity (noise-free). Rows correspond to MC-KAN (top) and [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Results for Experiment III, showing the true kernel [PITH_FULL_IMAGE:figures/full_fig_p013_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Results for Experiment III, showing the recovered kernel along the axis slices [PITH_FULL_IMAGE:figures/full_fig_p013_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: (a) Relative kernel error E(K) versus noise for three MC-KAN input normalizations (linear, learnable power-law, and fixed α1 = α2 = 0.75) and the Cheb-KAN baseline. The training data consist of seven temporal snapshots on a 32 × 32 spatial grid. Markers denote the mean over five noise realizations with fixed initialization, and shaded bands indicate ± one standard deviation. (b)–(c) Recovered kernel along … view at source ↗
Figure 8
Figure 8. Figure 8: Results for Experiment III, showing the recovered kernel along the axis slices [PITH_FULL_IMAGE:figures/full_fig_p015_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

71 extracted references · 8 linked inside Pith

  1. [1]

    Solution of a partial integro-differential equation arising from viscoelasticity.Interna- tional Journal of Computer Mathematics, 83(1):123–129, 2006

    Mehdi Dehghan. Solution of a partial integro-differential equation arising from viscoelasticity.Interna- tional Journal of Computer Mathematics, 83(1):123–129, 2006

  2. [2]

    Asymptotic stability in viscoelasticity.Archive for rational mechanics and analysis, 37(4):297–308, 1970

    Constantine M Dafermos. Asymptotic stability in viscoelasticity.Archive for rational mechanics and analysis, 37(4):297–308, 1970. 16

  3. [3]

    Diffusion with memory in two cases of biological interest.Journal of theoretical biology, 254(3):697–703, 2008

    Michele Caputo and Cesare Cametti. Diffusion with memory in two cases of biological interest.Journal of theoretical biology, 254(3):697–703, 2008

  4. [4]

    The random walk’s guide to anomalous diffusion: a fractional dynamics approach.Physics reports, 339(1):1–77, 2000

    Ralf Metzler and Joseph Klafter. The random walk’s guide to anomalous diffusion: a fractional dynamics approach.Physics reports, 339(1):1–77, 2000

  5. [5]

    Anomalous diffusion models and their properties: non-stationarity, non-ergodicity, and ageing at the centenary of single particle tracking

    Ralf Metzler, Jae-Hyung Jeon, Andrey G Cherstvy, and Eli Barkai. Anomalous diffusion models and their properties: non-stationarity, non-ergodicity, and ageing at the centenary of single particle tracking. Physical Chemistry Chemical Physics, 16(44):24128–24164, 2014

  6. [6]

    A general theory of heat conduction with finite wave speeds

    Morton E Gurtin and Allen C Pipkin. A general theory of heat conduction with finite wave speeds. Archive for Rational Mechanics and Analysis, 31(2):113–126, 1968

  7. [7]

    On heat conduction in materials with memory.Quarterly of Applied Mathematics, 29 (2):187–204, 1971

    Jace W Nunziato. On heat conduction in materials with memory.Quarterly of Applied Mathematics, 29 (2):187–204, 1971

  8. [8]

    Fast and oblivious convolution quadrature

    Achim Schädle, María López-Fernández, and Christian Lubich. Fast and oblivious convolution quadrature. SIAM Journal on Scientific Computing, 28(2):421–438, 2006

  9. [9]

    A kernel-independent sum-of-exponentials method.Journal of Scientific Computing, 93(2):40, 2022

    Zixuan Gao, Jiuyang Liang, and Zhenli Xu. A kernel-independent sum-of-exponentials method.Journal of Scientific Computing, 93(2):40, 2022

  10. [10]

    Optimal prediction and the mori–zwanzig rep- resentation of irreversible processes.Proceedings of the National Academy of Sciences, 97(7):2968–2973, 2000

    Alexandre J Chorin, Ole H Hald, and Raz Kupferman. Optimal prediction and the mori–zwanzig rep- resentation of irreversible processes.Proceedings of the National Academy of Sciences, 97(7):2968–2973, 2000

  11. [11]

    A priori estimation of memory effects in reduced- ordermodelsofnonlinearsystemsusingthemori–zwanzigformalism.Proceedings

    Ayoub Gouasmi, Eric J Parish, and Karthik Duraisamy. A priori estimation of memory effects in reduced- ordermodelsofnonlinearsystemsusingthemori–zwanzigformalism.Proceedings. Mathematical, Physical, and Engineering Sciences, 473(2205):20170385, 2017

  12. [12]

    Identification of a relaxation kernel using two boundary measures.arXiv preprint arXiv:1503.03883, 2015

    Luciano Pandolfi. Identification of a relaxation kernel using two boundary measures.arXiv preprint arXiv:1503.03883, 2015

  13. [13]

    Memory kernel reconstruction problems in the integro-differential equation of rigid heat conductor.Mathematical Methods in the Applied Sciences, 45 (14):8374–8388, 2022

    Durdimurod K Durdiev and Zhonibek Zh Zhumaev. Memory kernel reconstruction problems in the integro-differential equation of rigid heat conductor.Mathematical Methods in the Applied Sciences, 45 (14):8374–8388, 2022

  14. [14]

    Kernel determination problem for a integro–differential heat equation with a variable thermal conductivity

    Durdimurod Durdiev and Javlon Nuriddinov. Kernel determination problem for a integro–differential heat equation with a variable thermal conductivity. InAIP Conference Proceedings, volume 3004, page 040014. AIP Publishing LLC, 2024

  15. [15]

    Inverse problems for memory kernels by laplace transform methods.Zeitschrift für Analysis und ihre Anwendungen, 19(2):489–510, 2000

    Jaan Janno and Lothar von Wolfersdorf. Inverse problems for memory kernels by laplace transform methods.Zeitschrift für Analysis und ihre Anwendungen, 19(2):489–510, 2000

  16. [16]

    Identifying memory kernels in linear thermoviscoelasticity of boltzmann type.Mathematical Models and Methods in Applied Sciences, 4(06):807–842, 1994

    Cecilia Cavaterra and Maurizio Grasselli. Identifying memory kernels in linear thermoviscoelasticity of boltzmann type.Mathematical Models and Methods in Applied Sciences, 4(06):807–842, 1994

  17. [17]

    A survey of regularization methods for first-kind volterra equations

    Patricia K Lamm. A survey of regularization methods for first-kind volterra equations. InSurveys on solution methods for inverse problems, pages 53–82. Springer, 2000

  18. [18]

    Identification of the relaxation kernel in diffusion processes and viscoelasticity with memory via deconvolution.Mathematical Methods in the Applied Sciences, 40(7):2542–2549, 2017

    Luciano Pandolfi. Identification of the relaxation kernel in diffusion processes and viscoelasticity with memory via deconvolution.Mathematical Methods in the Applied Sciences, 40(7):2542–2549, 2017

  19. [19]

    Application of prony series to linear viscoelasticity

    JE Soussou, F Moavenzadeh, and MH Gradowczyk. Application of prony series to linear viscoelasticity. Transactions of the Society of Rheology, 14(4):573–584, 1970

  20. [20]

    Brunton, Joshua L

    Steven L. Brunton, Joshua L. Proctor, and J. Nathan Kutz. Discovering governing equations from data by sparse identification of nonlinear dynamical systems.Proceedings of the National Academy of Sciences, 113(15):3932–3937, 2016. DOI: 10.1073/pnas.1517384113

  21. [21]

    Neural ordinary differential equations.Advances in neural information processing systems, 31, 2018

    Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud. Neural ordinary differential equations.Advances in neural information processing systems, 31, 2018

  22. [22]

    Rudy, Steven L

    Samuel H. Rudy, Steven L. Brunton, Joshua L. Proctor, and J. Nathan Kutz. Data-driven discovery of partial differential equations.Science Advances, 3(4):e1602614, 2017. DOI: 10.1126/sciadv.1602614

  23. [23]

    Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators.Nature Machine Intelligence, 3(3):218–229, 2021

    Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis. Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators.Nature Machine Intelligence, 3(3):218–229, 2021. DOI: 10.1038/s42256-021-00302-5. 17

  24. [24]

    Symbolic–kan: Kolmogorov-arnold networks with discrete symbolic structure for interpretable learning.arXiv preprint arXiv:2603.23854, 2026

    Salah A Faroughi, Farinaz Mostajeran, Amirhossein Arzani, and Shirko Faroughi. Symbolic–kan: Kolmogorov-arnold networks with discrete symbolic structure for interpretable learning.arXiv preprint arXiv:2603.23854, 2026

  25. [25]

    Benjamin C Koenig, Suyong Kim, and Sili Deng. Kan-odes: Kolmogorov–arnold network ordinary dif- ferential equations for learning dynamical systems and hidden physics.Computer Methods in Applied Mechanics and Engineering, 432:117397, 2024

  26. [26]

    Minpo: Memory-informed neural pseudo- operator to resolve nonlocal spatiotemporal dynamics.arXiv preprint arXiv:2512.17273, 2025

    Farinaz Mostajeran, Aruzhan Tleubek, and Salah A Faroughi. Minpo: Memory-informed neural pseudo- operator to resolve nonlocal spatiotemporal dynamics.arXiv preprint arXiv:2512.17273, 2025

  27. [27]

    Physics informed neural networks for an inverse problem in peridynamic models.Engineering with Computers, 41(6):4003–4012, 2025

    Fabio V Difonzo, Luciano Lopez, and Sabrina F Pellegrino. Physics informed neural networks for an inverse problem in peridynamic models.Engineering with Computers, 41(6):4003–4012, 2025

  28. [28]

    Learning memory kernels in second-order volterra models: a data-driven approach for viscoelastic systems.Mechanics of Time-Dependent Materials, 30(2): 42, 2026

    Yacine Khaldi, Meriem Belaifa, and Amir Benzaoui. Learning memory kernels in second-order volterra models: a data-driven approach for viscoelastic systems.Mechanics of Time-Dependent Materials, 30(2): 42, 2026

  29. [29]

    Sparse identification of delay equations with distributed memory.arXiv preprint arXiv:2512.21070, 2025

    Dimitri Breda, Muhammad Tanveer, and Jianhong Wu. Sparse identification of delay equations with distributed memory.arXiv preprint arXiv:2512.21070, 2025

  30. [30]

    Jose A Carrillo, Gissell Estrada-Rodriguez, Laszlo Mikolas, and Sui Tang. Sparse identification of nonlocal interaction kernels in nonlinear gradient flow equations via partial inversion.Mathematical Models and Methods in Applied Sciences, 35(05):1073–1131, 2025

  31. [31]

    Physics-guided symbolic regression of nonlocal transport memory from sparse observations in aquatic systems.Preprint, available at SSRN 6404003, 2026

    Xiangnan Yu, Yong Zhang, HongGuang Sun, Zhibo Chen, and Yuntian Chen. Physics-guided symbolic regression of nonlocal transport memory from sparse observations in aquatic systems.Preprint, available at SSRN 6404003, 2026

  32. [32]

    Neural integro-differential equations

    Emanuele Zappala, Antonio H de O Fonseca, Andrew H Moberly, Michael J Higley, Chadi Abdallah, Jessica A Cardin, and David van Dijk. Neural integro-differential equations. InProceedings of the aaai conference on artificial intelligence, volume 37, pages 11104–11112, 2023

  33. [33]

    Learning integral operators via neural integral equations.Nature Machine Intelligence, 6(9):1046–1062, 2024

    Emanuele Zappala, Antonio Henrique de Oliveira Fonseca, Josue Ortega Caro, Andrew Henry Moberly, Michael James Higley, Jessica Cardin, and David van Dijk. Learning integral operators via neural integral equations.Nature Machine Intelligence, 6(9):1046–1062, 2024

  34. [34]

    ANODE: Unconditionally accurate memory-efficient gra- dients for neural ODEs

    Amir Gholami, Kurt Keutzer, and George Biros. ANODE: Unconditionally accurate memory-efficient gra- dients for neural ODEs. InProceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI), 2019

  35. [35]

    Discretize-optimize vs

    Derek Onken and Lars Ruthotto. Discretize-optimize vs. optimize-discretize for time-series regression and continuous normalizing flows.arXiv preprint arXiv:2005.13420, 2020

  36. [36]

    Heinz Werner Engl, Martin Hanke, and Andreas Neubauer.Regularization of inverse problems, volume

  37. [37]

    Springer Science & Business Media, 1996

  38. [38]

    A simple proof of a duality theorem with applications in scalar and anisotropic vis- coelasticity.arXiv preprint arXiv:1805.07275, 2018

    Andrzej Hanyga. A simple proof of a duality theorem with applications in scalar and anisotropic vis- coelasticity.arXiv preprint arXiv:1805.07275, 2018

  39. [39]

    On complete monotonicity of solution to the fractional relaxation equation with the n th level fractional derivative.Mathematics, 8(9):1561, 2020

    Yuri Luchko. On complete monotonicity of solution to the fractional relaxation equation with the n th level fractional derivative.Mathematics, 8(9):1561, 2020

  40. [40]

    Wave propagation in linear viscoelastic media with completely monotonic relaxation moduli.Wave Motion, 50(5):909–928, 2013

    Andrzej Hanyga. Wave propagation in linear viscoelastic media with completely monotonic relaxation moduli.Wave Motion, 50(5):909–928, 2013

  41. [41]

    Chebyshev polynomial-based kolmogorov-arnold networks: An efficient architecture for nonlinear function approximation.arXiv preprint arXiv:2405.07200, 2024

    Sidharth SS, Keerthana AR, Anas KP, et al. Chebyshev polynomial-based kolmogorov-arnold networks: An efficient architecture for nonlinear function approximation.arXiv preprint arXiv:2405.07200, 2024

  42. [42]

    Interpretable machine learning for science with PySR and SymbolicRegression.jl.arXiv preprint arXiv:2305.01582, 2023

    Miles Cranmer. Interpretable machine learning for science with PySR and SymbolicRegression.jl.arXiv preprint arXiv:2305.01582, 2023. DOI: 10.48550/arXiv.2305.01582

  43. [43]

    Khemraj Shukla, Juan Diego Toscano, Zhicheng Wang, Zongren Zou, and George Em Karniadakis. A comprehensive and fair comparison between mlp and kan representations for differential equations and operator networks.Computer Methods in Applied Mechanics and Engineering, 431:117290, 2024

  44. [44]

    Kolmogorov- arnold networks for data-driven, physics-informed, and deep-operator learning: a review, synthesis, and new analysis.Neural networks, page 108791, 2026

    Salah A Faroughi, Farinaz Mostajeran, Amin Hamed Mashhadzadeh, and Shirko Faroughi. Kolmogorov- arnold networks for data-driven, physics-informed, and deep-operator learning: a review, synthesis, and new analysis.Neural networks, page 108791, 2026. 18

  45. [45]

    Neural tangent kernel analysis to probe convergence in physics- informed neural solvers: Pikans vs

    Salah A Faroughi and Farinaz Mostajeran. Neural tangent kernel analysis to probe convergence in physics- informed neural solvers: Pikans vs. pinns.Computers & Mathematics with Applications, 215:155–191, 2026

  46. [46]

    Scaled-cpikans: Spatial variable and residual scaling in chebyshev-based physics-informed kolmogorov-arnold networks.Journal of Computational Physics, 537: 114116, 2025

    Farinaz Mostajeran and Salah A Faroughi. Scaled-cpikans: Spatial variable and residual scaling in chebyshev-based physics-informed kolmogorov-arnold networks.Journal of Computational Physics, 537: 114116, 2025

  47. [47]

    Kan or mlp: A fairer comparison.arXiv preprint arXiv:2407.16674, 2024

    Runpeng Yu, Weihao Yu, and Xinchao Wang. Kan or mlp: A fairer comparison.arXiv preprint arXiv:2407.16674, 2024

  48. [48]

    Salah A Faroughi, Nikhil M Pawar, Celio Fernandes, Maziar Raissi, Subasish Das, Nima K Kalantari, and Seyed Kourosh Mahjour. Physics-guided, physics-informed, and physics-encoded neural networks and operators in scientific computing: Fluid and solid mechanics.Journal of Computing and Information Science in Engineering, 24(4):040802, 2024

  49. [49]

    Input convex neural networks

    Brandon Amos, Lei Xu, and J Zico Kolter. Input convex neural networks. InInternational conference on machine learning, pages 146–155. PMLR, 2017

  50. [50]

    Monotonic networks.Advances in neural information processing systems, 10, 1997

    Joseph Sill. Monotonic networks.Advances in neural information processing systems, 10, 1997

  51. [51]

    Prakash Thakolkaran, Yaqi Guo, Shivam Saini, Mathias Peirlinck, Benjamin Alheit, and Siddhant Ku- mar. Can KAN CANs? Input-convex Kolmogorov-Arnold networks (KANs) as hyperelastic constitutive artificial neural networks (CANs).Computer Methods in Applied Mechanics and Engineering, 443:118089,

  52. [52]

    DOI: 10.1016/j.cma.2025.118089

  53. [53]

    Input convex kolmogorov arnold networks.arXiv preprint arXiv:2505.21208, 2026

    Thomas Deschatre and Xavier Warin. Input convex kolmogorov arnold networks.arXiv preprint arXiv:2505.21208, 2026

  54. [54]

    The bernstein polynomial basis: A centennial retrospective.Computer Aided Geometric Design, 29(6):379–419, 2012

    Rida T Farouki. The bernstein polynomial basis: A centennial retrospective.Computer Aided Geometric Design, 29(6):379–419, 2012

  55. [55]

    Kan: Kolmogorov–arnold networks

    Ziming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle, James Halverson, Marin Soljacic, Thomas Hou, and Max Tegmark. Kan: Kolmogorov–arnold networks. InInternational conference on learning representations, volume 2025, pages 70367–70413, 2025

  56. [56]

    Epi-ckans: Elasto-plasticity informed kolmogorov-arnold networks using chebyshev polynomials.arXiv preprint arXiv:2410.10897, 2024

    Farinaz Mostajeran and Salah A Faroughi. Epi-ckans: Elasto-plasticity informed kolmogorov-arnold networks using chebyshev polynomials.arXiv preprint arXiv:2410.10897, 2024

  57. [57]

    Physics-informed kolmogorov– arnold network with chebyshev polynomials for fluid mechanics.Physics of Fluids, 37(9), 2025

    Chunyu Guo, Lucheng Sun, Shilong Li, Zelong Yuan, and Chao Wang. Physics-informed kolmogorov– arnold network with chebyshev polynomials for fluid mechanics.Physics of Fluids, 37(9), 2025

  58. [58]

    Auto- matic differentiation in machine learning: a survey.Journal of machine learning research, 18(153):1–43, 2018

    Atilim Gunes Baydin, Barak A Pearlmutter, Alexey Andreyevich Radul, and Jeffrey Mark Siskind. Auto- matic differentiation in machine learning: a survey.Journal of machine learning research, 18(153):1–43, 2018

  59. [59]

    Pytorch: An imperative style, high- performance deep learning library.Advances in neural information processing systems, 32, 2019

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high- performance deep learning library.Advances in neural information processing systems, 32, 2019

  60. [60]

    Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980, 2014

    Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980, 2014

  61. [61]

    On the limited memory bfgs method for large scale optimization.Math- ematical programming, 45(1):503–528, 1989

    Dong C Liu and Jorge Nocedal. On the limited memory bfgs method for large scale optimization.Math- ematical programming, 45(1):503–528, 1989

  62. [62]

    Line search algorithms with guaranteed sufficient decrease.ACM Transactions on Mathematical Software (TOMS), 20(3):286–307, 1994

    Jorge J Moré and David J Thuente. Line search algorithms with guaranteed sufficient decrease.ACM Transactions on Mathematical Software (TOMS), 20(3):286–307, 1994

  63. [63]

    Cambridge University Press, 2017

    Hermann Brunner.Volterra integral equations: an introduction to theory and applications, volume 30. Cambridge University Press, 2017

  64. [64]

    Springer Science & Business Media, 2013

    Jim M Cushing.Integrodifferential equations and delay models in population dynamics. Springer Science & Business Media, 2013

  65. [65]

    World Scientific, 2022

    Francesco Mainardi.Fractional calculus and waves in linear viscoelasticity: an introduction to mathemat- ical models. World Scientific, 2022

  66. [66]

    Non-symmetrical dielectric relaxation behaviour arising from a simple empirical decay function.Transactions of the Faraday society, 66:80–85, 1970

    Graham Williams and David C Watts. Non-symmetrical dielectric relaxation behaviour arising from a simple empirical decay function.Transactions of the Faraday society, 66:80–85, 1970. 19

  67. [67]

    Detailed comparison of the williams–watts and cole– davidson functions.The Journal of chemical physics, 73(7):3348–3357, 1980

    Christopher P Lindsey and Gary D Patterson. Detailed comparison of the williams–watts and cole– davidson functions.The Journal of chemical physics, 73(7):3348–3357, 1980

  68. [68]

    Traveling waves in a convolution model for phase transitions.Archive for Rational Mechanics and Analysis, 138(2):105–136, 1997

    Peter W Bates, Paul C Fife, Xiaofeng Ren, and Xuefeng Wang. Traveling waves in a convolution model for phase transitions.Archive for Rational Mechanics and Analysis, 138(2):105–136, 1997

  69. [69]

    Numerical analysis for a nonlocal allen-cahn equation

    Peter W Bates, Sarah Brown, and Jianlong Han. Numerical analysis for a nonlocal allen-cahn equation. International Journal of Numerical Analysis and Modeling, 6(1):33–49, 2009

  70. [70]

    A nonlocal continuum model for biological aggregation.Bulletin of mathematical biology, 68(7):1601–1623, 2006

    Chad M Topaz, Andrea L Bertozzi, and Mark A Lewis. A nonlocal continuum model for biological aggregation.Bulletin of mathematical biology, 68(7):1601–1623, 2006

  71. [71]

    Swarm dynamics and equilibria for a nonlocal aggregation model.Nonlinearity, 24(10):2681–2716, 2011

    Razvan C Fetecau, Yanghong Huang, and Theodore Kolokolnikov. Swarm dynamics and equilibria for a nonlocal aggregation model.Nonlinearity, 24(10):2681–2716, 2011. 20 Appendix A. Theoretical Guarantee We now show that the kernel constraints introduced in Section 2.2.2 are preserved throughout the full depth of MC-KAN. Since the architecture enforces monoton...