{"id":"82c180b9-b02b-4a09-80a0-9ae6e27378e1","arxiv_id":"2504.12580","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"ChemKANs, a physics-structured KAN-ODE network, infer chemical kinetic models from noisy data and accelerate hydrogen-air combustion chemistry with 344 parameters at 2x speedup.","lead":"A new neural network architecture called ChemKAN is trained to replace the stiff chemistry calculations inside combustion simulations. In tests on hydrogen-air ignition and biodiesel kinetics, it matched the detailed chemistry with a compact network and a modest speedup, while resisting noisy data better than a standard deep learning baseline.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central unsupported leap is from 0-D homogeneous-reactor accuracy/speedup to operator-split use in multidimensional reacting flows; the paper's own Fig. 8(A) also shows the low-temperature generalization that would be needed is fragile.","rationale":"In good faith, the paper's 0-D results are internally consistent: the ChemKAN architecture encodes the kinetic/thermodynamic split of Eqs. 1-2, the biodiesel noise study uses an appropriate noise-free metric and shows a real robustness advantage over DeepONet, and the H2 case reports concrete parameter counts, a held-out condition, and a measured 2.0x homogeneous-reactor speedup. These results support a narrower claim: a 344-parameter ChemKAN can approximate GRI-Mech 3.0 source terms for homogeneous autoignition on the sampled initial-condition grid. The load-bearing gap is the leap from that narrow result to the abstract's 'generalizable to larger-scale turbulent flow simulations.' A source-term surrogate can in principle be lifted to operator-split CFD, but this requires the surrogate to remain accurate on states visited in flames, at the operating pressure actually used, and with transport effects present; none of that is demonstrated. The paper's own Fig. 8(A) shows degraded accuracy at 987.5 K, inside the nominal temperature range, which specifically threatens low-temperature ignition behavior in reacting-flow applications. Because the issue is a missing demonstration rather than a demonstrated contradiction, the reader's CONDITIONAL verdict remains appropriate; I would not move to ACCEPT or REJECT. The main adjustment I would impose is to soften or explicitly qualify the turbulence-generalization claim until a multidimensional or at least 1-D flame test is provided.","tokens_in":21119,"tokens_out":7197,"duration_ms":83869,"concrete_test":"Run a 1-D freely propagating laminar H2/air flame at 1 atm with the trained 344-parameter ChemKAN substituted for the GRI-Mech 3.0 source term under operator splitting, and compare laminar flame speed, temperature/species profiles, and cell-source-evaluation wall time against the detailed-chemistry solution. If the flame speed differs by more than about 5%, or if the source-term error on flame-sampled states is materially worse than the Fig. 8 training-condition errors, the transfer premise fails and the turbulence-generalization claim should be softened to a potential rather than a demonstrated capability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's most consequential premise is that a source-term surrogate trained on homogeneous-reactor trajectories can be coupled to flow solvers for multidimensional combustion. This is asserted in Sec. I ('In the operator splitting regime, such a surrogate can be directly coupled to existing CFD...') and again in Sec. III B ('generalizing to other simulation conditions when coupled to flow solvers, including simple laminar flames and complex 2-D and 3-D turbulent combustion conditions'), but it is never tested. The reported 2.0x speedup was measured only for 36 homogeneous reactor integrations in Arrhenius.jl; it does not establish cell-wise wall-clock behavior in a flame solver, where transport, operator splitting, and state distributions differ. The transfer precondition is that flame states lie on the thermochemical manifold covered by the 0-D training set, and the paper's own Fig. 8(A) weakens this: at 987.5 K, an unseen condition inside the stated 950-1200 K range, errors are roughly an order of magnitude worse than at nearby training conditions. The abstract's claim that the solver is 'generalizable to larger-scale turbulent flow simulations' is therefore an extrapolation, not a demonstrated result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ChemKANs, a variant of Kolmogorov-Arnold Network ordinary differential equations (KAN-ODEs) in which the gradient getter is split into a kinetic core and an optional thermodynamic superstructure, mirroring the structure of Eqs. (1) and (2). Two applications are studied: (i) inference of a biodiesel transesterification kinetic model from synthetic noisy data, compared against DeepONets, and (ii) acceleration of zero-dimensional hydrogen-air homogeneous reactor simulations, where a single 344-parameter ChemKAN is reported to reproduce all nine species and temperature profiles and to run about 2x faster than the detailed GRI-Mech-based chemistry in Arrhenius.jl. The paper claims robustness to added noise, resistance to overfitting, and generalizability of the learned source-term surrogate to larger-scale flow simulations.","tokens_in":21328,"tokens_out":3667,"duration_ms":40877,"significance":"If the results hold, the architecture is a meaningful contribution to scientific machine learning for combustion: it encodes known kinetic/thermodynamic coupling in the network structure, uses a single information-sharing network instead of ChemNODE's m+1 separate networks, and demonstrates strong parameter efficiency. The two-stage training procedure and forward sensitivity analysis are practical and clearly described. The paper also provides a useful demonstration that KAN-ODEs can be applied to stiff, multi-species systems rather than only small dynamical systems. However, the claimed generality to multidimensional turbulent combustion and the quantitative speedup are not demonstrated by the experiments, which are limited to homogeneous reactors; the paper's own Fig. 8(A) reveals fragile generalization at unseen low-temperature conditions.","major_comments":[{"comment":"The central acceleration claim is extrapolated from zero-dimensional homogeneous reactors to multidimensional operator-split flow solvers without any demonstration. The Introduction states that 'such a surrogate can be directly coupled to existing CFD or machine learning-based flow solvers', and Sec. III B asserts generalizability 'to other simulation conditions when coupled to flow solvers, including simple laminar flames and complex 2-D and 3-D turbulent combustion conditions', but no laminar flame, 2-D, or 3-D simulation is performed. The measured speedup is for 36 homogeneous reactor integrations only. This is a load-bearing gap because the abstract's claim that the solver is 'generalizable to larger-scale turbulent flow simulations' rests entirely on this untested transfer.","section":"Sec. I and Sec. III B"},{"comment":"Fig. 8(A) shows that at 987.5 K, an unseen condition lying inside the stated 950-1200 K initial-temperature range, reconstruction errors are roughly an order of magnitude worse than at neighboring training conditions (about 10^-3 versus 10^-4). This directly weakens the premise that the 35-condition training grid covers the thermochemical manifold needed for generalization to other conditions. The limitation paragraph in Sec. III C acknowledges that a non-uniform training grid with denser sampling in the cooler regions would be needed for CFD applications, but the abstract and Sec. III B still assert generalizability; the claims should be qualified to match this evidence.","section":"Sec. III B, Fig. 8(A)"},{"comment":"The reported 2.0x speedup is based on a single comparison of average solve times for 36 homogeneous reactor conditions in Arrhenius.jl. No error bars, no repeated trials, no wall-clock breakdown between gradient evaluation and integrator overhead, and no comparison at different tolerances or state distributions are reported. Since the speedup is a central quantitative claim and the paper explicitly connects it to 'potential for substantial acceleration unlocked by ChemKANs' in larger simulations, the measurement needs to be characterized more carefully, or the claim should be limited to the homogeneous-reactor configuration actually measured.","section":"Sec. III B, Table I"},{"comment":"The abstract's statement that ChemKANs exhibit 'no overfitting or model degradation in any of these training cases' is stronger than the evidence in Fig. 5(A). The noise-free testing MSE increases by roughly a factor of two from 0% to 15% added noise, which is a mild but real degradation, and the claim of 'no degradation' is therefore not supported by the reported metric. The narrative should be revised to say the degradation is small and much less than that of the DeepONet baseline.","section":"Sec. III A and Fig. 5(A)"}],"minor_comments":[{"comment":"Eq. (22) contains an extra closing parenthesis in 'u_pred_k(t_j, θ))' that should be removed.","section":"Sec. II C 4, Eq. (22)"},{"comment":"The caption says '405 additional testing locations between the initial 36', while the text says 441 total initial conditions with 406 unseen; please reconcile these counts explicitly (35 training plus 1 withheld plus 405 additional gives 441, but this accounting should be stated).","section":"Sec. III B, Fig. 8 caption"},{"comment":"The data availability statement says the data are 'available on request'; for a machine-learning paper whose main claims are empirical, releasing code, trained models, and data generation scripts would greatly strengthen reproducibility.","section":"Data availability"},{"comment":"The text refers to 'the H2/O2 mechanism from GRI-Mech 3.0 (9 species, 29 reactions)', but GRI-Mech 3.0 is a 53-species mechanism; please clarify that a subset is used and specify which reactions are retained.","section":"Sec. II D 2"},{"comment":"The bullet 'T raining stage 1' has an erroneous space in 'Training'; this is a typo that should be corrected.","section":"Sec. II C 5"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the leap from 0-D homogeneous reactors to multidimensional CFD lands: the paper's own Fig. 8(A) and limitation paragraph weaken the generalization claim, and the speedup is measured only in homogeneous reactors. The central architecture and in-domain results are plausible, but the abstract and Sec. III B overstate what has been demonstrated. I recommend major revision rather than rejection because the load-bearing issues are fixable by qualifying claims, adding uncertainty quantification on the speedup, and either removing or clearly labeling the CFD-generalization assertions as conjectures."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead ChemKANs. Bottom line: the core contribution is real and the paper is worth engaging with, but the abstract oversells the generalization claim and the release of artifacts is thin.\n\nWhat's actually new: Koenig et al. take their own KAN-ODE and LeanKAN results and wrap them in a chemistry-specific structure: a kinetic core that outputs species production rates plus a thermodynamic superstructure that is mostly a fixed linear map from those rates to dT/dt, with a small KAN correction. That structural inductive bias is sensible, and it lets them train a single network for all nine H2 species plus temperature. ChemNODE needed seven separate MLPs and omitted H, HO2, and H2O2. The 344-parameter network that gives a 2x speedup over detailed chemistry in homogeneous reactors is a plausible and genuinely useful result. The noisy-data study against DeepONet is also fair: they define a noise-free test loss and show ChemKAN degrades only mildly at 15% noise while DeepONet overfits. Good.\n\nThe paper is also honest about some of its own weaknesses: the limitations section admits the speedup is modest, and the low-temperature generalization difficulty in Fig. 8(A) is discussed in the text.\n\nSoft spots, in proportion: the biggest is the leap from 0-D homogeneous reactors to turbulent flow. The abstract says the solver is 'generalizable to larger-scale turbulent flow simulations,' but the only evidence is 36 homogeneous reactor integrations in Arrhenius.jl. That is not a demonstration of operator-split behavior in a flame solver. The transfer could fail because flame states are not all on the training manifold, and the paper's own Fig. 8(A) shows an order-of-magnitude error jump at 987.5 K, just below the training range. So the generalization claim needs either tempering or an additional test, even a 1D laminar flame. Also minor: no error bars on losses or speedup, no code released (data 'available on request'), and the DeepONet comparison is for one small biodiesel system.\n\nWho it's for: combustion modelers and people working on neural surrogates for stiff ODEs. It deserves a serious referee. I would send it to peer review, and ask for code/data deposit, error bars, and either a straightening of the abstract or a flame-coupled test.","headline":"A solid chemistry-structured KAN-ODE paper with real but modest results; the homogeneous-reactor demonstrations hold up, but the abstract's turbulent-flow generalization claim is untested and the paper's own temperature grid shows why that matters.","tokens_in":21925,"tokens_out":2107,"would_cite":true,"duration_ms":21687,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a 344-parameter physics-structured neural network can learn all hydrogen-air combustion chemistry and run stiff ODE integration twice as fast as a detailed mechanism.","keywords":["ChemKAN","Kolmogorov-Arnold networks","neural ordinary differential equations","combustion kinetics","hydrogen-air combustion","stiff chemistry surrogate","noise robustness","operator splitting"],"falsifier":"Replace detailed GRI-Mech 3.0 chemistry with the 344-parameter ChemKAN in a freely propagating hydrogen-air laminar flame calculation under operator splitting and compare the predicted laminar flame speed and extinction limits with the detailed mechanism; disagreement beyond the homogeneous-reactor error levels would refute the transferability claim, as would failure to ignite at the 987.5 K conditions where the paper already reports order-of-magnitude larger errors.","tokens_in":20848,"feed_emoji":"🔥","tokens_out":10639,"duration_ms":105108,"temperature":0.7,"pith_summary":"ChemKAN is a neural ordinary differential equation whose gradient getter is a Kolmogorov-Arnold network reshaped to mirror combustion kinetics: a kinetic core maps the current thermochemical state to species production rates, and a thermodynamic superstructure forms the temperature rate as a linear combination of those rates plus a small correction term. The paper's central claim is that this chemistry-specific inductive bias makes the network unusually parameter-lean and robust: a 344-parameter ChemKAN reproduces all nine species and temperature profiles for hydrogen-air homogeneous reactors across a range of initial temperatures and equivalence ratios, and it does so roughly twice as fast as the detailed GRI-Mech 3.0 mechanism in an ODE solver. In a separate model-inference task, the same ODE-coupled architecture recovers the underlying biodiesel transesterification kinetics from data with up to 15% added noise, showing no overfitting even when the network is overparameterized, while a DeepONet baseline overfits. The payoff, if the claims hold, is that kinetic surrogates can be trained on cheap homogeneous-reactor data and then replace the stiff chemistry term in larger reacting-flow simulations, where chemistry evaluation is typically the dominant computational cost.","feed_headline":"A 344-parameter neural net learns hydrogen combustion at 2x speed","feed_subtitle":"If it transfers to real flames, one trained network could replace the costliest part of combustion simulation.","key_machinery":"The load-bearing object is the ChemKAN architecture itself: a KAN-ODE whose learnable gradient function is split into a kinetic core, $\\mathrm{KAN}_{kin}$, mapping the full state $u=[Y_1,\\dots,Y_m,T]$ to the $m$ species production rates, and a thermodynamic superstructure that computes $dT/dt$ as a linear combination of those rates, mirroring the energy equation, plus a single-layer KAN correction for thermophysical parameter variation. The kinetic core stacks an additive KAN layer and a LeanKAN layer whose multiplicative sublayer captures the products of concentrations that appear in Arrhenius rate laws. Training proceeds in two stages, kinetics first and thermodynamics second, uses forward sensitivity analysis instead of adjoint differentiation to avoid stiffness-driven instability, and optionally adds an element-conservation penalty to the loss. This structure carries the argument by forcing temperature evolution to be built out of species production, sharing information across all species and outputs, and letting the outer ODE integrator smooth over noisy data.","core_discovery":"The central claim is that encoding the known kinetic-thermodynamic coupling into a KAN-ODE produces a source-term surrogate that is simultaneously accurate, sparse, and resistant to overfitting. For hydrogen-air combustion trained on 35 homogeneous-reactor conditions (one condition withheld), a single 344-parameter ChemKAN predicts temperature and mass fractions of every species, including the low-concentration radicals H, HO2, and H2O2 that the MLP-based ChemNODE baseline omitted, and it reproduces ignition delays across the studied range. The same network generalizes over most of a finer 441-condition grid, with the paper explicitly reporting degraded performance at colder, slower-igniting conditions around 987.5 K. For the biodiesel model-inference case, the kinetic core alone extracts smooth underlying profiles from sparse noisy data, and its test error stays nearly flat as parameters are added, where the DeepONet comparator's test error diverges. The paper describes the measured 2x ODE-integration speedup as a conservative lower bound and argues that the single-network full-state output makes the surrogate suitable for coupling to flow solvers.","pith_inferences":["A test the paper leaves open is to embed the trained 344-parameter ChemKAN in a laminar or turbulent reacting flow under operator splitting and compare flame speed, extinction limits, and ignition against detailed chemistry; the homogeneous-reactor 2x speedup could shrink or grow depending on per-cell evaluation cost.","The degraded 987.5 K results suggest the 35-condition grid is thinnest exactly where ignition chemistry is most temperature-sensitive, so a non-uniform or adaptive training grid, rather than a denser global grid, would likely recover accuracy in that regime.","The thermodynamic superstructure is not hydrogen-specific: any reacting system whose temperature equation is a linear combination of species production rates plus a thermophysical correction could reuse the same split, so the architecture should transfer to other fuels, pyrolysis, or thermal-runaway problems.","Because the outer ODE integrator is what smooths the noisy training data, noise robustness may scale with trajectory length or stiffness; varying the integration window in the biodiesel task would test that mechanism directly."],"forward_implications":["A single ChemKAN forward pass returns the source term for the entire thermochemical state, so surrogate models can keep all species, including minor radicals, rather than dropping them as the ChemNODE baseline did.","The 2x per-step ODE speedup should compound in multi-dimensional reacting flows, where chemical source-term evaluation is usually the dominant cost, if the homogeneous-reactor surrogate transfers under operator splitting.","The two-stage training recipe and the linear thermodynamic coupling generalize as a template for other stiff kinetic systems; the kinetic core alone already suffices for isothermal problems such as biodiesel transesterification.","The optional element-conservation penalty provides a physics-consistency check that should reduce mass-fraction drift in long-horizon integrations and in repeated operator-split calls."],"supporting_citations":[{"why":"Provides the MLP-based ChemNODE baseline: the parameter count, the 2.3x speedup, and the omitted H, HO2, and H2O2 species that ChemKAN is compared against.","marker":"[17]"},{"why":"Supplies the neural-ODE training mechanism of differentiating through an ODE solver to fit integrated profiles rather than raw gradients.","marker":"[19]"},{"why":"Is the DeepONet baseline in the biodiesel model-inference comparison, used to benchmark noise robustness and overfitting.","marker":"[20]"},{"why":"Introduces the Kolmogorov-Arnold network with learnable univariate activation functions, the base architecture ChemKAN augments.","marker":"[23]"},{"why":"Introduces KAN-ODEs, the gradient-getter-plus-ODE-solver framework that ChemKAN extends to chemistry.","marker":"[29]"},{"why":"Reports that KANs degrade under noise, the prior result ChemKAN's noise-robustness claims are measured against.","marker":"[32]"},{"why":"Provides the LeanKAN layer with multiplicative and additive sublayers used in the ChemKAN kinetic core.","marker":"[35]"},{"why":"Generates the homogeneous-reactor training data for the hydrogen-air case.","marker":"[50]"},{"why":"Defines the detailed H2/O2 mechanism that serves as ground-truth chemistry for the hydrogen case.","marker":"[51]"},{"why":"Is the combustion solver used for the 2x speedup comparison against detailed chemistry.","marker":"[52]"}],"fun_headline_variants":["KAN-ODE physics bias makes combustion surrogates noise-robust and 2x faster","344-param ChemKAN speeds hydrogen combustion ODEs without overfitting","Physics-coded KANs handle noisy sparse data, cut sim cost in half","ChemKANs: sparse, stiff-savvy surrogates that resist overfitting","ChemKAN with 344 weights doubles ODE speed, defies overfitting"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The promise of accelerating real combustion simulations depends on the untested premise that a network trained on zero-dimensional homogeneous-reactor data remains accurate in multi-dimensional flows under operator splitting, where the source term is evaluated on the local thermochemical state alone, and that the 35-condition training grid adequately covers the thermochemical manifold.","fun_headline_variants_meta":{"raw":{"variants":["KAN-ODE physics bias makes combustion surrogates noise-robust and 2x faster","344-param ChemKAN speeds hydrogen combustion ODEs without overfitting","Physics-coded KANs handle noisy sparse data, cut sim cost in half","ChemKANs: sparse, stiff-savvy surrogates that resist overfitting","ChemKAN with 344 weights doubles ODE speed, defies overfitting"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000733,"raw_usage":{"total_tokens":3337,"prompt_tokens":1064,"completion_tokens":2273,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":680,"completion_tokens_details":{"reasoning_tokens":2162}},"tokens_in":680,"tokens_out":2273,"duration_ms":19152,"temperature":1.0,"reasoning_tokens":2162,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:28:14.437812+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace detailed GRI-Mech 3.0 chemistry with the 344-parameter ChemKAN in a freely propagating hydrogen-air laminar flame calculation under operator splitting and compare the predicted laminar flame speed and extinction limits with the detailed mechanism; disagreement beyond the homogeneous-reactor error levels would refute the transferability claim, as would failure to ignite at the 987.5 K conditions where the paper already reports order-of-magnitude larger errors.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the DeepONet baseline in the biodiesel model-inference comparison, used to benchmark noise robustness and overfitting."},{"cited_title":"Retrieved from https://github.com/Blealtan/efficient-kan (2024)","cited_arxiv_id":null,"evidence_quote":"Is the combustion solver used for the 2x speedup comparison against detailed chemistry."}],"review_version":1}