{"id":"33b9a4cf-87e3-4ebf-aa3c-ef432970878a","arxiv_id":"2412.07514","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A hybrid mosquito model that replaces a fixed empirical development rate with a PINN-learned weather-dependent rate improves adult abundance and peak forecasts compared with the standard ODE model.","lead":"This paper trains a physics-informed neural network to learn the pupa development rate of Culex mosquitoes from weather data and plugs it into a classic population model. The hybrid model predicts adult mosquito peaks better than the standard model, which matters for timing vector-control interventions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported test period was used to select the architecture and activation function (Sec. 4, Tables 2-3), so the headline peak metrics are not a clean out-of-sample comparison; a nested selection protocol is needed.","rationale":"The paper's central contribution is empirical: a hybrid PINN parameterisation beats a fixed-parameter ODE on an eight-year out-of-sample period. For that claim to hold, the reported test metrics must come from a model whose choices did not use the test labels. The manuscript shows the opposite: Sec. 4 selects architecture by test-period RMSE/peak scores and Sec. 4.2 selects the SoftAbs epsilon the same way. This is a more proximal threat than the identifiability concern the reader flagged, because it is directly visible in the text and does not depend on assumptions about ODE sloppiness. I therefore partially agree with the reader: they listed test-set selection as a source of fragility, but their stated weakest assumption was identifiability/transferability. My concrete test would settle whether the selection actually matters. If the training-selected model still beats Dy PopMosq with the reported margins, the headline stands; if not, the validation advantage in Table 1 is inflated. The verdict remains CONDITIONAL because the paper is otherwise reproducible in intent (code/data links, detailed hyperparameters) and the RMSE advantage appears robust across architectures, but the peak-detection claim specifically is not yet trustworthy.","tokens_in":18739,"tokens_out":6205,"duration_ms":62266,"concrete_test":"Re-run the architecture and activation ablations with model selection performed only on the 2016-2017 training period (e.g., select architecture/epsilon/checkpoint by RMSE on a simulated training holdout, ideally using an inner split of the two training years), freeze that configuration, and then compute the 2000-2007 validation metrics. If the pre-selected Branched FourierMLP + SoftAbs(1e-4) still gives RMSE about 0.18 and peak F1 about 0.57, the concern is resolved; if a different configuration wins on training selection or the test gap shrinks, the reported Table 1 gains are partly a selection artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing weakness is not the inverse problem's identifiability but the way the 2000-2007 validation set is used twice. In Sec. 4.1 the paper compares four architectures (MLP, FourierMLP, Branched MLP, Branched FourierMLP) and in Sec. 4.2 compares output activations (Identity, ReLU, Softplus, Abs, SoftAbs with three epsilons), with metrics computed \"using the same experimental procedure\" as Sec. 2, i.e. on the test period. It then selects Branched FourierMLP and SoftAbs(eps=1e-4) as superior based on those test-period RMSE/peak scores (Tables 2 and 3), and Table 1 reports exactly this selected configuration as Hy PopMosq. Consequently the headline RMSE 0.18, peak recall 0.56, precision 0.63, F1 0.57 are not an out-of-sample evaluation of a pre-specified model; the eight-year test period has influenced which model is reported. The checkpoint-level selection described in Sec. 4.2/Appendix C (\"best RMSE when simulating with PINN-learned parameters\" on training data) does not repair this, because architecture and activation choices were made on the test set. Mitigating evidence: all four architectures in Table 2 have lower RMSE than Dy PopMosq's 0.26, so the RMSE advantage may survive re-selection. But the peak performance varies dramatically across architectures (recall 0.15-0.56) and activations (F1 0.06-0.57); the strongest claimed advantage is exactly the part most exposed to selection. This makes the central \"generally outperforms\" claim conditional on a clean nested evaluation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Hy PopMosq, a hybrid model that replaces the fixed empirical pupa development rate fP in the Dy PopMosq ODE model with a neural-network parameter function Θ(m; Wθ) learned via a physics-informed neural network from meteorological data and adult trap counts. The parameter network is trained on 2016-2017 Petrovaradin data and then used with the Dy PopMosq ODE to simulate adult abundance for 2000-2007. The paper reports lower RMSE (0.18 vs 0.26) and higher peak recall/precision/F1 (0.56/0.63/0.57 vs 0.13/0.09/0.11) for Hy PopMosq. Section 4 presents an ablation of architectures and output activations, motivating the choice of Branched FourierMLP and SoftAbs(ε=10^-4) that define the Hy PopMosq configuration.","tokens_in":19152,"tokens_out":5598,"duration_ms":49628,"significance":"If the comparison were clean, the paper would provide useful evidence that a weather-driven neural parameterisation of a single development rate can improve a mechanistic mosquito model's out-of-sample abundance and peak forecasts. The temporal holdout (train 2016-2017, test 2000-2007) is a genuine out-of-sample design, and the ablation is systematic. The authors should be credited for using a real multi-year dataset and for testing a hypothesis that is falsifiable. However, the reported test period was used to select the architecture and activation function, and the baseline comparison is not fully controlled because Dy PopMosq and Hy PopMosq use different initial conditions. These issues are fixable, but they currently make the headline quantitative claims conditional.","major_comments":[{"comment":"The Hy PopMosq configuration reported in Table 1 is the result of a model-selection process that used the 2000-2007 test-period metrics. In Sec. 4.1 the four architectures are compared \"using the same experimental procedure as in Sec. 2\", and Table 2 shows peak recall ranging from 0.146 (MLP) to 0.563 (Branched FourierMLP) and F1 from 0.101 to 0.567; Table 3 shows F1 ranging from 0.063 (ReLU) to 0.567 (SoftAbs with ε=10^-4). The chosen Branched FourierMLP and SoftAbs(ε=10^-4) are exactly the best performers on those test-period metrics, so Table 1's RMSE 0.18, recall 0.56, precision 0.63, and F1 0.57 are not an out-of-sample evaluation of a pre-specified model. The \"generally outperforms\" claim therefore needs a nested protocol: hold out a portion of 2000-2007 for architecture/activation selection, or select using training-period metrics, and only then evaluate the selected model on the remaining holdout. Reporting the test-period best as the headline result overstates the evidence.","section":"Section 4 (Tables 2 and 3) vs Table 1"},{"comment":"The comparison between Hy PopMosq and Dy PopMosq is not controlled with respect to initial conditions. The PINN simulations use initial conditions derived from the trained state network U, while Dy PopMosq uses a fixed initial condition of 300 for every state component, as stated in Appendix C and Sec. 2.2. Since the reported RMSE and peak metrics are computed on the simulated adult abundance, the hybrid model's advantage could partly reflect more favorable initialisation rather than the learned fP parameterisation. The authors should either run Dy PopMosq with the same data-derived initial conditions (or a range of initial conditions) and report the resulting metrics, or explicitly justify that the spin-up period (about 200 days) makes the initial condition irrelevant for the 7-8 year evaluation.","section":"Appendix C and Sec. 3.3"},{"comment":"The checkpoint-selection procedure is not applied on equal footing across activation functions. Models using ReLU and Softplus were selected at 11,000 and 12,000 training steps out of 300,000, when training loss had not been fully minimised, while Identity, Abs, and SoftAbs were selected at about 250,000 steps. This means the activation comparison in Table 3 confounds the activation choice with the number of training steps and the stopping criterion. Moreover, selecting the checkpoint by \"best RMSE when simulating with PINN-learned parameters\" on the training period is a training-based selection, which does not justify selecting the architecture or epsilon on the test period. The authors should report results for a fixed checkpoint rule (e.g., final step or best training loss) across all activations, and should clearly separate training-based selection from test-based evaluation.","section":"Sec. 4.2, final paragraph"},{"comment":"The peak-detection metrics are computed with a 7-day, 0.2-prominence detector, but no sensitivity analysis or alternative peak definition is provided. Because the strongest improvement over Dy PopMosq is in peak recall/precision/F1 (0.56/0.63/0.57 vs 0.13/0.09/0.11), and because these metrics also drive the architecture and activation selection in Tables 2-3, the authors should report how peak metrics vary with detector parameters (window, prominence threshold) and with the normalization used. This would establish that the peak advantage is not an artifact of a particular peak-finding configuration.","section":"Sec. 3.2 and Fig. 3 caption / Appendix C"}],"minor_comments":[{"comment":"The text says validation uses daily data from 2000-2007, but Table D.5 lists annual metrics only for 2001-2007; please reconcile or clarify whether 2000 is included in the average.","section":"Sec. 3.1 vs Table D.5"},{"comment":"The row 'No. Peaks' is not defined in the validation methodology; please state whether it is the average number of detected peaks per year and how it is computed.","section":"Table 1"},{"comment":"Equation (4) uses the symbol D (θ D Θ(m; Wθ)) which appears to be a typo for equality; also define all symbols in Eq. (5), where the letter 'e' appears in the Fourier feature definition.","section":"Eq. (4)"},{"comment":"Appendix C states the validation dataset has 7 years, while Sec. 3.1 and Table D.5 say the period is 2000-2007; please ensure the period counts are consistent throughout the manuscript.","section":"Appendix C"},{"comment":"The phrase 'No bias reduction is applied in the original model' is unclear; please specify whether this refers to bias correction of meteorological inputs or statistical bias in the ODE simulation.","section":"Sec. 2.2"}],"recommendation":"major_revision","confidential_remarks":"The main concern is selection leakage: the test period is used for architecture/activation selection and then reported as the evaluation. This is fixable by a nested holdout or by pre-specifying the model. The RMSE advantage appears relatively stable across architectures in Table 2, so the paper is likely salvageable, but the peak-metric claims need the most careful re-evaluation. I would also encourage releasing the training/evaluation code so that the checkpoint-selection procedure can be independently verified, since the current statement \"available from the corresponding author upon reasonable request\" limits reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things up front. The paper does something real: it replaces the fixed temperature-driven pupa development rate fP in a 10-stage Culex ODE model with a neural network that maps meteorological inputs plus day-of-year to that rate, learned via a PINN from two years of weekly adult trap counts. And the eight-year holdout (2000-2007) is a genuine temporal split, so the central comparison is not a self-fulfilling fit. The catch is that the same holdout also selected the architecture: the Sec. 4 ablation picks Branched FourierMLP and SoftAbs(1e-4) using test-period metrics, and Table 1 then reports that configuration as Hy PopMosq. The headline peak gains (recall 0.56 vs 0.13, F1 0.57 vs 0.11) are exactly the numbers that swing across architectures (recall 0.15-0.56), so treat them as conditional until a properly nested selection is run.\n\nWhat the paper does well. The sensitivity analysis in Appendix B identifies fP as the leverage point before fitting, which is the right way to justify what you learn. All four architectures in Table 2 beat the baseline on RMSE, so the abundance-tracking claim likely survives re-selection. Table D.5 gives full annual metrics rather than cherry-picking, and the data are on GitHub. The paper is also honest about the 2002/2007 failures and about the uneven checkpoint selection in Sec. 4.2. That transparency is real credit.\n\nSoft spots, in order. First, the test-set architecture selection is the load-bearing one, and it is not flagged: the Sec. 5 limitations do not mention it. This is fixable by pre-specifying the architecture or running a nested holdout. Second, the activation comparison is confounded with training length: ReLU and Softplus checkpoints were taken at 11k-12k of 300k steps while Identity/Abs/SoftAbs ran to about 250k, so some of the SoftAbs advantage may just be more training. Third, the meteorology for 2000-2007 is reconstructed by linear regression from a station some distance away, and there are no uncertainty estimates anywhere. With only two years of training data, identifiability of the parameter function is assumed, not demonstrated; a few ensemble runs would tell you how much of the validation gain hangs on a particular solution.\n\nMy overall read agrees with the conditional verdict: the core idea is sound, the holdout is real, and the main flaw is a selection-protocol issue rather than an internal contradiction. The paper deserves a serious referee. I would send it to review with instructions to fix the nested selection and add some uncertainty quantification. It is a solid case study for people building hybrid mechanistic-ML population models, not a breakthrough, and it should be cited with the selection caveat noted.","headline":"A genuine temporal holdout and a sensible choice of what to learn, but the same validation years also picked the architecture, so the headline peak metrics are conditional until a nested selection is run.","tokens_in":19695,"tokens_out":5102,"would_cite":true,"duration_ms":43176,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a neural-network replacement for the pupa-development parameter in a mechanistic mosquito model, learned from weather and two years of trap counts, beats the fixed-form model on an eight-year validation.","keywords":["mosquito population dynamics","physics-informed neural networks","inverse parameter estimation","mechanistic ODE models","pupa development rate","Culex pipiens","peak prediction","hybrid dynamic model"],"falsifier":"Train the PINN with several random initialisations and two or more different architectures, then compare the inferred $f_P$ functions: if functions that fit the two training years equally well diverge strongly outside the training period, the inverse problem is non-identifiable and the validation gains are not robust. Alternatively, hold out one of the eight validation years or a second trap site entirely, retrain only on the first site's two years, and check whether the RMSE and peak-F1 advantages persist in the held-out data.","tokens_in":18559,"feed_emoji":"🦟","tokens_out":7889,"duration_ms":69678,"temperature":0.7,"pith_summary":"This paper argues that a physics-informed neural network (PINN) can replace one fixed empirical parameter in a mechanistic mosquito-population model with a data-learned function of weather, and that the resulting hybrid model predicts adult Culex pipiens abundance and peaks better than the temperature-only mechanistic model. The authors train a neural network to infer the pupa development rate $f_P$—the parameter their sensitivity analysis identifies as the most influential—from two years of weekly adult trap counts together with daily temperature, humidity, and precipitation. On an eight-year out-of-sample simulation, the hybrid Hy PopMosq model achieves lower population RMSE (0.18 vs 0.26) and better peak detection (precision 0.63 vs 0.09, recall 0.56 vs 0.13) than the baseline Dy PopMosq model. An ablation study reports that a branched Fourier-feature MLP with a smoothed absolute-value output activation is the best-performing configuration. If correct, the result would make weather-driven, machine-learned parameterisation a practical route to more accurate vector-abundance forecasts.","feed_headline":"Neural network learns pupa rate from weather, beats fixed model","feed_subtitle":"Hybrid machine-learning model cuts population RMSE from 0.26 to 0.18 and lifts peak F1 from 0.11 to 0.57.","key_machinery":"The load-bearing object is the parameter network $\\Theta(m; W_\\theta)$, a neural network that maps a vector of daily and seven-day historical meteorological conditions plus the day of year to the pupa development rate $f_P$. It is trained jointly with a state network $U(t; W_U)$ by minimising a loss with a data term (weekly adult counts, with unobserved stages masked) and a physics term (residuals of the ten ODEs at collocation points). After training, $\\Theta$ is detached from $U$ and supplies $f_P$ to the original ODE system for forward simulation. The architecture is a two-branch FourierMLP with GELU hidden activations and a SoftAbs output activation, chosen by ablations that show this configuration best avoids trivial-zero convergence and best captures annual periodicity.","core_discovery":"The central claim is that the pupa development rate $f_P$ can be learned as a function of meteorological inputs by a PINN, and that substituting this learned parameter for the empirical temperature-only formula in the ten-stage ODE system gives better simulations of the adult blood-seeking population ($A_{b1}+A_{b2}$) over 2000–2007. The trained parameter network is the only change to the underlying Dy PopMosq model; once trained, it is frozen and used as a parameter-supplying function in forward ODE simulations. Across all years Hy PopMosq has lower RMSE for population and for growth rate, a standard deviation closer to the observed value, and higher peak recall and precision. The two worst years (2002 and 2007) remain difficult for both models, and the paper attributes those failures to environmental factors outside weather, such as temporary water retention. The architecture study attributes the hybrid's performance to Fourier features combined with a multi-branch structure, and to the SoftAbs activation preventing collapse to the trivial zero solution.","pith_inferences":["A direct test of transferability would be to train the parameter network at Petrovaradin and apply it, without retraining, to a second Culex site with its own trap counts; the current validation is at the same location across years.","The same PINN-inverse procedure could be applied to the larva development rate $f_L = 1.65 f_P$ or to mortality rates; the paper's sensitivity analysis suggests $f_P$ is the most influential, but a multi-parameter version would test whether learned interactions improve or destabilise the ODE system.","The 2002 and 2007 failures point to an implicit assumption that weather alone drives the learned parameter; incorporating flood-retention or land-cover proxies is a natural, testable extension suggested by the paper's own discussion."],"forward_implications":["Weather-driven learning of $f_P$ lowers population RMSE from 0.26 to 0.18 and raises peak F1 from 0.11 to 0.57 over the eight-year validation.","Because only the parameter function changes, the hybrid retains the interpretable stage-structured ODE description of mosquito biology.","The peak-detection gains imply the hybrid model is better positioned to time vector-control interventions than the temperature-only baseline.","The ablation results show that the two-branch Fourier feature structure, not simply network capacity, drives the advantage in capturing annual periodicity.","Both models miss sudden ecological events such as the 2002 and 2007 peaks, so the learned parameterisation does not remove the need for exogenous disturbance information."],"supporting_citations":[{"why":"Builds the base PINN mosquito model with ODE normalisation and gradient balancing that Hy PopMosq extends.","marker":"[28]"},{"why":"Supplies the ten-stage Dy PopMosq ODE model whose parameterisation the paper replaces.","marker":"[29]"},{"why":"Defines the PINN training objective (data loss plus physics loss) used for inverse parameter inference.","marker":"[33]"},{"why":"Introduces Fourier features that the architecture uses to capture high-frequency and periodic patterns.","marker":"[34]"},{"why":"Provides the Fourier-feature network design for PINNs and the rationale for its convergence benefits.","marker":"[35]"},{"why":"Supplies the GELU activation used in hidden layers.","marker":"[36]"},{"why":"Provides the automatic differentiation framework used to compute ODE residuals and gradients.","marker":"[38]"},{"why":"Sets the validation criterion that RMSE should be below the standard deviation of observations.","marker":"[41]"}],"fun_headline_variants":["PINN learns pupa rate from weather, bests fixed ODE","Neural net improves mosquito model, RMSE drops 31%","Weather-trained neural network sharpens vector forecasts","Physics-inspired net beats classic mosquito dynamics","Hybrid model learns pupa development from met data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that a unique, transferable pupa-development function can be recovered from two years of weekly adult trap counts at one site, with off-site weather data, and that this function stays valid for the eight-year validation period.","fun_headline_variants_meta":{"raw":{"variants":["PINN learns pupa rate from weather, bests fixed ODE","Neural net improves mosquito model, RMSE drops 31%","Weather-trained neural network sharpens vector forecasts","Physics-inspired net beats classic mosquito dynamics","Hybrid model learns pupa development from met data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000281,"raw_usage":{"total_tokens":1689,"prompt_tokens":995,"completion_tokens":694,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":611,"completion_tokens_details":{"reasoning_tokens":616}},"tokens_in":611,"tokens_out":694,"duration_ms":7020,"temperature":1.0,"reasoning_tokens":616,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:45:59.686110+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the PINN with several random initialisations and two or more different architectures, then compare the inferred $f_P$ functions: if functions that fit the two training years equally well diverge strongly outside the training period, the inverse problem is non-identifiable and the validation gains are not robust. Alternatively, hold out one of the eight validation years or a second trap site entirely, retrain only on the first site's two years, and check whether the RMSE and peak-F1 advantages persist in the held-out data.","supporting_citations":[{"cited_title":"Petric, Modelling the influence of meteorological conditions on mosquito vector population dynamics (diptera, culicidae","cited_arxiv_id":null,"evidence_quote":"Supplies the ten-stage Dy PopMosq ODE model whose parameterisation the paper replaces."},{"cited_title":"Raissi, P","cited_arxiv_id":null,"evidence_quote":"Defines the PINN training objective (data loss plus physics loss) used for inverse parameter inference."},{"cited_title":"Tancik, P","cited_arxiv_id":null,"evidence_quote":"Introduces Fourier features that the architecture uses to capture high-frequency and periodic patterns."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Fourier-feature network design for PINNs and the rationale for its convergence benefits."},{"cited_title":"Pielke, Mesoscale meteorological modeling, Academic Press, New York, N","cited_arxiv_id":null,"evidence_quote":"Sets the validation criterion that RMSE should be below the standard deviation of observations."}],"review_version":1}