{"id":"698a3f08-2940-4f5e-a32c-9cc25c33a03b","arxiv_id":"2506.13486","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"PALIRS combines active learning with MACE neural network potentials and a dipole moment model to predict infrared spectra of small organic molecules at a fraction of the DFT cost.","lead":"This paper introduces PALIRS, an open-source workflow that uses active learning to train a machine-learned force field and a dipole model, then computes infrared spectra from fast molecular dynamics. It reports accurate spectra for 24 small organic molecules at roughly 1/100th of the DFT data cost, with good agreement between machine-learned, ab initio, and experimental spectra.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dipole model coverage of production-trajectory configurations is asserted rather than demonstrated; a coupled validation is needed.","rationale":"I read the paper in good faith: it is a well-engineered workflow with open code and data, and the evidence (PCC 0.80 DFT-ML and 0.81 Exp.-ML) supports the claim for the 24-molecule in-distribution set at a meaningful 100x DFT data reduction. The reader's identified weakest assumption matches mine almost exactly, which is why I mark agreement as full. I nevertheless keep the verdict UNCHANGED because the reader's CONDITIONAL verdict already demands the validation I would request; my check is a sharpening of the same condition, not a new one. The concern is load-bearing in principle, but it is a missing-validation point rather than an observed contradiction: the paper even acknowledges elevated force uncertainty for 1,3-butadiene and still reports spectra, which is exactly why a coupled dipole/force check on production trajectories would settle it. I did consider whether the internal consistency of the workflow was fully sound and found no contradiction: Eq. 1 and the Wiener-Khinchin processing are standard, the active learning details in Sec. 5.4 are concrete, and the publicly available repository strengthens reproducibility. The one additional caveat worth noting is that the methodology does not explicitly verify that the dipole model's training data includes configurations at the higher temperatures used in production (100-900 K in Fig. 6), though the active learning data spans 300-700 K; this is best interpreted as part of the same production-trajectory coverage question rather than a separate failure. I therefore recommend no change: the conditional acceptance with the explicit requirement to validate dipole-model accuracy on independent production trajectories and on out-of-distribution molecules is the right scientific posture.","tokens_in":17756,"tokens_out":1910,"duration_ms":17382,"concrete_test":"Run a 50 ps MLMD production trajectory at 300 K for 1,3-butadiene (and for 2-3 in-distribution molecules) using the published final MLIP ensemble, compute dipole predictions with the published dipole model, then perform DFT single-point dipole calculations on 100 uniformly sampled trajectory frames (roughly one frame per 0.5 ps). Compare DFT dipole moments (magnitudes and time autocorrelation) against the ML dipole model on those frames. If the dipole MAE on production frames approaches or exceeds the test-set MAE of 7.62 mDebye by a large factor, or if the resulting IR spectrum from corrected dipoles differs from the published spectrum beyond the reported standard deviation, then the central claim needs qualification for out-of-distribution molecules.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that PALIRS reproduces DFT-AIMD and experimental IR spectra with fewer than 1,000 DFT single points per molecule. The IR spectrum is computed from dipole-moment trajectories (Eq. 1), so accuracy requires both the MACE force model and the separately trained MACE dipole model to remain accurate on the 50 ps production MLMD ensemble at each temperature. Active learning (Sec. 5.4) selects configurations by force uncertainty only; the dipole model is trained once on the final 16,067 selected structures with no dipole-specific acquisition or retraining. The dipole model's reported 7.62 mDebye MAE is measured on a test set generated in Sec. 5.6: a 300 K 100 ps MLMD trajectory using the first MLIP itself, clustered in MBTR space and labeled by DFT. This test set is therefore in-distribution for both the force model and the dipole model. The 1,3-butadiene transferability result (Sec. 3) directly exposes the gap: force uncertainties reach 10^-1 eV/A, and the paper reports poor Exp.-ML spectral agreement for that molecule, yet the stated explanation is limited C=C training representation. The reader's weakest assumption is exactly this load-bearing premise: force-disagreement-selected configurations sufficiently cover the dipole moment surface, including along production trajectories and for out-of-distribution molecules. The paper does not validate dipole-model error on a held-out production trajectory, nor on out-of-distribution molecules, nor does it show that predicted IR intensities are insensitive to dipole error. Because IR amplitudes are compared to experiment and DFT, this unvalidated coupling is the softest spot in the central efficiency claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces PALIRS, an active-learning workflow that trains a MACE ensemble-based machine-learned interatomic potential (MLIP) and a separate MACE dipole moment model, then uses machine-learning molecular dynamics (MLMD) trajectories and the dipole autocorrelation function (Eq. 1) to predict infrared spectra. The method is tested on 24 small organic molecules by comparing MLMD spectra against DFT-based AIMD spectra and NIST experimental spectra using Pearson correlation coefficient (PCC) and Wasserstein distance (WD). The authors report mean DFT-ML PCC 0.80 and Exp.-ML PCC 0.81, a speedup relative to AIMD, temperature-dependence studies for methanol and ethanol, and transferability tests on 8 additional molecules. The central claim is that PALIRS reproduces AIMD-quality IR spectra with fewer than 1,000 single-point DFT calculations per molecule, compared with roughly 100,000 DFT steps for AIMD.","tokens_in":18139,"tokens_out":6564,"duration_ms":71289,"significance":"If the central claim holds, PALIRS is a practically useful contribution to high-throughput IR spectroscopy for small organic molecules, and the open-source code and deposited datasets (Zenodo DOIs, GitLab repository) are valuable for reproducibility. The external comparison to NIST experimental spectra is a genuine benchmark, and the systematic analysis of trajectory length for spectral convergence is a useful practical guideline. The main limitation is that the dipole moment model is trained and evaluated on configurations selected by force-uncertainty active learning, so the accuracy of the dipole model on production trajectories and on out-of-distribution molecules is not directly demonstrated. This gap tempers the strength of the transferability claims but does not invalidate the core comparison for in-distribution molecules, which is supported by the DFT-ML spectral agreement.","major_comments":[{"comment":"The test set used for the MLIP and dipole moment errors is generated from a 100 ps MLMD trajectory produced by the first MLIP of the final ensemble, with structures clustered in MBTR space and labeled by DFT. This makes the test configurations in-distribution for the force model and likely also for the dipole model, since the training set was itself collected from active-learning MLMD runs at 300, 500, and 700 K. The dipole moment MAE of 7.62 mDebye in Table 1 therefore does not measure accuracy on out-of-distribution regions or on the full production ensemble. The central claim would be considerably strengthened by evaluating the dipole model on an independently generated test set (for example, from DFT-based AIMD trajectories or from MLMD trajectories of the final ensemble at multiple temperatures) and by reporting dipole errors on the actual production trajectories used for IR spectra.","section":"Section 5.6 and Figure 3"},{"comment":"Active learning selects configurations based on force uncertainty only; the dipole moment model is trained once on the final force-selected dataset without any dipole-specific acquisition or retraining. The assumption that force-disagreement-selected configurations also sufficiently cover the dipole moment surface is load-bearing for the IR spectrum prediction, but it is not validated. The paper reports force uncertainties up to 10^-1 eV/A for 1,3-butadiene but never reports the dipole model's error for that molecule or for the other transferability molecules. The statement in the Discussion that the predicted IR spectra \"remain consistent\" despite elevated force uncertainties is based on the standard deviation across the three ensemble trajectories, which measures sampling variability, not dipole-model accuracy. The authors should add dipole-model validation on held-out production configurations, or introduce dipole-uncertainty-based acquisition, or provide an explicit argument why force-uncertainty sampling guarantees dipole accuracy.","section":"Sections 5.4 and 3"},{"comment":"The claim that the ML models \"generalized well to larger, structurally similar molecules\" is only partially supported by the main text: 1,3-butadiene shows clear deviations in both frequency and intensity (Figure 7b), and the full set of 8 molecules is only presented in the Supplementary Information. The transferability metrics (PCC and WD) for all 8 molecules should be reported in the main text or at least in a table, and the \"generalizes well\" wording should be qualified in light of the butadiene result. This is not a fatal issue, but it affects the strength of the generalizability claim made in both the Results and the Conclusion.","section":"Section 2, 'Assessment of ML model transferability'"},{"comment":"The choice of 50 ps as the production trajectory length is based on a convergence study for a single molecule, methanol, comparing 20 ps and 50 ps runs (Figure 4). This length is then applied to all 24 training molecules and all 8 transferability molecules without per-molecule convergence checks. For molecules with slower conformational relaxation or low-frequency modes, 50 ps may not be sufficient for converged intensities. The authors should either provide per-molecule convergence evidence or discuss why the methanol result is expected to transfer to the other molecules in the dataset.","section":"Section 2, 'Infrared spectra calculation and length of dynamical simulation'"}],"minor_comments":[{"comment":"The text states that the ML predictions \"align even more closely with the experimental data than the DFT-based AIMD results,\" but Table 2 shows Exp.-DFT WD = 0.054 and Exp.-ML WD = 0.057, so the Wasserstein distance slightly favors DFT, while the PCC favors ML (0.80 vs 0.73). The claim should be qualified as metric-dependent or rephrased.","section":"Table 2 and Section 2, 'Performance of PALIRS in predicting infrared spectra'"},{"comment":"A maximum correlation depth of 1000 fs is used in the autocorrelation function, but no convergence study with respect to this parameter is reported. Since the correlation window length affects spectral resolution and statistical noise, a short sensitivity check or a justification of this value would improve the manuscript.","section":"Section 5.8"},{"comment":"DFT-based AIMD uses Berendsen equilibration followed by Nosé-Hoover thermostatting, whereas MLMD production runs use a Langevin thermostat with a friction coefficient of 0.01. The potential effect of different thermostats on the IR intensities is not discussed; a brief justification would be helpful.","section":"Section 5.7"},{"comment":"There is a typographical error: \"Additionaly\" should be \"Additionally.\"","section":"Declarations"},{"comment":"The MACE model architecture is described, but training hyperparameters such as learning rate, batch size, number of epochs, and loss weighting for energy, forces, and dipole moments are not reported. Providing these details would improve reproducibility.","section":"Section 5.1"},{"comment":"The text says peak positions converge by 20 ps, but the relative intensities differ between 20 and 50 ps and the 20 ps spectrum has an \"inverse trend\" compared with experiment. Clarify in the figure caption or text that convergence is primarily for peak positions, and that intensities require longer trajectories.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reports a useful workflow and a generally convincing in-distribution validation, but the validation of the dipole model on production and out-of-distribution configurations needs to be addressed before publication. I recommend that the revised version be reviewed by someone with expertise in uncertainty quantification for machine-learned potentials, and that the authors be asked to either provide the additional validation or substantially soften the transferability claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a solid, methods paper that delivers a genuinely useful open-source workflow (PALIRS) for MLIP-based IR spectra, with a real external benchmark against NIST and DFT-AIMD across 24 small molecules. The headline cost claim—under 1,000 DFT single points per molecule instead of roughly 100,000—is credible and consistent with the active-learning design. I would bet the central claim holds.\n\nWhat is actually new is not the individual ingredients: active learning for MLIPs and MLIP-based IR spectra both exist. The contribution is the integrated package, the systematic 24-molecule validation, the trajectory-convergence analysis, and the temperature-dependence tests. The code and datasets are public, which is real reproducible work and meaningfully strengthens the paper.\n\nThe soft spots, in proportion:\n\n1. The dipole model is trained once on force-uncertainty-selected structures, with no dipole-aware acquisition. The spectrum comes from the dipole autocorrelation, so this is the one load-bearing assumption that is asserted rather than demonstrated. The 7.62 mDebye MAE is measured on a test set generated by the final MLIP itself, so it is in-distribution for both models; it does not directly validate dipole accuracy on production trajectories or out-of-distribution molecules. The 1,3-butadiene case, with force uncertainty up to 1e-1 eV/A and poor Exp.-ML agreement, is exactly the regime where this matters. This is a genuine gap, but a fixable one: a few held-out production trajectories labeled with DFT dipoles would settle it.\n\n2. The 50 ps trajectory length was chosen by comparing methanol to experiment. That is a mild selection-on-outcome; it would be cleaner to present it as a pragmatic choice rather than a general convergence guarantee. Minor.\n\n3. The MLIP test set is produced by the model being tested (Methods 5.6). This is common practice in active-learning papers, but it means the reported model-error bars are probably optimistic. The external DFT/experimental spectral comparison partly compensates for this.\n\nOverall, the paper is honest about its limitations, the central formula is standard autocorrelation, and the citation pattern is appropriate. The work deserves a serious referee. I would send it to peer review and expect moderate revision, with the validation of dipole-model coverage being the main thing I would ask for.","headline":"A well-engineered, open-source workflow for MLIP-based IR spectra with credible 100x DFT-cost savings; the main gap is untested dipole-model coverage on production trajectories.","tokens_in":18613,"tokens_out":2456,"would_cite":true,"duration_ms":25954,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An active-learning workflow trains a neural-network potential and a dipole model on a few hundred DFT calculations per molecule and reproduces both ab initio and experimental infrared spectra.","keywords":["infrared spectroscopy","active learning","machine-learned interatomic potentials","data generation","equivariant neural networks","molecular dynamics","dipole moments","spectra prediction"],"falsifier":"For a molecule outside the training distribution, such as one with the C=C motif beyond ethene, run the production trajectories, compute DFT dipole moments at snapshots where the force ensemble disagreed most, and compare them with the dipole model; if the dipole errors are large and the spectrum changes when training data are instead selected by dipole disagreement, the force-uncertainty proxy is refuted.","tokens_in":17508,"feed_emoji":"🧪","tokens_out":13701,"duration_ms":131188,"temperature":0.7,"pith_summary":"This paper sets out to show that an active-learning workflow can train a machine-learned interatomic potential and a companion dipole-moment model from just a few hundred density functional theory (DFT) calculations per molecule. For 24 small organic molecules, roughly 600-800 selected structures per molecule (about 16,000 in total) were enough to push harmonic-frequency errors to roughly 4 cm$^{-1}$ and force errors to about 4 meV/Å. The spectra predicted from machine-learning molecular dynamics match DFT-based ab initio spectra and experimental gas-phase spectra in both peak positions and amplitudes, while cutting the number of DFT calculations by roughly a factor of 100. If correct, this makes high-throughput infrared screening of catalytically relevant molecules practical, including temperature-dependent spectra from 100 to 900 K.","feed_headline":"Active learning cuts the DFT cost of IR spectra 100-fold","feed_subtitle":"Fewer than 1,000 DFT calculations per molecule reproduce reference and measured IR spectra.","key_machinery":"The load-bearing mechanism is the active-learning loop built around an ensemble of equivariant message-passing neural-network potentials. Disagreement among ensemble members' force predictions marks each configuration's uncertainty; the most uncertain structures from short simulations at three temperatures are labelled by DFT and added to the training set, and the ensemble is retrained. A second network with a vector output learns the molecular dipole moment. The infrared spectrum is the Fourier transform of the autocorrelation of the dipole time derivative along 50 ps trajectories, with a 1000 fs correlation depth, a smoothing window, and averaging over three independent trajectories; similarity to reference spectra is quantified by a correlation coefficient and a transport distance.","core_discovery":"The central claim is that data selection by force uncertainty, not exhaustive sampling, is what makes machine-learned spectra affordable and accurate. The workflow starts from geometries sampled along harmonic normal modes, then repeatedly runs short dynamics with an ensemble of neural-network potentials and adds the configurations where ensemble force predictions disagree most, cycling at 300, 500, and 700 K until harmonic-frequency errors plateau. A second neural network is trained on the accumulated set to predict dipole moments, and the spectrum is computed from the autocorrelation of the dipole time derivative over three independent 50 ps trajectories. Across all 24 molecules the machine-learned spectra agree with the DFT reference spectra (mean correlation 0.80) and with experimental spectra (mean correlation 0.81), comparable to or better than the DFT-versus-experiment agreement of 0.68; the authors conclude that anharmonic effects and temperature dependence are captured implicitly by the dynamics.","pith_inferences":["If the force-disagreement selection rule is really a sufficient proxy for dipole-moment coverage, then a variant that actively selects configurations by dipole-model disagreement should not change the spectra; running that comparison for out-of-distribution molecules would test the workflow's weakest link.","The same acquisition logic could be transferred to other spectra that depend on a property surface—Raman intensities from polarizability, or vibrational circular dichroism from rotatory strength—with the property-specific uncertainty as the acquisition signal.","The better-than-DFT agreement with experiment may come substantially from averaging three independent trajectories rather than from the potential itself; separating sampling noise from model error by varying the number of trajectories would clarify where the accuracy gain originates."],"forward_implications":["IR spectra of small catalytically relevant molecules can be produced at DFT-level accuracy with roughly 100 times fewer single-point DFT calculations, making high-throughput spectral screening feasible.","Temperature-dependent spectra become practical by running the trained models at different temperatures, demonstrated here from 100 to 900 K, without new ab initio simulations.","Because the machine-learned dynamics scales roughly linearly with system size while DFT scales steeply, the same workflow should extend to larger molecules than ab initio dynamics can reach.","Transferability is bounded by chemical coverage: molecules with underrepresented bonding motifs, such as the C=C motif in 1,3-butadiene, show degraded spectra, so broader training sets or transfer learning are the natural route to wider applicability."],"supporting_citations":[{"why":"Supplies the central formula for IR spectra from the autocorrelation of the dipole time derivative and the trajectory-length guidance the workflow adopts.","marker":"[17]"},{"why":"The equivariant message-passing neural-network architecture used both for the potential-energy/force ensemble and for the dipole-moment model.","marker":"[36, 37]"},{"why":"Establishes committee force disagreement as the uncertainty signal that drives the active-learning data selection.","marker":"[56, 57]"},{"why":"The all-electron DFT code that computes the reference energies, forces, and dipole moments, including the ab initio molecular dynamics spectra used as benchmarks.","marker":"[59-62]"},{"why":"Provides the baseline-correction procedure and the correlation/transport-distance metrics used to quantify spectral similarity.","marker":"[65]"},{"why":"The experimental gas-phase reference spectra against which predicted peak positions and amplitudes are compared.","marker":"[66]"}],"fun_headline_variants":["Active learning cuts IR spectra cost 100-fold","IR spectra via active learning: 100x cheaper than DFT","Machine-learned potentials slash IR spectra simulation cost","Active learning predicts IR spectra at 1/100th DFT cost"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The workflow assumes that the configurations picked because the atomic-force predictions disagreed are also enough to train the separate dipole-moment model, so that model stays accurate along the trajectories that actually produce the spectrum.","fun_headline_variants_meta":{"raw":{"variants":["Active learning cuts IR spectra cost 100-fold","IR spectra via active learning: 100x cheaper than DFT","Machine-learned potentials slash IR spectra simulation cost","Active learning predicts IR spectra at 1/100th DFT cost"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001073,"raw_usage":{"total_tokens":4475,"prompt_tokens":912,"completion_tokens":3563,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":3497}},"tokens_in":528,"tokens_out":3563,"duration_ms":26447,"temperature":1.0,"reasoning_tokens":3497,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:59:50.824756+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a molecule outside the training distribution, such as one with the C=C motif beyond ethene, run the production trajectories, compute DFT dipole moments at snapshots where the force ensemble disagreed most, and compare them with the dipole model; if the dipole errors are large and the spectrum changes when training data are instead selected by dipole disagreement, the force-uncertainty proxy is refuted.","supporting_citations":[{"cited_title":"Journal of Chemical Theory and Computation17(2), 985–995 (2021) https://doi.org/10.1021/acs.jctc.0c01279","cited_arxiv_id":null,"evidence_quote":"Provides the baseline-correction procedure and the correlation/transport-distance metrics used to quantify spectral similarity."}],"review_version":2}