{"id":"d39258ec-63f0-435c-b615-8252aa5a2cf9","arxiv_id":"2412.12962","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Modified UNIFAC 2.0 fills all group-pair interaction parameters using low-rank matrix completion trained on over 500,000 data points, improving accuracy and scope over modified UNIFAC (Dortmund).","lead":"The paper merges a machine-learning matrix completion model into the modified UNIFAC group-contribution framework to fill all missing group-pair interaction parameters, producing a complete public parameter table for 63 main groups. A smart generalist might read it because complete, accurate mixture-property predictions would let chemical engineers simulate more vapor-liquid equilibria without new experiments or proprietary databases.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The complete-table claim rests on the rank-8 factorization of Eqs. (2)-(3), but the 594 group pairs with no training data are never validated because the Sec. 3.3 holdout covers only pairs that have data.","rationale":"The reader's conditional verdict already identifies the low-rank factorization in Eqs. (2)-(3) as the weakest assumption, and my read agrees. The added sharpening is that the Sec. 3.3 extrapolation study, while valuable, is structurally unable to validate the 594 completely unobserved group pairs: every one of the 100 selected test sets has data, so the test exercises interpolation within the observed matrix rather than the advertised completion of empty rows/columns. This is the load-bearing point because the paper's main practical contribution is the complete parameter table, and for those 594 entries there is no direct experimental check. The proposed rank-sensitivity and posterior-uncertainty test would give an inexpensive, concrete indication of whether the imputed values are stable or merely artifacts of K=8. Since the paper is otherwise technically sound and the deliverable is useful, the concern does not overturn the conditional recommendation; it strengthens the need for the requested validation before the tables are adopted as a default.","tokens_in":19700,"tokens_out":7497,"duration_ms":75391,"concrete_test":"Retrain the full model with K=4, K=8, and K=16 on the identical training set, then compare the imputed a_mn and b_mn values for the 594 group pairs that have no training data (Fig. S.1b). If these imputed parameters, or the predicted ln(gamma) and hE for representative binary mixtures requiring those pairs, change materially across K, the published point estimates are artifacts of the chosen rank rather than robust model predictions. Also report posterior predictive intervals for those 594 pairs from the existing variational posterior; intervals that cross zero or span a large fraction of the fitted parameter range would confirm that the imputation is unconstrained.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central promise of modified UNIFAC 2.0 is that its parameter tables are complete, so it can predict any mixture decomposable into the 125 subgroups. For the 594 group combinations (30% of the matrix) with no experimental data, the only mechanism producing parameters is the low-rank factorization a_mn = theta_a_m dot beta_a_n and b_mn = theta_b_m dot beta_b_n with K=8. This is an ad hoc structural assumption without physical derivation. The validation in Sec. 3.3 does not test the fully unobserved regime: the 100 randomly withheld group combinations listed in Tables S.1 and S.2 all have nonzero numbers of data points and mixtures, so they lie in the observed part of the matrix. The test demonstrates interpolation among data-supported pairs, not completion of unsupported entries. Some held-out results are already poor (e.g., hE MAE of 3.86 kJ/mol for pair 18-44 and 2.25 kJ/mol for pair 5-17), indicating uneven imputation quality even where data exist. Without evidence that K=8 imputations for the zero-data pairs are stable and physically sensible, the published 'complete' tables may be numerical artifacts, and the expanded-scope claim is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents modified UNIFAC 2.0, a group-contribution model that replaces the sequential parameter fitting of modified UNIFAC (Dortmund) with an end-to-end trained matrix completion approach. Pair-interaction parameters are modeled as low-rank latent feature products (Eqs. (2) and (3)), enabling simultaneous training of all a_mn and b_mn parameters and completion of the parameter table. The model is trained on more than 500,000 DDB data points for activity coefficients and excess enthalpies and is compared with the latest published modified UNIFAC version (called mod. UNIFAC 1.0). The authors report improved accuracy on the shared prediction horizon and a large expansion of applicability, and they provide complete parameter tables in the Supporting Information.","tokens_in":19945,"tokens_out":3576,"duration_ms":33933,"significance":"If the central claim holds, the contribution is significant and practically useful: it provides a freely available, complete parameterization of modified UNIFAC for 63 main groups, which can be dropped into existing process simulators. The paper also demonstrates a promising methodological template for embedding matrix completion into established thermodynamic group-contribution frameworks. The end-to-end training on a large experimental database, the provision of complete parameter tables as .csv files, and the inclusion of a held-out component test and a held-out pair-parameter test are concrete strengths that go beyond what is typical for parameter-fitting papers. The main scientific risk is that the completeness of the parameter table relies on a low-rank assumption that is not directly validated for the 30% of group pairs that have no training data, and the headline accuracy comparison is partly in-sample for the new model.","major_comments":[{"comment":"The central claim of complete predictive scope rests on the parameters imputed for the 594 group combinations (30% of the matrix) for which no experimental data are available (Supporting Information, Sec. S.1). The extrapolation test in Sec. 3.3 withholds 100 group combinations, but all of these, as listed in Tables S.1 and S.2, have nonzero numbers of data points and mixtures; they test interpolation among data-supported pairs, not the fully unobserved regime. To support the completeness claim, the authors should provide additional evidence for the zero-data entries, for example a stability analysis of the imputed parameters under retraining with different random seeds, a synthetic-data recovery test for a matrix with known low-rank structure, or a cross-validation scheme that removes entire rows/columns of the parameter matrix. Without such evidence, the imputed values for the 594 unsupported pairs may be numerical artifacts, and the claim of unlimited applicability within the 125 subgroups is not established.","section":"Sec. 3.3 and Sec. S.1"},{"comment":"The headline comparison showing improved accuracy of mod. UNIFAC 2.0 over mod. UNIFAC 1.0 is partly a comparison of an in-sample model with a model that was not trained on the evaluation data. Mod. UNIFAC 2.0 is trained on the full DDB data set, including data measured after 2016, whereas mod. UNIFAC 1.0 was published in 2016 and could not have used those data. The authors acknowledge this only by saying that it is 'reasonable to assume' the training sets are similar. This is not a substitute for a controlled evaluation: the results in Figs. 2 and 3 should be reported for a version of mod. UNIFAC 2.0 trained only on data available before 2016, or the claims should be explicitly restricted to the held-out experiments of Secs. 3.2 and 3.3, which are more convincing.","section":"Figs. 2 and 3 and Sec. 3.1"},{"comment":"Mod. UNIFAC 2.0 drops the c_mn parameters that appear in the original modified UNIFAC equation (Eq. (1)). The authors justify this by stating that c_mn was fitted for only very few group pairs in mod. UNIFAC 1.0, but those few fitted c_mn values can be important for the systems they describe. The comparison on the shared horizon therefore changes the model form as well as the fitting procedure. The paper should quantify how many group pairs in mod. UNIFAC 1.0 use nonzero c_mn and discuss the effect of setting them to zero on the comparison for those pairs.","section":"Sec. 2.1 and Eq. (1)"},{"comment":"The 100 test sets in the pair-extrapolation experiment are based on a single random selection, and many contain very few data points (e.g., one mixture with one data point for pairs 16-35, 17-28, 28-39, 33-34, 33-35, 36-40 in Table S.1). The aggregate box plots in Fig. 8 and Fig. S.6 obscure large per-pair failures, such as hE MAE of 3.86 kJ/mol for pair 18-44 and 2.25 kJ/mol for pair 5-17 in Table S.2. The authors should report the distribution of per-pair errors, relate them to the amount of held-out data, and ideally repeat the random selection several times to show that the aggregate result is robust. As it stands, the claim that 'true predictions achieve comparable accuracy' is based on a summary that is dominated by a few large test sets.","section":"Sec. 3.3 and Tables S.1/S.2"}],"minor_comments":[{"comment":"The latent dimension K=8 is stated to have been determined in preliminary studies, but no details or sensitivity analysis are provided. A short paragraph or a reference showing how K was chosen would help the reader assess the robustness of the low-rank assumption.","section":"Sec. 2.1"},{"comment":"The derivation of activity coefficients from VLE data uses the extended Raoult's law with an ideal gas phase and neglects the pressure dependence of the liquid-phase chemical potential. The authors note this limitation but do not quantify its impact for the 10 bar pressure limit. A brief estimate of the expected errors in ln gamma from these approximations would strengthen the data-preprocessing description.","section":"Sec. 2.2 and Eq. (8)"},{"comment":"The box plots are informative, but they do not show the number of mixtures per box. Since sample sizes vary strongly between the 'mod. UNIFAC 1.0 horizon' and 'mod. UNIFAC 2.0 only' sets, the numbers should be given in the captions or in the text.","section":"Figs. 2, 3, 7, 8, S.4, S.6"},{"comment":"The statement that the mean of the MAE is 'nearly halved' with mod. UNIFAC 2.0 is not supported by explicit numbers. Please report the exact mean and median MAE values for both models and for both the ln gamma and hE results, either in the text or in a table.","section":"Sec. 3.1"},{"comment":"The paper says mod. UNIFAC 2.0 outperforms mod. UNIFAC 1.0 for 'most group combinations' and reports 461 improved versus 267 deteriorated combinations. It should be stated what the total number of compared group combinations is, so the reader can see the fraction of cases where the new model is worse.","section":"Sec. 3.1 and Fig. 4b"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong candidate for this journal if the authors can address the validation of the fully unobserved parameter entries and qualify the in-sample comparison. The self-citation of the authors' prior MCM framework is appropriate and does not raise novelty concerns. The scope fits the journal well. I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The deliverable here is real: a complete set of modified UNIFAC pair-interaction parameters for 63 main groups, trained end-to-end on over 500,000 data points including excess enthalpies, released as CSV files. That is a genuine extension of the same group's UNIFAC 2.0, and for practitioners it removes the main blocker of the public modified UNIFAC. The paper deserves credit for shipping the parameter tables, for the unseen-component test, and for the pair-holdout test—even with its limits, it is more evidence than most UNIFAC papers provide.\n\nThe central soft spot is exactly what the stress-test note identifies. The 100 withheld group pairs in Section 3.3 all have nonzero data—they are interpolation tests. The 594 group pairs with no experimental data (30% of the matrix) are filled purely by the rank-8 factorization of Eqs. (2) and (3), and those imputed parameters are never validated. The paper's headline claim of expanded scope stands or falls on those entries, so this is a load-bearing gap, not a quibble. Some of the held-out pair results are also poor even where data exist (hE MAE of 3.86 kJ/mol for pair 18-44), suggesting the imputation quality is uneven. A repetition of the pair-holdout analysis with multiple random splits, plus sanity checks on the zero-data entries (e.g., comparing against the commercial UNIFAC Consortium table where available, or checking physical plausibility) would materially strengthen the paper.\n\nThe comparison to mod. UNIFAC 1.0 is also red-shifted: Figs. 2 and 3 evaluate a model trained on most of the same data against a 2016 model that was not. The authors acknowledge this and the unseen-component test partly compensates, but the headline \"nearly halved MAE\" should be read with that caveat. Minor issues: single random split in the pair test, no uncertainty quantification on the imputed parameters, and training data that are not publicly accessible (DDB is commercial). None of these are fatal, but they matter for a paper that recommends replacing the industrial standard.\n\nWho is this for? Anyone using UNIFAC in process simulation or phase-equilibrium work. It is not a conceptual reordering, but it is a practical step forward. It deserves a serious referee; I would send it to review, but the referee should push hard on the zero-data validation and ask for repeated splits and sensitivity on K.\n\nFor your reading group: maybe—good for a methods discussion, but the paper itself is straightforward. I would cite it if I needed a complete modified UNIFAC parameter set.","headline":"A useful, practical completion of modified UNIFAC, but the claim of full predictive scope rests on a low-rank prior that is never validated on the truly zero-data group pairs.","tokens_in":20532,"tokens_out":1253,"would_cite":true,"duration_ms":14542,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Modified UNIFAC 2.0 claims to complete all group-pair interaction parameters via machine-learned matrix completion and to nearly halve the mean error on the shared data horizon.","keywords":["modified UNIFAC","group-contribution method","matrix completion","activity coefficients","excess enthalpy","pair-interaction parameters","machine learning","vapor-liquid equilibrium"],"falsifier":"Measure activity coefficients and excess enthalpies for binary mixtures whose group pairs have no training data, roughly 594 of the 1953 group combinations, and compare the results with modified UNIFAC 2.0 predictions; if the mean absolute error on these physically measured unseen pairs substantially exceeds the model's reported error on seen pairs, the latent-space completion is filling gaps with artifacts rather than chemistry.","tokens_in":19413,"feed_emoji":"🧪","tokens_out":5557,"duration_ms":49185,"temperature":0.7,"pith_summary":"The paper tries to establish that the main limitation of the workhorse group-contribution method modified UNIFAC, namely missing pair-interaction parameters, can be removed by machine-learned matrix completion without changing the physical model equations. It presents modified UNIFAC 2.0, trained end-to-end on more than 500,000 experimental data points for activity coefficients and excess enthalpies, yielding a complete set of parameters for all 63 main groups. On mixtures both models can describe, the new model nearly halves the mean error in predicted logarithmic activity coefficients; on mixtures the older model cannot describe at all, the new model still performs well. If the claim holds, users get a drop-in replacement parameter table that extends prediction to any mixture decomposable into the model's 125 subgroups.","feed_headline":"Machine learning fills every missing UNIFAC interaction pair","feed_subtitle":"Trained on 500,000 data points, the updated model nearly halves average error and widens the scope of mixture predictions.","key_machinery":"The machinery is matrix completion by low-rank factorization. Each pair-interaction parameter is written as a dot product of two latent feature vectors of length $K = 8$, namely $a_{mn} = \\theta^a_m \\cdot \\beta^a_n$ and $b_{mn} = \\theta^b_m \\cdot \\beta^b_n$. Fitting all features simultaneously under a Bayesian likelihood with a heavy-tailed Cauchy error model lets every observed data point influence every parameter, so the model imputes values for the 594 group pairs that never appear in the training data. The standard modified UNIFAC equations then generate activity coefficients and excess enthalpies from the completed parameter tables.","core_discovery":"The central claim is that modified UNIFAC 2.0, trained end-to-end on more than 500,000 experimental data points for activity coefficients and excess enthalpies, produces complete pair-interaction parameter tables for all 63 main groups and achieves improved accuracy compared with the latest published modified UNIFAC (Dortmund) while significantly expanding the predictive scope. On the shared comparison horizon, the paper reports that the mean of the mixture-wise mean absolute error in logarithmic activity coefficients is nearly halved, and that on mixtures inaccessible to the older model the new model still reaches accuracy comparable to what the older model achieves where it applies. The temperature dependence is kept through the $a_{mn}$ and $b_{mn}$ parameters, while the rarely fitted $c_{mn}$ parameter is dropped from the model.","pith_inferences":["Beyond the paper, the same low-rank factorization could be applied to other group-contribution frameworks, because it treats the parameter matrix rather than the chemistry as the object to be completed.","A testable extension is to interpret the learned eight-dimensional latent vectors as coarse descriptors of group chemistry, checking whether distances in that space correlate with known functional-group similarity.","Because the Bayesian training yields posterior uncertainty for the features, one could flag group pairs whose imputed parameters carry high variance, guiding which new experiments would most reduce prediction risk."],"forward_implications":["Any binary or multicomponent mixture decomposable into the 125 subgroups can be passed through the model, because no missing group pair blocks the calculation.","Existing process simulators need only swap in the supplied complete parameter tables to gain the expanded predictive scope.","The mean mixture-wise mean absolute error in logarithmic activity coefficients on the shared horizon is nearly halved relative to the older model, with similar improvement seen for excess enthalpies.","Withheld-component and withheld-pair tests show only modest accuracy loss, so predictions for chemistries not directly fitted are plausible.","The training scheme can be rerun as new experimental data arrive, making the model updatable without changing its equations."],"supporting_citations":[{"why":"Defines the baseline modified UNIFAC (Dortmund) parameter set, group definitions, and the comparison horizon.","marker":"[5]"},{"why":"Establishes the parent UNIFAC 2.0 hybrid that first embedded matrix completion in UNIFAC.","marker":"[18]"},{"why":"Supplies the modified UNIFAC equations used to predict activity coefficients.","marker":"[21]"},{"why":"Supplies the modified UNIFAC parameter matrix and equations linking parameters to thermodynamic properties.","marker":"[22]"},{"why":"Shows that matrix completion can predict pair-interaction parameters in thermodynamic mixture models.","marker":"[16]"},{"why":"Provides the experimental activity coefficient and excess enthalpy data used for training.","marker":"[19]"},{"why":"Supplies the matrix-factorization method for recommendation systems that underlies the completion approach.","marker":"[10]"}],"fun_headline_variants":["ML completes UNIFAC tables, nearly halving prediction error","UNIFAC 2.0: hybrid ML fills all missing interaction pairs","500k data points power ML-completed UNIFAC parameter tables","Matrix completion ML expands UNIFAC predictive scope","Modified UNIFAC hybrid: ML predicts every interaction pair"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each missing pair-interaction parameter can be recovered from the dot product of two 8-dimensional latent feature vectors learned from other data; if group-pair interactions are not approximately low-rank in that space, the imputed parameters for the 594 unsupported group pairs are unconstrained artifacts.","fun_headline_variants_meta":{"raw":{"variants":["ML completes UNIFAC tables, nearly halving prediction error","UNIFAC 2.0: hybrid ML fills all missing interaction pairs","500k data points power ML-completed UNIFAC parameter tables","Matrix completion ML expands UNIFAC predictive scope","Modified UNIFAC hybrid: ML predicts every interaction pair"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000288,"raw_usage":{"total_tokens":1656,"prompt_tokens":879,"completion_tokens":777,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":691}},"tokens_in":495,"tokens_out":777,"duration_ms":7894,"temperature":1.0,"reasoning_tokens":691,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:32:58.360165+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure activity coefficients and excess enthalpies for binary mixtures whose group pairs have no training data, roughly 594 of the 1953 group combinations, and compare the results with modified UNIFAC 2.0 predictions; if the mean absolute error on these physically measured unseen pairs substantially exceeds the model's reported error on seen pairs, the latent-space completion is filling gaps with artifacts rather than chemistry.","supporting_citations":[{"cited_title":"Further Development of Modified UNIFAC (Dortmund): Revision and Extension 6","cited_arxiv_id":null,"evidence_quote":"Defines the baseline modified UNIFAC (Dortmund) parameter set, group definitions, and the comparison horizon."},{"cited_title":"Advancing Thermodynamic Group-Contribution Methods by Machine Learning: UNIFAC 2.0","cited_arxiv_id":"2408.05220","evidence_quote":"Establishes the parent UNIFAC 2.0 hybrid that first embedded matrix completion in UNIFAC."},{"cited_title":"A modified UNIFAC model","cited_arxiv_id":null,"evidence_quote":"Supplies the modified UNIFAC equations used to predict activity coefficients."},{"cited_title":"A modified UNIFAC model","cited_arxiv_id":null,"evidence_quote":"Supplies the modified UNIFAC parameter matrix and equations linking parameters to thermodynamic properties."},{"cited_title":"Making thermodynamic models of mixtures predictive by machine learning: matrix completion of pair interactions","cited_arxiv_id":null,"evidence_quote":"Shows that matrix completion can predict pair-interaction parameters in thermodynamic mixture models."},{"cited_title":"2023; www.ddbst.com","cited_arxiv_id":null,"evidence_quote":"Provides the experimental activity coefficient and excess enthalpy data used for training."},{"cited_title":"Matrix Factorization Techniques for Recommender Systems","cited_arxiv_id":null,"evidence_quote":"Supplies the matrix-factorization method for recommendation systems that underlies the completion approach."}],"review_version":1}