{"id":"ff644216-a36f-466d-b253-efdf04416650","arxiv_id":"2412.04993","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A nearest-neighbor-like method using sigma-profile similarity predicts infinite-dilution activity coefficients more accurately than modified UNIFAC and COSMO-SAC variants on a DDB data set.","lead":"This paper introduces a similarity-based imputation method (SBM) that predicts infinite-dilution activity coefficients of binary mixtures by averaging experimental values of mixtures with chemically similar components, using quantum-chemical sigma-profiles as the similarity measure. The authors report that the SBM beats established physical models such as modified UNIFAC and COSMO-SAC on a sparse experimental database, with accuracy within typical experimental uncertainty.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Selection of wσ, wP, and ξ on the full evaluation set makes the reported SBM-vs-benchmark gains in-sample estimates; nested resampling is required before the central claim is established.","rationale":"The reader's stated weakest_assumption is the similarity proxy premise (similar components imply similar γ∞). That is a real limitation for generalization, but it is not the main threat to the central benchmark claim, because the paper's headline is an empirical comparison on a specific DDB set. The decisive vulnerability is that the SBM's hyperparameters and threshold are selected using the same data that are then used to measure performance. The reader's rationale does mention hyperparameter selection on the evaluation set, but the weakest_assumption field points elsewhere, so my agreement is partial. This concern is load-bearing because every reported advantage (ξ=0.85 vs COSMO-SAC-dsp, ξ=0.87 vs modified UNIFAC, ξ=0.62 vs COSMO-SAC) is quoted from a curve generated after full-data model selection. Leave-one-out removes the target point for prediction, but the choice of wσ, wP, and ξ still sees all target values; therefore the quoted MAEs are optimistically biased estimates of the procedure's out-of-sample performance. The concrete nested-resampling test directly settles whether the advantage survives when hyperparameters are chosen without seeing the evaluated points. I do not see a more fundamental flaw: the method is clearly described, the leave-one-out mechanics are internally consistent, the benchmark comparison is in favor of the physical models (UNIFAC outliers are removed in a way that improves the benchmark), and the similarity heatmaps provide plausible face validity. The main fix is statistical: redo the model selection inside cross-validation, report paired comparisons on common sets, and add uncertainty quantification. If the nested test preserves the advantage, the central claim would be substantially strengthened; if not, the reported edge is largely selection-driven. Since the reader already returned CONDITIONAL, my stress-test does not change that verdict; it sharpens the specific condition that must be met: nested, selection-free evaluation on held-out data.","tokens_in":12903,"tokens_out":7881,"duration_ms":280726,"concrete_test":"Run a nested resampling benchmark. Split the 3,568 binary systems into 5 folds at the level of solute-solvent pairs. For each training fold, perform the same grid search over wσ∈{0,...,1}, wP∈{0,2}, and ξ∈{0.5,...,1} using only that fold's data (with an inner validation split to choose ξ), freeze the selected configuration, and then predict the held-out fold. Aggregate MAE and coverage across folds. Compare these held-out SBM numbers against modified UNIFAC (Dortmund), COSMO-SAC, and COSMO-SAC-dsp evaluated on exactly the same held-out points. If the aggregate held-out MAE/coverage no longer dominates the benchmarks at any ξ value, or only at substantially lower coverage, the headline claim is an artifact of selection on the test set. Report bootstrap 95% confidence intervals for the MAE difference at the chosen operating point.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reported comparison is not a valid out-of-sample evaluation. The SBM has three free choices: wσ (11 values), wP (2 values), and ξ (51 values). The authors first run the full leave-one-out grid over wσ and wP on the entire DDB set and select wσ=0.6, wP=2 as 'best' using MAE and scope computed from those same leave-one-out predictions (Fig. 3). They then scan ξ on the same full data and quote the favorable operating points ξ=0.85, 0.87, and 0.62 as evidence that an SBM variant outperforms COSMO-SAC-dsp and modified UNIFAC. Because every target value contributes to the model-selection criterion and to the ξ choice, the leave-one-out errors are selected, in-sample statistics; they are not estimates of how the procedure would perform when ξ and the weights are fixed before seeing the data to be predicted. The central claim 'one can always find a variant... that outperforms' is therefore not established by the numbers in Fig. 4. The similarity premise is a separate limitation; the more immediate threat to the headline is that the performance curve itself is selected from the evaluation set. The authors' note that physical models may have seen some of the data does not correct this in-sample selection; at best it makes the comparison conservative in a different direction.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a similarity-based method (SBM) for predicting infinite-dilution activity coefficients ln γ_ij^∞ of binary mixtures. The similarity between two components is computed from COSMO σ-profiles and surface areas via a weighted overlap measure (Eqs. 1–6), and the activity coefficient of a target mixture is predicted by arithmetically averaging the experimental values of mixtures that share one component with the target and whose other component has a similarity above a threshold ξ. Using leave-one-out validation on 3,568 DDB datapoints (221 solutes, 198 solvents at 298.15 K), the authors select hyperparameters w_σ, w_P, and ξ by grid search and report that some SBM variant outperforms modified UNIFAC (Dortmund), COSMO-SAC, and COSMO-SAC-dsp in both accuracy (MAE) and scope. The method is presented as generic and transferable to other binary-mixture properties.","tokens_in":13112,"tokens_out":4868,"duration_ms":55725,"significance":"If the reported performance survives a properly nested evaluation, the SBM would be a useful, conceptually simple baseline for mixture-property prediction: it requires only σ-profiles (available from the open-source Bell et al. database), involves no regression on the target values, and is transparent about the accuracy/scope trade-off controlled by ξ. The paper's strengths include a clearly specified algorithm, a leave-one-out protocol that genuinely removes the target value from the averaging step (so there is no direct leakage of the target into its own prediction), open quantum-chemical descriptors, and an honest acknowledgment that the physical benchmark models may have seen parts of the training data. The main weakness is that the hyperparameters and the threshold ξ are selected on the same data that are then used for the headline benchmark comparison, so the reported MAE values are in-sample statistics of a model-selection procedure; this is the load-bearing issue that must be addressed before the central claim can be accepted.","major_comments":[{"comment":"The procedure selects w_σ and w_P on the full DDB set using the leave-one-out MAE/scope computed from that same set (Fig. 3, 'Studied Model Variants', Eqs. 1 and 4), and then sweeps ξ on the same full set and quotes favorable operating points (ξ = 0.85, 0.87, 0.62) as evidence for the headline claim. Because every target value contributes both to model selection and to the reported error, the quoted MAE values are optimistic in-sample estimates of the whole procedure. The statement 'one can always find an SBM variant (by varying ξ) that outperforms it' is an existential claim over a parameter tuned on the evaluation set and is therefore not established. Please re-run the evaluation with a nested scheme: select (w_σ, w_P) and ξ on a training portion (or use an inner CV), then report errors on an untouched test portion; alternatively, report the performance of an a-priori fixed rule for choosing ξ (e.g., a target scope). At a minimum, report the distribution of test-set MAE over multiple random splits to show that the advantage is not an artifact of the selection step.","section":"Results and Discussion, 'Overall Performance of Different Similarity-Based Methods' and Fig. 4"},{"comment":"The headline comparisons are not made on a common subset of mixtures: at ξ = 0.85 the SBM covers N = 3,301 points while COSMO-SAC-dsp covers N = 3,199, and at ξ = 0.87 the SBM covers N = 3,115 versus N = 2,987 for modified UNIFAC (Dortmund). The MAE values are computed on different, only partially overlapping subsets, so a direct comparison of accuracy at different scopes is not apples-to-apples. The histograms later restrict all methods to the 1,748 common points, but that matched analysis is only shown for ξ = 0.93, not for the quoted operating points. Please report, for each ξ, the MAE of all methods on the intersection of their predictable sets (or on a fixed common set), and state which mixtures are excluded in each comparison; otherwise the combined claim of 'higher accuracy and broader scope' is not fully supported by the displayed numbers.","section":"Results and Discussion, 'Comparison to Physical Benchmark Models', Fig. 4"},{"comment":"No uncertainty quantification is provided for any MAE or scope value. The quoted margins are small (e.g., 0.27 vs 0.33 against modified UNIFAC at ξ = 0.87, and 0.30 vs 0.61 against COSMO-SAC-dsp at ξ = 0.85), and the MAE values vary smoothly with ξ (Fig. 5a), so without confidence intervals (bootstrap over the leave-one-out residuals, or across repeated splits) the reader cannot assess whether the differences are significant. This is especially important because the model-selection bias discussed above and the differing subsets discussed in the previous comment both affect the magnitude of the reported gains. Please add standard errors or bootstrap confidence intervals to the MAE values in Figs. 3–5 and to the quoted numbers.","section":"Results and Discussion, Fig. 5 and all reported MAE values"},{"comment":"The removal of eight outliers from modified UNIFAC (Dortmund) is statistically consequential: the MAE decreases from 0.6477 to 0.3340 after their removal. This filtering is reasonable and it actually makes the comparison harder for the SBM, but the criterion for outlier removal should be specified a priori (e.g., defined by a deviation threshold or by a documented DDB quality flag) rather than being applied only to the benchmark. Please also show the sensitivity of the comparison in Fig. 4 to including or excluding these eight points, and clarify whether any analogous outlier handling is applied to the SBM or to the COSMO variants. Without this, the benchmark comparison is not fully reproducible.","section":"Supporting Information, Table S.1 and 'Outliers of Modified UNIFAC (Dortmund)'"}],"minor_comments":[{"comment":"The caption contains the typo 'COMSO-SAC' (twice) for COSMO-SAC; please correct it, as it may confuse readers searching for the model name.","section":"Fig. 4 caption"},{"comment":"The moving-average window width (2 bins, corresponding to 0.002 e/Å^2) is a fixed hyperparameter that is not varied in the grid search; the paper should state explicitly that this choice was not optimized, or, if it was tested in preliminary studies, give the range explored.","section":"Similarity-Based Method, Eq. (6)"},{"comment":"Please clarify how predictions are formed when the target mixture has both similar solvents (same solute) and similar solutes (same solvent) with scores above ξ: is the final prediction a simple average over the union of all available neighboring mixtures, or is one side preferred? The description of the averaging step is ambiguous and affects the leave-one-out protocol.","section":"Prediction of Activity Coefficients"},{"comment":"The DDB data are available only under license, so the full dataset cannot be shared. Please state whether the preprocessed data matrix (the 221 × 198 matrix used in the study) can be made available in a form that does not violate the DDB license, or provide a synthetic demonstration dataset so that the pipeline can be re-run by readers.","section":"Data Availability Statement"},{"comment":"In Fig. 5b the legend labels 'SBM(ξ)' and 'SBM(ξ = 0.93)' are redundant because the same line is plotted in both panels; consider labeling the highlighted operating point directly in the caption to avoid confusion.","section":"Results and Discussion, Fig. 5"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The central problem is a model-selection-on-the-evaluation-set issue. The per-point leave-one-out prediction itself is honest (the target is removed before averaging), but the grid search over w_σ and w_P and the sweep over ξ are performed on the full dataset, and the quoted operating points are chosen from the resulting performance curves. This makes the headline SBM-vs-benchmark numbers in-sample. I would require a nested resampling evaluation (or a fixed a-priori ξ-selection rule) as a condition for publication. The matched-subset comparison and uncertainty intervals are secondary but important. The approach is otherwise clearly presented and the idea is sound; the authors' own note about benchmark models having seen some data is a valid point in the other direction, but it does not address the SBM's own selection bias. I did not find evidence of direct circularity in the averaging step, and I found the disclosure of the modified UNIFAC outlier removal to be transparent, though the removal criterion needs to be made more precise."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper introduces a similarity-based method (SBM) for infinite-dilution activity coefficients. The idea is genuinely simple and appealing: compute a similarity between two components from their COSMO sigma-profiles (with polar weighting and smoothing) and their surface-area ratio, then predict an unknown gamma_infinity by averaging experimental values from mixtures sharing a similar solute or solvent above a threshold xi. What's new is the specific similarity score and its use as a neighbor kernel for leave-one-out imputation. Prior work used fingerprints or matrix completion; I haven't seen this exact construction. The method is clearly specified, the LOO protocol is sound in itself, and the authors are transparent about data preprocessing and about removing modified UNIFAC outliers in a way that favors the benchmark, which is honest.\n\nThe soft spot is the validity of the headline comparison. The weights w_sigma and w_P are tuned on the full dataset via grid search, and then xi is scanned on the same dataset, with the reported favorable operating points (xi=0.85, 0.87) selected from that same curve. Every target value contributes to the model-selection criterion, so the MAE/N curves are in-sample statistics. The claim that 'one can always find an SBM variant that outperforms' is not established by these numbers; it is a post-hoc selection. This is fixable: nested cross-validation or a genuinely held-out test set would settle it. I would also want error bars on the MAE differences, and ideally a public benchmark subset of the data since the full DDB data are licensed.\n\nA separate, more fundamental limitation is the similarity premise itself: that similar components give similar gamma_infinity in any partner. That will break down for specific interactions like hydrogen bonding, and the authors implicitly acknowledge this by noting water has no similar components. This doesn't kill the method for many practical cases, but it bounds its generality.\n\nOverall, the paper is worth serious refereeing. The idea is simple, the implementation is clean, and the intended use (gap-filling and experiment planning) is real. But the performance claim needs a proper out-of-sample test before I'd believe the comparison to UNIFAC/COSMO-SAC. My recommendation: send to peer review, require nested resampling and error bars.","headline":"A clean, simple imputation idea that is plausibly useful, but the headline comparison to UNIFAC/COSMO-SAC is not out-of-sample because hyperparameters are tuned on the same data; needs nested resampling before the central claim is established.","tokens_in":13699,"tokens_out":1710,"would_cite":false,"duration_ms":18637,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Similarity-based imputation with quantum-chemical descriptors predicts infinite-dilution activity coefficients more accurately than modified UNIFAC and COSMO-SAC-dsp on the same dataset.","keywords":["activity coefficients","infinite dilution","similarity-based method","sigma-profiles","COSMO","imputation","UNIFAC","quantum-chemical descriptors"],"falsifier":"A direct test would be to pick a solute-solvent pair whose most similar replacement in the database scores above $\\xi$ but is known to interact via a specific mechanism the $\\sigma$-profile overlap underweights, such as a hydrogen-bond donor/acceptor mismatch or steric shielding, and compare the SBM's averaged prediction with a fresh measurement at 298.15 K. More systematically, one could remove all alkanes from the training matrix and ask whether SBM predictions for alkane-containing mixtures degrade; if they stay accurate, the similarity score transfers, and if they collapse, the apparent performance is carried by dense similar neighbors rather than by the descriptor.","tokens_in":12656,"feed_emoji":"🧪","tokens_out":10301,"duration_ms":86777,"temperature":0.7,"pith_summary":"The paper introduces the similarity-based method (SBM), which predicts a mixture property by treating the experimental database as a sparse matrix and imputing missing entries from mixtures with similar component pairs. Component similarity is measured from quantum-chemical $\\sigma$-profiles and surface areas, and the paper applies the method to infinite-dilution activity coefficients $\\gamma^\\infty_{ij}$ at 298.15 K. The central claim is that, after tuning a threshold $\\xi$ and two weighting hyperparameters, the SBM predicts $\\ln\\gamma^\\infty_{ij}$ more accurately than modified UNIFAC (Dortmund), COSMO-SAC, and COSMO-SAC-dsp while covering at least as many data points. If correct, this gives a simple, descriptor-based route to thermodynamic data for unstudied binaries and a generic recipe for other mixture properties.","feed_headline":"Similarity imputation beats physics-based activity-coefficient models","feed_subtitle":"Using sigma-profiles and measured data, the method matches or beats modified UNIFAC and COSMO-SAC-dsp.","key_machinery":"The load-bearing object is the similarity score $S_{mn}$ (Eq. 1), a number between 0 and 1 built from two pieces: $S^\\sigma_{mn}$, the bin-wise overlap $\\sum_k \\min(\\bar p_m(\\sigma_k), \\bar p_n(\\sigma_k))$ of modified $\\sigma$-profiles with polar regions optionally emphasized by $w_P$, and $S^A_{mn}$, the ratio of the smaller to the larger cavity surface area. These are combined with weight $w_\\sigma$. A threshold $\\xi$ selects which solvent or solute replacements count as similar, and the prediction is the arithmetic mean of the corresponding experimental $\\ln\\gamma^\\infty_{ij}$ values from a leave-one-out training set. The score's role is to decide, for each unstudied pair, which measured mixtures are relevant enough to average.","core_discovery":"The paper's central claim is that pairwise similarity of components, computed as a weighted overlap of polar-enhanced $\\sigma$-profiles combined with surface-area ratio, is enough information to predict infinite-dilution activity coefficients by imputation. For a target solute $i$ and solvent $j$, the SBM collects experimental $\\ln\\gamma^\\infty_{ij}$ values from mixtures where the other partner is replaced by a component with similarity above a threshold $\\xi$, and averages them. In leave-one-out testing on the Dortmund Data Bank set of 3,568 points (221 solutes, 198 solvents), the method with $w_\\sigma=0.6$, $w_P=2$ can be tuned so that, e.g., at $\\xi=0.85$ it covers 3,301 points with MAE 0.30 versus COSMO-SAC-dsp's 3,199 points with MAE 0.61, and at $\\xi=0.87$ it covers 3,115 points with MAE 0.27 versus modified UNIFAC (Dortmund)'s 2,987 points with MAE 0.33. The paper states that for every physical benchmark there is a threshold choice at which the SBM is both more accurate and broader in scope. The approach is presented as transferable to any binary mixture property for which a partially filled data matrix exists.","pith_inferences":["A testable extension of the paper's logic is to run the SBM on a different property, such as excess enthalpy or infinite-dilution selectivity, with the same $w_\\sigma=0.6$, $w_P=2$ weights; if the score is genuinely generic, the Pareto front should still beat group-contribution baselines.","The paper's benchmark comparison leaves a gap for components with no close neighbor, with water as the highlighted example; a natural complement would be a hybrid that falls back on COSMO-SAC or UNIFAC when the nearest similarity score falls below threshold.","Because leave-one-out only removes one mixture at a time, the reported accuracy tests interpolation within the existing database coverage; performance on entirely new chemical classes, where similarity to known components is low, is not yet measured and is the main extrapolation risk.","The similarity matrix itself could be used to select the next measurements that most increase matrix coverage, so the method couples naturally to an active-learning loop."],"forward_implications":["For the considered database and temperature, a user can choose $\\xi$ to obtain a model that is simultaneously more accurate and broader in coverage than modified UNIFAC (Dortmund) or COSMO-SAC-dsp.","Because only $\\sigma$-profiles, surface areas, and an experimental data matrix are needed, the same imputation recipe applies to other binary properties such as excess enthalpies or vapor-liquid equilibria.","At $\\xi=0.93$, over half the database is predictable with most deviations within $\\pm 0.1$ in $\\ln\\gamma^\\infty_{ij}$, i.e. within common experimental uncertainty.","The few highly similar neighbors per mixture mean targeted measurement of one representative system can yield accurate predictions for a cluster of systems, supporting proxy-substance strategies and design-of-experiments planning."],"supporting_citations":[{"why":"Supplies the experimental infinite-dilution activity coefficients that form the imputation matrix and the benchmark evaluation points.","marker":"[11]"},{"why":"Provides the open-source $\\sigma$-profiles and COSMO-SAC/COSMO-SAC-dsp implementations used for descriptors and benchmarks.","marker":"[35]"},{"why":"The modified UNIFAC (Dortmund) benchmark that the SBM is claimed to outperform.","marker":"[17]"},{"why":"The COSMO-SAC-dsp benchmark used as a second physical comparison.","marker":"[22]"},{"why":"The COSMO-SAC benchmark with the widest scope that the SBM matches at $\\xi=0.62$ while reaching a lower MAE.","marker":"[21]"},{"why":"Introduces $\\sigma$-profiles, the quantum-chemical descriptors underlying the similarity score.","marker":"[7]"},{"why":"Defines the leave-one-out validation procedure that produces the reported MAE and scope values.","marker":"[36]"}],"fun_headline_variants":["Similarity imputation beats UNIFAC and COSMO-SAC","Sigma-profile similarity predicts activity coefficients better","Data-driven imputation outperforms classic activity models","Pairwise similarity imputation wins over physics-based methods"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that two components with similar charge-density profiles and surface areas will behave similarly enough in any given partner that their measured activity coefficients can be averaged; if that similarity is not a reliable proxy, the imputed prediction inherits errors from chemically incompatible mixtures.","fun_headline_variants_meta":{"raw":{"variants":["Similarity imputation beats UNIFAC and COSMO-SAC","Sigma-profile similarity predicts activity coefficients better","Data-driven imputation outperforms classic activity models","Pairwise similarity imputation wins over physics-based methods"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000293,"raw_usage":{"total_tokens":1751,"prompt_tokens":1035,"completion_tokens":716,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":651,"completion_tokens_details":{"reasoning_tokens":653}},"tokens_in":651,"tokens_out":716,"duration_ms":9077,"temperature":1.0,"reasoning_tokens":653,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:00:20.514272+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be to pick a solute-solvent pair whose most similar replacement in the database scores above $\\xi$ but is known to interact via a specific mechanism the $\\sigma$-profile overlap underweights, such as a hydrogen-bond donor/acceptor mismatch or steric shielding, and compare the SBM's averaged prediction with a fresh measurement at 298.15 K. More systematically, one could remove all alkanes from the training matrix and ask whether SBM predictions for alkane-containing mixtures degrade; if they stay accurate, the similarity score transfers, and if they collapse, the apparent performance is carried by dense similar neighbors rather than by the descriptor.","supporting_citations":[{"cited_title":"2023; www.ddbst.com","cited_arxiv_id":null,"evidence_quote":"Supplies the experimental infinite-dilution activity coefficients that form the imputation matrix and the benchmark evaluation points."},{"cited_title":"H.; Mickoleit, E.; Hsieh, C.-M.; Lin, S.-T.; Vrabec, J.; Breitkopf, C.; J \\\"a ger, A","cited_arxiv_id":null,"evidence_quote":"Provides the open-source $\\sigma$-profiles and COSMO-SAC/COSMO-SAC-dsp implementations used for descriptors and benchmarks."},{"cited_title":"Considering the dispersive interactions in the COSMO-SAC model for more accurate predictions of fluid phase behavior","cited_arxiv_id":null,"evidence_quote":"The COSMO-SAC-dsp benchmark used as a second physical comparison."},{"cited_title":"I.; Lin, S.-T","cited_arxiv_id":null,"evidence_quote":"The COSMO-SAC benchmark with the widest scope that the SBM matches at $\\xi=0.62$ while reaching a lower MAE."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the leave-one-out validation procedure that produces the reported MAE and scope values."}],"review_version":1}