{"id":"cf0db563-04c2-4715-9b7a-9b2dd5748a63","arxiv_id":"2502.05576","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Human-guided analysis of AI-sampled stacking sequences yields three descriptors and a compact formula for the thermal conductivity of graphene-WS2 heterostructures.","lead":"AI sampling and human inspection jointly produce three stacking rules (Pa, Pb, Pc) that predict thermal conductivity in graphene-WS2 heterostructures with a simple formula. The work is a case study in human-AI collaboration for materials design, potentially guiding the creation of low-thermal-conductivity layered materials.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. 4's predicted range (about 0.0256 to 0.0289 W/m-K) excludes the paper's reported optimal thermal conductivity of 0.018 W/m-K, so the central claim that the model predicts low-κ structures fails.","rationale":"I considered the reader's concern about empirical potentials. That is a legitimate limitation, but it is an external validity issue that applies to all AGF studies and cannot be settled from the paper alone. The model-range problem is a direct internal inconsistency: the paper's own Equation 4 and its own reported optimal thermal conductivity are mutually incompatible. A model whose output range is bounded between about 0.0256 and 0.0289 W/m-K cannot be said to predict a thermal conductivity of 0.018 W/m-K; the discrepancy is 42% for the very structure the paper highlights as the theoretical optimum. This undermines the paper's central claim that Eq. 4 'predicts thermal conductivity of graphene-WS2 heterostructures' and that it can guide design of low-thermal-conductivity materials. The lack of a defined train/test split or error bars prevents ruling out that the reported accuracy metrics are computed in a way that hides this failure. Therefore, the verdict remains conditional, but the conditions must explicitly require demonstrating that the model captures the low-κ regime, or the claim should be revised. I partially agree with the reader: the potential accuracy issue is real, but the model-range inconsistency is more immediately decisive and checkable from the paper's own data.","tokens_in":11124,"tokens_out":11091,"duration_ms":94524,"concrete_test":"Compute Eq. 4 for the structure '10000100110011' and compare to the reported AGF value of 0.0180 W/m-K; then evaluate Eq. 4 on all 16,384 stacking sequences and count how many have AGF κ below the model's minimum of 0.0256. If this count is nonzero, the model cannot predict the lowest-κ structures. Additionally, re-run the RF and SR fitting with a random 80/20 train/test split, report R² and mean absolute error on the held-out test set, and plot residuals versus κ to check whether the low-κ tail is systematically overpredicted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"For the structure '10000100110011' identified as the SR model's theoretical optimum, counting the binary sequence gives Pa=(8+1)/(8+1)=1, Pb=(8+1)/(8+1)=1, and Pc=(2×4+1)/(6+1)=9/7≈1.286. Substituting into Eq. 4 yields κ_pred=0.178/(9.89+1+1+0.143×9/7)+0.0109≈0.0256 W/m-K. The paper reports the AGF value for this structure as 0.0180 W/m-K, 42% lower. Because the denominator of Eq. 4 is bounded (Pa,Pb≤1 and Pc typically below about 1.3), the model output is confined to roughly 0.0256–0.0289 W/m-K for all 16,384 structures; it cannot represent any structure with κ below 0.0256. The SLEPA-identified optimal structure also has κ=0.018, and low-κ candidates are precisely the design targets the model is meant to identify. The reported R²=0.70 for the RF model and the undefined '64% accuracy' for SR are not accompanied by a train/test split or residual analysis, so they may mask a systematic failure in the low-κ tail. This is an internal inconsistency in the paper's own reported numbers, independent of whether the empirical potentials are physically accurate.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a human-AI collaboration workflow for 14-layer graphene-WS2 heterostructures: SLEPA combined with mode-resolved atomistic Green's function (AGF) generates a small thermal-conductivity dataset; human inspection defines three sequence descriptors Pa, Pb, and Pc; random forest and symbolic regression produce a compact formula κ = 0.178/(9.89 + Pa + Pb + 0.143Pc) + 0.0109 (Eq. 4); and mode-resolved AGF attributes each descriptor to specific frequency and incidence-angle phonon-suppression mechanisms. The paper claims that this interpretable model predicts thermal conductivity across the full space of 16,384 stacking sequences and yields actionable design rules for ultralow thermal conductivity.","tokens_in":11494,"tokens_out":5273,"duration_ms":53845,"significance":"If the central quantitative claim were sound, the work would provide a useful template for combining active learning, human feature construction, and transparent symbolic regression in phonon engineering. The full N=10 enumeration (1024 AGF calculations), the use of mode-resolved transmission data to interpret each descriptor, and the explicit design rules are valuable contributions. However, the central predictive claim is currently undermined by an inconsistency between Eq. 4 and the reported AGF values in the low-κ regime, and the reported accuracy metrics lack the protocol needed to support generalizability. As a computational study, all conclusions are also conditional on the empirical interatomic potentials used for the AGF ground truth, which the paper does not discuss critically.","major_comments":[{"comment":"The model's stated theoretical minimum of 0.0256 W/m-K is inconsistent with the reported AGF value of 0.0180 W/m-K for the same structure '10000100110011'. Because Eq. 4's denominator is bounded for the defined descriptors, the model output cannot reach the low-κ values that SLEPA identifies and that the paper explicitly targets. This is not a small extrapolation error but a systematic failure in the region of interest. The authors should refit the symbolic expression, add a residual analysis for the low-κ tail, or restrict the predictive claims to the range actually supported by Eq. 4.","section":"Construction of a predictive model, Eq. (4) and Fig. 8(c)"},{"comment":"The reported R2 = 0.70 for the random forest and the '64% accuracy' for symbolic regression are not accompanied by a train/test split, cross-validation, error bars, or residual plots. Since the descriptors Pa, Pb, and Pc were selected by human inspection of the same SLEPA-generated dataset that is later used to fit the models, these metrics are at risk of being in-sample and cannot establish out-of-sample predictive performance. A clear data-splitting or resampling protocol and a residual-vs-κ plot are needed, with particular attention to the low-κ regime.","section":"Construction of a predictive model, Fig. 8"},{"comment":"The claim that SLEPA 'mimics the original large dataset' is supported only by visual comparison of histograms. No quantitative distribution-distance metric (e.g., Kolmogorov-Smirnov or Hellinger distance) is reported, and no repeated-run statistics are provided to show that the 100-case SLEPA outcome is robust. Adding such metrics for the 10%, 20%, 30%, and 40% sample sizes would make the validation conclusion load-bearing rather than qualitative.","section":"Validation of SLEPA, Figs. 2–3"},{"comment":"The feature-construction loop uses the SLEPA dataset both to discover Pa, Pb, and Pc and to fit the RF/SR models; this creates an in-sample selection effect that the paper does not address. The independent mode-resolved AGF mechanism analysis in Fig. 9 provides useful external grounding for the physical interpretation, but it does not validate the numerical accuracy of Eq. 4. The authors should clarify the chronology of descriptor selection and model fitting and evaluate the final model on held-out structures outside the SLEPA training pool.","section":"Identification of meaningful features and Construction of a predictive model"}],"minor_comments":[{"comment":"The caption lists '(b) SLEPA, (b) Bayesian optimization', duplicating the label for two different panels; the second panel should be labeled (c).","section":"Figure 3 caption"},{"comment":"The heading contains the typo 'frequences'; it should read 'frequencies'.","section":"Table 1"},{"comment":"The text says the SLEPA-optimal structure '11000000101101' has a thermal conductivity of 0.018 W/m-K, while the SR-optimal structure is later said to be 0.0180 W/m-K and 'only slightly higher'. The rounding and the comparison should be made consistent so the reader can see whether these are the same value.","section":"Results, SLEPA for 14-layer heterostructures"},{"comment":"The sentence 'the left and right leads consist of two layers of graphene or graphite' is ambiguous; it should specify whether the leads are graphene, graphite, or both depending on the terminal layers of the central heterostructure.","section":"Methods, Mode-resolved AGF"},{"comment":"The data availability statement only offers data 'from the corresponding author on reasonable request'; given the reproducibility emphasis of the study, a persistent repository for the 1024 and 1300 AGF datasets would strengthen the paper.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision rather than rejection because the principal inconsistency (Eq. 4 cannot represent the reported low-κ optimum) is potentially fixable by refitting the symbolic expression and adding a proper evaluation protocol. If the authors cannot repair the low-κ behavior, the central predictive claim should be withdrawn or substantially narrowed. The paper would also benefit from a more measured framing: the current abstract and conclusions claim a general human-AI collaboration framework, while the evidence is limited to one empirical-potential-based computational case study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the human-AI pipeline is a nice idea and the mechanism analysis is careful, but the final SR model is internally inconsistent with the AGF data — it cannot represent any structure with κ below about 0.025 W/m-K, even though the optimal structure it nominates has κ = 0.018. That undermines the paper's main claim that the model identifies low-κ design rules.\n\nWhat's new: using SLEPA to serve a small, representative dataset to human experts, then having those experts hand-extract stacking descriptors (Pa, Pb, Pc) and testing them with modal AGF. The descriptor definitions are simple and interpretable, and Figure 9's mode-resolved transmission is a solid way to tie each descriptor to specific phonon frequencies and incidence angles. That part is worth taking seriously.\n\nThe serious problem: Eq. 4 is bounded. With Pa, Pb ≤ 1 and Pc ≈ 1.3 for the nominated optimum, the denominator is about 12.07 and κ = 0.0256. For any structure, the model output stays in a narrow band around 0.025–0.029, while the AGF data includes 0.018. The paper even acknowledges that the \"theoretical minimum\" (0.0256) is not the actual AGF value (0.018) for that structure, but doesn't register the contradiction. So the model doesn't predict low-κ structures; it systematically overestimates them. This is not a subtle statistical issue; it is a fundamental range mismatch. Also, the R² = 0.70 for RF and \"64% accuracy\" for SR are reported without train/test splits or residual analysis, and the SLEPA validation is visual histogram comparison. The descriptors were chosen by inspecting the same SLEPA dataset used to fit the model, so the fit is partly in-sample. The AGF ground truth rests on empirical potentials, which is a standard caveat, not the main issue.\n\nWho is this for? People working on materials informatics workflows, especially human-in-the-loop methods. The workflow idea may be worth pursuing, but as a paper the predictive model is not reliable. It deserves a serious referee, not a desk reject, because the flaw is technical and fixable — but it needs major revision: proper train/test splits, error bars, residual/range analysis, and either a model that can cover the low-κ tail or a more honest statement of what Eq. 4 can do.\n\nRecommendation: send to peer review with a strong request for major revision; the referee should verify the dynamic range of Eq. 4.","headline":"The workflow is appealing, but Eq. 4 cannot produce the low-κ values the paper is after, so the central prediction claim fails on the paper's own numbers.","tokens_in":12003,"tokens_out":4829,"would_cite":false,"duration_ms":42477,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that three human-identified stacking-order parameters, surfaced by an entropy-based sampler and embedded in a symbolic-regression formula, predict thermal conductivity in graphene-WS2 heterostructures.","keywords":["thermal conductivity","graphene-WS2 heterostructures","phonon transmission","human-AI collaboration","symbolic regression","entropic population annealing","atomistic Green's function","interpretable machine learning"],"falsifier":"Recompute thermal conductivities of the same 14-layer stacking sequences with an ab initio or experimentally benchmarked phonon method and check whether Pa, Pb, and Pc still separate low- from high-conductivity structures; even a single sequence with large Pa, Pb, and Pc that conducts better than Eq. (4) predicts would break the claimed design rule.","tokens_in":10957,"feed_emoji":"🔥","tokens_out":6540,"duration_ms":60148,"temperature":0.7,"pith_summary":"Human involvement in materials machine learning is usually limited to setting up and monitoring the algorithm; this paper asks what happens when human expertise sits in the loop, interpreting what the AI proposes before a second AI builds the final model. Using heat conduction in 14-layer graphene-WS2 heterostructures as the test bed, the authors show that a sampling method called SLEPA can produce a small dataset that reproduces the full distribution of thermal conductivities, a human then extracts three stacking-order parameters Pa, Pb, and Pc from the extremes of that distribution, and symbolic regression compresses the parameters into a compact formula. The paper's central claim is that this human-AI sequence yields an interpretable predictive model—κ = 0.178/(9.89 + Pa + Pb + 0.143Pc) + 0.0109—that captures the dependence of thermal conductivity on stacking order and, through mode-resolved atomistic Green's function analysis, identifies the specific frequencies and incidence angles at which each parameter suppresses phonon transmission. If true, it would mean that expert intuition can be injected into materials discovery in a data-efficient way, producing design rules rather than black-box predictions.","feed_headline":"Three stacking parameters predict heat flow in 2D stacks","feed_subtitle":"A human-AI loop distills 14-layer graphene-WS2 designs into an equation and pinpoints the phonons each parameter blocks.","key_machinery":"The carrying mechanism is an AI-human-AI pipeline whose load-bearing components are three human-identified stacking parameters computed from the binary layer sequence of the heterostructure (0 = graphene, 1 = WS2). Pa = (n0+1)/(sum0+1) measures how strongly graphene is buried away from the outer WS2 layers; Pb = (n>00+1)/(sum0+1) measures the share of graphene that appears in runs of length two or more; Pc = (n1 n11 +1)/(sum1+1) measures the product of single- and double-layer WS2 runs. SLEPA, a sampling method combining entropic sampling with a surrogate Gaussian-process model, supplies a small dataset that reproduces the full thermal-conductivity distribution, and the human step converts that distribution into these three descriptors. Symbolic regression then fits κ = 0.178/(9.89 + Pa + Pb + 0.143Pc) + 0.0109. The descriptors carry the argument because they are discrete-structure statistics that both correlate with conductivity and map onto specific phonon-suppression windows in frequency-incidence space.","core_discovery":"On its own terms, the paper discovers three physically interpretable descriptors of stacking order in graphene-WS2 heterostructures and shows that they control the thermal conductivity through distinct phonon-suppression channels. Pa captures whether graphene layers are concentrated between outer WS2 layers, Pb captures whether graphene forms runs of two or more consecutive layers, and Pc captures the product of single- and double-layer WS2 runs. The final model, κ = 0.178/(9.89 + Pa + Pb + 0.143Pc) + 0.0109, predicts thermal conductivity from these three numbers, and the mode-resolved AGF analysis attributes each descriptor to a specific suppression regime: Pa suppresses normally incident low-frequency phonons, Pb suppresses normally incident mid-frequency phonons, and Pc suppresses both normally incident mid-frequency and obliquely incident high-frequency phonons. The paper argues that this mechanism-resolved, closed-form model can guide nanostructure design directly.","pith_inferences":["Outside the paper's own claims, the same SLEPA-to-human-features-to-symbolic-regression loop could be applied to other interface-controlled properties, with the human step identifying local structural motifs rather than global stack order.","Because Pa, Pb, and Pc are simple run-length statistics, they may generalize to other layered heterostructures and to longer layer counts, though the model may need an explicit layer-number dependence.","The model is trained on zero-temperature AGF conductances; an extension to finite-temperature or anharmonic effects would test whether the same parameters preserve their ranking of structures.","The near-minimum structures found by SLEPA and by the symbolic-regression formula are not identical, so an exhaustive check over all 16,384 candidates would directly quantify how much the human-selected descriptors miss."],"forward_implications":["The three stacking parameters can rank any 14-layer graphene-WS2 heterostructure by thermal conductivity without running a full phonon calculation.","The design rules translate into fabrication guidance: terminate with WS2 layers, keep graphene in multi-layer blocks, limit WS2 to one or two contiguous layers, and balance the two materials.","Because the model is an explicit formula, it can be inverted to search for stacking sequences that achieve a target thermal conductivity.","The frequency and incidence-angle map in Table 1 identifies which phonon populations to engineer, such as tuning Pa to suppress low-frequency normal-incidence phonons."],"supporting_citations":[{"why":"Introduces the SLEPA algorithm used to generate the small representative dataset of stacking configurations.","marker":"[26]"},{"why":"Provides the mode-resolved atomistic Green's function formalism used to compute phonon transmission and the spectral decomposition behind Table 1.","marker":"[35]"},{"why":"Supplies the optimized Tersoff potential for graphene, on which the ground-truth thermal conductivities depend.","marker":"[33]"},{"why":"Supplies the Stillinger-Weber potential for WS2 used in the AGF ground-truth calculations.","marker":"[34]"},{"why":"LAMMPS is used to calculate the force constants from which AGF transmission is obtained.","marker":"[32]"},{"why":"Random Forest is used to validate the predictive power of the three human-selected features against the AGF targets.","marker":"[28]"},{"why":"Symbolic regression is used to distill the features into the closed-form conductivity formula.","marker":"[29]"}],"fun_headline_variants":["Three descriptors unlock thermal conductivity in 2D stacks","Human-AI loop distills heat flow into three parameters","Three stacking knobs control heat in graphene-WS2","AI plus human insight pins phonon-blocking parameters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ground truth for every prediction is the thermal conductivity computed by atomistic Green's function with empirical interatomic potentials (Tersoff for graphene, Stillinger-Weber for WS2, and Lennard-Jones for van der Waals contacts); if those potentials misrepresent phonon transmission at the interface, the three parameters and the fitted formula would be artifacts of the force field rather than physical rules.","fun_headline_variants_meta":{"raw":{"variants":["Three descriptors unlock thermal conductivity in 2D stacks","Human-AI loop distills heat flow into three parameters","Three stacking knobs control heat in graphene-WS2","AI plus human insight pins phonon-blocking parameters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000603,"raw_usage":{"total_tokens":2820,"prompt_tokens":958,"completion_tokens":1862,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":1797}},"tokens_in":574,"tokens_out":1862,"duration_ms":12731,"temperature":1.0,"reasoning_tokens":1797,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T18:45:26.777982+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute thermal conductivities of the same 14-layer stacking sequences with an ab initio or experimentally benchmarked phonon method and check whether Pa, Pb, and Pc still separate low- from high-conductivity structures; even a single sequence with large Pa, Pb, and Pc that conducts better than Eq. (4) predicts would break the claimed design rule.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the SLEPA algorithm used to generate the small representative dataset of stacking configurations."},{"cited_title":"Ong, Phys","cited_arxiv_id":null,"evidence_quote":"Provides the mode-resolved atomistic Green's function formalism used to compute phonon transmission and the spectral decomposition behind Table 1."},{"cited_title":"Lindsay, D.A","cited_arxiv_id":null,"evidence_quote":"Supplies the optimized Tersoff potential for graphene, on which the ground-truth thermal conductivities depend."},{"cited_title":"Mobaraki, A","cited_arxiv_id":null,"evidence_quote":"Supplies the Stillinger-Weber potential for WS2 used in the AGF ground-truth calculations."},{"cited_title":"Plimpton, J","cited_arxiv_id":null,"evidence_quote":"LAMMPS is used to calculate the force constants from which AGF transmission is obtained."},{"cited_title":"Breiman, Mach","cited_arxiv_id":null,"evidence_quote":"Random Forest is used to validate the predictive power of the three human-selected features against the AGF targets."},{"cited_title":"Cranmer, A","cited_arxiv_id":null,"evidence_quote":"Symbolic regression is used to distill the features into the closed-form conductivity formula."}],"review_version":1}