{"id":"06e7f590-4dc8-444c-802a-7bc3c8befd54","arxiv_id":"2607.21246","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A TabPFN model trained only on 25 quaternary Mo-W-S-Se-Te alloys reconstructs DFT dielectric spectra of held-out quaternaries with R²>0.98 and generalizes zero-shot to quinary alloys with R²>0.97.","lead":"The paper trains a tabular machine-learning model on computer-simulated light-absorption spectra of 25 four-element TMD alloys and predicts held-out spectra with very high accuracy. It also adds a physics-informed energy-sampling trick that fits the model's context limit and lets it guess spectra for alloys it never saw.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Configurational averaging is untested: one random 4x4x1 supercell per composition means the reported R2/MAE may describe single arbitrary draws rather than the composition's true optical response.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the random-substitution procedure yields one representative configuration per composition, and no evidence establishes that this configuration is typical. I agree that this is the most load-bearing issue because it sits directly under the central claim. If the dielectric spectrum depends strongly on the random atomic arrangement, then a composition-only feature vector is insufficient in principle, and every reported R2/MAE number is evaluating agreement with a single stochastic draw rather than with the composition's mean optical response. That would not merely weaken the zero-shot narrative—it would undermine the fundamental mapping the paper claims to have learned. The other issues listed by the reader (the four unaccounted quaternary alloys, inconsistent MAE units, and overstated zero-shot language) are real but secondary: they affect reproducibility and interpretation, whereas the configurational question affects whether the target itself is well defined. I do not think the concern warrants rejection: high held-out accuracy is indirect evidence that configurational variance is probably not overwhelming, and the missing test is cheap and easy to run. The appropriate response is to keep the verdict conditional pending this configurational check. No adjustment to the reader's verdict is needed.","tokens_in":19564,"tokens_out":10378,"duration_ms":120848,"concrete_test":"Pick 2–3 compositions spanning the dataset—e.g., one quaternary training composition such as Mo6W10S13Te19, one quinary test composition such as Mo3W13S11Se17Te4, and one binary such as Mo4S8. For each, build 3–5 independent random 4x4x1 supercells with identical stoichiometry but different random seeds and recompute the PBE-GGA dielectric spectra using the same CASTEP settings and k-grid. Quantify the inter-configuration spread (mean absolute deviation in εx2/εz2, peak-energy shifts) and compare with the reported MAE (<0.10). If the spread is much smaller than the MAE, the single-configuration approximation is adequate; if it is comparable or larger, rerun the TabPFN evaluation against configurational averages and report error bars. This is decisive because the central claim predicts the spread should be negligible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Computational Details ('Supercell Construction and Alloying') states that for each target composition atoms are 'randomly substituted over the available metal and chalcogen sites' and that a single 'representative atomic configuration' is used 'without explicitly enumerating all possible symmetrically inequivalent arrangements.' Thus each composition contributes exactly one DFT spectrum. The central claim—that composition fractions plus photon energy suffice for TabPFN to emulate DFT-level spectra—requires that this one draw is effectively the composition's optical response. No configurational convergence check is provided; Appendix A verifies only k-point convergence. If random Mo/W and S/Se/Te arrangements shift peak positions or alter transition-matrix elements by amounts comparable to the reported MAE (<0.10), then the DFT labels are stochastic, the R2/MAE metrics are computed against arbitrary single configurations, and the 'zero-shot' success could be partly an artifact of the particular supercells chosen. This concern applies equally to the quinary generalization claim, whose labels are also one random supercell per composition. The paper's own Appendix C concedes degraded zero-shot predictions for binary/quinary systems, but the configurational issue is more fundamental: it questions the determinism of the target itself. The concern is not that the authors are wrong, but that a load-bearing premise is unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper builds a dataset of 99 monolayer Mo–W–S–Se–Te TMD alloy structures (binary, ternary, quaternary, quinary) and computes their DFT-level dielectric functions along in-plane and out-of-plane directions. It then trains TabPFN, a tabular foundation model, using only element fractions and photon energy as features, with a physics-informed energy subsampling strategy that allocates 330 points per material (10 sub-gap, 250 active, 70 tail) to fit within TabPFN's in-context limit. Trained exclusively on 25 quaternary alloys, the model is reported to reconstruct held-out quaternary spectra with R2 > 0.98 and MAE < 0.10 for all four components, outperforming Extra Trees and XGBoost, and to generalize zero-shot to binary, ternary, and quinary alloys, with quinary predictions reaching R2 > 0.97. The paper also derives refractive index, extinction coefficient, and absorption coefficient from the predicted dielectric functions.","tokens_in":19869,"tokens_out":4957,"duration_ms":56342,"significance":"If the results hold, the paper makes a useful contribution: it demonstrates that a composition-only feature representation plus photon energy, fed to a pretrained tabular transformer, can emulate DFT-level dielectric spectra across a broad TMD alloy space without task-specific training. The holdout design is a genuine generalization test, not circular: the model weights are frozen and the test spectra are held-out DFT calculations. The paper also ships a transparent subsampling budget and compares against standard tree-based baselines. The main significance is the zero-shot claim: a model trained only on quaternary compositions predicting quinary alloys with R2 > 0.97 would be practically valuable for screening. However, the single-configuration-per-composition dataset and the inconsistent reporting of MAE units currently leave the strength of the quantitative claims uncertain.","major_comments":[{"comment":"The dataset uses one random 4x4x1 supercell per composition. The text states atoms are 'randomly substituted' and a single 'representative atomic configuration' is used without enumerating inequivalent arrangements. No check is reported that different random configurations yield similar dielectric spectra; Appendix A verifies only total-energy convergence with k-point density, not configurational convergence of optical properties. Since the regression targets are these single-configuration DFT spectra, the reported MAE (<0.10) may be smaller than the configurational scatter of the target. This is load-bearing for the claim that composition fractions plus photon energy suffice. Please compute multiple independent configurations for a representative subset (e.g., 3–5 compositions across quaternary and quinary), report spectral standard deviations, and either average labels or qualify the c","section":"Computational Details, 'Supercell Construction and Alloying'; Appendix A"},{"comment":"There is a direct conflict about the units of MAE for the imaginary components. Figure 6's caption says MAE for εx2 and εz2 is 'reported in log-transformed units,' while Table 3's caption says MAE values are 'in the original linear scale.' The abstract states MAE <0.10 for all four dielectric components without qualification. A log-scale MAE of 0.10 is not comparable to a linear-scale MAE; it corresponds approximately to multiplicative errors in ε2. Report all MAE values in a single identifiable unit, ideally back-transformed to the physical ε2 scale, and correct the abstract and table.","section":"Figure 6 caption; Table 3 caption; Abstract"},{"comment":"The zero-shot claim should be qualified. Table 3 shows binary and ternary R2 values of 0.82–0.90 and MAE 0.15–0.25, which is substantially weaker than the quaternary interpolation performance; Appendix C concedes degraded predictions for binary and quinary materials. The abstract highlights only quinary (R2>0.97), which is accurate, but the unqualified sentence in the text 'generalized in a zero-shot manner to binary, ternary, and quinary alloys' may be misread as uniformly high accuracy. Please state the performance by family and temper the interpretive claim that the model 'learned transferable, physically meaningful compositional representations' based on this evidence.","section":"Results and Discussion, 'Generalization across compositions beyond training data'; Appendix C"}],"minor_comments":[{"comment":"Typographical issues: the title contains 'T e' instead of 'Te'; 'xGBoost' appears with inconsistent capitalization; several figure labels contain the literal string 'uni00A0' instead of a space. Please run a final proofreading pass.","section":"Title and figure labels"},{"comment":"The text says 'R2 score ranges from 0 to 1,' but R2 can be negative for predictions worse than the mean. The formula is standard, but the prose should be corrected.","section":"Evaluation Metrics, Eq. (24)"},{"comment":"The phrase 'Zero set uses data of only one composition for initializing' is unclear. What is initialized? In TabPFN there is no training, so 'initializing' needs explanation or removal.","section":"Figure 2 caption"},{"comment":"The energy subsampling budget (10/250/70 points) and the boundary max(10 eV, 3Eg) are presented as fixed choices without sensitivity analysis. A short ablation showing that nearby budgets give similar results would make the method appear less ad hoc.","section":"Subsampling Strategy, Table 1"},{"comment":"The claim that 'no prior work has predicted the optical properties of TMD alloys' is strong. The cited Barhoumi et al. work predicts absorption spectra, and there is a broader ML-for-spectra literature. Please sharpen the novelty statement to distinguish spectral prediction as a function of composition from the current work.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The configurational sampling issue is the key uncertainty. The holdout design and external comparison to held-out DFT are strengths, and the claim is testable: a few additional DFT calculations with multiple random supercells per composition would establish whether the target is deterministic enough for the reported MAE to be meaningful. I would not reject, but the manuscript should not be accepted until this is addressed and the MAE-unit conflict is corrected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick read: this is a useful, honest ML/DFT paper that does something genuinely new—predicting full dielectric spectra of Mo-W-S-Se-Te alloys from composition plus photon energy using TabPFN with a physics-informed energy subsampling scheme. The main claim is plausible, and the holdout design is the right kind: train only on quaternary alloys, then test on held-out quaternaries plus binary, ternary, and quinary compositions. The quinary generalization (R2 > 0.97) is impressive, and the authors are transparent that binary and ternary performance is more modest (R2 0.82-0.90 in Table 3).\n\nWhat it does well: the new 99-structure DFT dataset, the physically sensible subsampling strategy (dense sampling in the active absorption window), and a clean comparison against Extra Trees and XGBoost. The paper is readable, and the distinction between in-context learning and traditional training is explained carefully.\n\nSoft spots, in descending order:\n\n1. Configurational averaging. Each composition gets one random 4x4x1 supercell, and the paper provides no check that different random arrangements yield similar spectra. If configurational disorder shifts peaks by more than the reported MAE, the DFT labels are noisy and every R2/MAE number is about a single arbitrary draw. This is the biggest gap. It may not matter—composition effects may dominate—but it's unverified. A few independent supercells per composition would settle it.\n\n2. No code or data repository. That blocks verification and reuse, which is essential for a dataset-and-model paper.\n\n3. MAE units are inconsistent. The abstract says MAE < 0.10 for all four components, but for ε2 that's in log-transformed space. Table 3 gives linear-scale MAE, but the main text doesn't flag the distinction. This overstates accuracy.\n\n4. Count discrepancy: 36 quaternary alloys, 25 for training, 7 in the test set—leaving 4 unexplained. Minor but sloppy.\n\nNone of these, taken together, refute the central result. The external holdout gives it real weight. But the paper isn't fully verifiable as written. I'd send it to peer review and ask for the code/data, a few repeat-supercell checks, and a consistent statement of errors. Not a desk reject.\n\nThis is the kind of work I'd want in the literature after those fixes.","headline":"Useful, plausible ML-DFT screening paper for TMD alloys, but the one-supercell-per-composition assumption and a few reporting gaps need fixing before I'd trust the numbers.","tokens_in":20366,"tokens_out":3882,"would_cite":true,"duration_ms":41856,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A tabular model, trained only on quaternary Mo–W–S–Se–Te alloys, reconstructs DFT-grade dielectric spectra (R²>0.98, MAE<0.10) and predicts binary, ternary, and quinary compositions with no in-context examples.","keywords":["two-dimensional TMD alloys","dielectric function","optical-property prediction","TabPFN","in-context learning","physics-informed sampling","density functional theory","composition-property mapping"],"falsifier":"Compute DFT dielectric spectra for at least three independently randomized 4×4×1 supercells of one quaternary composition (for example Mo10W6S13Te19). If the spread among those single-configuration spectra is comparable to or larger than the claimed MAE<0.10, then the reported accuracy is measuring agreement with one random configuration, not with the alloy's representative optical response, and the central claim would not survive.","tokens_in":19461,"feed_emoji":"⚛️","tokens_out":7467,"duration_ms":74802,"temperature":0.7,"pith_summary":"The paper sets out to show that the frequency-dependent optical response of two-dimensional Mo–W–S–Se–Te transition-metal dichalcogenide (TMD) alloys can be predicted from composition alone, without a fresh DFT calculation for every new stoichiometry. Using a tabular foundation model that learns in-context from a small table of DFT-computed dielectric spectra, plus a non-uniform energy sampling scheme that concentrates points where interband absorption is strongest, the authors report that quaternary-trained models reproduce held-out quaternary spectra with R²>0.98 and MAE<0.10, and generalize with zero in-context examples to binary, ternary, and quinary alloys, with quinary R²>0.97. If correct, this replaces combinatorial first-principles optical screening with a single forward pass, yielding refractive index, extinction coefficient, and absorption coefficient for device design across a broad alloy space.","feed_headline":"TabPFN predicts TMD alloy optical spectra from composition alone","feed_subtitle":"Trained on quaternary alloys only, it reaches R2 > 0.98 and transfers to five-element mixes.","key_machinery":"The central object is TabPFN, a prior-data fitted transformer that performs in-context learning: rather than gradient-training on the target dataset, it takes the DFT-computed training rows as a context table and produces predictions for query rows in a single forward pass, with attention across samples and across features. Because its context is capped near 10,000 rows, the paper introduces a physics-informed energy subsampling scheme: each material contributes 330 energy points, with 10 below the band gap, 250 in the optically active window from the gap to max(10 eV, 3×gap), and 70 in the high-energy tail. With 25 quaternary training materials this yields about 8,250 rows. The four dielect","core_discovery":"On its own terms, the paper's discovery is that a composition-only feature vector—the atomic fractions of Mo, W, S, Se, and Te plus the photon energy—is sufficient input for a pre-trained tabular transformer to emulate DFT-level dielectric spectra across a five-element TMD alloy space. Training exclusively on 25 quaternary compositions, subsampled to 330 energy points per material via a physics-informed grid, the model reproduces the real and imaginary parts of the in-plane and out-of-plane dielectric functions of held-out quaternary alloys with R²>0.98 and MAE<0.10, and—without ever seeing binary, ternary, or quinary rows in context—predicts those families as well, with quinary predictions","pith_inferences":["An implication the paper leaves implicit: because each composition is represented by one random alloy configuration, the reported R² and MAE conflate learning error with configurational variance; a decisive follow-up is to average DFT spectra over several random configurations and retrain, checking whether accuracy improves or the error bars absorb the current gap.","A testable extension of the subsampling argument: hold the 330-point budget fixed but replace the physics-informed grid with uniform sampling over 0–33 eV; if R² degrades markedly, the case for physics-informed sampling is confirmed; if not, most of the credit belongs to the model prior.","The paper's own Appendix C notes that binary and quinary zero-shot spectra sometimes fail to separate closely spaced peaks; a natural next step is to seed the context with a handful of binary and quinary rows and measure how much fine-structure fidelity improves.","Because the descriptor is strictly compositional plus photon energy, the same recipe is portable to other 2D alloy families; a quick check would be to apply it to a held-out chalcogen or metal pair not in Mo–W–S–Se–Te."],"forward_implications":["A composition-only descriptor plus a physics-informed energy grid is sufficient for a tabular foundation model to reconstruct DFT-quality dielectric spectra of quaternary TMD alloys (R²>0.98, MAE<0.10).","Zero-shot transfer works across alloy order: binary and ternary predictions reach R²≈0.82–0.90 and quinary predictions exceed R²=0.97, so combinatorial screening of higher-order alloys becomes feasible without new DFT data.","Refractive index, extinction coefficient, and absorption coefficient follow directly from the predicted dielectric components, giving device-relevant quantities without additional first-principles runs.","Deployment reduces to one forward pass with no gradient-based training or hyperparameter search, making cheap optical screening repeatable as new compositions are proposed.","The authors' proposed workflow—small reference set for a target family, in-context prediction across candidates, DFT validation only on the short list—would concentrate DFT cost on the most promising compositions."],"fun_headline_variants":["Composition-only TabPFN predicts TMD alloy dielectric spectra","Zero-shot: quaternary-trained model predicts all TMD alloy optics","Physics-informed sampling lets TabPFN predict TMD alloy spectra","Composition alone suffices to predict TMD alloy optics","TabPFN maps alloy fractions to dielectric spectra zero-shot"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that one randomly substituted 4×4×1 supercell per composition is a faithful representative of that alloy's optical response; if different random configurations of the same composition give materially different dielectric spectra, the DFT labels and every reported accuracy figure are noise around an uncharacterized configurational average.","fun_headline_variants_meta":{"raw":{"variants":["Composition-only TabPFN predicts TMD alloy dielectric spectra","Zero-shot: quaternary-trained model predicts all TMD alloy optics","Physics-informed sampling lets TabPFN predict TMD alloy spectra","Composition alone suffices to predict TMD alloy optics","TabPFN maps alloy fractions to dielectric spectra zero-shot"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000787,"raw_usage":{"total_tokens":3362,"prompt_tokens":855,"completion_tokens":2507,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":2436}},"tokens_in":599,"tokens_out":2507,"duration_ms":18697,"temperature":1.0,"reasoning_tokens":2436,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T08:01:13.152668+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute DFT dielectric spectra for at least three independently randomized 4×4×1 supercells of one quaternary composition (for example Mo10W6S13Te19). If the spread among those single-configuration spectra is comparable to or larger than the claimed MAE<0.10, then the reported accuracy is measuring agreement with one random configuration, not with the alloy's representative optical response, and the central claim would not survive.","supporting_citations":[],"review_version":1}