{"id":"1a8900b4-dc65-45d3-8433-608812c221f3","arxiv_id":"2412.09917","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Using an improved ab initio starting point and a train/test protocol, the authors construct the p35-i3 shell-model Hamiltonian for N=50 isotones with 100-130 keV spectrum agreement.","lead":"Prior empirical shell-model Hamiltonians are fit to sparse nuclear data and can overfit. This paper combines an improved ab initio starting Hamiltonian with a train/test fitting protocol, yielding the p35-i3 interaction that reproduces N=50 isotone spectra to about 100-130 keV and gives an extrapolation accuracy of about 300 keV.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 35-VLC cutoff is selected from the same validation curve used to quote 100–130 keV RMSD; until a nested or truly frozen holdout is reported, the headline 'predictive power' is optimistic.","rationale":"I read the paper as a methodological proposal: SVD fitting from an improved IMSRG(3f2) starting point, with a train/test split used to choose the number of varied linear combinations, should give trustworthy extrapolations to exotic nuclei. The strongest quantitative support is the validation curve in Fig. 3 and the Z ≤ 36 test in Fig. 5. The main internal weakness is that the same validation information is doing double duty: it sets Nc and then defines the reported error. That is a textbook model-selection optimism, not a claim about the physics. The plateau over 15–35 VLC reduces but does not remove the bias, because the quoted 100–130 keV is still the minimum of the curve. The Z ≤ 36 test goes in the right direction, but it is not fully independent because 35 was selected using splits that included Z ≤ 36 levels; moreover, some 99In constraints are extrapolated levels, so part of the boundary agreement is built in. I do not see an internal mathematical inconsistency in the SVD formalism; the issue is the statistical protocol. The proposed nested or frozen holdout is a feasible, decisive check. Since the reader already flagged this same weak point and asked for conditional acceptance, my stress-test does not move the verdict.","tokens_in":13919,"tokens_out":9207,"duration_ms":103504,"concrete_test":"Run a nested cross-validation version of Fig. 3: for each candidate VLC, split all levels into outer train/test, use a further inner split of the outer training set to select VLC by minimizing inner validation error, and evaluate the selected Hamiltonian on the untouched outer test set. Report the mean and spread of outer-test RMSD over repeated splits. Separately, for the Z ≤ 36 extrapolation, restrict all model selection to Z > 36 data (for example, choose VLC by leave-one-isotone-out on Z > 36) and evaluate once on the complete, frozen Z ≤ 36 set. If the outer-test RMSD remains within the 100–130 keV band, the reported estimate survives; if it shifts toward the roughly 300 keV Z ≤ 36 value, the headline 'predictive power' must be revised downward.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the choice of Nc = 35 in Sec. V. The authors take the minimum of the green (validation) RMSD curve in Fig. 3 to fix the number of varied linear combinations, then quote the resulting RMSD (100–130 keV) as the predictive performance of p35–i3. Because the validation set has already been consulted to select model complexity, that number is an optimistically biased estimate of error on genuinely new data: the minimum over roughly 35 candidate cutoffs is expected to lie below the true generalization error even when the true model is not overfit. The Z ≤ 36 extrapolation test does not fully repair this, because the same 35-VLC choice was made using random splits that include Z ≤ 36 levels, so the target region has influenced hyperparameter selection. A secondary aggravator is that some low-Z boundary 'data' (the 3/2− and 5/2− states of 99In, Sec. IV) are not direct measurements but systematics-based extrapolations used with 50 keV uncertainties, so part of the agreement near 100Sn is not a fresh prediction. The central claim should be re-evaluated with a genuinely independent holdout.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents two developments intended to improve the predictive power of empirical shell-model Hamiltonians when calibration data are sparse. First, it uses an improved VS-IMSRG(3f2) starting Hamiltonian for the πj4 model space (protons in 0f5/2, 1p3/2, 1p1/2, 0g9/2), arguing that this starting point requires fewer phenomenological adjustments than the previous IMSRG(2) approximation. Second, it introduces a protocol based on singular-value decomposition (SVD) fitting combined with random 80/20 training/testing partitions, repeated 2000 times, to select the number of varied linear combinations (VLC) and to estimate predictive error. The authors select Nc = 35 varied linear combinations, producing the p35-i3 Hamiltonian, and report that it predicts a subset of experimental spectra for all nuclei in the πj4 space to within a 100-130 keV RMSD, with a neutron-rich extrapolation test (Z ≤ 36) reaching approximately 300 keV. The paper concludes that these developments enable more reliable extrapolation to exotic isotopes.","tokens_in":14156,"tokens_out":2962,"duration_ms":33514,"significance":"If the claims are correct, the paper would make a useful methodological contribution to the calibration of shell-model Hamiltonians in data-sparse regions, relevant for rare-isotope facilities and r-process studies. The use of a modern ab initio starting point, IMSRG(3f2), is timely, and the SVD formalism is standard and clearly presented. The 2000-split training/validation protocol is a sensible internal check and is a strength of the paper. The paper also provides the final Hamiltonian parameters and results in supplemental material, which aids reproducibility. However, the headline claim of predictive power is weakened by a selection-bias issue: the same validation data used to choose the 35-VLC cutoff are also used to quote the 100-130 keV RMSD. The extrapolation test in Sec. V is an important attempt, but it is contaminated because the hyperparameter choice was made using random splits that include Z ≤ 36 levels. These issues are fixable with a nested or frozen holdout procedure, but as presented the central quantitative claim is not yet fully supported.","major_comments":[{"comment":"The choice of Nc = 35 is made by taking the minimum of the green validation RMSD curve in Fig. 3, and the value of that curve (100-130 keV) is then quoted in Sec. VI as the predictive power of p35-i3. Because the validation data have already been used to select the model complexity, this error estimate is optimistically biased: the minimum over roughly 35 candidate cutoffs is expected to lie below the true generalization error even without overfitting. To support the claim, the authors should report a genuinely independent test error, for example by partitioning the data once into training, validation, and test sets, choosing Nc on the validation portion, and evaluating the final Hamiltonian on the untouched test portion (or by a nested cross-validation that accounts for the model-selection step).","section":"Sec. V, Fig. 3"},{"comment":"The Z ≤ 36 extrapolation test does not fully remedy the selection-bias problem, because the same Nc = 35 was selected using random partitions that include Z ≤ 36 levels, so the target region has influenced hyperparameter selection. To claim extrapolative reliability, the cutoff should be fixed using only Z > 36 data — for example, by choosing Nc from a validation curve built exclusively from Z > 36 levels — and then applied to the Z ≤ 36 data as a true holdout. As written, the 300 keV figure is not an unbiased measure of extrapolation error.","section":"Sec. V, Fig. 5"},{"comment":"The manuscript blends measured and model-dependent input in the data used for fitting and validation. The caption of Fig. 1 states that the 9/2+ energy shown in the 'experimental' panel was obtained from the p35-i3 Hamiltonian itself, which is circular if that panel is presented as experimental. In addition, the 99In 3/2− and 5/2− excited states are not direct measurements but extrapolations based on 131In systematics, and Sec. IV states that they were assigned 50 keV uncertainties without the 150 keV theoretical error. These artificially small uncertainties make those states act as hard constraints, so agreement near 100Sn is not a fresh prediction. The authors should exclude such model-dependent points from the test and validation splits, or at minimum use realistic uncertainties and clearly flag them as extrapolated input.","section":"Sec. IV and Fig. 1"}],"minor_comments":[{"comment":"The heading 'SVD Proceedure' contains a typo; it should read 'SVD Procedure'.","section":"Sec. III heading"},{"comment":"The sentence 'The 9/2+ energy shown in the experimental panel was obtained from our p35–i3 Hamiltonian' is confusing; the panel should be relabeled so that theory-generated points are not presented as experimental data.","section":"Fig. 1 caption"},{"comment":"The data selection is described narratively but the actual list is only in the supplemental material; a table in the main text listing the included and excluded levels with their adopted Jπ values and uncertainties would improve transparency.","section":"Sec. IV"},{"comment":"The text says 'The calculations presented in FIG. 5 sampled from the full range of data 28 ≤ Z ≤ 50', but Fig. 5 actually shows the partition with training restricted to A > 86 (Z > 36); the wording should be clarified to distinguish the interpolation experiment (Fig. 3) from the extrapolation experiment (Fig. 5).","section":"Sec. V, Fig. 5"},{"comment":"In the text after Eq. (2), the notation is dense: the matrix G is used for the fit matrix and also referred to as the error matrix without a derivation; a short paragraph linking G^{-1} to parameter covariances would help readers not familiar with the SVD fitting literature.","section":"Sec. II, Eq. (1)"},{"comment":"There is a minor capitalization inconsistency: 'states for 99in' should be '99In'.","section":"Sec. II"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the methodological idea is timely, but the central quantitative claim of predictive power rests on validation-based model selection on the same data used to quote the error. A nested holdout or a truly frozen Z ≤ 36 test set would strengthen the paper substantially. The inclusion of the 9/2+ state in the 'experimental' panel of Fig. 1 and the use of extrapolated 99In states with reduced uncertainties are also concerns that should be addressed before publication. I would not recommend rejection, as the issues are fixable within the scope of the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the paper’s real contribution is modest but concrete. It shows that the IMSRG(3f2) starting Hamiltonian is close enough to experiment that fitting fewer SVD linear combinations gets you to a good energy RMSD, and it packages that into a usable p35-i3 interaction for the πj4 space. Section III is a clean, standard exposition of the SVD fitting machinery. The idea of splitting data into training/testing to choose how many varied linear combinations to keep is borrowed from ML but new to this fitting context, and the 2000-split scatter plot in Fig. 3 is a useful visual.\n\nThe strongest part is the extrapolation test in Sec. V: train on Z>36, test on Z≤36, get about 300 keV at 35 VLC. That is a genuinely out-of-sample check and it supports the overfitting narrative.\n\nNow the soft spots, in proportion. The headline '100–130 keV RMSD' is not an honest estimate of predictive error, because the 35-VLC cutoff was selected from the minimum of the green validation curve and then the same curve's minimum is quoted. That is selection on the validation set. The minimum over ~35 candidate cutoffs will sit below the true generalization error. The Z≤36 test doesn't repair this, since the random splits used to pick 35 already contain Z≤36 levels. So the central quantitative claim needs a nested holdout or a pre-registered frozen validation set.\n\nTwo secondary issues. Fig. 1 places the p35-i3 9/2+ energy in the 'experimental' panel, and some 99In levels are systematics-based extrapolations assigned 50 keV uncertainties rather than direct measurements. That is disclosed in the text, so it's not deceptive, but it means part of the agreement near 100Sn is not a fresh prediction. Also, the promised supplemental files don't exist yet, and there is no baseline comparison to JUN45 or other older Hamiltonians under the same train/test split, so we can't judge how much of the improvement comes from the protocol vs. the starting Hamiltonian.\n\nThat said, the paper is honest on its own terms and the math is right. I would send it out; the referee should push for an honest held-out evaluation and for the data release. The p35-i3 interaction will be useful to the community even if the methodology claim needs tightening.","headline":"Useful demonstration that the IMSRG(3f2) starting point plus a train/test-guided SVD cutoff improves shell-model calibration in a sparse region, but the quoted 100–130 keV predictive error is optimistic because the same validation curve selects the cutoff.","tokens_in":14747,"tokens_out":2330,"would_cite":true,"duration_ms":25617,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A better ab initio starting Hamiltonian and a train/test validation protocol yield the p35-i3 shell-model Hamiltonian, which reproduces $\\pi j4$ spectra to 100–130 keV RMSD and extrapolates to neutron-rich nuclei with $Z \\leq 36$ at…","keywords":["shell model","empirical Hamiltonians","IMSRG","singular value decomposition","overfitting","nuclear spectra","exotic isotopes","πj4 model space"],"falsifier":"Measure the 100Sn binding energy and 99In excited states, whose values p35-i3 predicts; if the deviations exceed the roughly 300 keV found in the $Z \\leq 36$ extrapolation test, the paper's claim that the validation protocol yields reliable extrapolation would be falsified.","tokens_in":13657,"feed_emoji":"⚛️","tokens_out":9632,"duration_ms":81751,"temperature":0.7,"pith_summary":"This paper tries to show that empirical shell-model Hamiltonians can be made genuinely predictive even when calibration data are sparse. The authors combine a more accurate ab initio starting Hamiltonian, obtained with the IMSRG(3f2) approximation, with a training-and-testing protocol that chooses how many Hamiltonian parameter combinations to vary. Fitting to proton configurations in the $\\pi j4$ model space between 78Ni and 100Sn, they obtain the p35-i3 Hamiltonian, which reproduces known spectra to a 100–130 keV root-mean-square deviation and extrapolates to neutron-rich nuclei with $Z \\leq 36$ to about 300 keV. These tools matter because nuclei relevant to rare-isotope facilities and the r-process sit in model spaces where calibration data are far sparser than in the well-studied sd shell.","feed_headline":"Validation protocol yields 100–130 keV shell-model predictions","feed_subtitle":"A train/test-calibrated Hamiltonian extrapolates reliably to neutron-rich exotic isotopes.","key_machinery":"The carrying mechanism is the singular-value decomposition (SVD) fit. The fit matrix $G$ — the symmetric product of wavefunction overlaps weighted by inverse data variances — is diagonalized to produce uncorrelated linear combinations of the Hamiltonian parameters, each with an error $d_i$; only combinations with error below a cutoff are varied, while the rest retain their ab initio values. The number of varied linear combinations (VLC) is chosen from the minimum of the validation RMSD curve obtained by randomly splitting known energies into 80% training and 20% testing sets, repeated 2000 times. The improved starting point is the IMSRG(3f2) factorization, which incorporates intermediate three-body operators in nested commutators at the same computational cost as the earlier IMSRG(2). The final product is the p35-i3 Hamiltonian with 35 varied linear combinations.","core_discovery":"The central claim is that the p35-i3 Hamiltonian—built by averaging 78Ni- and 100Sn-referenced IMSRG(3f2) valence-space Hamiltonians and then adjusting 35 of the best-determined linear combinations of single-particle energies and two-body matrix elements by fits to energy data—predicts a subset of the experimental spectra for all nuclei in the $\\pi j4$ space to within a 100–130 keV RMSD. A companion extrapolation test, training only on $Z > 36$ and evaluating $Z \\leq 36$, reaches about 300 keV RMSD before overfitting sets in beyond 35 varied linear combinations. The paper argues that the IMSRG(3f2) starting point is closer to experiment than earlier ab initio Hamiltonians, needing far fewer adjusted linear combinations, and that the validation protocol protects against overfitting.","pith_inferences":["Beyond the paper: because the same held-out data are used both to choose the 35-VLC cutoff and to report validation error, the 100–130 keV figure is likely optimistic for truly unmeasured nuclei; the $Z \\leq 36$ test is a more honest estimate, though if that region contributed to cutoff selection it is not fully independent.","Beyond the paper: a stricter evaluation would reserve an entire region, such as all $Z \\leq 36$ data, from both fitting and cutoff selection, then evaluate the chosen Hamiltonian on that region exactly once.","Beyond the paper: applying the same protocol in the sd shell, where abundant data exist, would let one simulate sparse calibration by withholding random subsets and check whether validation-selected cutoffs beat fixed cutoffs on truly withheld data."],"forward_implications":["The protocol should allow empirical Hamiltonians to be calibrated with considerably less data than the sd-shell cases needed, without losing predictive accuracy.","The IMSRG(3f2) starting point should reduce the number of parameters requiring adjustment, making the unadjusted remainder of the Hamiltonian more physically trustworthy.","For the $\\pi j4$ region, the paper makes concrete predictions for unmeasured states, including the binding energy of 100Sn, excited states of 99In, and a 9/2+ state in 79Cu, which upcoming experiments can check.","The same train/test selection procedure could be applied to larger model spaces where full optimization is computationally heavy, improving extrapolation to exotic isotopes relevant to the r-process."],"supporting_citations":[{"why":"Supplies the factorized IMSRG(3f2) approximation whose three-body corrections produce the improved starting Hamiltonian.","marker":"[3]"},{"why":"Establishes the USD-style empirical shell-model fit and the roughly 200 keV RMSD benchmark that this paper aims to improve upon.","marker":"[1]"},{"why":"Introduces the SVD-based fitting with a singular-value cutoff that the paper's varied-linear-combination protocol extends.","marker":"[2]"},{"why":"Provides the VS-IMSRG framework used to derive the valence-space effective Hamiltonians from the ab initio interaction.","marker":"[6]"},{"why":"Gives the JUN45 Hamiltonian in the same jj44 model space, used for comparison of level ordering and fits.","marker":"[8]"},{"why":"Provides the 79Cu low-lying level data that constrain single-particle energies near 78Ni.","marker":"[22]"},{"why":"Provides the 99In level data that pin down hole states near the 100Sn closed shell.","marker":"[23]"},{"why":"Supplies the EM 1.8/2.0 NN+3N interaction used as input to the IMSRG calculations.","marker":"[25]"},{"why":"ENSDF is the source of the experimental energies and uncertainties used in the fits.","marker":"[27]"}],"fun_headline_variants":["Shell-model fit hits 100–130 keV for exotic nuclei","Validation protocol boosts shell-model extrapolation to 100 keV","35-parameter shell-model predicts exotic isotopes reliably","New shell-model Hamiltonian improves predictive power","Overfitting check sharpens shell-model predictions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the validation error measured on randomly held-out known levels, even after those same held-out levels helped choose the 35-combination cutoff, is an unbiased measure of how well the Hamiltonian will predict unmeasured exotic isotopes.","fun_headline_variants_meta":{"raw":{"variants":["Shell-model fit hits 100–130 keV for exotic nuclei","Validation protocol boosts shell-model extrapolation to 100 keV","35-parameter shell-model predicts exotic isotopes reliably","New shell-model Hamiltonian improves predictive power","Overfitting check sharpens shell-model predictions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1409,"prompt_tokens":778,"completion_tokens":631,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":394,"completion_tokens_details":{"reasoning_tokens":558}},"tokens_in":394,"tokens_out":631,"duration_ms":6728,"temperature":1.0,"reasoning_tokens":558,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:35:24.074347+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the 100Sn binding energy and 99In excited states, whose values p35-i3 predicts; if the deviations exceed the roughly 300 keV found in the $Z \\leq 36$ extrapolation test, the paper's claim that the validation protocol yields reliable extrapolation would be falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the factorized IMSRG(3f2) approximation whose three-body corrections produce the improved starting Hamiltonian."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the USD-style empirical shell-model fit and the roughly 200 keV RMSD benchmark that this paper aims to improve upon."},{"cited_title":"Explicitly: yi = dici = di pX l=1 el AT il Simultaneously, linear combinations of abinitio Hamiltonian parameters are determined from equation (4)","cited_arxiv_id":null,"evidence_quote":"Introduces the SVD-based fitting with a singular-value cutoff that the paper's varied-linear-combination protocol extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the VS-IMSRG framework used to derive the valence-space effective Hamiltonians from the ab initio interaction."},{"cited_title":"In the single-particle model this gives a 1 p1/2 − 1p3/2 spin-orbit splitting of 0.855 MeV","cited_arxiv_id":null,"evidence_quote":"Provides the 79Cu low-lying level data that constrain single-particle energies near 78Ni."},{"cited_title":"Talmi and I","cited_arxiv_id":null,"evidence_quote":"Provides the 99In level data that pin down hole states near the 100Sn closed shell."},{"cited_title":"Auerbach and I","cited_arxiv_id":null,"evidence_quote":"Supplies the EM 1.8/2.0 NN+3N interaction used as input to the IMSRG calculations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ENSDF is the source of the experimental energies and uncertainties used in the fits."}],"review_version":1}