{"id":"97e3bfbc-c11c-4e09-a4dc-b5908b02a497","arxiv_id":"2412.20356","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Minimizing a neural-network-estimated beam emittance with Bayesian optimization tunes electron microscope aberrations in minutes, but the real-microscope comparison is scored by the same network.","lead":"Bayesian optimization of a neural-network-predicted beam emittance can automatically tune an electron microscope's aberration correctors in about four minutes, based on simulation and two real instruments. The experimental speed advantage is real only if the network's emittance estimate is trustworthy, because the network is both the optimizer's target and the final scorekeeper.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Experimental comparison is circular: the CNN serves as both optimization objective and evaluation metric, and no independent optical-state measurement is reported, so the claimed superiority over Zemlin tuning is not established.","rationale":"The central claim has two parts: the conceptual equivalence of emittance minimization to aberration correction, and the empirical demonstration that BO with a CNN objective outperforms conventional Zemlin tuning. The simulation section (III.A) is convincing because emittance is computed directly from GPT ray-traced phase space, and BO improves the probe compared to Nelder-Mead. The weak point is the experimental section (III.B). There, the CNN from Part I is the sole quantitative oracle. It provides the BO's initial labels, is minimized by the acquisition function, and is then used to score the final states of both BO and Zemlin corrections. No independent measurement of the final optical state is reported. If the CNN generalizes imperfectly to real microscopes, the reported numeric comparison is unreliable, and the qualitative Ronchigram flat areas are insufficient to quantify a 3.7x improvement. The missing support is a calibration or independent verification step. This is precisely the weakest assumption the reader identified. A straightforward check is to measure the residual aberrations of the BO-final state with the same Zemlin tableau method used as the baseline. If the BO state truly has lower residual aberrations, the CNN is trustworthy and the claim stands; if not, the experimental comparative claim collapses to a simulation-only result. I considered the defocus-invariance issue as an alternative, but it is secondary because defocus can be set separately, and the paper does not state whether defocus is included in the optimized controls. Therefore I recommend keeping the CONDITIONAL verdict: simulation accepted, experimental comparative claim conditional on independent verification.","tokens_in":11100,"tokens_out":10643,"duration_ms":109480,"concrete_test":"On the Spectra 300, after BO converges, measure the residual aberrations with the same Zemlin tableau routine used for the comparison, and compare the coefficient magnitudes and probe size to the Zemlin-tuned state. If the BO state does not show residual aberrations at least as small as the Zemlin result, the CNN-based objective was biased and the comparative claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The experimental demonstration hinges on the CNN from Part I as both the optimization objective and the only quantitative evaluator. In Section III.B, ten Ronchigrams are scored by this CNN to seed the GP, the acquisition function minimizes CNN-predicted emittance, and the final comparison to the Zemlin tableau (0.0266±0.0037 vs 0.0977±0.0041) is made by re-scoring Ronchigrams with the same CNN. No independent ground truth—residual aberration coefficients, probe size, or image resolution—is reported for the BO-final state. If the CNN's mapping from real Ronchigrams to emittance is biased (due to sample thickness, camera gain, aperture mismatch, or its trained invariance to defocus), the optimizer is chasing a corrupted target and the 'better optical state' claim could be an artifact. The qualitative Ronchigram flat-area improvement is supporting evidence, but it does not establish the quantitative superiority claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a Bayesian optimization (BO) framework for automated aberration correction in scanning transmission electron microscopy (STEM). The objective is a scalar beam-emittance-growth metric predicted by a convolutional neural network (CNN) from a single Ronchigram, developed in Part I. The authors propose a deep-kernel Bayesian optimization (DKBO) variant to capture inter-channel correlations. They validate the method in General Particle Tracer (GPT) simulations of a simplified hexapole-corrected microscope, showing that BO outperforms the Nelder-Mead simplex baseline in 2D and 6D tuning tasks, and report online experiments on a ThermoFisher Titan Cryo-S/TEM and a ThermoFisher Spectra 300 (plus a Nion UltraSTEM in Appendix A), where the method converges in about 50 iterations and 4 minutes. They compare the final state to the standard Zemlin tableau method on the Spectra 300 and claim their approach achieves a better optical state (normalized emittance 0.0266 vs. 0.0977) with a higher convergence rate.","tokens_in":11258,"tokens_out":4304,"duration_ms":43933,"significance":"If the claims hold, the method is a substantial practical advance: it replaces a 2–3 hour manual tuning procedure with a fully automated, roughly 4-minute routine that uses a single scalar metric and does not require explicit measurement of aberration coefficients. The simulation results are credible and provide independent support: GPT ray tracing shows a large FWHM reduction (from 0.385 μm to 0.0064 μm in one example), a visibly enlarged Ronchigram flat area, and a reduced phase-space area. The convergence benchmarks over 20 repetitions are a solid feature. However, the central experimental claim that BO 'outperforms conventional approaches' is currently supported only by evaluations using the same CNN that serves as the optimization objective, which is a circular comparison. The paper would be significantly strengthened by an independent optical-state measurement (e.g., residual aberration coefficients, probe size, or image resolution).","major_comments":[{"comment":"The quantitative experimental comparison is circular. In Section III.B, the CNN emittance predictor is used as (i) the objective for Bayesian optimization, (ii) the metric for the final BO state (0.0266 ± 0.0037), and (iii) the metric for the Zemlin-tuned state (0.0977 ± 0.0041). No independent measurement of the optical state—such as residual aberration coefficients, direct probe size, or a resolution test—is reported for the BO-final states. Therefore the headline claim that the proposed method 'outperforms conventional approaches' in experiments is not established by the reported data. The authors should either provide an independent evaluation of the BO-final states or explicitly limit the experimental claim to a qualitative Ronchigram improvement.","section":"III.B"},{"comment":"The comparison with the Zemlin tableau is not a controlled experiment. The text states that the Spectra 300 was tuned 'several times' using the Zemlin method, but gives no details on the number of tuning iterations, the stopping criterion, the operator variance, or whether the same experimental conditions (sample, aperture, camera settings) were used as in the BO runs. The BO results are averaged over 10 runs, whereas no distribution is reported for the Zemlin results. This makes the quantitative margin (0.0266 vs. 0.0977) hard to interpret even if the circularity concern were resolved.","section":"III.B"},{"comment":"The dramatic FWHM reduction (0.3850 μm to 0.0064 μm) is reported for a single illustrative run, not averaged over the 20 repetitions used for the convergence benchmarks in Figure 6. The text claims that 'the optimization of aberration correctors according to the minimization of beam emittance can effectively eliminate lower order aberrations,' but this claim is supported by only one example. Please report the distribution of final FWHM or emittance over the 20 runs and assess statistical significance.","section":"III.A / Figure 4"},{"comment":"The claim that the deep kernel 'eliminates the chance of being trapped at local optima' is not quantitatively supported. Figure 7 shows qualitative scatter plots of queried points, but no metric for exploration/exploitation balance or local-optimum avoidance is provided, and the standard error bars in Figure 6 overlap in several regions. Please quantify exploration (e.g., coverage of the parameter space, GP predictive variance, or final emittance variance across repetitions) or temper the claim.","section":"II.D / Figure 7"}],"minor_comments":[{"comment":"There is a duplicated heading: 'C. Physics-informed kernel' appears twice, immediately after 'C. Bayesian optimization'. Also, equation numbers are reused (Equation (5) appears twice, and Equation (2) is used for both the emittance definition and the Matern kernel). Please renumber consistently.","section":"II.C"},{"comment":"Typographical errors: 'Baysian' should be 'Bayesian', 'shruk' should be 'shrunk', and 'posess' should be 'possess'.","section":"III.A"},{"comment":"The paper's title advertises 'physics-informed' optimization, but the implemented DKBO method uses a deep neural network kernel and the Hessian-derived correlation kernel described in Section II.C is not used in the actual experiments or benchmarks. Please clarify how 'physics-informed' applies to the implemented method, or adjust the title/abstract to avoid overclaiming.","section":"II.C"},{"comment":"The text states the budget of iterations for the GPT-6D runs is 300, while the right panel of Figure 6 is labeled 'within 100 iterations.' Please clarify the final budget and whether the plotted results are truncated or the budget was changed.","section":"III.A / Figure 6"},{"comment":"The CNN emittance predictor from Part I is used as the objective for all experiments, but the paper gives only a reference to Part I and does not summarize its architecture, training data, or validation on experimental Ronchigrams. Since the experimental results depend entirely on this predictor, a brief description or at least a validation summary (especially regarding defocus invariance and transfer to different microscopes) should be included.","section":"III.B"},{"comment":"No data availability statement is provided. Please include a statement on whether the GPT simulation scripts, BO code, and experimental datasets will be made available.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The circular evaluation is the main technical obstacle to acceptance. The simulation results are credible and provide strong independent support for the core idea, so the paper is salvageable through major revision. I would ask the authors to either add an independent experimental validation (e.g., residual aberration coefficients from the corrector software, or a resolution measurement) or substantially weaken the experimental superiority claim. Also, the comparison to Zemlin should be better controlled. The paper would benefit from a data availability statement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague – this is Part II of the emittance-based aberration correction work. The new contribution is the complete BO pipeline: CNN-predicted emittance from Ronchigrams as the objective, deep kernel BO to handle coupled lens channels, and validation on two (three including Nion) real microscopes with 50 iterations in about 4 minutes. The simulation work is genuinely useful: GPT ray tracing with 6 tunable elements, comparison against Nelder-Mead, and a clear demonstration that BO beats the simplex baseline; DKBO shows faster long-run convergence at the cost of early overfitting. The probe FWHM drop from 0.385 μm to 0.0064 μm is a direct ray-traced measurement, independent of the CNN, so the feasibility argument is not circular there.\n\nThe soft spot is the experimental comparison to the Zemlin tableau. The same CNN from Part I is the optimization objective and the evaluation metric: Section III.B seeds the GP with CNN-predicted emittance, the acquisition minimizes CNN-predicted emittance, and the final BO vs Zemlin numbers (0.0266±0.0037 vs 0.0977±0.0041) come from re-scoring Ronchigrams with that same CNN. No independent ground truth—residual aberration coefficients, measured probe size, or image resolution—is reported for the BO-final state. If the CNN's simulated-to-real transfer is biased (camera gain, sample thickness, aperture mismatch, defocus invariance), the optimizer is chasing a corrupted target and the 'better optical state' claim is not established quantitatively. The qualitative flat-Ronchigram-area improvements support the method but do not nail the comparison.\n\nMinor issues: hyperparameters and CNN training details live in Part I/Appendix D, so reproducibility is only partial; it's unclear whether the online experiments used DKBO or vanilla BO with an RBF kernel; and the claim that this outperforms conventional approaches is stronger than the evidence warrants, given the self-referential evaluation.\n\nProportionally: the simulation is solid, the engineering is real, and the workflow is a genuine step toward autonomous operation. The experimental circularity is fixable—add an independent probe-size or aberration-coefficient measurement for the BO state, and release the CNN weights and BO code. This deserves a serious referee, not a desk reject. If I were refereeing, I'd recommend major revision with a required independent experimental metric, but I'd expect the central feasibility claim to survive.\n\nFor a reading group: maybe bring it for the BO/microscopy intersection, but the circularity makes it a good discussion piece on evaluation pitfalls.","headline":"Solid simulation-backed automation paper whose headline experimental claim rests on a circular metric; worth refereeing for the simulation and the method, with a demand for independent experimental validation.","tokens_in":11812,"tokens_out":3760,"would_cite":true,"duration_ms":36453,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that aligning an aberration-corrected scanning transmission electron microscope can be reduced to optimizing a single scalar—the beam emittance growth predicted from a single Ronchigram by a neural network—using Bayesian…","keywords":["Bayesian optimization","beam emittance","aberration correction","scanning transmission electron microscopy","Ronchigram","deep kernel learning","Gaussian process surrogate","autonomous microscope tuning"],"falsifier":"Take the states the optimizer declares optimal on a real microscope and measure the probe directly—by imaging a known atomic lattice, by 4D-STEM ptychographic reconstruction, or by reconstructing the phase-space emittance from the Wigner distribution—and check whether the CNN's emittance ranking matches the physical probe quality; if the CNN says a state is good while the probe is demonstrably aberrated, the central equivalence between the learned metric and aberration correction fails.","tokens_in":10894,"feed_emoji":"🔬","tokens_out":9533,"duration_ms":90744,"temperature":0.7,"pith_summary":"This paper argues that tuning an aberration-corrected scanning transmission electron microscope—normally a 2-to-3-hour expert procedure—can be recast as a single-number optimization. The authors' key move is to replace the standard practice of measuring many individual aberration coefficients with a quantity called beam emittance growth, which they show is mathematically equivalent to the quality of the aberration correction and is convex in the aberration coefficients (one bowl-shaped landscape to descend). A neural network trained on simulated Ronchigrams—the diffraction shadow patterns formed from an amorphous sample—estimates this emittance from one experimental Ronchigram, and a Bayesian optimizer (with a deep neural-network kernel) then searches the corrector controls to minimize it. In simulations and on three real microscopes, the loop converges in about 50 iterations—roughly 4 minutes—to a lower normalized emittance than the conventional Zemlin-tableau method, suggesting that automated, machine-driven alignment is practical.","feed_headline":"Bayesian loop aligns electron microscopes in 4 minutes","feed_subtitle":"It replaces hour-long expert alignment with one beam-quality number and about 50 probe images.","key_machinery":"The load-bearing object is the beam emittance growth $\\varepsilon_{rms}^2$, a single scalar obtained from the second moments of the Wigner distribution of the 2D electron wave function at the aperture. For an aberration function $\\chi(\\vec{\\alpha})$ expanded in Krivanek notation, it can be computed from $\\langle|\\nabla\\chi|^2\\rangle$, $\\langle|\\vec{\\alpha}|^2\\rangle$, and $\\langle \\vec{\\alpha}\\cdot\\nabla\\chi\\rangle$, which the authors show reduces to an integral over the squared aperture brightness and the gradient of $\\chi$. Its useful properties are that it is independent of defocus and convex in the aberration coefficients, so the Hessian with respect to aberration coefficients is positive semidefinite. That convexity makes it a well-behaved black-box objective, and a CNN trained on simulated Ronchigrams supplies the mapping from experimental images to emittance. The optimizer is Bayesian: a Gaussian-process surrogate with a kernel (RBF, Matern, or a deep-kernel network $k(g(\\mathbf{x}_1,\\mathbf{w}), g(\\mathbf{x}_2,\\mathbf{w}))$) is updated after each Ronchigram, and the next corrector setting is chosen by an acquisition function (upper confidence bound). The deep-kernel variant is what lets the surrogate learn couplings between separate control channels without measuring the aberration coefficients explicitly.","core_discovery":"The central claim is that minimizing beam emittance growth is equivalent to performing aberration correction, so the whole tuning problem collapses to optimizing a single scalar objective rather than estimating individual Zernike coefficients. Building on the Part I derivation, the authors take the root-mean-square emittance growth $\\varepsilon_{rms}^2$ defined from the Wigner distribution of the electron wave function, note that it depends only on the gradient of the aberration function $\\chi(\\vec{\\alpha})$ and is convex in the aberration coefficients, and use a deep convolutional network trained on simulated Ronchigrams to predict it from a single experimental Ronchigram. Around this predictor they build a Bayesian optimization loop, testing generic RBF and Matern kernels and a deep-kernel surrogate that learns correlations between control channels. The reported results are that Bayesian optimization outperforms the Nelder-Mead simplex baseline in simulation, that the deep kernel converges to lower emittance than isotropic kernels given enough iterations, and that on two ThermoFisher microscopes the automated loop reaches a CNN-predicted normalized emittance of $0.0266 \\pm 0.0037$ and $0.0264 \\pm 0.0045$ in about 50 iterations and 4 minutes, versus $0.0977 \\pm 0.0041$ after Zemlin-tableau tuning on the Spectra 300.","pith_inferences":["A natural extension, not spelled out in the paper, is to use the same Wigner-based emittance metric for other charged-particle optical systems—accelerator beamlines or electron sources—wherever a CNN predictor can be trained.","The paper does not compare the CNN's emittance scores with a direct experimental phase-space measurement; doing so (for example, with 4D-STEM or a Wigner reconstruction of the probe) would independently confirm that the optimized state is genuinely better.","The periodic 'deGauss' reset hints that magnetic hysteresis is a confound for the surrogate; an extension would be to model that history explicitly in the kernel or state, potentially removing the need for resets and speeding convergence.","Because the CNN was trained on simulation and applied without per-instrument recalibration, a practical extension is to fine-tune it with the first few online runs on each microscope, making the objective robust to instrument-specific optics and detectors."],"forward_implications":["Aberration correction no longer requires measuring individual Zernike aberration coefficients; a single Ronchigram plus one scalar objective suffices, which removes the multi-image tilt-series bottleneck of the Zemlin tableau.","A commercial microscope can be aligned from a random starting state in about 50 iterations and 4 minutes, compared with hours for a human expert or several multi-minute Zemlin measurements, making frequent re-tuning practical.","Because the CNN can score every candidate state from a single Ronchigram, the loop can run on an amorphous sample area, allowing re-alignment in the middle of an experiment rather than only at the start.","Deep-kernel surrogates that learn input couplings reduce the chance of getting trapped near local optima and reach lower emittance than isotropic RBF/Matern kernels in the long run, according to the simulation benchmarks.","The same emittance-minimization-plus-Bayesian-optimization framework is claimed to generalize to other high-dimensional, expensive scientific-instrument tuning tasks."],"supporting_citations":[{"why":"Supplies the Wigner-Weyl transform derivation that connects the aberration function to phase-space emittance.","marker":"[3]"},{"why":"Part I conference report that establishes the CNN mapping from Ronchigrams to emittance and the convexity properties used as the objective.","marker":"[4]"},{"why":"Gives the Bayesian optimization framework the method builds on for sample-efficient black-box optimization.","marker":"[5]"},{"why":"The BEACON baseline that optimizes an image-contrast metric with Bayesian optimization, which the paper compares against.","marker":"[7]"},{"why":"Provides the Hessian-based correlation transformation used to construct a physics-informed kernel from the emittance Hessian.","marker":"[8]"},{"why":"The General Particle Tracer code used to simulate trajectories and generate the simulated microscope objective for benchmarking.","marker":"[14]"},{"why":"Introduces deep kernel learning, the basis of the deep-kernel Gaussian-process surrogate used by DKBO.","marker":"[15]"},{"why":"Supports the combination of deep kernel learning with Bayesian optimization as an approximate posterior sampling method.","marker":"[17]"}],"fun_headline_variants":["Bayesian optimizer tunes electron microscope in 4 minutes","Emittance metric cuts microscope alignment to minutes","Automated aberration correction in 4 minutes via Bayesian optimization","Bayesian emittance optimization aligns electron microscopes fast"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole demonstration rests on the assumption that the neural network trained on simulated Ronchigrams predicts true emittance on real, imperfect microscopes closely enough that minimizing its output also minimizes actual aberrations.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian optimizer tunes electron microscope in 4 minutes","Emittance metric cuts microscope alignment to minutes","Automated aberration correction in 4 minutes via Bayesian optimization","Bayesian emittance optimization aligns electron microscopes fast"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00078,"raw_usage":{"total_tokens":3490,"prompt_tokens":1033,"completion_tokens":2457,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":2395}},"tokens_in":649,"tokens_out":2457,"duration_ms":17876,"temperature":1.0,"reasoning_tokens":2395,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:23:34.121492+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the states the optimizer declares optimal on a real microscope and measure the probe directly—by imaging a known atomic lattice, by 4D-STEM ptychographic reconstruction, or by reconstructing the phase-space emittance from the Wigner distribution—and check whether the CNN's emittance ranking matches the physical probe quality; if the CNN says a state is good while the probe is demonstrably aberrated, the central equivalence between the learned metric and aberration correction fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Part I conference report that establishes the CNN mapping from Ronchigrams to emittance and the convexity properties used as the objective."},{"cited_title":"BEACON -- Automated Aberration Correction for Scanning Transmission Electron Microscopy using Bayesian Optimization","cited_arxiv_id":"2410.14873","evidence_quote":"The BEACON baseline that optimizes an image-contrast metric with Bayesian optimization, which the paper compares against."},{"cited_title":"The go-to choice of kernel is the radial basis function (RBF) kernel","cited_arxiv_id":null,"evidence_quote":"Provides the Hessian-based correlation transformation used to construct a physics-informed kernel from the emittance Hessian."},{"cited_title":"Krivanek, N","cited_arxiv_id":null,"evidence_quote":"The General Particle Tracer code used to simulate trajectories and generate the simulated microscope objective for benchmarking."},{"cited_title":"De Loos, S.B","cited_arxiv_id":null,"evidence_quote":"Introduces deep kernel learning, the basis of the deep-kernel Gaussian-process surrogate used by DKBO."}],"review_version":1}