{"id":"23cbb699-37dc-4f9e-bb89-94c2852e3a1c","arxiv_id":"1909.02372","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"An element-wise ultrasound simulator, with CT-density to acoustic-property curves fitted and then validated against MRI thermometry in 72 patients, predicts transcranial focused ultrasound focal heating with validation R²=0.71.","lead":"This paper models transcranial focused ultrasound treatments element by element, then tunes how CT skull density maps to sound speed and attenuation so simulated heating matches MRI temperature measurements. The tuned model predicts focal heating across 72 patients with validation R² about 0.7, which could improve planning of noninvasive brain surgery.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validation R² is not fixed-model: the 'after skull change' density/sound speed curve is fitted to deviant sonications selected by the optimized attenuation model; the reported out-of-sample R² depends on this ad hoc regime switch.","rationale":"The reader's weakest assumption identifies the same concern: the second sound speed relationship is ad hoc and the deviant split is model-dependent. My analysis confirms this is the most load-bearing issue because the final R² reported in the strongest claim is computed after applying that second curve to sonications that were selected, in both training and validation, by a criterion based on the optimized attenuation model's residuals. This makes the supposed out-of-sample validation less independent than it appears. However, the concern does not warrant rejection: the primary element-wise simulation framework is demonstrated on a large clinical cohort, the validation cohort is genuinely new patients, and the shape/obliquity correlations (R² 0.62, 0.74) do not rely on the debatable second curve. The paper also discloses its own limitations candidly. A single computational check can quantify how much of the R² depends on the ad hoc regime switch; until that is reported, the conditional verdict stands. I therefore leave the reader's verdict unchanged.","tokens_in":24250,"tokens_out":5407,"duration_ms":56189,"concrete_test":"Re-run the 40-patient validation simulations using only the optimized attenuation model with the pre-change density/sound speed relationship, applying no deviant classification and no second curve; compute the R² for all 396 sonications. If this single-curve R² is within ~0.05 of the reported 0.71, the second curve is not load-bearing. If it drops by more than 0.1, the headline R² depends on the ad hoc regime switch. As a second check, apply the second curve to validation sonications selected by a fixed accumulated-energy threshold (e.g., >17.6 kJ, the mean from training) rather than the residual-based rule; a large R² drop would confirm the classification is not predictive.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the element-wise model with fitted CT-density/attenuation and CT-density/sound speed relationships predicts relative peak temperature rise (R² 0.74 training, 0.71 validation) and focal geometry. The weak link is Section II.D's ad hoc second density/sound speed curve for 'after skull change' sonications. In Section II.C, the attenuation relationship is optimized on sonications before 'deviation,' and the deviants are selected automatically when two or more sonications exceed two SD above the mean difference 'for the optimized attenuation model' (Section III.B). Those same deviants are then used to fit the second curve (Section II.D), and the final R² includes them. Thus the model is not a single fixed predictor; it uses a regime switch keyed to residuals of the same model. The paper itself acknowledges (Discussion) that the change may not be in sound speed, may not be uniform/binary, and that the interpretation 'may not be correct for every patient' (patients 27 and 30). If the deviation is actually caused by unmodeled physics (dynamic bone changes, nonlinear propagation, thermal property errors), the second curve is an overfit to residual patterns and the validation R² will not reproduce under a pre-specified rule. The authors' own finding that measured heating is more diffuse than simulated is consistent with missing physics rather than a sound-speed change.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an element-wise simulation framework for transcranial MRI-guided focused ultrasound (TcMRgFUS). Each element of a hemispherical phased array is simulated separately with the k-Wave toolbox, using the same CT-derived skull model and the phase/magnitude values used during treatment, and the pressure fields are combined and fed into a bioheat solver to estimate temperature rise. Using MRTI data from 32 patients (431 sonications), the authors optimize polynomial relationships between CT-derived skull density and acoustic attenuation and sound speed, including a second sound-speed relationship intended to apply after a presumed irreversible change in skull properties during treatment. They validate the optimization on 40 additional patients (396 sonications), reporting R2 of 0.74 for the training cohort and 0.71 for the validation cohort for relative peak temperature rise, and R2 of 0.62 and 0.74 for focal dimensions and obliquity. The paper also reports that measured heating is more spatially diffuse than simulated and that the acoustic energy required to reach an ablative thermal dose varies by an order of magnitude.","tokens_in":24592,"tokens_out":5685,"duration_ms":59074,"significance":"If the central claim holds, this is a substantial clinical modeling contribution: a large (72-patient) dataset with per-element simulations, MRTI comparison, and an independent validation cohort is rare in this field, and the element-wise architecture is genuinely useful for rapid iteration over phase/magnitude schemes. The paper ships explicit optimized coefficients (Table 3), uses the open-source k-Wave toolbox, and reports falsifiable comparisons to measured heating shape, size, and obliquity. The independent 40-patient validation is a real strength. However, the main predictive claim is weakened by the post-hoc, residual-defined regime switch for the second sound-speed curve: the reported R2 figures do not correspond to a single fixed predictive model, and the paper's own Discussion acknowledges that the assumed physical mechanism may be incorrect. The approach and the optimization framework are nevertheless valuable, and the identified limitations appear addressable with additional analysis.","major_comments":[{"comment":"The 'after skull change' density/sound speed curve is fitted to sonications that were selected automatically as 'deviants' using the optimized attenuation model (two or more sonications exceeding two standard deviations above the mean difference in Fig. 7d-f). This selection is a function of the model's own residuals, so the second polynomial in Table 3 is fit to the residual patterns of the first model rather than to independently measured physical changes. Because the reported R2 values (0.74 and 0.71) include sonications simulated with this second curve, the paper does not report the predictive performance of a single fixed predictor. Please report, for both the training and validation cohorts, the R2, slope, and intercept for non-deviant sonications under the 'before' curve alone, for the deviant sonications under the 'after' curve, and for a pre-specified two-regime model in which the switch point is fixed in advance (e.g., based on accumulated acoustic energy) rather than selected by residual inspection.","section":"Section II.D and III.B (Eq. 6, Fig. 7)"},{"comment":"The validation on patients 33-72 inherits the same residual-based regime switch: deviant sonications in the validation patients are identified using the optimized attenuation model's residual plot (Fig. 11c-d) and then simulated with the 'after' curve. Thus the out-of-sample R2 does not independently test the substantive claim that a change in skull sound speed causes the deviation. The Discussion itself states that the change may not be spatially uniform or binary, that sound speed 'may not even be correct' as the dominant factor, and that patients 27 and 30 deviate without prior accumulated energy. These admissions indicate that the second curve could be an overfit to residual patterns from missing physics (e.g., shear-mode transmission, nonlinear propagation, incorrect thermal properties). Please re-analyze the validation cohort with a pre-specified switch rule and report the performance difference between the two-curve model and a single-curve model; also report how many validation sonications would be classified as deviant by the fixed rule and whether the 'after' curve actually improves out-of-sample prediction.","section":"Section III.E and Discussion (Fig. 11)"},{"comment":"The abstract and results sections report R2 of 0.74 for patients 1-32 and 0.71 for patients 33-72 without flagging that the first number is in-sample: the attenuation polynomial (Eq. 3) and both sound-speed polynomials (Eq. 6) were optimized on the MRTI temperatures of patients 1-32, and the deviant-selection threshold was also chosen from these data. Presenting 0.74 as the 'agreed well' number conflates training and validation performance. Please label the 0.74 as training performance in the abstract, results, and conclusion, and lead with the out-of-sample result for the pre-specified fixed model, including separate metrics for non-deviant and deviant subsets.","section":"Section III.B and Abstract"}],"minor_comments":[{"comment":"The typesetting of several equations is garbled, with missing subscripts and an unclear grid-size variable, which makes the polynomial coefficient definitions and summation limits hard to follow; please revise the notation for publication.","section":"Equations (2)-(5) and (6)-(11)"},{"comment":"The visual comparison uses 50% contours for MRTI and 25% contours for simulated heating; this asymmetry should be stated in the main text as well as the caption, and the systematic underestimation of focal dimensions (R2 0.62 with simulated widths smaller than measured) should be discussed as a limitation rather than only as a correlation.","section":"Fig. 5a"},{"comment":"The optimized coefficients are presented without uncertainties or a statement of the density range over which they were fit; please provide confidence intervals or a plot of the curves overlaid on the observed density histogram.","section":"Table 3"},{"comment":"The candid acknowledgment that the deviation 'may not even be correct to assume that the sound speed is the dominant factor' is appreciated, but the conclusion still states the model predicted focal temperature rise 'on average'; please qualify the conclusions to reflect the unresolved mechanism.","section":"Section IV (Discussion)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's principal risk is the model-selection circularity in the two-curve sound-speed model, which is load-bearing for the headline R2 values. I recommend that the editor seek a statistical review of the residual-based regime switch. If the authors can provide a pre-specified switch rule and out-of-sample metrics for fixed models, the paper would be substantially strengthened. The data are extensive and the element-wise methodology is a genuine contribution, so I view these concerns as addressable within a major revision rather than grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's the short version: this paper shows that an element-wise simulation approach for transcranial focused ultrasound can predict relative peak heating across 72 patients with usable accuracy, and it does so with a genuinely separate validation cohort. The method itself isn't new—it's essentially Clement et al. 2000—but the clinical-scale empirical fitting of CT-density to attenuation and sound speed is new, and the validation is the real contribution.\n\nWhat it does well: 825 sonications, 40 patients held out. The R2 of 0.71 on the held-out patients is the number that matters, not the training R2. The shape and obliquity comparisons are also a plus; the example of the internal capsule involvement is clinically meaningful. And the authors are candid about the limitations, which is more than you usually get.\n\nThe soft spots: the 'after skull change' sound speed curve is fit to sonications that were flagged as deviant using the already-optimized attenuation model. That is a circularity, and the paper admits the interpretation is ad hoc. If the true cause is something else—dynamic bone changes, nonlinear propagation, thermal property error—the second curve is a residual fit and the reported R2 won't generalize under a pre-specified rule. The measured heating being systematically more diffuse than simulated points in the same direction. That said, the validation cohort was unseen when the curves were fixed, so the 0.71 is still an honest out-of-sample number for the full two-curve model, even if the physical story is shaky.\n\nMy take: the central engineering claim—that you can use this approach to predict relative focal heating and obliquity—holds up, with the caveat that the regime switch weakens the physical interpretation but not the empirical validation. The paper is worth a serious referee. I'd send it out.","headline":"A clinically valuable and honestly reported validation of a known simulation strategy, with one data-driven regime switch that complicates the physical story but not the out-of-sample result.","tokens_in":25067,"tokens_out":1656,"would_cite":true,"duration_ms":19094,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that separately simulated phased-array elements, combined with CT-derived skull acoustic properties optimized against MR thermometry, predict focal heating in transcranial ultrasound to an $R^2$ of 0.74 in 32…","keywords":["transcranial MRI-guided focused ultrasound","thermal ablation","element-wise simulation","phased array transducer","CT-derived skull density","acoustic attenuation","sound speed","MR thermometry"],"falsifier":"Measure skull sound speed directly before and after high-intensity sonication, for example by through-skull time-of-flight or transmission measurements on ex vivo human skulls exposed to comparable acoustic energies; if the sound speed does not shift in the way the fitted second density/sound speed curve describes, the two-curve model and the reported validation $R^2$ would not generalize.","tokens_in":2012,"feed_emoji":"🧠","tokens_out":2335,"duration_ms":74813,"temperature":0.7,"pith_summary":"The paper is trying to show that transcranial MRI-guided focused ultrasound thermal ablation can be modeled by simulating each transducer element in isolation, storing the resulting pressure fields, and then combining them with the phase and amplitude settings used during treatment. After fitting two empirical relationships between CT-derived skull density and acoustic attenuation and sound speed, the model matched measured peak temperature rises across 825 sonications in 72 patients, with $R^2 = 0.74$ in the optimization cohort and $R^2 = 0.71$ in an independent validation cohort. The same simulations captured the size, shape, and tilt of the heated region, although the measured heating was more spatially diffuse than predicted. This matters because a fast, validated simulation could help plan treatments, screen patients who need very high acoustic energy, and reveal oblique or off-target heating that single-plane MR thermometry might miss.","feed_headline":"Element-wise simulation predicts focused ultrasound heating","feed_subtitle":"Peak temperatures matched across 825 sonications in 72 patients, and focal shape and tilt were captured too.","key_machinery":"The central object is the element-wise simulation library: each transducer element is treated as a 1 cm flat circular piston and simulated separately with the k-Wave toolbox in a $44 \\times 44 \\times 492$ voxel grid at 0.325 mm spacing, then the resulting complex pressure fields are stored so that any phase and magnitude distribution can be combined in memory. The argument is carried by two fitted polynomial relationships connecting CT-derived skull density to acoustic properties: attenuation vs. density, $\\alpha(\\rho) \\approx \\sum_m A_m \\rho^m$, and inverse sound speed vs. density, $1/c(\\rho) \\approx \\sum_m B_m \\rho^m$. A simplified attenuation model and a transmission-coefficient model based on acoustic impedances and Snell's law allowed the authors to test 10,000 candidate density/attenuation and density/sound speed curves per sonication, and the bioheat equation was then used to estimate temperature rise from the resulting pressure fields.","core_discovery":"The central claim is that an element-wise simulation approach, in which each of the 993 active phased-array elements is modeled separately and the complex pressure fields are rotated and interpolated into a common frame, can predict the focal temperature rise and heating geometry of clinical transcranial focused ultrasound treatments. Using the manufacturer's phase and magnitude corrections, the simulations reproduced the relative peak temperature rise with $R^2 = 0.74$ in the first 32 patients and $R^2 = 0.71$ in 40 additional patients, and the simulated focal dimensions and obliquity correlated with MR thermometry at $R^2 = 0.62$ and $R^2 = 0.74$, respectively. The paper further claims that the growing mismatch between simulated and measured heating after accumulated acoustic energy reflects an irreversible change in skull sound speed, and that a second density/sound speed polynomial, fitted to the deviant sonications, improves predictions in those cases. The energy required to reach an ablative temperature of 55°C varied by an order of magnitude across patients, from 3.3 to 36.1 kJ, indicating wide patient-to-patient variability in skull transmission.","pith_inferences":["A likely next step, not demonstrated in the paper, is to use MRTI feedback during a treatment to update the assumed skull sound speed curve between sonications, turning the two-curve model into an adaptive correction scheme.","The systematically more diffuse heating seen in MRTI compared with simulation points to unmodeled acoustic pathways such as shear-wave conversion at oblique incidence, reverberation, or nonlinear propagation; adding these to the element-wise framework is a testable extension that could close the gap in focal width and in the observed cooling-rate mismatch.","Because the deviant sonications used to fit the second density/sound speed curve were themselves selected using the optimized attenuation model, the second curve should be validated against direct measurements of skull sound speed before and after sonication, for example through-skull time-of-flight in ex vivo bone, rather than inferred only from temperature mismatch.","The worst-outlier patients, who needed more than 20 kJ to reach ablative temperature without being predicted by any single skull metric, hint that factors beyond density-based attenuation, such as microstructural scattering or localized skull heterogeneity, need to enter the model; a concrete test would be whether per-element phase-error or scattering terms improve prediction in exactly those pati"],"forward_implications":["If the element-wise library is built once, any phase or magnitude correction scheme can be evaluated rapidly, enabling near-real-time exploration of beam steering and aberration correction without rerunning full three-dimensional simulations.","The optimized CT-density to attenuation and sound speed relationships, fitted on 32 patients and validated on 40 more, give other transcranial focused ultrasound modeling efforts a directly usable parameterization for human skull bone.","Because the simulations reproduce focal obliquity, including cases where heating tilts in two planes and extends toward the internal capsule, they can flag off-target heating risks that single-plane MR thermometry may miss.","The finding that ablative energy varies from 3.3 to 36.1 kJ across patients, with simulations predicting this energy at $R^2 = 0.45$, suggests that simulation-based screening could supplement the currently used skull density ratio, though the most difficult patients remain poorly predicted.","The two-curve density/sound speed model implies that skull acoustic properties can change during a treatment, and that the observed loss of treatment efficiency at high accumulated acoustic energy is at least partly a sound speed effect that can be modeled."],"supporting_citations":[{"why":"Supplies the baseline empirical density/attenuation and density/sound speed relationships from excised human skulls that the optimization starts from and refines.","marker":"[15]"},{"why":"Provides the k-Wave MATLAB toolbox used for the element-wise pressure field simulations.","marker":"[21]"},{"why":"Supplies the non-invasive focusing method and the transmission coefficient formula used in the simplified sound speed optimization.","marker":"[12]"},{"why":"Supplies the bioheat equation used to estimate temperature rise from the simulated absorbed acoustic power.","marker":"[20]"},{"why":"Establishes the proton resonance frequency shift method that underlies the MR thermometry data used as the measurement ground truth.","marker":"[23]"},{"why":"Documents the reduction in treatment efficiency at high acoustic powers, motivating the paper's interpretation of mid-treatment deviations as a change in skull properties.","marker":"[11]"},{"why":"Provides the numerical study of oblique focus in transcranial focused ultrasound that frames the paper's analysis of focal obliquity and its clinical risk.","marker":"[8]"},{"why":"Offers a recent alternative rapid beam simulation framework for transcranial focused ultrasound, providing context for the element-wise approach's claims of speed and flexibility.","marker":"[18]"}],"fun_headline_variants":["Element-wise simulation matches focused ultrasound heating","Per-element simulation predicts ultrasound ablation heat","Modeling each array element forecasts brain ablation heating","Element-wise model predicted heating in 72 patients"],"cache_read_input_tokens":27136,"weakest_assumption_plain":"The paper's second fitted curve rests on the assumption that the growing mismatch between simulated and measured heating after accumulated acoustic energy is an irreversible change in skull sound speed that a second density/sound speed polynomial can capture, even though the sonications used to fit that curve were themselves selected as deviations from the optimized attenuation model, making the data split model-dependent.","fun_headline_variants_meta":{"raw":{"variants":["Element-wise simulation matches focused ultrasound heating","Per-element simulation predicts ultrasound ablation heat","Modeling each array element forecasts brain ablation heating","Element-wise model predicted heating in 72 patients"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000859,"raw_usage":{"total_tokens":3804,"prompt_tokens":1097,"completion_tokens":2707,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":713,"completion_tokens_details":{"reasoning_tokens":2652}},"tokens_in":713,"tokens_out":2707,"duration_ms":22297,"temperature":1.0,"reasoning_tokens":2652,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:51:36.443460+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure skull sound speed directly before and after high-intensity sonication, for example by through-skull time-of-flight or transmission measurements on ex vivo human skulls exposed to comparable acoustic energies; if the sound speed does not shift in the way the fitted second density/sound speed curve describes, the two-curve model and the reported validation $R^2$ would not generalize.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the k-Wave MATLAB toolbox used for the element-wise pressure field simulations."},{"cited_title":"Pichardo, V","cited_arxiv_id":null,"evidence_quote":"Supplies the baseline empirical density/attenuation and density/sound speed relationships from excised human skulls that the optimization starts from and refines."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the non-invasive focusing method and the transmission coefficient formula used in the simplified sound speed optimization."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the bioheat equation used to estimate temperature rise from the simulated absorbed acoustic power."},{"cited_title":"Ishihara, A","cited_arxiv_id":null,"evidence_quote":"Establishes the proton resonance frequency shift method that underlies the MR thermometry data used as the measurement ground truth."},{"cited_title":"Hughes, Y","cited_arxiv_id":null,"evidence_quote":"Documents the reduction in treatment efficiency at high acoustic powers, motivating the paper's interpretation of mid-treatment deviations as a change in skull properties."},{"cited_title":"Hughes, Y","cited_arxiv_id":null,"evidence_quote":"Provides the numerical study of oblique focus in transcranial focused ultrasound that frames the paper's analysis of focal obliquity and its clinical risk."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Offers a recent alternative rapid beam simulation framework for transcranial focused ultrasound, providing context for the element-wise approach's claims of speed and flexibility."}],"review_version":1}