{"id":"89ff1972-4cd9-45d4-a960-4c271ddc49f9","arxiv_id":"2507.15035","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"OpenBreastUS provides a large-scale, anatomically realistic benchmark of 16 million breast ultrasound simulations and demonstrates neural-operator-based full-waveform inversion on clinical in vivo breast data.","lead":"OpenBreastUS is a new dataset with 8,000 realistic breast models and over 16 million simulated ultrasound wavefields, built to train and test neural networks that solve wave equations. The authors benchmark six neural operators on forward simulation and image reconstruction, and apply these models to reconstruct living human breast tissue in vivo.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The in vivo headline is not supported as written: §IV-C trains models on 3 frequencies and 64 sources, while §IV-D claims models trained on 8 frequencies and 256 transmitters, with no documented retraining.","rationale":"I read the paper's central claim as the demonstration that neural-operator surrogates trained on OpenBreastUS can generalize to real clinical USCT data and enable efficient in vivo imaging. The strongest available evidence is the BFNO-based FWI reconstruction in §IV-D, but that section conflicts with the benchmark setup in §IV-C. This is more concrete and more easily checkable than the phantom-realism concern, which is why I focus on it. The dataset itself is a real contribution: 8,000 VICTRE-derived phantoms, 16 million CBS-simulated wavefields, a broad benchmark of six forward and four inverse operators, and the OOD generalization study in §IV-E are all valuable and plausibly executed. I also credit the detailed dataset descriptions and the explicit limitations section. The concern does not undermine the dataset's utility or the benchmark's internal ranking; it targets the headline in vivo claim. Because the issue is one of missing or contradictory reporting rather than an obviously wrong method, the appropriate verdict remains conditional: the authors should clarify whether the §IV-D models were retrained, provide the training configuration, and add at least one quantitative or same-configuration baseline comparison to the clinical reconstructions. The reader's weakest-assumption statement identified the phantom-realistic premise and the frequency/source mismatch as separate fragilities; I agree with the latter and regard the former as subordinate, so my agreement is partial. No ad hominem or language beyond the evidence is intended: the fix is a reporting correction, not a redesign.","tokens_in":20725,"tokens_out":3892,"duration_ms":48044,"concrete_test":"Download the released code and weights from the URL in §IV-A and inspect the training configuration for the models used in §IV-D. Determine whether they were trained on 3 frequencies and 64 sources, as §IV-C states, or on 8 frequencies and 256 sources. Then rerun the two in vivo reconstructions (malignant tumor and benign cyst) under both configurations and report a quantitative comparison against the FDFD reference, such as SSIM, PSNR, or tumor-boundary error. If the clinical result is only produced by a model trained on 8 frequencies and 256 sources, the retraining must be described and the benchmark tables updated; if the 3-frequency, 64-source model reproduces Figures 5–6, the mismatch is cosmetic and the paper only needs clarifying text.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of efficient in vivo imaging rests on the experiments in §IV-D, but the training configuration of the models used there is ambiguous and apparently inconsistent with §IV-C. Section IV-C states that the forward and inverse baselines were trained and tested using a subset of three frequencies (300, 400, and 500 kHz) and 64 uniformly sampled sources out of 256. Section IV-D then says 'Since our neural operators were trained on data at 8 frequencies (300–650 kHz), we evaluated the baseline models using clinical data restricted to these 8 frequencies,' and the Figure 5 caption specifies BFNO with 8 frequencies and 256 transmitters. No retraining step is described anywhere. If the §IV-D models are the same 3-frequency, 64-source models from §IV-C, applying them to 8 frequencies and 256 transmitters is a substantial distribution shift, and the clean tumor delineation in Figures 5–6 would be surprising and unexplained. If the models were retrained on the full configuration, the benchmark tables and experimental section omit that fact, making the reported results non-reproducible and the 'first in vivo imaging' claim unverifiable. This is an internal inconsistency, not merely a missing baseline: it determines whether the headline experiment was actually run in the configuration claimed. The phantom-realism concern raised by the reader is downstream of this issue, because even a perfectly realistic phantom distribution cannot rescue an experiment whose inputs are not the ones described. A second, compounding gap is that the in vivo evaluation is purely qualitative, with no same-configuration classical FWI baseline and no quantitative comparison against the FDFD reference, so there is no way to tell whether the shown tumor and cyst are accurately localized or merely plausible structures produced by the iterative procedure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"OpenBreastUS is a large-scale dataset of 8,000 anatomically realistic breast phantoms (VICTRE-based) together with 16.4 million frequency-domain Helmholtz wave simulations using a 256-transducer ring array at 8 frequencies. The paper benchmarks five forward neural operators (U-Net, FNO, AFNO, BFNO, MgNO), four inverse approaches (DeepONet, InversionNet, NIO, and optimization-based FWI with neural surrogates), and reports that BFNO-based gradient inversion reconstructs a malignant tumor and a benign cyst on two in vivo cases from the Karmanos Cancer Institute. The central claim is that models trained on OpenBreastUS generalize to clinical USCT data, enabling efficient in vivo imaging with neural operator solvers.","tokens_in":20881,"tokens_out":4975,"duration_ms":49538,"significance":"If the central claim holds, OpenBreastUS would be a substantial resource for neural operator research in wave imaging, combining realistic anatomy with a credible CBS reference solver, and the in vivo transfer would be a notable first. The paper's strengths include the scale and realism of the dataset, the use of an external clinical dataset never seen in training, the held-out phantom evaluation, and the comparison against a CBS numerical solver rather than a weaker reference. The OOD experiments (source locations, breast types, GRF media) also provide useful evidence about generalization. However, the internal inconsistency in training configuration and the purely qualitative in vivo evaluation mean the headline claim is not yet established.","major_comments":[{"comment":"§IV-C states that forward baselines were evaluated on a subset of three frequencies (300, 400, 500 kHz) and 64 of 256 sources, and that inverse baselines were trained and tested on the same three frequencies. §IV-D then says the neural operators 'were trained on data at 8 frequencies (300–650 kHz)' and the Figure 5 caption specifies BFNO with 8 frequencies and 256 transmitters. No retraining step is described anywhere in the paper. If the §IV-D models are the same as the §IV-C models, the 8-frequency/256-source clinical input is a large out-of-distribution shift; if they were retrained, the training setup is undocumented. Either way, the in vivo experiment as reported is not reproducible, and the headline claim depends on it.","section":"§IV-C, §IV-D, Fig. 5"},{"comment":"Table II includes a 600 kHz row, while §IV-C says only 300, 400, and 500 kHz were used, and §IV-D says 8 frequencies (300–650 kHz). This inconsistency makes the actual training frequencies for the benchmark models ambiguous. The paper must state exactly which frequencies and source counts were used for each model in Tables II–III and for the Section IV-D in vivo experiments.","section":"Table II, §IV-C, §IV-D"},{"comment":"The in vivo demonstration is entirely qualitative. There are no quantitative metrics (e.g., SSIM/PSNR against the FDFD reference, lesion contrast, or boundary error) and no statistical assessment. The claim that BFNO 'successfully reconstructed a clear malignant tumor' is supported only by visual inspection of two clinical cases. Since the abstract's 'first in vivo imaging' claim rests on these figures, quantitative evaluation should be added.","section":"§IV-D, Figs. 5–6"},{"comment":"The benchmark tables report a single point estimate per model and condition, with no standard deviation across training seeds or test folds. The paper's stated purpose is to benchmark neural operators, so the reported rankings (e.g., MgNO at RRMSE 0.0028 vs. BFNO at 0.0113 in Table II) cannot be assessed for statistical significance. At least three seeds with error bars and, where possible, significance tests should be reported.","section":"Tables II, III, V, VI, VII"}],"minor_comments":[{"comment":"Section III-B refers to 'OpenWaves dataset' instead of OpenBreastUS.","section":"§III-B"},{"comment":"In Appendix B, 'WIn this paper' appears to be a typo for 'In this paper'.","section":"Appendix B, InversionNet"},{"comment":"Section IV-C contains the typo 'neural operatos', which should be 'neural operators'.","section":"§IV-C"},{"comment":"Figures 5(a) and 6(a) are labeled 'Ground Truth' but are themselves reconstructions produced by an FDFD solver; relabel them as 'Reference reconstruction' to avoid implying true ground truth.","section":"Figs. 5–6"},{"comment":"The phantom generation relies on hand-set VICTRE parameters (a1b, a1t, a2l, a2r, a3, targetFatFrac ranges) with no sensitivity analysis; a brief discussion of how these choices affect the benchmark would help users interpret the dataset.","section":"§III-C2, Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The main issue is the ambiguous training configuration for the in vivo experiments. If the authors can confirm that the models were retrained on 8 frequencies and 256 sources with full documentation, and add quantitative in vivo metrics, the paper would meet the bar. The 'first in vivo imaging' claim should also be verified against the existing USCT literature to avoid overclaiming novelty."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper is worth referee time for the dataset alone, but the in vivo headline is not supported as written. The main benchmark is fine; the clinical transfer section has a load-bearing reporting gap.\n\nWhat is new and good: OpenBreastUS is a real contribution. 8,000 VICTRE-based breast phantoms across four density classes, 16M frequency-domain wavefields with realistic ring-array geometry (256 transducers, 300–650 kHz), and detailed documentation of the generation pipeline. The CBS solver is a credible reference for the ground-truth fields. The benchmark of five forward solvers and four inverse pipelines is thorough, and the ranking (MgNO best forward; surrogate-based FWI beats direct inversion) is plausible and consistent with the physics argument that direct inversion memorizes priors. The OOD source-location and breast-type tables add useful evidence. The GRF comparison is a nice demonstration that simplified media overestimate generalization.\n\nThe soft spots, in order of severity. First, the in vivo experiment is internally inconsistent with the benchmark. Section IV-C trains the baselines on three frequencies and 64 sources; Section IV-D says the models were trained on eight frequencies and uses 256 transmitters in the figures, with no retraining step described. That is a qualitative jump in input dimension and distribution. If the same models were applied outside their training configuration, the clean tumor delineation is inexplicable; if they were retrained, the paper must say so and report the setup. Either way the abstract's 'for the first time' claim is unverifiable as written. Second, the in vivo evaluation is purely visual. No same-configuration classical FWI baseline, no quantitative metric, no error bars. The FDFD reference from [15] is at 20 frequencies and 1024 sensors, so it is not a matched comparison. Third, there are no seed variations anywhere in the paper, so it is impossible to tell whether the benchmark gaps are significant. Fourth, the 'first in vivo imaging using neural operator solvers' is not reconciled with the authors' own prior Neural Born Series paper [30]. Minor: the dataset URL has no commit hash or checksum.\n\nThe phantom-realism worry (VICTRE + random perturbations) is real but downstream of the training-configuration problem. The benchmark numbers on held-out phantoms stand on their own; the in vivo transfer is the only place where realism matters, and that section is currently too thin to validate it.\n\nRecommendation: send this to peer review. The dataset and benchmark are valuable enough to warrant serious referee time, but the paper needs major revision: clarify/retrain for the in vivo configuration, add quantitative in vivo evaluation with a matched baseline, add error bars, and resolve the prior-work claim. I would not cite the in vivo result until that is fixed, but I would likely cite the dataset.","headline":"A genuinely useful large-scale USCT dataset and benchmark, undermined by an internally inconsistent report of the in vivo experiment that needs fixing before the headline claim can be trusted.","tokens_in":21659,"tokens_out":2761,"would_cite":true,"duration_ms":28963,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"OpenBreastUS lets neural wave solvers image real breast tissue for the first time.","keywords":["ultrasound computed tomography","neural operators","full-waveform inversion","breast imaging","Helmholtz equation","benchmark dataset","in vivo validation","Fourier neural operator"],"falsifier":"A decisive check would be to run the same BFNO-driven FWI pipeline on a held-out clinical USCT dataset from a different scanner, with tumor locations confirmed by biopsy or pathology; if reconstructed lesion boundaries do not match pathology within the expected margin, or if image quality collapses on this second dataset, the claimed transfer is not general.","tokens_in":1549,"feed_emoji":"🩻","tokens_out":2756,"duration_ms":92276,"temperature":0.7,"pith_summary":"OpenBreastUS is a large-scale dataset of anatomically realistic breast phantoms paired with frequency-domain wavefields, built to test whether neural operator surrogates can replace classical Helmholtz solvers in ultrasound computed tomography. The paper argues that existing PDE datasets are too simplified, so it generates 8,000 virtual breasts across four density classes and 16.4 million simulations using the geometry and frequencies of a real 256-transducer ring system. On these data, it benchmarks five forward solvers and four inverse-imaging approaches, finding that multigrid and Born-type Fourier operators give the most accurate wavefields, and that surrogate-driven gradient optimization beats direct inversion networks. The central claim is that a Born Fourier Neural Operator trained on OpenBreastUS, embedded in iterative full-waveform inversion, reconstructs a malignant tumor and a benign cyst on clinical in vivo data, which is the first demonstration of in vivo breast imaging with neural operator solvers.","feed_headline":"Neural wave solvers image live breast tumors","feed_subtitle":"A 16-million-simulation breast dataset makes learned solvers transfer from synthetic tissue to real patients.","key_machinery":"The load-bearing object is the frequency-domain heterogeneous Helmholtz equation, solved by the Convergent Born Series algorithm to create ground-truth wavefields for 8,000 virtual breast anatomies under a real annular-array USCT configuration. Neural operator baselines, including FNO, BFNO, AFNO, MgNO, and U-Net, learn the map from sound-speed map and source to complex wavefield; the best forward models are then plugged into the adjoint-state gradient, replacing the two Helmholtz solves per iteration with one network evaluation. The mixture-of-experts frequency-specific operators, together with the adjoint-gradient formula $-2(\\omega^j)^2 \\lambda u / c^3$, are what make the in vivo inversion tractable.","core_discovery":"The paper's central claim is that a neural surrogate trained purely on synthetic breast anatomy can replace the numerical Helmholtz solver inside full-waveform inversion and still produce clinically useful reconstructions of real patient breasts. Using gradient-based optimization with a Born Fourier Neural Operator (BFNO) as the forward map, the authors reconstruct a malignant tumor with irregular boundaries and a benign cyst with smooth boundaries from clinical USCT data restricted to the dataset's eight frequencies, while direct inversion networks (NIO, InversionNet, DeepONet) fail to recover interior structure. The result is presented as evidence that the OpenBreastUS phantoms capture the scattering statistics of real tissue closely enough for transfer, and that forward operators, which learn only the conditional mapping from medium to wavefield, are more robust than inverse operators that must also learn the anatomical prior.","pith_inferences":["A likely consequence the paper leaves implicit is that the same phantom-generation approach could synthesize training pairs for almost any ring-array USCT geometry, making scanner-specific retraining cheap once a forward surrogate is trusted.","Because the in vivo test used one dataset and two patients, the strongest next test is external validation on a different device and population; positive results there would convert the demonstration into a general method.","The frequency and source mismatch between training (three frequencies, 64 sources) and application (eight frequencies, 256 transmitters) suggests the model transfers more broadly than the benchmark explicitly trains for, and isolating whether that generalization is physics or coincidence would be a clean ablation.","Extending the forward model to include acoustic attenuation, which the paper lists as future work, would likely improve tumor characterization and is a natural test of whether the dataset's realism bottleneck is tissue structure or missing physics."],"forward_implications":["In iterative full-waveform inversion, replacing each pair of Helmholtz solves with a forward neural operator cuts per-gradient cost to a single network pass, making quasi-real-time USCT reconstruction a realistic target instead of a numerical bottleneck.","Forward surrogate models outperform direct inversion networks on realistic tissue because they learn the likelihood rather than the posterior, so they are less prone to memorizing the training anatomy distribution.","OpenBreastUS-trained operators generalize to unseen breast types, unseen source locations, and even to media like Gaussian random fields, while models trained on simplified random media fail on breast tissue.","Higher breast density and higher frequency systematically degrade all baselines, so future neural solvers will need to target dense-tissue scattering and high-wavenumber behavior.","In the two clinical cases, gradient-based optimization with BFNO distinguishes malignant from benign lesions by boundary regularity, indicating a path toward computer-assisted breast-disease screening."],"supporting_citations":[{"why":"Supplies the clinical USCT data from human breasts and the FDFD solver-based reference reconstructions used as ground truth for the in vivo validation.","marker":"[15]"},{"why":"Supplies the virtual breast phantom generation method that produces the 8,000 anatomically realistic 3D breast models sliced into 2D tissue maps.","marker":"[9]"},{"why":"Supplies the Convergent Born Series algorithm used to simulate the 16.4 million ground-truth frequency-domain wavefields.","marker":"[29]"},{"why":"Supplies the Born Fourier Neural Operator architecture that serves as the forward surrogate in the successful in vivo reconstructions.","marker":"[18]"},{"why":"Supplies the gradient-based optimization framework with neural operators that replaces numerical Helmholtz solves in the adjoint method.","marker":"[30]"}],"fun_headline_variants":["16M simulations teach neural solvers to image live tumors","Synthetic-to-real transfer enables in vivo breast imaging","OpenBreastUS: benchmarking neural operators for USCT","First in vivo breast reconstruction with learned solvers","Neural wave solvers trained on phantoms image real breasts"],"cache_read_input_tokens":23552,"weakest_assumption_plain":"The whole transfer result rests on the assumption that synthetic breast anatomies, after hand-set scaling and small random sound-speed perturbations, have the same spatial statistics as real breast tissue; the in vivo validation also assumes models trained on three frequencies and 64 sources apply to eight frequencies and 256 transmitters without explicit retraining.","fun_headline_variants_meta":{"raw":{"variants":["16M simulations teach neural solvers to image live tumors","Synthetic-to-real transfer enables in vivo breast imaging","OpenBreastUS: benchmarking neural operators for USCT","First in vivo breast reconstruction with learned solvers","Neural wave solvers trained on phantoms image real breasts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001088,"raw_usage":{"total_tokens":4536,"prompt_tokens":924,"completion_tokens":3612,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":3532}},"tokens_in":540,"tokens_out":3612,"duration_ms":30641,"temperature":1.0,"reasoning_tokens":3532,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:42:17.485061+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check would be to run the same BFNO-driven FWI pipeline on a held-out clinical USCT dataset from a different scanner, with tumor locations confirmed by biopsy or pathology; if reconstructed lesion boundaries do not match pathology within the expected margin, or if image quality collapses on this second dataset, the claimed transfer is not general.","supporting_citations":[{"cited_title":"2-d slicewise waveform inversion of sound speed and acoustic attenuation for ring array ultrasound tomography based on a block lu solver,","cited_arxiv_id":null,"evidence_quote":"Supplies the clinical USCT data from human breasts and the FDFD solver-based reference reconstructions used as ground truth for the in vivo validation."},{"cited_title":"3-d stochastic numerical breast phantoms for enabling virtual imaging trials of ultrasound com- puted tomography,","cited_arxiv_id":null,"evidence_quote":"Supplies the virtual breast phantom generation method that produces the 8,000 anatomically realistic 3D breast models sliced into 2D tissue maps."},{"cited_title":"3-d stochastic numerical breast phantoms for enabling virtual imaging trials of ultrasound computed tomography,","cited_arxiv_id":null,"evidence_quote":"Supplies the Convergent Born Series algorithm used to simulate the 16.4 million ground-truth frequency-domain wavefields."},{"cited_title":"A convergent born series for solving the inhomogeneous helmholtz equation in arbitrarily large media,","cited_arxiv_id":null,"evidence_quote":"Supplies the gradient-based optimization framework with neural operators that replaces numerical Helmholtz solves in the adjoint method."}],"review_version":1}