{"id":"bb818499-180e-4f4d-89fe-823a35598735","arxiv_id":"2507.01152","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"SonoGym provides parallel, realistic ultrasound simulation for training RL and imitation-learning agents on robotic orthopedic tasks including navigation, reconstruction, and surgery.","lead":"SonoGym is a new simulation platform that generates realistic ultrasound images from CT scans and lets robots practice three surgical tasks: navigating to a target, reconstructing bone surfaces, and guided drilling. It benchmarks reinforcement learning and imitation learning agents, and reports which methods train successfully and where they fail.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ODT evidence for sim-to-real is not out-of-domain: Section 5 varies GAN seeds, not imaging domain; the only true held-out test (Table 6) shows safe ratio 52.86%, so the central transfer claim is unsupported.","rationale":"Reading in good faith, the paper delivers a useful simulation platform with code and data release, and it demonstrates that PPO, A2C, ACT, and DP can learn all three tasks in simulation. The strongest claim, however, is not just that policies train, but that the learning-based simulator's ODT results indicate sim-to-real potential. That inference is undercut by the ODT protocol. Varying random seeds of a GAN trained on the same ex-vivo data (Fig. 14) changes speckle texture but does not change anatomy, probe, tissue properties, or ultrasound physics; it is a within-distribution perturbation. The one test that does cross a real boundary, new patient anatomy, shows a collapse consistent with the reader's weakest assumption. The paper's own limitations list real-robot validation as future work, so this is a missing-evidence problem rather than an internal contradiction; the platform itself remains credible. Therefore the reader's CONDITIONAL verdict is appropriate: platform claims stand, but the sim-to-real generalization claim needs either a true out-of-domain evaluation or removal. This stress-test refines the reader's concern from 'transfer is untested' to 'the provided ODT evidence does not test the claimed transfer,' hence partial agreement.","tokens_in":18111,"tokens_out":6603,"duration_ms":83413,"concrete_test":"Redefine ODT to an actual domain shift: acquire paired CT-US from a held-out in-vivo specimen or use a public US-CT dataset, then evaluate the trained navigation and surgery policies on the real US images under the same closed-loop control loop. If the safe ratio and position errors resemble Table 6 rather than the IDT columns of Table 1, the sim-to-real claim fails. As a cheaper image-level check, compute FID or LPIPS between SonoGym learning-based outputs on TotalSegmentator CT slices and real in-vivo US images; a large gap would invalidate interpreting seed-averaged ODT as domain generalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of sim-to-real potential rests on the 'Out-of-Domain Test' in Section 5, but the paper defines ODT as training on four pix2pix networks and testing on a fifth, all trained on the same paired ex-vivo CT-US dataset with different random seeds. This varies texture noise, not the actual domain gap: the training and test images share the same anatomy distribution, transducer, and acquisition physics. The only genuinely out-of-domain evaluation, surgery on a held-out TotalSegmentator patient with learning-based simulation (Table 6), shows large degradation: safe ratio 52.86% for PPO and side error 27.3 mm versus 5.42 mm in the in-domain test in Table 5. The paper acknowledges cross-patient difficulty but still concludes 'potential of sim-to-real transfer over the ultrasound imaging domain' from the seed-variation ODT results. This is an invalid operationalization of ODT and does not support the platform's usefulness for real-world policy transfer.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents SonoGym, a parallel robotic ultrasound simulation platform that provides both physics-based and learning-based (pix2pix) ultrasound image generation from CT-derived patient models. It proposes three surgical tasks—navigation, bone surface reconstruction, and ultrasound-guided pedicle screw drilling—and formulates them within MDP, submodular MDP, and state-wise constrained MDP frameworks. The authors benchmark PPO, A2C, a SafeRPlan safety-filtered PPO, ACT, and Diffusion Policy across these tasks, reporting learning curves, quantitative task metrics, and timing measurements. The main claims are that the platform enables stable policy learning across tasks, is efficient enough for parallel training, and that generalization experiments suggest the potential of sim-to-real transfer over the ultrasound imaging domain while cross-patient generalization remains challenging.","tokens_in":18357,"tokens_out":7290,"duration_ms":180347,"significance":"SonoGym addresses a genuine gap: existing surgical simulation platforms focus on laparoscopic or soft-tissue manipulation and rarely include patient-specific intraoperative imaging modalities such as ultrasound. The open-source release, expert demonstration datasets, and benchmarks for multiple RL/IL algorithms are valuable community assets. The honest reporting of unsuccessful baselines (SAC, PPO-Lagrangian, Decision Transformer on surgery) strengthens the paper's credibility. The computational efficiency results support the title's 'high performance' claim. If the simulation realism and transfer claims are appropriately qualified, the platform is a useful contribution to robot learning for robotic ultrasound.","major_comments":[{"comment":"The condition labeled 'Out-of-Domain Test (ODT)' for the learning-based ultrasound simulation is not an out-of-domain test as described in Section 5 (Experiment setup): it trains agents on four pix2pix networks and tests on a fifth, where all five networks share the same paired ex-vivo CT-US dataset and differ only by random seed. This varies generator noise/texture, not the imaging domain (anatomy distribution, transducer, acquisition physics). Therefore the conclusions that the gaps between LB_ODT and LB 'demonstrate the potential of sim-to-real transfer over the ultrasound imaging domain' (Fig. 6 caption) and that training with multiple networks 'address the sim-to-real gap between images' (Section 5.2 Q4) are unsupported. The only genuinely held-out domain evaluation, the new-patient surgery test in Table 6, shows large degradation (PPO safe ratio 52.86%, side error 27.3 ± 36.4 mm vs. 5.42 ± 4.9 mm in the in-domain result in Table 5). The authors should re-label this condition (e.g., 'generator randomization') or, ideally, evaluate on real ultrasound images or a distinct data distribution before claiming evidence for sim-to-real transfer.","section":"Section 5.2 Q4; Fig. 6; Table 1"},{"comment":"The quantitative evaluation of ultrasound realism reports LPIPS 0.2415, SSIM 0.3940, and PSNR 15.96 as point estimates without standard deviations, without variance across the five trained networks, and without a quantitative comparison to the model-based simulator or to existing baselines such as [8], despite the claim that the values are 'close' to [8]. Because the realism of the learning-based simulation is a load-bearing premise for the transfer claims, the authors should report variance across networks and include a quantitative comparison with the model-based approach and prior ultrasound simulation methods.","section":"Section 5.1, Q1"}],"minor_comments":[{"comment":"The arrow in the table header 'safe ratio ↓[%]' is inverted; higher safe ratio is better, so it should be '↑[%]'.","section":"Table 1"},{"comment":"The platform is consistently called 'SonoGym' except for the spelling 'Sonogym' in the abstract; please unify the spelling throughout.","section":"Abstract"},{"comment":"The rotation error term in the navigation reward appears without the weight w1, although the text states that w1 balances position error (in mm) and rotation error (in rad); please clarify whether w1 multiplies both terms.","section":"Section 4.1, reward equation"},{"comment":"The reported timings of 0.0089 s and 0.1107 s for 100 environments do not specify whether the measurement includes slice extraction plus rendering or only the ultrasound renderer, nor the GPU batch size; please state the measurement boundary.","section":"Section 5.1, timing results"},{"comment":"The caption states 'comparable performance between ODT and IDT for LB' without significance tests or standard deviations; please provide error bars or statistical tests to support this claim, since the table otherwise omits standard deviations.","section":"Table 1, caption"},{"comment":"The abstract refers to 'vision transformers' for IL agents, but the paper trains ACT and Diffusion Policy; ACT uses a transformer within a CVAE but is not a vision transformer in the standard sense, so more precise terminology is recommended.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid systems contribution with clear value to the community, and the major issues are re-labeling and claim-limiting rather than fundamental flaws in the platform. I recommend major revision; with the ODT claim corrected and realism metrics properly quantified, the paper could be acceptable. The authors should also ensure the 'modified SafeRPlan' is described in enough detail for reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it ships: code, assets, expert datasets, and multiple robot/patient models are public, and the core claim that you can train DRL and IL agents in a GPU-parallel simulator across three ultrasound-guided orthopedic tasks is backed by learning curves and tables. Second, the \"out-of-domain\" generalization evidence is weaker than the label suggests, and the paper's sim-to-real conclusion rests on that weak evidence.\n\nWhat is actually new is the integration: physics-based (ray-traced) and learning-based (pix2pix) ultrasound synthesis in one parallel platform, with three tasks—navigation, bone surface reconstruction, and ultrasound-guided drilling—and benchmarks for PPO/A2C, submodular RL, safe RL, ACT, and diffusion policy. No prior work cited combines all of this; the reconstruction task with submodular rewards and the surgery task with a safety filter are new to the robotic-ultrasound literature. The released code and data are a real contribution and should be credited.\n\nThe main soft spot is the Out-of-Domain Test in Section 5. What they call ODT is training agents on four pix2pix networks and testing on a fifth, all trained on the same paired ex-vivo CT-US dataset. That varies GAN seeds, hence texture noise; it does not change anatomy, transducer, or acquisition physics. The only genuinely held-out evaluation—surgery on a new TotalSegmentator patient (Table 6)—drops safe ratio to 52.86% for PPO and raises side error from 5.42 mm in-domain to 27.3 mm. The authors acknowledge cross-patient difficulty but still say the seed-variation results show \"potential of sim-to-real transfer over the ultrasound imaging domain.\" That is an overclaim. The realism metrics are also thin: SSIM 0.394, PSNR 15.96, no variance or direct comparator, and the \"not significant\" claim for LB_ODT vs LB has no statistical test behind it. These issues are fixable.\n\nI also note the GAN is trained on seven ex-vivo spine specimens and transferred to in-vivo CT volumes; the paper honestly lists real-robot validation as future work. So the platform claim is credible; the transfer claim is not yet established.\n\nWho gets value: researchers working on robotic ultrasound planning or surgical robot learning who need a fast training environment with realistic image feedback. It deserves a serious referee because the shipped environment and reproducible benchmarks are genuinely useful. I would accept with revision: rename or reframe ODT as network-randomization robustness, add statistics for the realism metrics and significance claims, and temper the sim-to-real sentence in the abstract and conclusion.\n\nRecommendation: send it to peer review. It will be a solid contribution after those revisions.","headline":"Genuinely useful open platform for robotic ultrasound research, but the 'out-of-domain' test varies only GAN seeds and the one true held-out patient test shows large drops, so the sim-to-real conclusion is overclaimed.","tokens_in":18940,"tokens_out":2940,"would_cite":true,"duration_ms":28537,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SonoGym is a scalable robotic-ultrasound simulator that trains deep RL and imitation-learning policies for navigation, surface reconstruction, and ultrasound-guided spine drilling.","keywords":["robotic ultrasound","simulation platform","deep reinforcement learning","imitation learning","ultrasound image simulation","spine surgery","pedicle screw placement","submodular reinforcement learning"],"falsifier":"Run a policy trained in SonoGym on a real robotic ultrasound system with a cadaver or volunteer and measure the same metrics used in the paper, such as navigation position and rotation error or the surgery safe ratio; if the safe ratio falls near the paper's own held-out-patient value of 52.86% or navigation error exceeds the reported band of 16 mm, the central sim-to-real premise fails. A cheaper first check is to compute LPIPS or SSIM between SonoGym's generated images and real ultrasound images from a specimen not used in training: if the generative outputs are no closer to real ultrasound than the model-based outputs are, the learning-based branch loses its claimed advantage.","tokens_in":1933,"feed_emoji":"🤖","tokens_out":4782,"duration_ms":130338,"temperature":0.7,"pith_summary":"Robotic ultrasound could make spinal procedures more reproducible, but deep reinforcement learning and imitation learning have not been used much there because no fast, realistic simulator existed. SonoGym attacks that gap: it renders ultrasound images from CT-derived patient models two ways, with a physics-based ray-tracing model and with a generative network, and it runs tens to hundreds of environments in parallel so agents can collect experience quickly. The paper's central claim is that with this combination, learning-based agents actually train well: PPO reaches close-to-expert performance on navigation and surgery, and submodular PPO and A2C beat a heuristic scanning path on bone-surface reconstruction. The authors also quantify where the approach is not yet reliable, especially when a policy trained on several patients is tested on a new patient, where the safe ratio in the drilling task drops to 52.86%. That mix of demonstrated training success and honest remaining gaps is the contribution: a platform that lets the community work on the hard parts instead of rebuilding simulators.","feed_headline":"Simulator trains robotic-ultrasound agents for spine surgery in hours","feed_subtitle":"PPO, ACT, and diffusion policies train on navigation, reconstruction, and drilling in dozens of parallel simulated environments.","key_machinery":"The load-bearing mechanism is the ultrasound image-rendering pipeline: for each of dozens of parallel environments, the current end-effector pose defines an ultrasound image plane inside the patient coordinate frame, and the simulation slices the 3D CT volume and its segmentation map along that plane to produce a 2D CT slice and a label slice. From those inputs, either the model-based branch computes reflection and backscattering with a convolutional ray-tracing model, with acoustic impedance set from CT intensity, or the learning-based branch feeds the CT slice to a generative image-translation network. This batch pipeline renders a 200 by 150 image for 100 environments in 0.0089 seconds for the model-based branch and 0.1107 seconds for the learning-based branch on an RTX 3090 Ti, which is what makes PPO training feasible in about 2 hours for model-based and 10 hours for learning-based simulation. The task-specific MDP formulations are the second piece: a partially observable MDP for navigation, a submodular MDP with marginal-gain rewards for reconstruction, and a state-wise constrained MDP with a safety filter for surgery.","core_discovery":"SonoGym's thesis is that the missing piece for robot learning in robotic ultrasound is not the learning algorithm but the training environment. The paper shows that with real-time parallel ultrasound simulation—both a convolutional ray-tracing model and a generative network trained on paired CT-ultrasound data from seven ex-vivo spine specimens—PPO agents train stably and achieve close-to-expert performance, outperforming A2C on navigation and surgery, while submodular PPO and A2C surpass the heuristic open-loop trajectory used in prior reconstruction work. The environment encodes the three tasks as specialized MDPs: navigation as a partially observable MDP with ultrasound images as observations, reconstruction as a submodular MDP whose reward is the marginal gain in covered bone-surface area, and surgery as a state-wise constrained MDP with an unsafe-region cost. On generalization, the paper claims that training against several ultrasound noise networks keeps navigation errors within acceptable ranges (below 16 mm in position and 12 degrees in rotation) and keeps surgical performance roughly level, while generalization to a held-out patient remains a clear failure mode, with the safe ratio falling to 52.86%.","pith_inferences":["If the out-of-domain imaging results carry to a real machine, sim-to-real transfer could be achieved by training against several simulated ultrasound styles rather than collecting real ultrasound data; the paper does not make this claim.","Adding anatomical variability to the training patients may improve held-out-patient performance more than improving image fidelity, because the sharpest reported failure is inter-patient, not inter-noise, generalization.","A testable extension is to train on both model-based and learning-based images at once, randomizing over simulation branches, which could improve generalization beyond either branch alone.","We infer that the submodular-MDP coverage formulation applies beyond ultrasound, for example to laparoscopic surface scanning, where coverage and path length compete in the same way."],"forward_implications":["PPO agents trained from ultrasound image observations can reach near-expert performance on navigation and surgery, so image-based policies are viable when the true anatomical pose is unknown.","Submodular rewards based on the marginal gain in covered bone-surface area let reconstruction agents beat the heuristic open-loop scanning path, producing higher coverage with lower rotation and path length.","Training with several ultrasound-generator models keeps navigation and surgery performance roughly in-domain, which the authors read as evidence that the imaging-domain sim-to-real gap is addressable.","Imitation-learned ACT and diffusion policies train successfully on navigation and ACT on surgery, but PPO with a safety filter is more safety-aware in the drilling task, while ACT achieves better insertion accuracy, so the two families trade off.","Inter-patient generalization remains unsolved: on a held-out sixth patient, PPO's safe ratio drops to 52.86% and side error rises to 27.3 mm, so patient diversity in the training set is a key bottleneck."],"supporting_citations":[{"why":"Supplies the convolutional ray-tracing model that produces the model-based ultrasound images.","marker":"[38]"},{"why":"Provides the CT-intensity-based acoustic impedance and reflection treatment that refines the model-based simulation.","marker":"[21]"},{"why":"Is the generative adversarial image-translation pipeline used to learn the CT-slice-to-ultrasound translation.","marker":"[15]"},{"why":"Provides the paired CT-ultrasound dataset from seven ex-vivo spine specimens used to train the learning-based simulator.","marker":"[53]"},{"why":"Supplies the CT volumes and segmentations from which the patient models in the simulator are derived.","marker":"[52]"},{"why":"Introduces submodular MDPs whose marginal-gain rewards define the reconstruction task.","marker":"[37]"},{"why":"Defines state-wise constrained MDPs used to formulate the safety constraint in the surgery task.","marker":"[57]"},{"why":"Provides the heuristic open-loop scanning path and reconstruction setup that the learned reconstruction policies are compared against.","marker":"[24]"},{"why":"Contributes the Action Chunking Transformer used as one of the imitation-learning baselines.","marker":"[56]"},{"why":"Contributes the Diffusion Policy used as one of the imitation-learning baselines.","marker":"[6]"}],"fun_headline_variants":["Simulator trains robotic-ultrasound agents for surgery in hours","Parallel ultrasound sim enables fast robot learning for spine surgery","SonoGym: real-time robotic ultrasound training for complex tasks","Simulating ultrasound for robot-guided orthopedic surgery in parallel","Rapidly train robotic ultrasound agents for navigation and drilling"],"cache_read_input_tokens":20992,"weakest_assumption_plain":"The load-bearing premise is that ultrasound images generated from CT volumes—in particular those from a generative network trained on seven ex-vivo spine specimens—are realistic enough that policies trained in SonoGym would behave the same way on real ultrasound and real patients, an assumption the paper does not test outside simulation.","fun_headline_variants_meta":{"raw":{"variants":["Simulator trains robotic-ultrasound agents for surgery in hours","Parallel ultrasound sim enables fast robot learning for spine surgery","SonoGym: real-time robotic ultrasound training for complex tasks","Simulating ultrasound for robot-guided orthopedic surgery in parallel","Rapidly train robotic ultrasound agents for navigation and drilling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1347,"prompt_tokens":1044,"completion_tokens":303,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":660,"completion_tokens_details":{"reasoning_tokens":220}},"tokens_in":660,"tokens_out":303,"duration_ms":4415,"temperature":1.0,"reasoning_tokens":220,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:58:30.466151+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a policy trained in SonoGym on a real robotic ultrasound system with a cadaver or volunteer and measure the same metrics used in the paper, such as navigation position and rotation error or the surgery safe ratio; if the safe ratio falls near the paper's own held-out-patient value of 52.86% or navigation error exceeds the reported band of 16 mm, the central sim-to-real premise fails. A cheaper first check is to compute LPIPS or SSIM between SonoGym's generated images and real ultrasound images from a specimen not used in training: if the generative outputs are no closer to real ultrasound than the model-based outputs are, the learning-based branch loses its claimed advantage.","supporting_citations":[{"cited_title":"Patient-specific 3d ultrasound simulation based on convolutional ray-tracing and appearance optimization","cited_arxiv_id":null,"evidence_quote":"Supplies the convolutional ray-tracing model that produces the model-based ultrasound images."},{"cited_title":"Visualization and gpu-accelerated simulation of medical ultrasound from ct images","cited_arxiv_id":null,"evidence_quote":"Provides the CT-intensity-based acoustic impedance and reflection treatment that refines the model-based simulation."},{"cited_title":"UltraBones100k: A reliable automated labeling method and large-scale dataset for ultrasound-based bone surface extraction","cited_arxiv_id":"2502.03783","evidence_quote":"Provides the paired CT-ultrasound dataset from seven ex-vivo spine specimens used to train the learning-based simulator."},{"cited_title":"Totalseg- mentator: robust segmentation of 104 anatomic structures in ct images","cited_arxiv_id":null,"evidence_quote":"Supplies the CT volumes and segmentations from which the patient models in the simulator are derived."},{"cited_title":"Submodular rein- forcement learning","cited_arxiv_id":null,"evidence_quote":"Introduces submodular MDPs whose marginal-gain rewards define the reconstruction task."},{"cited_title":"State-wise constrained policy optimization, 2024","cited_arxiv_id":null,"evidence_quote":"Defines state-wise constrained MDPs used to formulate the safety constraint in the surgery task."},{"cited_title":"Robot-assisted ultrasound reconstruction for spine surgery: from bench-top to pre-clinical study","cited_arxiv_id":null,"evidence_quote":"Provides the heuristic open-loop scanning path and reconstruction setup that the learned reconstruction policies are compared against."}],"review_version":1}