{"id":"a629f7d6-a49a-4040-adcb-3b9b7b70c571","arxiv_id":"2502.03360","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A 3D MedNeXt network predicts all 180 VMAT fluence maps from beam's-eye-view projections of a 3D dose map, with DVHs close to the target dose and inference under 20 ms.","lead":"Researchers trained a 3D neural network to turn a 3D radiation dose map into all 180 beam intensity patterns, or fluence maps, needed for a full VMAT radiotherapy arc in a single pass under 20 milliseconds. If the results hold, this could make prostate cancer treatment planning much faster and could later extend to other tumor sites.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central DVH validation skips leaf sequencing: predicted fluence maps are never converted to deliverable MLC segments, so the ultra-fast VMAT planning claim is unproven.","rationale":"The reader's CONDITIONAL verdict is appropriate, and the weakest assumption about the forward model and the well-posedness of the inverse mapping is real. However, the single most load-bearing gap is more specific and even more decisive: the predicted quantity is an idealized fluence map, not a deliverable MLC sequence, and the paper never integrates leaf sequencing into its validation. The title, abstract, and conclusion promise 'ultra fast VMAT planning' and 'small DVH difference validates the approach,' yet no experiment shows that the predicted fluence maps can be realized by the MLC and still produce the claimed dose distribution. This is an omission in the argument, not an internal inconsistency in the architecture: the PSNR/SSIM improvements over U-Net baselines are plausible, the dataset scaling result is interesting, and the BEV preprocessing is a reasonable contribution. But the clinical conclusion depends on the unverified leap from predicted fluence maps to a deliverable plan. A conditional acceptance requiring the deliverability validation, or an explicit statement that the network predicts MLC positions directly, would be consistent with the evidence. Thus the reader's verdict stands unchanged.","tokens_in":8305,"tokens_out":5562,"duration_ms":54171,"concrete_test":"Run the prediction-to-delivery loop on the validation set: convert each predicted fluence map to MLC leaf positions and MUs using a Varian-compatible leaf-sequencing algorithm (or the authors' inverse model), compute dose with Acuros AXB, and compare target versus predicted DVHs plus dose metrics (D98, Dmean, D2%, V95%) with confidence intervals. If the predicted fluence maps are already intended as deliverable MLC positions, provide the exact conversion and validate it. If after sequencing the DVH differences exceed clinical tolerances (e.g., more than 2% in D98), the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that a single 20 ms network pass can produce VMAT fluence maps whose DVHs are 'very close' to the target, enabling ultra-fast inverse planning. This claim rests on a validation chain that is not described and not demonstrated. The network outputs 2D fluence maps, but VMAT delivery is defined by MLC leaf positions and MUs at control points; a fluence map is an intermediate, non-unique representation. The paper never runs the predicted fluence maps through leaf sequencing, nor does it specify how the 'Acuros AXB dose calculation function' was applied to the predicted fluence maps. Acuros AXB is a clinical dose engine for beams defined by MLC apertures and weights; converting arbitrary predicted fluence maps to such apertures requires an additional algorithm (the inverse of the model used to create targets) that is absent. If the DVHs in Fig. 4 were computed directly from idealized fluence maps, they bypass the deliverability constraints of the MLC (maximum leaf speed, tongue-and-groove, interdigitation), and the comparison is not clinically meaningful. The reported PSNR/SSIM are on the same idealized fluence domain, so they do not close this gap. The forward model used to build targets ignores tongue-and-groove and may not match the Eclipse/Acuros dose used as input, making the training targets and the input dose mutually inconsistent. The missing MAE in Gy and the absence of DVH metrics or error bars further prevent assessing whether the 'close DVH' claim is real. Therefore the 'ultra-fast VMAT planning' conclusion is not supported by the experiments as presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a deep learning method to predict all 180 fluence maps of a single-arc VMAT plan from the 3D dose distribution in one network pass. The dose map is preprocessed into beam's eye view projections at each control point and fed to a 3D MedNeXt architecture. The authors generate additional VMAT plans using Eclipse to enlarge the dataset from 117 to 2266 prostate plans, train with L1+L2 losses, and report PSNR/SSIM improvements over 2D and 3D U-Nets. They also claim that DVHs calculated from the predicted fluence maps are very close to the target DVHs. The central claims are the architectural contribution (BEV transform + 3D convolutions) and dataset scaling, with a downstream goal of ultra-fast VMAT planning.","tokens_in":8581,"tokens_out":6993,"duration_ms":55738,"significance":"If the dosimetric validation were rigorous, the work would be a useful contribution to automated VMAT planning by showing that fluence maps can be predicted in a single fast forward pass and that dataset size is a key driver of accuracy. The paper provides a clear comparison against 2D and 3D U-Net baselines and consistently reports PSNR gains; the large-scale Eclipse-generated dataset is a practical resource. However, the current manuscript does not provide quantitative DVH metrics or a deliverability assessment, so the clinical significance of the reported gains is not established.","major_comments":[{"comment":"The manuscript states that MAE in Gy and DVHs were computed for the best method, yet the Results section reports no MAE value and no quantitative DVH statistics (e.g., D95, V95, mean or max dose differences, gamma pass rates). The only evidence for the central claim that 'the resulting DVHs are very close' is four qualitative examples in Fig. 4 and a statement that all DVHs looked similar. This does not substantiate the claim. Please report the numerical MAE/Gy and DVH metrics across the validation set, ideally with confidence intervals.","section":"II.4 / Results"},{"comment":"The dose calculation from the predicted fluence maps is not described. Acuros AXB takes MLC leaf positions and MU weights as input, not arbitrary 2D fluence maps; therefore a leaf-sequencing or aperture conversion step is required. The manuscript neither specifies the conversion algorithm nor states whether deliverability constraints (leaf speed, tongue-and-groove, interdigitation) were respected. If the DVHs were computed directly from idealized fluence maps, they bypass the very constraints that define VMAT delivery, and the comparison is not clinically meaningful. This is a load-bearing gap for the ultra-fast planning claim.","section":"II.4 / Fig. 4"},{"comment":"The training targets are fluence maps computed with a forward model that includes leaf leakage and inter-control-point motion but ignores tongue-and-groove, while the input dose is computed by Eclipse/Acuros from the MLC positions. The manuscript does not verify that the simplified fluence-to-dose forward model is consistent with the Eclipse/Acuros dose used as input. A large discrepancy would make the supervised task ill-posed (the network would be asked to map a dose that cannot be exactly reproduced by the target fluence representation). Please report a gamma analysis between the dose computed from the target fluence maps and the original Eclipse dose for the training targets.","section":"II.1"},{"comment":"The results are reported as single point estimates with no error bars, no multi-seed runs, and no statistical significance tests. This makes it impossible to judge whether the smaller differences (e.g., 3D MedNeXt trained on Ecl. 500 vs Ecl. full: 26.01 vs 28.72 dB on Ecl. 500) are meaningful. Additionally, the SSIM values for the 3D U-Net (e.g., 0.2664 in Table 3) appear anomalously low relative to its PSNR (~22 dB) compared with the 2D U-Net (SSIM ~0.89 at similar PSNR), which suggests a possible issue in the SSIM computation or in the 3D U-Net training; please verify.","section":"II.4 / Tables 2-3"}],"minor_comments":[{"comment":"The number of generated plans is given as 'around 2266' in the text and 2266 in Table 1; please use exact numbers consistently.","section":"II.1"},{"comment":"The statement 'we are using for the first time a 3D convolutional network architecture' conflicts with the later mention that previous studies (refs 10, 11) used a 3D U-Net to predict fluence maps. Please clarify what is new relative to those works.","section":"II"},{"comment":"There are several typographical errors: 'the propose 3D network', 'fund that they all looked similar', and 'litterature' should be corrected.","section":"III"},{"comment":"The training details lack the number of epochs, learning rate schedule, and data augmentation, which are needed for reproducibility.","section":"II.4"},{"comment":"The DVH figure is not labeled with the structures; it is unclear which curves correspond to PTV or OARs. Please add a legend or caption explanation.","section":"Fig. 4"},{"comment":"The inference time of 20 ms is reported without specifying the GPU/CPU hardware, input resolution, and batch size; please provide these for reproducibility.","section":"III"}],"recommendation":"major_revision","confidential_remarks":"The paper relies heavily on refs 6, 17, and 22 from the same group for the dose prediction, data generation, and leaf sequencing modules; the current manuscript is essentially a module paper for a larger pipeline. The editor may wish to consider whether the journal's scope requires end-to-end validation. Also, the missing quantitative DVH metrics and the vague 'around 2266 plans' phrasing give an impression of incompleteness that a revision should address."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a legitimate extension of prior VMAT fluence prediction work. Two genuine contributions: the BEV projection that turns gantry rotation into a translation, and the Eclipse-based dataset scaling from about 100 to over 2200 prostate plans. The MedNeXt backbone shows clear PSNR/SSIM gains over 2D and 3D U-Nets, and the ablation showing ~4.8 dB gain from dataset size is believable. The paper does what it claims on the fluence-map reconstruction task.\n\nThe soft spot is the validation chain. The Methods say they computed dose with Acuros AXB and generated DVHs, but they never report the MAE in Gy, and they never run the predicted fluence maps through leaf sequencing to get deliverable MLC segments. Acuros expects MLC apertures, not arbitrary 2D fluence maps, so it is unclear how the DVHs in Fig. 4 were actually produced. If those DVHs come from idealized fluence, they bypass leaf speed, tongue-and-groove, and interdigitation constraints, so 'very close' DVHs may not survive delivery. The forward model used to build targets ignores tongue-and-groove, and the paper does not discuss whether the input 3D dose is consistent with that forward model. That said, the Discussion is careful to call the network a module within a larger framework, not a complete planner, so the overclaim is modest but real.\n\nAlso no error bars, no multi-seed runs, no statistical tests. The PSNR/SSIM tables are plausible but would be stronger with variance estimates. No code or data released, and the self-citations for the generation pipeline (ref 17) and leaf sequencing (ref 22) are heavy, but those are the relevant prior works from the same group, so I do not count that against the science.\n\nThis paper is for researchers working on deep learning for radiotherapy planning, especially inverse mapping from dose to fluence or MLC parameters. It deserves a serious referee: the idea is sound, the dataset construction is useful, and the architecture comparison is clearly presented. The revision needs to either provide the missing dose metrics and a deliverability check or soften the clinical framing.\n\nRecommendation: send to peer review, but expect major revision on the validation and claims.","headline":"Good architecture study for fast VMAT fluence prediction, with a real dataset-scaling contribution, but the dosimetric validation skips the deliverability step, so the 'ultra-fast planning' claim is not yet proven.","tokens_in":9186,"tokens_out":1944,"would_cite":true,"duration_ms":17710,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single 3D neural network can reconstruct all 180 fluence maps of a one-arc VMAT plan directly from a 3D dose map in under 20 milliseconds, with dose-volume histograms close to those of the input plan.","keywords":["VMAT","fluence map prediction","beam's eye view","3D convolutional network","MedNeXt","inverse treatment planning","dose-volume histogram","machine learning radiotherapy"],"falsifier":"Recompute the dose delivered by the predicted fluence maps through a full leaf-sequencing and delivery simulation that includes tongue-and-groove and real MLC speed and acceleration constraints, then compare per-voxel dose and DVHs with the original Eclipse plan on an independent validation set; a clinically meaningful disagreement would falsify the claim that the predicted maps are directly deliverable and dose-equivalent.","tokens_in":8099,"feed_emoji":"🎯","tokens_out":9266,"duration_ms":81026,"temperature":0.7,"pith_summary":"This paper argues that the time-consuming iterative optimization step in VMAT treatment planning can be replaced by one feedforward neural network. The network takes a 3D dose distribution, projects it into beam's-eye views of the 180 control points of a single gantry arc, and predicts all 180 fluence maps simultaneously in less than 20 ms. It is trained in a supervised way, with Eclipse-generated prostate plans as targets and a mixture of L1 and L2 losses. On validation data the predicted fluence maps yield dose-volume histograms very close to those of the target dose, and the authors conclude that this makes ultra-fast inverse planning feasible, either as a standalone module or as a warm start for an iterative optimizer.","feed_headline":"Beam's-eye 3D network outputs a full VMAT arc's fluence maps in 20 ms","feed_subtitle":"It turns a 3D dose map into 180 fluence maps in one pass, with DVHs close to the target plan.","key_machinery":"The load-bearing machinery is the beam's-eye-view (BEV) transform: the 3D dose is projected into the coordinate frame of each of the 180 control points, producing 180 2D dose projections that are stacked into a 3D tensor whose depth axis is the gantry angle. A 3D MedNeXt encoder-decoder with ConvNeXt residual blocks and circular padding over the gantry dimension processes this tensor and regresses all 180 fluence maps at once. The fluence-map targets are computed from Eclipse-optimized MLC positions and monitor units with a forward model that includes leaf leakage and inter-control-point motion but omits tongue-and-groove, and the training loss combines L1 and L2 terms. The depth-direction convolutions let the network exploit translation equivariance along the rotation axis, while the ConvNeXt blocks supply the long-range mixing needed because the total dose is a sum over all control points.","core_discovery":"The paper's central claim is that the inverse mapping from a planned 3D dose to the full set of per-control-point fluence maps of a single-arc VMAT plan can be learned and executed in one forward pass. The evidence is a 3D MedNeXt network that ingests 180 beam's-eye-view dose projections stacked along the gantry-angle dimension and jointly regresses the 180 corresponding fluence maps. On the validation split of 2,266 Eclipse-generated prostate plans the method reaches 28.71 dB PSNR and 0.9712 SSIM, and dose recomputed from the predicted fluence maps with the Acuros dose engine produces DVHs that closely track the target plan's DVHs. The authors take this as support for using the network as a module in ultra-fast inverse planning or as an initialization for an iterative VMAT optimizer.","pith_inferences":["The dose-to-fluence inverse is likely non-unique, so DVH closeness does not by itself prove the predicted MLC sequence is the deliverable sequence the Eclipse plan would execute; a delivery-simulation or QA test would settle machine compatibility.","The BEV-stacked representation turns gantry rotation into translation along the depth axis, a transformation that should transfer to multi-arc VMAT, IMRT, or other rotating delivery modalities if training data with the same control-point structure are available.","The reported lack of benefit from adding CT or contours as extra input channels suggests the 3D dose map alone carries most of the planning information this network can use, at least for prostate cases and this dataset size.","Because all plans are Varian single-arc prostate, the model's machine specificity is an open question: testing on other linac MLC models, collimator angles, or anatomical sites would reveal how much retraining is needed."],"forward_implications":["One forward pass, under 20 ms excluding data loading, produces all 180 fluence maps of a single-arc VMAT plan from a dose map, so inverse planning no longer needs an iterative leaf-sequencing loop for these cases.","Because the training targets encode MLC constraints from real plans, the predicted maps should be rapidly leaf-sequenceable, making the network a plug-in module in an end-to-end planning pipeline.","Dataset size is a first-order driver: tripling the training set from about 100 to 400 plans adds 2.1 dB PSNR, and going from 400 to about 1,900 adds another 2.7 dB, so collecting more clinically planned cases is a direct route to better predictions.","The same network can initialize an established VMAT optimizer such as Varian's Photon Optimization algorithm, potentially reducing the iterations needed to reach a clinical plan."],"supporting_citations":[{"why":"Defines VMAT optimization and the fluence-map/leaf-sequencing loop that this work proposes to shortcut.","marker":"[1]"},{"why":"Prior deep-learning inverse mapping from dose to fluence maps that this work extends from per-map 2D prediction to joint 3D prediction.","marker":"[10]"},{"why":"Closest precedent: 3D dose-driven automatic VMAT machine-parameter generation that this method builds on and compares with.","marker":"[12]"},{"why":"Introduces the ConvNeXt block the 3D network uses for long-range feature mixing.","marker":"[13]"},{"why":"Provides the 3D U-Net baseline architecture used in the ablation study.","marker":"[14]"},{"why":"Supplies the original prostate VMAT plans and the seed data later replanned with Eclipse.","marker":"[16]"},{"why":"Describes the Eclipse Scripting API replanning pipeline used to scale the dataset from 117 to 2,266 plans.","marker":"[17]"},{"why":"Defines the MedNeXt architecture that the proposed 3D network is built on.","marker":"[18]"},{"why":"Provides the 2D U-Net baseline that processes the 180 control points as input and output channels.","marker":"[21]"},{"why":"Documents the Eclipse dose calculation and Photon Optimizer algorithms used for DVH validation and as the proposed initialization target.","marker":"[23]"}],"fun_headline_variants":["3D network maps dose to 180 VMAT fluence maps in 20 ms","Beam's-eye view AI predicts full VMAT arc fluence maps in 20 ms","Single-pass deep learning speeds VMAT planning to 20 ms","Network from dose to fluence maps in under 20 ms for VMAT"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the fluence maps computed from the Eclipse-optimized MLC positions and monitor units, via a forward model that includes leaf leakage and inter-control-point motion but ignores tongue-and-groove, together with the Eclipse/Acuros dose engine, accurately represent a deliverable VMAT plan and that the input 3D dose map is exactly the dose produced by those same fluence maps.","fun_headline_variants_meta":{"raw":{"variants":["3D network maps dose to 180 VMAT fluence maps in 20 ms","Beam's-eye view AI predicts full VMAT arc fluence maps in 20 ms","Single-pass deep learning speeds VMAT planning to 20 ms","Network from dose to fluence maps in under 20 ms for VMAT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1385,"prompt_tokens":1099,"completion_tokens":286,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":715,"completion_tokens_details":{"reasoning_tokens":201}},"tokens_in":715,"tokens_out":286,"duration_ms":3288,"temperature":1.0,"reasoning_tokens":201,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T04:58:17.259385+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the dose delivered by the predicted fluence maps through a full leaf-sequencing and delivery simulation that includes tongue-and-groove and real MLC speed and acceleration constraints, then compare per-voxel dose and DVHs with the original Eclipse plan on an independent validation set; a clinically meaningful disagreement would falsify the claim that the predicted maps are directly deliverable and dose-equivalent.","supporting_citations":[{"cited_title":"Otto, Volumetric modulated arc therapy: IMRT in a single gantry arc, Medical physics 35 , 310--317 (2008)","cited_arxiv_id":null,"evidence_quote":"Defines VMAT optimization and the fluence-map/leaf-sequencing loop that this work proposes to shortcut."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior deep-learning inverse mapping from dose to fluence maps that this work extends from per-map 2D prediction to joint 3D prediction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Closest precedent: 3D dose-driven automatic VMAT machine-parameter generation that this method builds on and compares with."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the ConvNeXt block the 3D network uses for long-range feature mixing."},{"cited_title":"C i c ek, A","cited_arxiv_id":null,"evidence_quote":"Provides the 3D U-Net baseline architecture used in the ablation study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the original prostate VMAT plans and the seed data later replanned with Eclipse."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the MedNeXt architecture that the proposed 3D network is built on."},{"cited_title":"Ronneberger, P","cited_arxiv_id":null,"evidence_quote":"Provides the 2D U-Net baseline that processes the 180 control points as input and output channels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the Eclipse dose calculation and Photon Optimizer algorithms used for DVH validation and as the proposed initialization target."}],"review_version":1}