{"id":"35ec4e46-ec2b-434a-89c2-5ce4955a3434","arxiv_id":"2504.13240","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The authors defend the topological gap protocol's false discovery rate estimate against two recent comments, adding robustness checks and disclosing a bias-range cropping bug in the published supplementary information.","lead":"This paper is Microsoft Quantum's point-by-point rebuttal of two external comments that criticized the statistical test used to claim topological superconductivity in their nanowire devices. The authors stand by their earlier papers and argue that the critics did not find any real flaw in the false discovery rate estimate, while also disclosing a data-processing bug in one of the original papers.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The FDR transfer hinges on an untested distributional match between simulations and experiments; the rebuttal's robustness tables do not validate it, so the 'no flaws' claim is not established.","rationale":"The reader's verdict is CONDITIONAL, and the weakest assumption identified is exactly the distributional match. I agree. The paper deserves credit for releasing code and simulated data, for the new robustness tables, and for explicitly addressing each point. However, the FDR is an estimate computed from simulations; its validity for real devices depends on the simulated data being representative of actual devices. The paper's response to Ref. 3's range-sensitivity point is an in-model stress test: Table II and Fig. 3 show that within the simulation ensemble, cropping the field range by 0.5 T changes the FDR by about one percentage point (e.g., from <7.9% to <9.0% for SLG-β at n2D,int=4.0). That is evidence about the protocol's sensitivity to hyperparameters, not evidence that the simulation ensemble matches experiment. Likewise, the bug admitted in the published SI for Ref. 2 changes TGP outcomes for device B (new SOI2, larger ROI2); the paper argues this does not affect the FDR estimate, which is correct because the FDR is a simulation-based quantity. But the bug does show that 'no flaws' must be read narrowly. The single most load-bearing condition is the untested distributional match; a concrete two-sample test or a re-fitting of simulation parameters to experimental data would settle whether the FDR bound transfers. Since this condition is not currently demonstrated, the conditional verdict is appropriate; my concern does not change the reader's verdict.","tokens_in":13437,"tokens_out":9827,"duration_ms":85230,"concrete_test":"Train a classifier on TGP-relevant features (extracted transport gap values, zero-bias peak heights and widths, nonlocal conductance magnitudes, and their dependence on B and Vp) using the simulated datasets in the public repository [14] and the experimental datasets from Ref. 2 (available on Zenodo). If the classifier can distinguish simulation from experiment with accuracy significantly above chance (e.g., cross-validated AUC > 0.7), the distributional-match assumption fails and the FDR transfer is unsupported. A null result would strengthen the claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that no flaws have been identified in the FDR estimate and that the external objections are unfounded—rests on the assumption, quoted from Ref. 1, that 'the simulated data is drawn from the same probability distribution as the data produced by real devices.' The new robustness analysis (Tables I and II, Fig. 3) varies parameters inside the simulation model; it does not test whether the simulation's joint distribution of conductance features matches the experimental devices. The cited localization-length measurements and device characterization constrain disorder and normal-state transport, but they do not establish that TGP-relevant features—zero-bias peak statistics, nonlocal gap shapes, Andreev-enhanced subgap conductance, coupling asymmetries—are drawn from the same distribution. If the model is misspecified in any such dimension, the FDR bound of <8% could be biased low, and the claim that the objections are 'unfounded' would not follow. This is the least secure link in the argument, and the paper provides no quantitative test of it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a point-by-point rebuttal by the Microsoft Quantum collaboration to two comments by H. F. Legg (arXiv:2502.19560 and arXiv:2503.08944) on the topological gap protocol (TGP) as used in Phys. Rev. B 107, 245423 (2023) and in the Supplementary Information of Nature 638, 651-655 (2025). The rebuttal addresses seven points from Ref. 3 and several from Ref. 4, covering the definition of the transport gap, the threshold parameter Gth, the difference between simulated and experimental analysis (average_over_cutter), measurement ranges and their effect on TGP outcomes, the definition of 'topological' via the scattering invariant, and the interpretation of conductance data in the parity-readout devices. The manuscript provides new simulation statistics in Tables I and II, a stability analysis of ROI2s under magnetic-field-range cropping in Fig. 3, and a reproducibility statement with code links. The central claim is that no flaws have been identified in the FDR estimate (below 8%) and that the objections raised in Refs. 3 and 4 are unfounded.","tokens_in":13652,"tokens_out":5774,"duration_ms":51422,"significance":"If the rebuttal is correct, it defends the reliability of the TGP as a statistical tuning tool, preserving the conclusions of two high-profile experimental papers. The manuscript's strengths include the provision of reproducible code and package environment files, new simulation tables showing FDR bounds under changed analysis settings, and a direct stability analysis of individual ROI2s under magnetic-field-range modifications. The point-by-point format is useful for the community. However, the significance is tempered by two load-bearing caveats identified in the manuscript itself: the admitted cropping bug in the published SI for the Nature paper, and the explicitly quoted but not quantitatively tested assumption that simulated data are drawn from the same probability distribution as experimental data. These caveats mean that the strong 'no flaws' claim in the abstract is not fully established as written.","major_comments":[{"comment":"The manuscript admits a data-processing bug in the published Supplementary Information for Ref. 2: the bias range was incorrectly cropped, and correcting it changes TGP outcomes for device B, including the appearance of a new SOI2 for one cutter and an increase in ROI2 size. This admission is difficult to reconcile with the abstract's claim that 'no flaws have been identified in our estimate of the FDR' and that the objections are 'unfounded.' At least one objection (Ref. 4) identified a genuine error in the published data analysis. The rebuttal should explicitly narrow the 'no flaws' claim to the FDR estimate itself and acknowledge that the published SI contained a flaw, or provide a detailed argument for why an error that changes TGP outcomes does not affect the FDR bound. As written, the strong claim overreaches the manuscript's own findings.","section":"Technical Response to Ref. 4, paragraph beginning 'We note that the published version of the TGP data...'"},{"comment":"The FDR transfer from simulation to experiment relies on the assumption, quoted in the manuscript, that 'the simulated data is drawn from the same probability distribution as the data produced by real devices.' The new robustness analysis in Tables I and II and Fig. 3 varies parameters inside the simulation model (average_over_cutter, magnetic field range) but does not test the distributional match between simulated and experimental conductance features. Independent localization-length measurements (Fig. 5) constrain disorder but do not establish that TGP-relevant features--zero-bias peak statistics, nonlocal gap shapes, Andreev-enhanced subgap conductance, junction asymmetries--are drawn from the same distribution. The rebuttal should either provide a quantitative comparison between simulated and experimental conductance distributions or explicitly state that the FDR bound is conditional on this untested assumption. Without such a test, the claim that 'no flaws have been identified' is not fully supported.","section":"Quotation from Ref. 1 in 'Technical Response to Ref. 3', point 2"},{"comment":"The manuscript states that 'changes in Gth of less than 20% result in only minor variations in the FDR (i.e. a few percent).' This statement is load-bearing for the robustness of the FDR estimate, but no quantitative evidence is provided for this specific claim. Figure 1 shows the presence/absence of ROI2s for a single measurement as Gth varies from 0.04 to 0.06 Gmax, but it does not report FDR values or statistics over the simulation ensemble. If the FDR is to be claimed robust against Gth variations, the manuscript should include a simulation table or plot showing FDR as a function of Gth, or should rephrase the statement to describe the observed stability without assigning a precise few-percent bound.","section":"Technical Response to Ref. 3, point 1 (threshold Gth)"}],"minor_comments":[{"comment":"The phrase 'the objections in arXiv:2502.19560 and arXiv:2503.08944 are unfounded' is too broad given the admitted cropping bug in the published SI. Recommend softening to 'the objections do not affect the FDR estimate' or similar.","section":"Abstract"},{"comment":"The text says the experimental stage 2 ranges are 'well within' the simulated distribution shown in Fig. 2, but the figure does not overlay the experimental values. Adding such an overlay would make the claim directly verifiable.","section":"Section 'Technical Response to Ref. 4', paragraph on measurement ranges"},{"comment":"The explanation that 'most of the ROI2s are at high fields for the DLG-ε stack' is helpful but could be expanded to state whether this is a property of the simulated model or a consequence of the chosen parameter ranges.","section":"Table II footnote [15]"},{"comment":"The discussion of the scattering invariant as a finite-size criterion is clear, but the manuscript refers to 'Fig. 32 in our paper' without a self-contained reproduction; since this is a response, a brief description of the relevant panels would aid readers who do not have Ref. 1 open.","section":"Section 'Technical Response to Ref. 3', point 4"},{"comment":"The manuscript is written as a collective reply under 'Microsoft Quantum' with a long author list in a footnote. For reproducibility and transparency, it would help to indicate which authors produced the new simulations and tables.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a rebuttal that will likely attract wide attention. The admitted cropping bug in the published SI of the Nature paper is a substantive issue; the editor may wish to consider whether this should be accompanied by an erratum or a clarification to the original publication. The rebuttal's strong 'unfounded' language is not well matched to the manuscript's own admissions, and I would recommend that the revised version either narrow the claims or provide the missing distributional-match and Gth-robustness evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read on this rebuttal. If you're following the Majorana nanowire debate, it's worth a skim, but it does not settle the underlying question.\n\nWhat's actually new: three robustness checks on the TGP false discovery rate — Table I (average_over_cutter=True), Table II (B≤2.5 T), and Fig. 3 (ROI2 survival under field-range cropping). These are parameter scans over the same simulation dataset used in the original PRB paper, so they are a defense rather than a new result. The authors also admit, for the first time, a bug in the published Nature SI: the bias range was cropped asymmetrically by one pixel. They quantify the effect as minor (<5 μeV for 96% of pixels) but concede it changes TGP outcomes for device B, producing a new SOI2 and enlarging an ROI2. That's an honest disclosure, and they protect the referee correspondence by noting it referred to a fixed-field slice.\n\nThe point-by-point refutation is largely coherent. They show the gap-extraction code matches the paper text, that the average_over_cutter difference does not change the FDR, and that the '37% topological fraction' in the comment misreads Stage 1 conditioning — the raw fraction is ~11%. The distinction between 'sensitivity to hyperparameters' and 'unreliable FDR' is well stated.\n\nThe soft spots. First, the central claim that 'no flaws have been identified in our FDR estimate' leans on the original caveat that simulated data is drawn from the same distribution as real devices. All the new robustness tables vary parameters inside the simulation; none test the simulation-to-experiment transfer. The localization-length data constrain disorder but do not validate the joint distribution of zero-bias-peak statistics and nonlocal gap shapes. That remains the load-bearing untested link. Second, the Gth sensitivity claim — 'changes of less than 20% result in only minor variations in the FDR' — is stated without supporting data; Fig. 1 is one experimental example, not a simulation scan. Third, the SI bug, while small, nonetheless changes the published TGP output for device B; the authors argue it doesn't alter the main claims, but it does mean the Nature SI as published was not the actual TGP result.\n\nOverall: this is a competent defense that goes some way toward answering Legg's comments, but it does not close the distributional-realism gap. The 'no flaws' language is overconfident. If you work on topological wire tune-up or on validation of statistical protocols against device physics, it is worth reading; if not, it is skippable.\n\nFor peer review: yes, I'd send it to a serious referee. Responses to comments on high-profile papers deserve scrutiny, and the new tables plus the SI-bug admission should be formally checked. My expectation is that referee reports would ask for a quantitative Gth scan and a more careful statement about the simulation-device distributional match.","headline":"A competent rebuttal that adds useful robustness checks but overclaims 'no flaws' while leaving the simulation-to-device distributional match untested.","tokens_in":14847,"tokens_out":3559,"would_cite":false,"duration_ms":31951,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that two external critiques fail to identify any flaw in the topological gap protocol's false discovery rate estimate.","keywords":["topological gap protocol","false discovery rate","Majorana zero modes","InAs-Al hybrid devices","topological superconductivity","rebuttal","transport measurements","quantum capacitance parity measurement"],"falsifier":"Re-run the released TGP code on the published simulated datasets with every implementation detail set to the experimental pipeline — cutter averaging on, corrected symmetric bias cropping, and magnetic-field ranges cropped to $B_{\\max} \\leq 2.5$ T as in Table II — and count false positives under the paper's own topological criterion; if the fraction of false positives among roughly 700 regions of interest exceeds 8%, the central claim fails. The code repository and data paths named in the paper make this check directly executable.","tokens_in":13269,"feed_emoji":"⚛️","tokens_out":7025,"duration_ms":61439,"temperature":0.7,"pith_summary":"This paper is a point-by-point rebuttal of two external comments on the topological gap protocol (TGP), a statistical test that decides whether a measured InAs-Al nanowire region is in a topological superconducting phase. The central claim is that neither comment identifies any flaw in the protocol's false discovery rate (FDR), the probability that a region flagged as topological is actually trivial, previously bounded below 8%. The authors argue the critics confuse a statistical tuning tool with a smoking-gun test: the TGP may be sensitive to its thresholds and measurement ranges, but that does not mean it frequently misclassifies trivial regions. Concretely, they show the released code matches the paper's gap-extraction description, that the one-parameter difference between simulated and experimental analysis does not change the FDR, and that cropping the magnetic-field range by half a tesla leaves the FDR low. If correct, the conclusions of the original device papers stand, including the parity-measurement result that underpins the topological qubit program.","feed_headline":"Topological gap protocol's low error rate survives both critiques","feed_subtitle":"A rebuttal says no flaw was found in the <8% false discovery rate, keeping the qubit tune-up claims intact.","key_machinery":"The load-bearing object is the false discovery rate (FDR): the probability that the TGP flags a trivial region as topological, estimated by running the protocol on simulated transport data calibrated to independently measured device disorder (localization length over 1 $\\mu$m). The rebuttal's machinery is the transfer argument: because the simulations are drawn from the same distribution as real devices, a low simulated FDR bounds the experimental FDR. The two new tables are the operative mechanism: Table I shows that using the experimental setting `average_over_cutter=True` leaves the FDR statistically unchanged, and Table II shows that cropping $B_{\\max}$ to 2.5 T keeps the FDR below 9% — together with a threshold-sensitivity analysis showing that changes in $G_{\\mathrm{th}}$ of less than 20% move the FDR by only a few percent.","core_discovery":"On the paper's own terms, the claim to be defended is that no flaw has been shown in the TGP's false discovery rate: re-running the protocol with the experimental setting `average_over_cutter=True` yields at most one false positive in roughly 700 regions of interest (Table I), and cropping the magnetic-field range to $B_{\\max} \\leq 2.5$ T changes the FDR bound only slightly (Table II). The paper upholds the original claims that a TGP pass indicates a topological region with probability above $1-8\\%=92\\%$ for the SLG-$\\beta$ stack and above $94\\%$ for the DLG-$\\epsilon$ stack, that measurement-range variations are explained by cooldown-to-cooldown disorder changes and stage-1 cluster sizes, and that the acknowledged bias-cropping bug in the Nature supplement alters extracted gap values by less than 5 $\\mu$eV in more than 96% of pixels without changing the parity result. It concludes that the comments attack a 'smoking gun' methodology the papers never used.","pith_inferences":["One could press further than the paper does: the robustness checks vary one parameter at a time, so a joint worst-case sweep over $G_{\\mathrm{th}}$, cutter averaging, bias cropping, and $B_{\\max}$ would directly test whether the FDR bound remains below 8% under combined stress.","The same calibration logic could be exported to other material platforms: if simulations matched to measured localization length yield a similarly low FDR for other nanowire systems, the protocol would be a general tune-up tool rather than a device-specific one.","The debate implicitly raises a question the paper leaves open: the true-positive criterion uses the scattering invariant together with stable zero-bias peaks and gap reopening, but the false-negative rate is explicitly unquantified; a reader wanting to use the TGP for discovery, not just tune-up, would want that number."],"forward_implications":["If the rebuttal is correct, the <8% FDR bound (and <6% for the DLG-$\\epsilon$ stack) survives, so regions that pass the TGP remain high-confidence operating points for topological qubit tune-up.","The distinction between statistical and smoking-gun tests becomes the operative frame: hyperparameter sensitivity of the TGP is not evidence of bias, and critiques must engage with the FDR itself.","The acknowledged bias-cropping bug in the Nature supplement is contained: gap values shift by under 5 $\\mu$eV in more than 96% of pixels, and the flux-dependent bimodal random-telegraph-signal parity result is not affected.","Measurement-range differences between devices and cooldowns are presented as natural consequences of stage-1 cluster sizes and device drift, with the simulation range distribution in Fig. 2 containing the experimental ranges."],"supporting_citations":[{"why":"Supplies the topological gap protocol, the simulated datasets, the <8% FDR estimate, and the device characterization (localization length) that the entire rebuttal defends.","marker":"[1]"},{"why":"The parity-measurement paper whose tune-up procedure (Supplementary S4.3) is the second target of the comments; the rebuttal upholds its main claim.","marker":"[2]"},{"why":"The first target comment; its claims about gap extraction, simulation-experiment differences, data ranges, and the topological definition are rebutted point by point.","marker":"[3]"},{"why":"The second target comment; its claims that the TGP reports gapped and gapless regions and that conductance data show a gapless, disordered system are rebutted.","marker":"[4]"},{"why":"The original protocol proposal that the TGP extends and whose statistical logic the paper restates as the framing for the rebuttal.","marker":"[5]"},{"why":"The released code; the paper argues it matches the published gap-extraction description and enables validation of Tables I and II.","marker":"[14]"},{"why":"Supplies the Andreev-enhancement mechanism used to argue that strong local subgap conductance does not imply an abundance of low-energy states.","marker":"[18]"},{"why":"Supplies the symmetry relations for local and nonlocal conductance used to rebut the claim of particle-hole symmetry breaking.","marker":"[19]"}],"fun_headline_variants":["Rebuttal: no flaws found in topological gap protocol's error rate","Topological qubit tune-up protocol stands firm against critiques","False discovery rate defense: TGP critics shown unfounded","No flaw in gap protocol's error estimate, rebuttal argues","Topological gap protocol errors survive scrutiny, authors counter"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the assumption that the simulated datasets used to calibrate the protocol are drawn from the same probability distribution as the real device data; if the simulations overstate how representative the devices are, the claimed FDR bound could be too low.","fun_headline_variants_meta":{"raw":{"variants":["Rebuttal: no flaws found in topological gap protocol's error rate","Topological qubit tune-up protocol stands firm against critiques","False discovery rate defense: TGP critics shown unfounded","No flaw in gap protocol's error estimate, rebuttal argues","Topological gap protocol errors survive scrutiny, authors counter"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000572,"raw_usage":{"total_tokens":2746,"prompt_tokens":1031,"completion_tokens":1715,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":1631}},"tokens_in":647,"tokens_out":1715,"duration_ms":10439,"temperature":1.0,"reasoning_tokens":1631,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:13:48.606635+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the released TGP code on the published simulated datasets with every implementation detail set to the experimental pipeline — cutter averaging on, corrected symmetric bias cropping, and magnetic-field ranges cropped to $B_{\\max} \\leq 2.5$ T as in Table II — and count false positives under the paper's own topological criterion; if the fraction of false positives among roughly 700 regions of interest exceeds 8%, the central claim fails. The code repository and data paths named in the paper make this check directly executable.","supporting_citations":[{"cited_title":"This is incorrect","cited_arxiv_id":null,"evidence_quote":"Supplies the topological gap protocol, the simulated datasets, the <8% FDR estimate, and the device characterization (localization length) that the entire rebuttal defends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The parity-measurement paper whose tune-up procedure (Supplementary S4.3) is the second target of the comments; the rebuttal upholds its main claim."},{"cited_title":"The variations in experimental data parameters are explained [1], and TGP outcomes are not the consequence of measurement choices","cited_arxiv_id":null,"evidence_quote":"The first target comment; its claims about gap extraction, simulation-experiment differences, data ranges, and the topological definition are rebutted point by point."},{"cited_title":"There is no redefinition of topological within Ref","cited_arxiv_id":null,"evidence_quote":"The second target comment; its claims that the TGP reports gapped and gapless regions and that conductance data show a gapless, disordered system are rebutted."},{"cited_title":"This claim is based on point 3 above, according to which variations in experimental parameter ranges determine whether a region is classified as gapped or gapless","cited_arxiv_id":null,"evidence_quote":"The original protocol proposal that the TGP extends and whose statistical logic the paper restates as the framing for the rebuttal."}],"review_version":1}