{"id":"c096bbe1-0225-49fb-bf4f-ac79e09fa023","arxiv_id":"2507.21222","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Moderate quantum randomness, and in some cases physical noise, improved MNIST classification accuracy on trapped-ion and IBM hardware compared with the classical limit.","lead":"This paper runs an interpolation-tunable quantum neural network on trapped-ion and superconducting hardware to classify MNIST images. It reports that moderate quantum randomness, and in some cases hardware noise, can push borderline images from wrong to right classifications.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'noise can be beneficial' claim rests on a within-error-bar difference in Fig. 2a and a single 10-shot NY image, with no direct noise characterization; the result is not yet established.","rationale":"The paper's most defensible contribution is the tunable BQNN architecture and the broad agreement of hardware validation with prior simulation. The load-bearing weak point is the inference that physical noise is beneficial. The reader's weakest_assumption targets a possible systematic-vs-incoherent confound; I agree that the paper lacks direct noise characterization, but I would broaden the concern: the evidence is underpowered before the mechanism question is even reached. Fig. 2a's hardware excess is admitted to be within error bars, so it cannot support the statement that physical noise is the only explanation. Fig. 2b/c and Fig. 3 use a single preselected NY image with no stated shot count or error bars, so the observed improvements are consistent with sampling variation. A replication with more shots/images and direct noise benchmarking would settle whether the headline effect is real. Because the authors disclose limitations and hedge their broader claims, the appropriate verdict remains CONDITIONAL pending this statistical and calibration evidence; I would not reject the paper on the current evidence, but I also would not accept the noise-benefit claim as established.","tokens_in":8484,"tokens_out":8444,"duration_ms":106599,"concrete_test":"Replicate the NY-image and 55-image scans with at least 100 shots/image on fresh MNIST test images using the same IBM and trapped-ion platforms; immediately before each run, measure single-qubit gate fidelity, readout confusion matrices, and calibration drift (e.g., randomized benchmarking and repeated calibration). Then compute bootstrap 95% confidence intervals for the hardware-vs-QASM validation-rate differences at a=0 and a=0.5, and test whether the NY-image a=0 success rate is reproduced and consistent with a noisy simulation built from the measured error model rather than from a deterministic calibration bias. If the hardware-vs-simulation excess is no longer significant, or the NY success is explained by systematic miscalibration, the 'noise is beneficial' claim should be withdrawn; if the excess survives and is captured by measured incoherent noise, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract asserts that noise 'is shown to improve' BQNN performance, but the supporting evidence is underpowered. In the main scan (55 images, 10 shots/image), the authors state that hardware validation rates 'slightly exceed' QASM simulation for a≤0.5 yet 'the difference is within error bars'; they nevertheless conclude 'its presence can only be attributed to physical noise.' A within-error-bar excess cannot establish the presence of a noise benefit. The only direct evidence is the single NY image 6929: at a=0, the trapped-ion device returns 1/10 correct and the IBM device 5/10. With 10 shots these proportions have standard errors of roughly 9 and 16 percentage points, and the two-platform difference is not significant; the paper reports no readout-error rates, gate fidelities, or calibration data for these runs. The observed a=0 successes could therefore be sampling fluctuations, systematic angle miscalibration, or readout bias that happens to favor the true label, rather than beneficial incoherent noise. The noise-injection scans (Fig. 2c, Fig. 3) use the same single image and show non-monotonic, platform-dependent patterns without error bars. Because the headline claim explicitly invokes physical noise as an improvement mechanism, this is a load-bearing gap: if the a=0 'noise benefit' is a calibration artifact or sampling noise, that part of the central claim loses its empirical support. The more modest claim that quantum uncertainty at moderate a agrees with Ref. [5] and QASM simulation is better supported, but the hardware demonstration of the noise-benefit mechanism is not.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper implements a quantum neural network (BQNN) on IBM superconducting and trapped-ion hardware, with a tunable classical-to-quantum interpolation parameter a. The network is trained classically but inference is run on hardware. The central claims are: (i) at moderate a, hardware validation rates exceed those of the classical a=0 network; (ii) for an ambiguous MNIST image (index 6929) that fails classically, hardware at a=0 sometimes classifies correctly due to physical noise; and (iii) injected single- and two-qubit gate pairs can, in some regimes, improve rather than degrade performance. The paper also presents a noise-injection benchmarking method. The experimental results are compared with QASM simulation and with the prior simulations of Ref. [5].","tokens_in":8789,"tokens_out":4503,"duration_ms":53411,"significance":"If the claims are established, this would be one of the few hardware demonstrations of a quantum neural network outperforming its classical counterpart on a standard image classification task, and it would provide a concrete example of noise-assisted quantum machine learning. The paper is also useful as a multi-platform benchmark (superconducting versus two trapped-ion gate implementations) with publicly available code. However, the statistical basis is thin: the main aggregate scan uses only 55 images with 10 shots each, the detailed noise-benefit analysis relies on a single image with 10 shots, and the paper provides no direct noise characterization. The result is therefore a promising proof-of-concept rather than a definitive demonstration, and the strongest claims in the abstract go beyond what the data currently support.","major_comments":[{"comment":"The claim that the experimental validation rate 'slightly exceeds' QASM simulation for a ≤ 0.5 and that 'its presence can only be attributed to physical noise' is internally inconsistent with the same sentence's acknowledgement that 'the difference is within error bars'. A within-error-bar excess cannot establish the presence of a noise benefit, and it certainly cannot support the abstract's statement that physical noise 'is shown to improve' BQNN performance. Please provide the number of shots per point, confidence intervals, and either a direct noise measurement or a statistical test that quantifies the evidence for a hardware advantage.","section":"Fig. 2a and following paragraph"},{"comment":"The a=0 hardware successes on NY image 6929 (1/10 for the trapped-ion device and 5/10 for the IBM device) are each based on only 10 shots. The binomial confidence intervals for 1/10 and 5/10 overlap zero and overlap each other, so these data are statistically indistinguishable from sampling fluctuation. No calibration data (readout error rates, gate fidelities, or qubit coherence times) are reported for the specific runs, so the attribution to 'physical noise' cannot be distinguished from deterministic calibration bias or readout bias. The 'two nearby minima' picture is plausible but currently unfalsified by the reported data.","section":"Fig. 2b and 'Perhaps the most interesting observation' paragraph"},{"comment":"The noise-injection scans that support the 'noise can be beneficial' conclusion are performed on a single image (NY 6929 and one YY image) and are plotted without error bars, shot counts, or replicates. The non-monotonic behavior in Fig. 3 (e.g., the trapped-ion jump after one two-qubit gate pair and the IBM improvement with additional pairs) is presented as evidence for a mechanism, but with no statistical uncertainty it is impossible to assess whether these features are reproducible or are fluctuations. These data need error bars, repeated runs, or a per-image statistical analysis before the conclusion can be accepted.","section":"Fig. 2c and Fig. 3"},{"comment":"The statement that 'The NY images are the source of the boost of performance upon quantization' is a causal claim about the aggregate improvement in Fig. 2a, but Fig. 2a shows only the aggregate validation rate over 55 images. No per-image breakdown is provided to demonstrate that the improvement is concentrated in NY images rather than spread across the full test set. Please provide a per-image analysis or at least a comparison of accuracy on NY vs. non-NY images at different a values.","section":"First paragraph after Fig. 2b"}],"minor_comments":[{"comment":"Equation (4) appears to have a formatting error: the k=1 case reads 'W 1I' which should likely be 'W^1 I', and the parenthesis structure is unclear. Please correct the equation and ensure the argument of the activation function is written unambiguously.","section":"Equation (4)"},{"comment":"The phrase 'arbitrary angels of rotation' is a typo; it should be 'arbitrary angles of rotation'.","section":"After Eq. (5)"},{"comment":"The sentence 'This predictably hinders classification performance and we observe validation rates dropping with with a' contains a duplicated 'with'.","section":"Second paragraph of Section after Fig. 2b"},{"comment":"The x-axis label '# U U†' does not define what U is or how it relates to the parameter n described in the text; please clarify in the caption that U denotes a native single-qubit or two-qubit gate and that n pairs are inserted after each layer.","section":"Fig. 2c caption"},{"comment":"The abstract states that 'Increasing [a] ... is shown to improve network performance', but the demonstration relies on a small sample (55 images, 10 shots each) with error bars that are not shown in Fig. 2a. Please add error bars or confidence intervals to the aggregate validation-rate plots, or soften the wording to reflect the statistical strength of the evidence.","section":"Abstract and Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-motivated hardware benchmark with useful multi-platform data, but its headline claims about noise-induced improvement rest on statistically underpowered measurements. The central intermediate-a result is consistent with the authors' prior simulation (Ref. [5]), so the paper is partly a confirmatory study; that framing should be made more explicit. I would be willing to reconsider after the authors add more shots/images, report calibration data, provide per-image breakdowns, and include error bars on the noise-injection scans. With those additions, the paper could be a solid contribution to the QML benchmarking literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the real news: the intermediate-a effect predicted in the authors' earlier paper does show up on actual hardware. A 55-image scan on IBM machines reproduces the simulated bump at a≈0.5, and the trapped-ion and IBM results for the borderline image track the same qualitative curve. That is a legitimate cross-platform benchmark, and they ship code. The paper also does not oversell: it says outright that the circuit is separable and classically simulable, and it frames the 'advantage' as a regularizer, not quantum supremacy.\n\nThe soft spot is the 'noise is helpful' headline. In Fig. 2a, the hardware validation rates exceed QASM simulation for a≤0.5—but the authors admit the difference is within error bars and then say it 'can only be attributed to physical noise.' That inference does not follow. The direct evidence is one NY image, index 6929, with 10 shots per platform: 1/10 on trapped ion, 5/10 on IBM at a=0. With 10 shots, those are sampling fluctuations. The paper does not report gate fidelities, readout error rates, or calibration data for the runs, so a deterministic bias in the rotation angles or a readout asymmetry that happens to favor the true label is not ruled out.\n\nThe noise-injection scans are also single-image and lack error bars; they are interesting as exploratory device tests, but not enough to support the general claim that 'certain types of physical noise can be beneficial.' The two-minima heuristic is plausible, but it is an explanation of an unestablished effect.\n\nNone of this is fatal to the paper's core value. The moderate-a result stands. The fix is straightforward: more images, more shots, or at least honest confidence intervals and a rewording that distinguishes 'consistent with noise' from 'caused by noise.' If a referee asks for that, the paper can get there.\n\nWho it is for: people doing small-scale QML hardware benchmarking will want this. It deserves to be reviewed seriously, but it should not be accepted as-is. My vote would be major revision with a request for either stronger statistics or a more careful statement of what the data can support.","headline":"Hardware confirmation of the intermediate-a effect is genuine, but the headline 'noise helps' claim rests on a single 10-shot image and a within-error-bar excess.","tokens_in":9334,"tokens_out":2414,"would_cite":true,"duration_ms":27526,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.Lx"],"model":"deepseek-v4-flash","headline":"This paper demonstrates a tunable quantum neural network whose intermediate quantum uncertainty improves handwritten-digit classification over its classical limit on trapped-ion and superconducting hardware, with physical noise sometimes…","keywords":["quantum neural networks","binary neural network","MNIST classification","trapped-ion processor","superconducting qubits","quantum activation function","noise-assisted inference","mid-circuit measurement"],"falsifier":"Run the same NY-image inference at $a=0$ on a device while independently logging gate fidelities and readout error rates; if the classification boost persists after error mitigation removes the incoherent component, or if it tracks a known systematic rotation or readout bias rather than random errors, the paper's attribution of the boost to physical noise is refuted.","tokens_in":8290,"feed_emoji":"⚛️","tokens_out":7215,"duration_ms":74124,"temperature":0.7,"pith_summary":"The paper reports a quantum generalization of a binary neural network, called the BQNN, and runs its inference on trapped-ion and superconducting processors to classify handwritten-digit images. The network has a knob, the interpolation parameter $a$, that tunes it continuously between a classical deterministic network and a quantum network with measurement uncertainty. The core result is that at moderate $a$, quantum uncertainty in the measurement-based activations improves classification accuracy over the classical limit, and that on an intermediate set of ambiguous images the improvement appears even at $a=0$ because physical hardware noise pushes outputs toward the correct label. The paper therefore argues that certain physical noise, normally viewed as harmful, can be beneficial for some quantum machine learning tasks. Their experiments also show that injecting additional gate noise eventually degrades performance toward random guessing.","feed_headline":"Quantum uncertainty lifts neural-net accuracy on real hardware","feed_subtitle":"A tunable network beats its classical limit on borderline images, and device noise can help.","key_machinery":"The load-bearing object is the BQNN circuit: a three-layer feedforward network whose neurons are qubits, whose activations are projective measurement outcomes, and whose single-qubit rotation angles are functions of the previous layer's measurements. The rotation angle for neuron $i$ in layer $k$ is $\\theta_i^k = \\frac{\\pi}{2}(1-\\phi_a((W^1 I)_i))$ for the first layer and $\\theta_i^k = \\frac{\\pi}{2}(d^{k-1}_i - \\phi_a((W^k d^{(k-1)})_i))$ for later layers, with $\\phi_a(x)=\\operatorname{htanh}(x/a)$ reducing to $\\operatorname{sgn}(x)$ as $a\\to0$. At $a=0$ the rotations are multiples of $\\pi$ and the feedforward is deterministic and classically equivalent, while at nonzero $a$ the rotations produce superposition states and hence random measurement outcomes. Training uses a clipped straight-through gradient estimator in simulation, and inference uses repeated runs with majority voting. The noise-injection procedure, inserting $n$ pairs of native gates $U$ and $U^\\dagger$ after each layer, is the instrument used to benchmark how physical gate noise affects classification.","core_discovery":"On the paper's own terms, the central discovery is that a network whose feedforward pass consists of qubit rotations followed by projective measurements, with rotation angles set by previous measurement outcomes, displays a 'quantumness' window in which real-device inference outperforms the same network's classical $a=0$ limit. The boost comes from borderline images that the classical network deterministically misclassifies; at $a>0$, measurement randomness lets the output fluctuate between two nearby minima of the network's effective energy landscape, and on hardware the same fluctuation is caused partly by physical noise at $a=0$. For clearly classified images the outcome is stable, so the noise sensitivity is specific to ambiguous inputs. The paper verifies this picture across superconducting and two trapped-ion implementations, and shows that injecting extra single-qubit or two-qubit gate pairs eventually destroys performance, but that small or moderate noise can improve an NY image's validation rate rather than hurt it.","pith_inferences":["If the NY-image sensitivity to noise is generic rather than peculiar to this digit dataset, quantum inference could serve as a data-quality probe: entries whose classifications are unstable under small hardware perturbations are exactly the ambiguous ones a classical network hides.","A natural testable extension, not performed in the paper, is to train the BQNN under small injected physical noise rather than only in clean simulation; the paper expects a systematic advantage from such noise-aware training, and that expectation could be checked by comparing validation rates.","Because the demonstrated circuit is separable and uses only mid-circuit measurements followed by classical feedback, scaling it to entangled partial-mid-circuit-measurement architectures would connect it to measurement-induced phase transitions; the authors flag this direction but do not demonstrate it, so the quantum-advantage claim remains prospective.","If the noise-injection metric, the number of $UU^\\dagger$ pairs needed to reach random-guess performance, is reliable, it offers a neural-network-specific alternative to circuit-level error rates for comparing quantum platforms."],"forward_implications":["At moderate $a$, the BQNN on real superconducting and trapped-ion devices matches or exceeds the classical network's fraction of correctly classified images on the tested subset, with the gain concentrated on ambiguous borderline images.","NY images, misclassified classically but recovered quantumly, respond to hardware noise at $a=0$ in a way that clear YY images do not, suggesting that noisy quantum inference can flag ambiguous entries in a dataset.","Inserted single-qubit $UU^\\dagger$ pairs can raise the NY-image validation rate on the trapped-ion device before enough noise drives all classifications toward random guessing.","Two-qubit gate-pair injection produces a sharp initial improvement on the trapped-ion NY image followed by decay to the random-guess rate, while the superconducting platform shows more resilience.","The qualitative agreement between simulation and three hardware implementations supports the BQNN as a portable benchmark for comparing device noise and for studying quantum-versus-classical behavior in neural networks."],"supporting_citations":[{"why":"Defines the BQNN architecture, the quantum activation function, and the simulated intermediate-quantumness accuracy benefit that this paper experimentally tests.","marker":"[5]"},{"why":"Supplies the spin-model energy-landscape picture of trained neural networks used to explain why noise switches NY images between nearby minima.","marker":"[4]"},{"why":"Provides the straight-through gradient estimator used to train the binarized network despite the discontinuous sign activation.","marker":"[23]"},{"why":"Supplies the MNIST handwritten-digit test images used as the classification benchmark.","marker":"[25]"},{"why":"Establishes the circuit-QED transmon platform underlying the superconducting hardware used for inference.","marker":"[26]"},{"why":"Describes the trapped-ion system and gate implementations used for the microwave and Raman experiments.","marker":"[30]"},{"why":"Supports the claim that stochastic noise during training acts as a regularizer, the classical analogue of the observed quantum-noise benefit.","marker":"[31]"},{"why":"Documents high measurement-error rates on superconducting devices, used to explain why that platform already shows noise-induced improvement at $a=0$ without injected noise.","marker":"[32]"}],"fun_headline_variants":["Quantum uncertainty improves MNIST accuracy on real chips","Noise can help quantum neural networks classify borderline images","Quantum neural net edges out classical on tricky images","Moderate noise boosts quantum network performance","Quantum randomness aids classification on trapped-ion and IBM hardware"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's noise-benefit claim assumes that hardware deviations from ideal simulation are random errors rather than a fixed bias that happens to favor the correct label, and it does not directly measure the noise.","fun_headline_variants_meta":{"raw":{"variants":["Quantum uncertainty improves MNIST accuracy on real chips","Noise can help quantum neural networks classify borderline images","Quantum neural net edges out classical on tricky images","Moderate noise boosts quantum network performance","Quantum randomness aids classification on trapped-ion and IBM hardware"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000189,"raw_usage":{"total_tokens":1347,"prompt_tokens":968,"completion_tokens":379,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":308}},"tokens_in":584,"tokens_out":379,"duration_ms":4987,"temperature":1.0,"reasoning_tokens":308,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:58:50.162211+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same NY-image inference at $a=0$ on a device while independently logging gate fidelities and readout error rates; if the classification boost persists after error mitigation removes the incoherent component, or if it tracks a known systematic rotation or readout bias rather than random errors, the paper's attribution of the boost to physical noise is refuted.","supporting_citations":[{"cited_title":"Natural Quantization of Neural Networks","cited_arxiv_id":"2503.15482","evidence_quote":"Defines the BQNN architecture, the quantum activation function, and the simulated intermediate-quantumness accuracy benefit that this paper experimentally tests."},{"cited_title":"NIST special database 19: Handprinted forms and charac- ters database,","cited_arxiv_id":null,"evidence_quote":"Supplies the MNIST handwritten-digit test images used as the classification benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the claim that stochastic noise during training acts as a regularizer, the classical analogue of the observed quantum-noise benefit."},{"cited_title":"Tomesh, P","cited_arxiv_id":null,"evidence_quote":"Documents high measurement-error rates on superconducting devices, used to explain why that platform already shows noise-induced improvement at $a=0$ without injected noise."}],"review_version":1}