{"id":"08b84294-9e12-4451-a23a-6208e5dcd441","arxiv_id":"2607.03337","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A full kill-chain reconstructs QNN structure from simulated power traces then injects timed crosstalk to approximate adversarial inputs on AQT trapped-ion hardware.","lead":"Researchers chained power-trace reconnaissance, constrained adversarial examples, and timed single-qubit crosstalk to attack a QNN on real trapped-ion hardware. The work shows how multi-tenant quantum cloud services can be hit by linked hardware-model attacks that isolated defenses miss.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Hardware crosstalk stage fails to induce the intended misclassification; success is redefined via MAE of expectation vectors rather than attack impact.","rationale":"The reader correctly flags the simulated power traces (IV-D) as a weak assumption for Stage 1. That concern is real but secondary: even if the simulator perfectly matched AQT laser electronics, the decisive Stage-3 result on real hardware still fails to produce the expected adversarial impact (Table VIII). The paper’s own numbers show crosstalk accuracy higher than clean, so the multi-stage chain does not culminate in a successful physical attack under the natural success criterion (misclassification). Redefining success via MAE of expectation vectors (Table IX) is a post-hoc metric that does not restore the claim. I therefore agree the verdict should stay CONDITIONAL, but the more load-bearing soft spot is the missing accuracy drop on hardware rather than the simulated recon alone. A single re-evaluation of labels/accuracy on the same 50 circuits would settle whether the physical link holds.","tokens_in":22476,"tokens_out":581,"duration_ms":5855,"concrete_test":"Re-run the 50-sample Ibex experiment (or the full 338) reporting per-sample argmax labels and accuracy for clean / analytical-adv / crosstalk under identical shot counts and the same [3,5,7,9] layout; if crosstalk accuracy remains ≥ clean accuracy (or label-flip rate stays near zero), the physical realization claim fails and the kill-chain is incomplete.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is an end-to-end kill-chain that realizes an adversarial perturbation on real AQT hardware via timed neighbor RY crosstalk. Stage 1 (power-trace recon) is simulated (IV-D) and Stage 2 (masked PGD) succeeds only on a minority of samples (Table VII: 338/885). On the 50 samples run on Ibex, Table VIII shows accuracies clean=0.64, adversarial=0.56, crosstalk=0.68 — crosstalk is no worse than clean and better than the analytical adversarial input. The paper then redefines success by the smallest MAE between hw_adv and hw_crosstalk expectation vectors (Table IX, 0.1236). That metric shows the induced state is closer to the adversarial hardware state than to clean, but does not establish that the physical attack actually flips labels or degrades accuracy as an adversarial example must. Without a demonstrated accuracy drop or label-flip rate under the physical schedule, the final link of the kill-chain remains unproven; the claim reduces to “we can move expectation vectors a little via crosstalk,” which is weaker than the abstract’s end-to-end attack demonstration.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The manuscript presents a multi-stage kill-chain attack against a re-uploading variational QNN on AQT trapped-ion hardware. Stage 1 uses simulated power traces to identify the victim architecture among 24 benchmark circuits via timing filters and L2 distance. Stage 2 generates adversarial examples with (masked) PGD under white-box access, restricting perturbations to RY-encoded features to match hardware crosstalk constraints. Stage 3 schedules neighbor RY rotations timed to the inferred encoding gates, using a measured crosstalk factor f≈0.01, and evaluates the induced effect on 50 selected samples on AQT Ibex. Superconducting power-trace and Active SWAP experiments appear in the appendix. The authors claim an end-to-end demonstration that side-channel reconnaissance enables a physical crosstalk attack approximating adversarial examples, and discuss QaaS implications and mitigations.","tokens_in":22797,"tokens_out":1522,"duration_ms":15788,"significance":"Framing QML security as a multi-stage kill chain (reconnaissance → constrained adversarial construction → timed physical crosstalk) is a useful systems-level contribution beyond isolated attack papers. Demonstrating the pipeline on real trapped-ion hardware, with explicit mapping from input-space δ to neighbor RY schedules and measured crosstalk factors (Appendix B, Table IV), is novel and relevant to multi-tenant QaaS. The constrained-PGD construction and the honest reporting of hardware–simulator MAE gaps (Table IX) are strengths. If the final physical stage were shown to reliably degrade classification, the work would be a strong reference for hardware and cloud providers. As written, the significance is tempered by the fact that the decisive hardware result does not establish attack impact via accuracy or label flips.","major_comments":[{"comment":"Section V-C, Table VIII: On the 50 hardware samples the reported accuracies are clean=0.64, adversarial=0.56, crosstalk=0.68. The physical crosstalk schedule does not reduce accuracy relative to clean and is slightly better than clean; it therefore fails the standard definition of an adversarial example (inducing misclassification). The abstract and introduction claim an end-to-end attack that “realizes the adversarial perturbation on the device.” That claim is not supported by classification outcomes. Redefining success via the smallest MAE between hw_adv and hw_crosstalk expectation vectors (Table IX, 0.1236) shows that the induced state moves toward the adversarial hardware state, but does not establish that the kill-chain’s final stage achieves the stated attack goal. Either additional experiments that produce a clear accuracy drop / label-flip rate under the physical schedule, or a","section":null},{"comment":"Section IV-D and V-A: The reconnaissance stage that supplies the timing map for gate placement rests entirely on a hand-crafted rectangular-pulse simulator (fixed 10 µs / 200 µs pulses, 2 µs idles, ad-hoc amplitude corrections {a_i}). No real AQT power traces are used. The L2 ranking that selects benchmark circuit 7 (and thereby the architecture and schedule) is therefore only as valid as this simulator. Because subsequent stages depend on that timing map, the manuscript should either (i) validate the simulator against real control electronics / published AQT pulse data, or (ii) clearly demote Stage 1 to a simulated reconnaissance assumption and state that the end-to-end hardware claim begins after architecture is known. As written, the “full attack chain on ion traps” overstates what was executed on the device.","section":null},{"comment":"Section V-C / sample selection: The 50 hardware circuits are drawn from the 338/885 training samples for which masked PGD succeeded on the statevector surrogate (Table VII). Accuracies are then reported on hardware where even the clean inputs drop to 0.64. The paper notes the sim–hardware mismatch but still presents the pipeline as a successful end-to-end attack. A load-bearing revision is needed: report label-flip rates and confusion matrices for clean / adversarial / crosstalk on the same 50 samples (analogous to the superconducting confusion matrices in Fig. 11), and quantify how often crosstalk actually changes the argmax relative to clean hardware, not only MAE of expectation vectors.","section":null}],"minor_comments":[{"comment":"Table II threat model: “explicit timing information about the victim circuit is required” for the crosstalk stage; the text should state more clearly whether this timing is assumed known a priori or is claimed to be recovered from the (simulated) power-trace stage, and under what multi-tenant scheduling assumptions concurrent placement is realistic.","section":null},{"comment":"Section IV-F: The uniform f_j=0.01 approximation is reasonable given Table IV, but the amplitude-saturation / gate-splitting rule (γ capped at π) should be checked for the largest |Δθ_i| that actually occur; a short histogram of required γ would help the reader assess how often multi-π sequences are needed.","section":null},{"comment":"Figure 4 and PCA inverse: The visual adversarial examples are in the reconstructed 16×16 domain; a brief note that the attack itself operates in the 16-dimensional PCA feature space (not pixel space) would avoid confusion.","section":null},{"comment":"Appendix A Active SWAP results (Table XIV): Individual CNOT disturbances raise accuracy (0.87, 0.88 vs 0.80 baseline) while the pair lowers it to 0.75. The discussion correctly cautions against over-interpretation; a one-sentence statement that this is not claimed as a reliable sabotage attack would align the appendix with the main-text caution.","section":null},{"comment":"Notation: The expanded input dimension is written as R^48 and D=3·16 in IV-F, while the architecture description earlier uses 16 PCA features with three re-uploads; a single consistent symbol for the expanded encoding vector would improve readability.","section":null},{"comment":"References: The kill-chain taxonomy is self-cited as [4]; for a journal audience, a short self-contained summary of the five stages (already present in III-A) is fine, but ensure the arXiv version of [4] is stable or cite the formal venue if available.","section":null}],"recommendation":"major_revision","confidential_remarks":"The work is a genuine multi-stage systems paper and is more ambitious than typical single-vector QML security notes. The main risk for the journal is overclaim in the abstract relative to Table VIII. If the authors reframe the hardware result honestly (crosstalk moves expectation vectors toward the adversarial hardware state; classification impact remains inconclusive under current noise) and either strengthen the physical attack or demote Stage 1 to simulation, the paper becomes a solid contribution. I would not reject on novelty grounds; the trapped-ion end-to-end framing and the constrained RY crosstalk construction are worth publishing after the claim is aligned with the data."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new piece is the end-to-end pipeline on AQT Ibex: simulated power-trace architecture recovery, masked PGD that only touches RY-encoded features (respecting virtual RZ), then timed neighbor RY gates that try to realize those angles via measured crosstalk. Prior work treated the stages separately or only in simulation; this is the first report that actually schedules the physical disturbance on a real ion chain against a trained re-uploading QNN.\n\nWhat they do well is the system-level framing and the honesty about hardware constraints. The kill-chain taxonomy is useful, the RY-only mask is a clean adaptation to AQT physics, the crosstalk-factor measurement (Appendix B) is careful, and they report the awkward accuracy numbers instead of burying them. Table IX showing the smallest MAE between hw_adv and hw_crosstalk is real evidence that the induced state moved in the intended direction. The superconducting appendix and the defensive discussion for QaaS operators are also solid.\n\nSoft spots, in proportion: Stage 1 rests entirely on a hand-crafted rectangular-pulse simulator (Section IV-D) because real power traces were unavailable; if the amplitude/timing model is off, the timing map used for gate placement is wrong. Masked PGD succeeds on only 338/885 clean samples. On the 50 hardware runs, Table VIII gives clean 0.64, adversarial 0.56, crosstalk 0.68 — the physical attack does not lower accuracy relative to clean. They redefine success via MAE of expectation vectors rather than label flips. That is a genuine weakening of the abstract’s claim, but it is not hidden; the paper itself flags the counter-intuitive numbers and the simulator-to-hardware gap. Free parameters (f_j frozen at 0.01, PGD schedule) exist but are secondary.\n\nThis is for people who care about multi-tenant QaaS isolation and hardware-level QML security. The math and citation pattern look fine; the data are limited but honestly presented. I would send it to peer review — the system-level demonstration is worth referee time even if the final link needs tightening. Worth reading and citing if you work on quantum side-channels or QML robustness; not a must-read outside that niche.","headline":"First real-hardware chaining of power-trace recon + constrained adversarial examples + timed ion-trap crosstalk on a QNN; the physical stage moves expectation vectors but does not deliver the promised accuracy drop.","tokens_in":23398,"tokens_out":572,"would_cite":true,"duration_ms":5525,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A multi-stage kill-chain can reverse-engineer a quantum neural network on real trapped-ion hardware and then flip its predictions by timed crosstalk.","keywords":["quantum neural networks","kill-chain attack","power side-channel","crosstalk","adversarial examples","trapped ions","QaaS security","re-uploading encoding"],"falsifier":"Capture genuine analogue power traces from the same trapped-ion controller while the victim circuit runs; if those traces no longer uniquely match the correct benchmark architecture, or if the subsequent crosstalk schedule fails to move the hardware expectation vectors closer to the adversarial ones, the claimed end-to-end chain collapses.","tokens_in":23375,"feed_emoji":"⚛️","tokens_out":635,"duration_ms":5481,"temperature":0.7,"pith_summary":"This paper shows that an attacker can chain side-channel reconnaissance, adversarial-example construction, and physical noise injection into a single campaign against a variational quantum classifier running on a real trapped-ion processor. Power-trace signatures (simulated from the pulse schedule) first identify the victim circuit’s qubit count, depth, and entangling pattern. That structural knowledge is then used to craft input perturbations that are restricted to the rotation axes the hardware can actually disturb via nearest-neighbour crosstalk. Finally the attacker places carefully timed single-qubit rotations on adjacent ions so that the resulting crosstalk approximates the intended adversarial perturbation. On the device the crosstalk configuration moves the measured expectation values closest to those of the analytically adversarial inputs, demonstrating that hardware-level effects can realise a classical-style evasion attack. The result matters because cloud quantum services already expose multi-tenant scheduling and pulse-level diagnostics; the work therefore supplies a concrete, end-to-end threat model rather than isolated vulnerabilities.","feed_headline":"Kill-chain flips a quantum neural net via timed ion crosstalk","feed_subtitle":"Power traces recover the circuit; neighbour rotations then realise the adversarial input on real hardware","key_machinery":"The multi-stage kill-chain itself: reconnaissance yields a temporal map of the victim’s data-encoding RY gates; masked PGD then produces adversarial angles only on those RY components; a linear crosstalk model converts the required angle offsets into neighbour rotations that are scheduled at the recovered times.","core_discovery":"An end-to-end multi-stage attack—power-trace architecture recovery, constrained projected-gradient adversarial examples, and timed neighbour RY crosstalk—can be executed against a trained re-uploading quantum neural network on real trapped-ion hardware, with the crosstalk configuration producing the smallest mean-absolute error to the adversarial hardware expectation vectors.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Multi-stage kill-chain hits QNN via power traces and ion crosstalk","Side-channel recon plus timed RY crosstalk flips trapped-ion QNN","Power-trace recovery and neighbour crosstalk realize QNN adversarial input","End-to-end attack chain: recon to crosstalk on re-uploading QNN hardware","Trapped-ion QNN defeated by timed neighbour rotations after circuit recovery"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The reconnaissance stage rests entirely on a hand-crafted rectangular-pulse simulator whose amplitudes and fixed durations may not match the real laser-drive electronics of the device.","fun_headline_variants_meta":{"raw":{"variants":["Multi-stage kill-chain hits QNN via power traces and ion crosstalk","Side-channel recon plus timed RY crosstalk flips trapped-ion QNN","Power-trace recovery and neighbour crosstalk realize QNN adversarial input","End-to-end attack chain: recon to crosstalk on re-uploading QNN hardware","Trapped-ion QNN defeated by timed neighbour rotations after circuit recovery"]},"model":"grok-4.5","effort":"low","cost_usd":0.003352,"raw_usage":{"total_tokens":1040,"prompt_tokens":626,"num_sources_used":0,"completion_tokens":106,"cost_in_usd_ticks":33520000,"prompt_tokens_details":{"text_tokens":626,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":308,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":626,"tokens_out":106,"duration_ms":3013,"temperature":1.0,"reasoning_tokens":308,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T03:10:49.614823+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Capture genuine analogue power traces from the same trapped-ion controller while the victim circuit runs; if those traces no longer uniquely match the correct benchmark architecture, or if the subsequent crosstalk schedule fails to move the hardware expectation vectors closer to the adversarial ones, the claimed end-to-end chain collapses.","supporting_citations":[],"review_version":1}