REVIEW 2 major objections 6 minor 34 references
A quantum belief-update service on IBM Heron can run sequential Tiger POMDP steps without changing the planner’s action.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-10 21:53 UTC pith:IEHEKWKI
load-bearing objection Honest Heron case study: all-step FPAA keeps sequential Tiger posteriors planner-safe on shallow oracles, with action agreement and boundary BIQAE calibration. the 2 major comments →
QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On present IBM Heron hardware, an all-step fixed-point amplitude amplification loop, paired with boundary-aware Bayesian amplitude estimation, can run sequential Tiger POMDP belief updates and return planner-facing posteriors that stay close to exact Bayes (max Hellinger 0.009 on the 8-step run and 0.021 on the 12-step run) and that select the same immediate action as exact Bayes under the standard Tiger reward rule, with zero measured cumulative value loss on the reported decision checks.
What carries the argument
QANTIS as a calibrated belief-update service: a shallow belief oracle plus all-step fixed-point amplitude amplification (softer phases that avoid overshoot so every listen step can be amplified) and boundary-aware BIQAE (a coarse scan that chooses a near-zero, near-one, or interior prior before fine estimation), feeding an ordinary classical Bayes update whose posterior becomes the next prior.
Load-bearing premise
The claim rests on shallow two-qubit Tiger oracles remaining a faithful planner-facing service; the paper itself treats circuit depth per belief update as the binding limit when encodings get richer.
What would settle it
Rerun the same sequential Tiger trajectory with all-step FPAA and boundary-aware BIQAE on Heron, compute hardware versus exact Bayes posteriors at every step, and check whether any step produces a different immediate action under the stated Tiger reward rule or pushes max Hellinger clearly outside the reported operating band.
If this is right
- A classical planner can treat the quantum step as a drop-in rare-evidence posterior service and still receive ordinary probability distributions.
- All-step fixed-point amplification removes the need for a skip guard that otherwise decides when amplification is safe.
- Boundary-aware estimation is required infrastructure for sequential belief tracking, not a cosmetic fix, because loops repeatedly hit amplitudes near zero and one.
- The useful operating regime is low-probability observations where classical sampling becomes expensive, not a claim that the whole autonomy stack becomes quantum.
- Longer horizons and Heron transfer runs stay inside the same band only while compiled belief oracles remain shallow.
Where Pith is reading between the lines
- If depth, not state-space size, is the first bottleneck, the next useful hardware milestone is shallower problem-specific belief encodings rather than more qubits alone.
- The same service contract could be stress-tested on other small POMDP unit cells with closed-form Bayes references to separate amplification fidelity from Tiger-specific structure.
- Once reliable queue-exclusive timing is logged, the natural next comparison is sample cost per accepted posterior against classical rare-event sampling at the same evidence rates.
- Boundary-aware priors may transfer to other iterative quantum estimators that repeatedly visit near-zero or near-one amplitudes in closed loops.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a controlled IBM Heron hardware case study of QANTIS as a calibrated sequential belief-update service for a two-state Tiger POMDP. The quantum step estimates a rare-event evidence amplitude (via amplification plus boundary-aware BIQAE); an ordinary classical Bayes update then returns a planner-facing posterior. On the same observation trajectory the authors compare no amplification, guarded Grover-AA, and all-step fixed-point AA (FPAA). Headline all-step FPAA reports max Hellinger 0.009 (8-step) and 0.021 (12-step) versus exact Bayes; 20- and 32-step controls stay in a similar band; and under a fixed Tiger immediate-reward rule every reported decision check shows action agreement with zero cumulative value loss. Supporting material includes same-backend BIQAE boundary calibration, a rare-event logical amplification envelope, Heron R3 transfer rows, and explicit non-claims on wall-clock speedup and end-to-end autonomy.
Significance. If the scoped result holds, the paper supplies a rare, auditable hardware operating envelope for sequential POMDP belief updates on present superconducting devices rather than a hardware-advantage claim. Strengths that should be credited include: (i) same-trajectory no-AA / guarded-Grover / all-step-FPAA controls with exact Bayes Hellinger references (Tables III–IV); (ii) planner-facing action and value-loss checks under an external Tiger reward rule (Table VI); (iii) a concrete two-phase boundary-aware BIQAE protocol with same-backend Pittsburgh pairs that collapse near-zero/near-one error (Table VII, Fig. 6); (iv) an explicit claim hierarchy (Table II) and resource accounting that separate primary 8/12-step results from longer-horizon and scaling probes; and (v) public code and artifact manifests. The contribution is narrow—two-qubit Tiger oracles at ISA depth ~15–18—but the systems framing (posterior service, not full autonomy stack) is appropriate and useful for the hybrid quantum–classical planning community.
major comments (2)
- [§III-D, Table IV, Abstract, Contribution 1] §III-D and Table IV: the primary Hellinger numbers used in the abstract and Contribution 1 (max 0.009 on the 8-step run) come from the 32 768-shot FPAA headline row, roughly 3× the ~10k per-step budget of the No-AA and guarded-Grover controls. The matched-shot FPAA control at 10k shots reports max Hellinger 0.033 / mean 0.019, which is not better than No-AA (0.029 / 0.015). The manuscript already notes that Table IV is not a fully budget-matched accuracy proof, but the abstract and contribution list still lead with 0.009 without the shot context. Because the central claim is that all-step FPAA can be reused without corrupting the planner-facing posterior, the abstract, Contribution 1, and the opening of §III-D should state the shot budget next to each headline Hellinger figure and more sharply separate “stability under all-step amplification (action preservation)” from “accuracy improvem
- [§III-C, §III-D, Table V] §III-C–D and Table V: the sequential results are largely single-trajectory (or few-repeat) runs on fixed observation sequences. Table V gives some Fez repeats, but there are no reported run-to-run distributions, confidence intervals, or multi-seed summaries for the primary 8-step Kingston Hellinger series or for the action-agreement checks. For a hardware case study whose product is an operating envelope, at least a short multi-execution summary (e.g., max/mean Hellinger over N independent submissions of the same 8-step path, or bootstrap intervals on the BIQAE estimates) is load-bearing for the claim that the service “preserves” the posterior across a horizon. Without it, the 0.009 / 0.021 figures remain point estimates that are hard to distinguish from favorable calibration windows.
minor comments (6)
- [Fig. 4, §III-D1] Fig. 4 / §III-D1: the rare-event “971832× logical amplification” row is useful as a sample-complexity envelope, but the accompanying text already notes that transpilation collapses the circuit to ≤6 single-qubit gates. Consider moving that caveat into the figure caption so the plot cannot be misread as a deep-circuit fidelity result.
- [Table III, Fig. 3] Table III vs Fig. 3: the 8-step audit table and the Hellinger plot are complementary; a one-sentence pointer in the table caption to Fig. 3 (and vice versa) would help readers who land on only one of them.
- [§IV-B] §IV-B: the noise-aware visibility is described as a session-level plug-in from short calibration circuits. State explicitly whether that visibility is frozen for the whole sequential trajectory or re-estimated per step, and whether the same plug-in is used for the No-AA baseline (which does not need BIQAE).
- [§V-A, Table VIII] §V-A / Table VIII: queue-inclusive and queue-exclusive timings are correctly declared out of scope, but the table still lists “resource accounting.” Renaming the table (e.g., “Shot and circuit accounting”) would avoid implying wall-clock comparison.
- [Table III] Notation: “listen(0)” / “listen(1)” in Table III are never defined in the main text (left vs right growl). A short footnote or parenthetical would remove ambiguity.
- [Related Work] Related work: the positioning against QANTIS v1 is clear; a single sentence on how the present all-step FPAA + boundary BIQAE stack differs from the multi-step OAA distortion analyses of Zecchi et al. and the QBRL simulation of Cunha et al. would further pin the empirical gap.
Circularity Check
No significant circularity: hardware posteriors are scored against independent exact classical Bayes and fixed external Tiger decision rules; self-citation to QANTIS v1 is background, not a load-bearing definition of the new claims.
full rationale
The paper is an empirical hardware case study, not a first-principles derivation that could close on its own inputs. The planner-facing product is an ordinary Bayes posterior; the quantum step only estimates the evidence amplitude, which is then plugged into the classical update and compared to the closed-form exact Bayes posterior on the same Tiger trajectory (Tables III–VI, Hellinger and action/value checks). Action agreement uses a fixed external immediate-reward rule (open above 90% / below 10%, listen in between), so agreement is not forced by construction of the estimator. All-step FPAA is taken from Yoder–Low–Chuang; BIQAE from Li et al.; rare-event logical amplification is measured against analytic Grover–Brassard targets, not fitted from the same counts. Visibility is explicitly a session-level plug-in from known-amplitude calibration circuits, with no joint-identifiability or prediction claim. Self-citation to QANTIS v1 supplies the guarded baseline, platform context, and deferred full derivations of standard Bayes/AA primitives; it does not define the new all-step FPAA, boundary-aware BIQAE, sequential Hellinger, or decision-check results. Those are independently measured on Heron against classical oracles and no-AA controls. No self-definitional loop, fitted-input-as-prediction, uniqueness import, or renamed known result is present. Score 0 is the honest finding.
Axiom & Free-Parameter Ledger
free parameters (5)
- per-step shot budget (headline FPAA) =
32768
- adaptive v1 shot allocation =
~15% reallocation; 8k–16k/step
- BIQAE coarse-scan boundary threshold =
0.1
- noise-aware visibility plug-in =
session-level estimate (not a single global number)
- Tiger observation accuracy =
0.85
axioms (6)
- domain assumption POMDP beliefs update by Bayes rule: predict, weight by observation likelihood, normalize by evidence.
- standard math Amplitude amplification reduces rare-event sampling cost from ~1/p toward ~1/√p under ideal reflections (Brassard–Høyer–Mosca).
- standard math Yoder–Low–Chuang fixed-point phase schedules keep amplification monotone and reduce overshoot for concentrated amplitudes.
- domain assumption BIQAE maintains a Bayesian posterior over amplitude angle and updates with sinusoidal (or noise-aware) likelihoods after chosen Grover depths.
- domain assumption Standard Tiger immediate-reward rule: open right above 90% belief tiger is left, open left below 10%, listen otherwise; listen cheap, wrong door costly.
- ad hoc to paper Hellinger distance is an appropriate bounded metric for planner-facing posterior fidelity in this study.
invented entities (3)
-
QANTIS belief-update service contract
no independent evidence
-
Boundary-aware two-phase BIQAE calibration protocol
no independent evidence
-
All-step FPAA sequential belief-tracking policy
no independent evidence
read the original abstract
Autonomous systems under partial observability act on beliefs, not raw sensor events. QANTIS treats the quantum processor as a calibrated belief-update service in that loop: it receives a prior and an observation model, estimates the rare-event evidence term, and returns an ordinary posterior to a classical planner. This paper asks whether that service can be reused across a sequential Tiger POMDP horizon on present IBM Heron hardware without corrupting the planner-facing posterior. We answer with a controlled hardware case study rather than an end-to-end autonomy or wall-clock speedup claim. The study compares no amplification, guarded Grover amplification, and all-step fixed-point amplification on the same trajectory, then checks whether the returned posterior would change the downstream action. All-step FPAA preserves the Tiger posterior across the reported 8-step and 12-step primary runs, and the 20-step and 32-step controls remain inside the same operating band. In every reported decision check, the hardware posterior and the exact Bayes posterior select the same immediate action. Boundary-aware BIQAE stabilizes amplitude estimation near zero and near one, while a rare-event sweep maps the logical sample-complexity envelope for one-in-a-million evidence. The result is an operating envelope for a hardware-calibrated belief-update primitive, not a standalone hardware-advantage claim.
Figures
Reference graph
Works this paper leans on
-
[1]
Quantum amplitude amplification and estimation,
G. Brassard, P. Høyer, M. Mosca, and A. Tapp, “Quantum amplitude amplification and estimation,” inQuantum Computation and Information, ser. Contemporary Mathematics. American Mathematical Society, 2002, vol. 305, pp. 53–74
work page 2002
-
[2]
QANTIS: A hardware-validated quantum platform for POMDP planning and multi-target data association,
B. Y . Eker, S. S. Arslan, O. Nazlı, M. S. Demirgil, and F. Delig ¨oz, “QANTIS: A hardware-validated quantum platform for POMDP planning and multi-target data association,” 2026, arXiv:2603.00785
-
[3]
S. Aaronson, “Quantum POMDPs,” 2014, arXiv:1406.2858
work page internal anchor Pith review Pith/arXiv arXiv 2014
-
[4]
Quantum Bayesian Networks Can Speed up Reinforcement Learning in Partially Observable Environments
G. Cunha, A. Ram ˆoa, A. Sequeira, M. de Oliveira, and L. Barbosa, “Hy- brid quantum-classical algorithm for near-optimal planning in POMDPs,” 2025, arXiv:2507.18606
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[5]
P. Cintio, A. Michelangeli, and A. Tsutskov, “Complexity scaling and optimal policy degeneracy in quantum reinforcement learning via unitary- control-then-measure models,” Apr. 2026, arXiv:2604.13096
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[6]
Projected Dynamic Programming for Sequential Quantum State Discrimination
J. Jeong, D. Ji, H. Jang, and K. Jeong, “Projected dynamic pro- gramming for sequential quantum state discrimination,” Apr. 2026, arXiv:2604.15393
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[7]
A. A. Zecchi, C. Sanavio, S. Perotto, and S. Succi, “Improved amplitude amplification strategies for the quantum simulation of classical transport problems,” 2025, arXiv:2502.18283
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[8]
Grover’s algorithm is an approximation of imaginary-time evolution,
Y . Suzuki, M. Gluza, J. Son, B. H. Tiang, N. H. Y . Ng, and Z. Holmes, “Grover’s algorithm is an approximation of imaginary-time evolution,” 2025, arXiv:2507.15065
-
[9]
Benchmarking quantum reinforcement learning,
N. Meyer, C. Ufrecht, G. Yammine, G. Kontes, C. Mutschler, and D. D. Scherer, “Benchmarking quantum reinforcement learning,” inProceedings of the 42nd International Conference on Machine Learning (ICML), 2025
work page 2025
-
[10]
Harnessing Bayesian statistics to accelerate iterative quantum amplitude estimation,
Q. Li, A. Vidwans, Y . Wang, and M. B. Soley, “Harnessing Bayesian statistics to accelerate iterative quantum amplitude estimation,”Quantum, vol. 10, p. 1962, 2026
work page 1962
-
[11]
Bayesian quantum amplitude estimation,
A. Ram ˆoa and L. P. Santos, “Bayesian quantum amplitude estimation,” Quantum, vol. 9, p. 1856, 2025, arXiv:2412.04394
-
[12]
On the bias in iterative quantum amplitude estimation,
K. Miyamoto, “On the bias in iterative quantum amplitude estimation,” EPJ Quantum Technology, vol. 11, no. 1, p. 42, 2024
work page 2024
-
[13]
Quantum annealing applied to de-conflicting optimal trajectories for air traffic management,
T. Stollenwerk, B. O’Gorman, D. Venturelli, S. Mandra, O. Rodionova, H. K. Ng, B. Sridhar, E. G. Rieffel, and R. Biswas, “Quantum annealing applied to de-conflicting optimal trajectories for air traffic management,” 2021
work page 2021
-
[14]
Enhancing multiple object tracking accuracy via quantum annealing,
Y . Ihara, “Enhancing multiple object tracking accuracy via quantum annealing,”Scientific Reports, vol. 15, p. 24294, 2025
work page 2025
-
[15]
Quantum annealing vs. QAOA: 127 qubit gate-model IBM quantum computer and D-Wave advantage,
E. Pelofske, A. B ¨artschi, and S. Eidenbenz, “Quantum annealing vs. QAOA: 127 qubit gate-model IBM quantum computer and D-Wave advantage,”npj Quantum Information, vol. 10, p. 18, 2024
work page 2024
-
[16]
Planning and acting in partially observable stochastic domains,
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra, “Planning and acting in partially observable stochastic domains,”Artificial Intelligence, vol. 101, no. 1–2, pp. 99–134, 1998
work page 1998
-
[17]
Monte-Carlo planning in large POMDPs,
D. Silver and J. Veness, “Monte-Carlo planning in large POMDPs,” inAdvances in Neural Information Processing Systems 23 (NeurIPS). Curran Associates, Inc., 2010, pp. 2164–2172
work page 2010
-
[18]
M. A. Nielsen and I. L. Chuang,Quantum Computation and Quantum Information, 10th ed. Cambridge University Press, 2010
work page 2010
-
[19]
The Hungarian method for the assignment problem,
H. W. Kuhn, “The Hungarian method for the assignment problem,”Naval Research Logistics Quarterly, vol. 2, no. 1–2, pp. 83–97, 1955
work page 1955
-
[20]
Y . Bar-Shalom and X.-R. Li,Multitarget-Multisensor Tracking: Principles and Techniques. Storrs, CT: YBS Publishing, 1995
work page 1995
-
[21]
An algorithm for tracking multiple targets,
D. B. Reid, “An algorithm for tracking multiple targets,”IEEE Transac- tions on Automatic Control, vol. 24, no. 6, pp. 843–854, 1979
work page 1979
-
[22]
Quantum approximate optimization algorithm with fixed number of parameters,
S. Saavedra-Pino, R. Quispe-Mendiz´abal, G. Alvarado Barrios, E. Solano, J. C. Retamal, and F. Albarr ´an-Arriagada, “Quantum approximate optimization algorithm with fixed number of parameters,” 2025
work page 2025
-
[23]
A quantum approximate optimization algorithm,
E. Farhi, J. Goldstone, and S. Gutmann, “A quantum approximate optimization algorithm,” 2014
work page 2014
-
[24]
Fixed-point quantum search with an optimal number of queries,
T. J. Yoder, G. H. Low, and I. L. Chuang, “Fixed-point quantum search with an optimal number of queries,”Physical Review Letters, vol. 113, p. 210501, 2014
work page 2014
-
[25]
Grover’s quantum searching algorithm is optimal,
C. Zalka, “Grover’s quantum searching algorithm is optimal,”Physical Review A, vol. 60, no. 4, pp. 2746–2751, 1999
work page 1999
-
[26]
Dequantizing Short-Path Quantum Algorithms
F. Le Gall and S. Tamaki, “Dequantizing short-path quantum algorithms,” Apr. 2026, arXiv:2604.12131. TABLE X OPERATIONAL READING OF THEQANTISVALIDATION RESULTS. THE TABLE SUMMARIZES WHAT EACH EXPERIMENT ESTABLISHES FOR AN AUTONOMOUS DECISION STACK WHILE KEEPING DEPLOYMENT LIMITS EXPLICIT. Validation component Evidence in this paper Operational reading Her...
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[27]
A Nested Amplitude Amplification Protocol for the Binary Knapsack Problem
L. Demmler and M. Hess, “A nested amplitude amplification protocol for the binary knapsack problem,” Apr. 2026, arXiv:2604.05776
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[28]
S. Asmussen and P. W. Glynn,Stochastic Simulation: Algorithms and Analysis, ser. Stochastic Modelling and Applied Probability. New York: Springer, 2007, vol. 57
work page 2007
-
[29]
Confidence Intervals for Rate Estimation with Importance Sampling in Autonomous Vehicle Evaluation
A. Chen, R. R. Zhou, J. J. Lee, N. Chamandy, and H. Hohnhold, “Confidence intervals for rate estimation with importance sampling in autonomous vehicle evaluation,” Apr. 2026, arXiv:2604.03827
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[30]
Efficient tensor network simulation of IBM's Eagle kicked Ising experiment
J. Tindall, M. Fishman, E. M. Stoudenmire, and D. Sels, “Efficient tensor network simulation of IBM’s Eagle kicked Ising experiment,” PRX Quantum, vol. 5, p. 010308, 2024, arXiv:2306.14887
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[31]
T. Beguˇsi´c, J. Gray, and G. K.-L. Chan, “Fast and converged classical simulations of evidence for the utility of quantum computing before fault tolerance,”Science Advances, vol. 10, p. eadk4321, 2024
work page 2024
-
[32]
L. Mauron and G. Carleo, “Challenging the quantum advantage frontier with large-scale classical simulations of annealing dynamics,” Mar. 2025, arXiv:2503.08247
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[33]
T. Midha, G.-M. Sommers, J. Tindall, and D. Abanin, “Belief propagation and tensor network expansions for many-body quantum systems: Rigorous results and fundamental limits,” Apr. 2026, arXiv:2604.03228
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[34]
Extreme Quantum Advantage for Rare-Event Sampling
C. Aghamohammadi, S. P. Loomis, J. R. Mahoney, and J. P. Crutch- field, “Extreme quantum advantage for rare-event sampling,” 2017, arXiv:1707.09553; Phys. Rev. X 8, 011025 (2018)
work page internal anchor Pith review Pith/arXiv arXiv 2017
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.