{"id":"f87c5f07-ff92-4969-943a-49b8c80d4631","arxiv_id":"2608.09625","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"For diffusion QST at 3 qubits, the Bloch representation reaches 0.907 fidelity while Hermitian direct reaches 0.394, reversing the 2-qubit ranking and showing that local conditioning alone does not predict reconstruction quality.","lead":"This paper compares seven ways of writing down a quantum state, or parameterizations, for diffusion-based quantum state tomography at 2 and 3 qubits. It finds that the parameterization with the best local geometry often reconstructs the state worst, because unconstrained coordinates produce invalid states that are distorted by projection after sampling.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 3-qubit central claim is not testable as written: no measurement basis, shot allocation, or conditioning representation is specified, so the Bloch-vs-Hermitian ranking could be a measurement-protocol artifact.","rationale":"The reader's weakest assumption is indeed the load-bearing one: the measurement model is absent, and Table 9's shot-level ranking is the primary evidence for the central claim. My stress-test agrees with the CONDITIONAL verdict: the paper makes an interesting and possibly correct claim, and it is commendably honest about learning-rate confounds, but the unspecified measurement protocol makes the headline result non-reproducible. The calibration-table inconsistency (Table 3 vs Table 14) is a real additional defect, but it is secondary to the missing measurement specification because even a fully consistent set of κ values would not permit independent verification of the end-to-end ranking. I am not identifying a mathematical error in the geometric framework itself; the Jacobian Gram matrix definitions are clear and the analytic benchmarks (Appendix B.2) are appropriate. The concern is empirical reproducibility: 'shots' has no operational meaning without a POVM and shot-allocation rule, and the diffusion model's conditioning representation is a core part of the pipeline that is never defined. A concrete replication with explicit measurement models would settle whether the Bloch advantage is physics or protocol. Therefore I recommend keeping the reader's CONDITIONAL verdict unchanged rather than escalating or downgrading it.","tokens_in":21835,"tokens_out":5971,"duration_ms":63789,"concrete_test":"Independently re-run the 3-qubit end-to-end evaluation under two explicitly specified, informationally complete measurement models, e.g.: (A) all 3^n Pauli settings with equal shots per setting and shot allocation uniform over settings; (B) a SIC-POVM with shots randomized over projectors. In both cases, encode the measurement outcome frequencies as a normalization-invariant conditioning vector, set λ_meas=0.0 as in the paper, and compute Table 9 (100 test states, K=20, no CFG). If Bloch does not beat Hermitian direct at 300 shots by a comparable margin (≈0.5) under both models, or if the Hermitian-direct versus Cholesky ordering flips, then the missing measurement specification is a load-bearing artifact rather than a parameterization-geometry effect. The check should also recompute Table 3 zero-noise Cholesky κ_spec to decide between 1265× and 1793×.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract and §4.7 rests entirely on Table 9, yet the evaluation protocol in §4.1 specifies only \"100 test states × 6 shot levels [10, 20, 30, 50, 100, 300] × K = 20 samples\" with CFG w=4.0 for the 2-qubit protocol, and \"no CFG\" for the 3-qubit Table 9. The measurement model is never defined: no POVM or measurement basis, no rule for how shots are distributed across settings, no description of how measurement statistics are encoded as a conditioning vector for the diffusion model, and no definition of the measurement-consistency loss λ_meas beyond stating λ_meas=0.0. This is load-bearing because the shot-level ranking is the entire evidence for the geometric-conditioning claim; a different POVM (e.g., Pauli versus SIC-POVM versus random Haar) or a different shot-allocation rule changes the posterior information content at each shot count and can plausibly reverse the order of Bloch and Hermitian direct. The paper itself shows that the ranking reverses under changes to learning rate and CFG (§4.5, Appendix E), so the protocol details cannot be treated as incidental. In addition, an internal inconsistency remains unresolved: Table 3 reports 3-qubit Cholesky κ_spec = 1265×, while Table 14 reports 1793× at p=0.0 (and 99× vs 33× at 2 qubits), so the calibration tables are not self-consistent. The missing measurement specification is the primary blocker: without it, the headline result cannot be reproduced or adjudicated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a geometric framework for comparing density-matrix parameterizations in diffusion-based quantum state tomography. It defines two metrics on the Jacobian Gram matrix J^T J: spectral dynamic range (SDR) and diagonal anisotropy (DA); calibrates seven parameterizations at 2- and 3-qubit scales; and trains three of them (Bloch/Gell-Mann, Hermitian direct, Cholesky) end-to-end at 2 and 3 qubits. The main claims are that geometric conditioning alone does not predict end-to-end fidelity, that bounded-domain parameterizations (Bloch) suffer less projection-induced information loss at 3 qubits (0.907 vs 0.394 for Hermitian direct at 300 shots), and that fix-trace provides the best conditioning-constraint tradeoff. The paper also presents matched-learning-rate controls, multi-seed validation, and a coordinate-aware sigma_data calibration recipe.","tokens_in":22175,"tokens_out":11032,"duration_ms":96660,"significance":"The J^T J diagnostics are simple, inexpensive, and do not require fitting to the end-to-end results; the use of analytical Jacobians for Cholesky, Hermitian direct, and Expmap strengthens the calibration. The matched-learning-rate experiments and global-sigma_data control in Appendix E are the right kinds of ablations for isolating the geometric claim. If the 3-qubit end-to-end result survives a fully specified measurement protocol, the paper would establish parameterization as a first-order design axis and provide practical guidance for diffusion QST. However, the current manuscript does not specify the measurement model, contains inconsistent calibration numbers, and frames a learning-rate artifact as a 'reversal'; these issues must be addressed before the central claim can be accepted.","major_comments":[{"comment":"The evaluation protocol does not specify the measurement model. Section 4.1 states only '100 test states × 6 shot levels [10, 20, 30, 50, 100, 300] × K = 20 samples' with CFG w=4.0, and Section 4.7 adds 'no CFG' for Table 9, but nowhere is the POVM or measurement basis defined, the rule for distributing shots across settings specified, or the encoding of measurement statistics as the denoiser conditioning vector described; the measurement-consistency loss lambda_meas is only stated as 0.0. Since the central quantitative claims (Bloch 0.907 vs Hermitian direct 0.394 at 300 shots in Table 9, and the comparison with the MLE baseline in Table 6) depend on how much posterior information each shot count provides, the headline ordering could be an artifact of the unspecified measurement scheme. Please provide the complete measurement protocol, including basis, shot allocation, conditioning representation, and the exact role of lambda_meas, or explicitly state that 'shots' enter only through a fixed conditioning vector with no measurement-consistency term.","section":"§4.1, §4.7, Tables 6 and 9"},{"comment":"The zero-noise Cholesky calibration is internally inconsistent. Table 2 reports the 2-qubit SDR as 33x and Table 3 reports the 3-qubit SDR as 1265x (repeated in Table 15), while Table 14 reports the p=0.0 values as 99x (2q) and 1793x (3q), and Section 5.4 uses the Table 14 values to quote an 87% (2q) and 98.8% (3q) degradation under depolarizing noise. Both tables are described as medians over the same 30-state protocol, so the discrepancy cannot be attributed to a different sampling scheme. Please reconcile these numbers; the calibration atlas in Section 3 is a central contribution and cannot contain two mutually inconsistent versions of the same quantity.","section":"Tables 2, 3, 14, 15; §5.4"},{"comment":"The 'reversal' of the 2-qubit ranking is not supported by the controlled experiments. The abstract and Section 4.7 compare the 2-qubit result of Table 7 (Hermitian direct above Bloch) with the 3-qubit result of Table 9 (Bloch above Hermitian direct), but Section 4.5 and Appendix E state that Table 7 used parameterization-specific learning rates and extra tuning, and Table 23 shows that at matched learning rate with no CFG, Bloch already outperforms Hermitian direct at every shot level at 2 qubits (+0.26 to +0.38). Thus the ranking does not reverse under a controlled comparison; what changes is the size of the Bloch advantage (roughly +0.38 at 300 shots at 2q vs +0.51 at 3q). The abstract and the Section 3.8 decision-tree text should be rewritten to describe a growing advantage rather than a reversal, and the Table 7 result should not be presented as the 2-qubit baseline.","section":"Abstract; §4.4; §4.7; Appendix E, Table 23"}],"minor_comments":[{"comment":"The running header states 'Accepted in Quantum 2017-05-09' for a manuscript dated 2026; this is impossible and should be corrected or removed.","section":"Running header"},{"comment":"Figures 1–4 are referenced in Sections 3.6–3.10 but no figure images appear in the supplied text; the published version must include them.","section":"§3.6–§3.10"},{"comment":"C.3 reports Bloch sigma_data,diag=0.1546 and Herm-direct sigma_data,diag=0.2341, while Appendix E says the per-coordinate values are sigma_data,diag=0.2341 and sigma_data,off=0.2183 for both parameterizations; please make the calibration values consistent.","section":"Appendix C.3 vs Appendix E"},{"comment":"Section 2.3 contains an unresolved cross-reference '(see §??)' immediately after the EDM noise-schedule discussion; fix the reference.","section":"§2.3"},{"comment":"The phrase '3 measurement repeats' appears only in the two-by-two cross-validation; the main evaluation protocol in §4.1 does not define repeats and should be harmonized.","section":"§4.5"}],"recommendation":"major_revision","confidential_remarks":"I am not aware of a novelty conflict, but the submission has several formatting artifacts (impossible header date, missing figures, unresolved cross-reference) that suggest the source was not prepared in final journal format; please verify the provenance before sending the manuscript back to the authors."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Shuangju Chang's design-space study is worth a look, but with a big caveat. The geometric framework—J^T J calibration of seven parameterizations—is a solid, useful contribution. The scaling tables are informative, the analytical Jacobians are a nice touch, and the paper is unusually honest about its own confounds, going out of its way to show that the 2-qubit Hermitian advantage was a learning-rate artifact. The matched-lr and global-sigma ablations are good practice.\n\nThe problem is the end-to-end validation. The headline 3-qubit result—Bloch 0.907 vs. Hermitian-direct 0.394 at 300 shots—rests on an evaluation protocol that is not actually specified. We never learn the measurement basis (Pauli? SIC-POVM?), how shots are distributed, or how measurement statistics are encoded as a conditioning vector. The measurement-consistency loss is set to 0.0, and that's about it. Without that, the ranking could be driven by the measurement protocol rather than parameterization geometry. The paper itself shows the ranking flips under changes to learning rate and CFG, so protocol details matter.\n\nThere's also a numerical inconsistency: Table 3 says 3-qubit Cholesky SDR is 1265x, but Table 14 reports 1793x at zero noise; and the 2-qubit values are 33x vs 99x. The tables disagree on the very quantities being calibrated. And at 3 qubits there is no classical tomography baseline to contextualize the diffusion results.\n\nThe core geometric claims—that Bloch is isotropic, that Cholesky degrades badly, that fix-trace is a good tradeoff—are independent of the end-to-end experiments and likely correct. The J^T J framework is a standard conditioning object newly applied. But the central empirical finding as stated is not testable from the manuscript as written, so I'd treat the performance ranking as provisional.\n\nWho's this for? Practitioners working on generative QST or anyone choosing a parameterization for constrained diffusion. The calibration tables alone are worth a skim. But I wouldn't rely on the 3-qubit fidelity numbers until the measurement model is disclosed.\n\nMy recommendation: send it to review, but expect major revision. The framework is worth refereeing; the empirical section needs to be either completed or substantially weakened. The authors need to specify the measurement protocol, reconcile the SDR tables, and ideally add a classical baseline.","headline":"Useful geometric calibration study, but the headline 3-qubit ranking is not reproducible because the measurement model is never specified—and the SDR tables disagree with each other.","tokens_in":22746,"tokens_out":5137,"would_cite":false,"duration_ms":44105,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A parameterization's geometric conditioning does not predict end-to-end diffusion quantum state tomography quality, and bounded representations win at three qubits.","keywords":["diffusion quantum state tomography","density matrix parameterization","Jacobian Gram matrix","geometric conditioning","projection loss","Bloch/Gell-Mann representation","Cholesky parameterization","noise schedule calibration"],"falsifier":"Train the three 3-qubit models under two or more explicitly specified informationally complete measurement schemes (e.g., Pauli 6-basis measurements with different shot allocations, and a SIC-POVM) with identical training protocols; if Hermitian direct outperforms Bloch at 300 shots under any such scheme, the claimed geometric explanation is incomplete. A cheaper check: compute the correlation between each model's raw-output projection cost F_proj − F_raw and its end-to-end fidelity across shot levels; the paper's mechanism predicts the projection cost ordering fully explains the fidelity ordering.","tokens_in":21587,"feed_emoji":"⚛️","tokens_out":8658,"duration_ms":67968,"temperature":0.7,"pith_summary":"This paper establishes a design-space principle for diffusion-based quantum state tomography (QST): the coordinate system used to represent a density matrix is a first-order design axis, not an implementation detail. It introduces a geometric toolbox based on the Jacobian Gram matrix J⊤J, calibrates seven parameterizations at 2 and 3 qubits, and trains identical diffusion models on the three most distinct ones. The headline result is that local conditioning does not predict reconstruction quality: at 3-qubit scale the near-isotropic Hermitian direct parameterization (κ = 2.0×) scores 0.394 fidelity at 300 shots, worse than Cholesky (κ = 27×) at every shot level, while the bounded Bloch/Gell-Mann parameterization reaches 0.907. The paper's explanation is that unbounded coordinates routinely leave the set of valid density matrices, and the ensuing projection destroys measurement information, whereas Bloch's maximally mixed state sits at the center of its valid region. A sympathetic reader should care because the field's default (Cholesky) may be the wrong baseline, and the paper offers a principled selection rule that favours bounded or constraint-preserving representations as system size grows.","feed_headline":"Isotropy alone fails at 3-qubit quantum state tomography","feed_subtitle":"A bounded representation hits 0.907 fidelity where a near-perfect isotropic one stops at 0.394.","key_machinery":"The load-bearing object is the Jacobian Gram matrix G = J^T J, where J = ∂ρ/∂y maps infinitesimal coordinate changes to changes in the density matrix. Its eigenvalues define two diagnostics: the spectral dynamic range κ_spec = λ_max/λ_min, measuring overall coordinate isotropy, and the diagonal anisotropy κ_diag = max_i G_ii / min_i G_ii, measuring per-coordinate scale variation. These two numbers are used to rank seven parameterizations (Cholesky, Hermitian direct, trace-norm, fix-trace, Bloch/Gell-Mann, exponential map, log-Cholesky) and to derive selection heuristics, but the paper's own end-to-end experiments show the metrics are necessary yet not sufficient: projection behaviour at the valid-domain boundary, not local conditioning, governs the final fidelity.","core_discovery":"The central discovery is that geometric conditioning alone does not predict end-to-end performance in diffusion QST, because the projection that enforces physical constraints acts as a non-linear, many-to-one information bottleneck. At 3-qubit scale, Hermitian direct (κ_spec = 2.0×) performs worse than Cholesky (κ_spec = 27×) at all shot levels, a 13.5× isotropy advantage turning into a fidelity disadvantage of up to +0.51; the Bloch/Gell-Mann representation (κ = 1.0×) achieves 0.907 while Hermitian direct reaches only 0.394 at 300 shots. The geometric mechanism identified is that the positive semidefinite constraint couples diagonal and off-diagonal coordinates through |ρ_ij|² ≤ ρ_ii ρ_jj, a coupling an unconstrained model cannot respect, so its outputs violate the constraint and the projection annihilates spectral components carrying the measurement signal. In the Bloch representation the maximally mixed state lies at the centre of a bounded valid ball, providing a buffer against such violations; this also explains why training validation fidelity, measured before projection, reverses the ranking (Hermitian direct 0.799 vs Bloch 0.451) relative to end-to-end reconstruction (0.394 vs 0.907).","pith_inferences":["We infer that the projection-loss mechanism will strengthen at 4+ qubits: the paper's pure geometric calibration shows Cholesky degrading to κ~10³ and log-Cholesky to κ~10⁹, and the codimension of the valid manifold grows with d, so unbounded parameterizations have more directions in which to violate the PSD constraint.","We infer the same design-space logic applies to other constrained-manifold generative problems, e.g., correlation matrices, SPD tensors in imaging, and quantum process tomography; a bounded or constraint-satisfying coordinate system should systematically beat an unbounded one whenever a projection step is applied to enforce the constraint.","We infer that the paper's comparison would be hardened by a fully specified measurement model; if a different informationally complete POVM or shot allocation changed the ordering, the geometric story would need revision, since the measurement posterior interacts with the parameterization through the consistency loss."],"forward_implications":["Parameterization choice should be a reported and controlled experimental variable in diffusion QST; at 3 qubits, the best and worst parameterizations differ by up to 0.51 fidelity at 300 shots.","The community default Cholesky parameterization becomes a poor choice as qubit count grows (κ 33×→1265×), and its 3-qubit end-to-end fidelity (0.535 at 300 shots) trails Bloch (0.907).","Bounded-domain representations (Bloch/Gell-Mann) and constraint-preserving ones (fix-trace) are recommended for n≥3 qubits; the paper's decision tree guides when automatic constraint satisfaction is preferred over isotropy.","Training/validation fidelity in parameter space is not a reliable proxy for reconstruction quality; the projection step must be included, otherwise rank reversals (0.799 training vs 0.394 reconstruction for Hermitian direct) appear paradoxical.","Per-coordinate noise-schedule calibration (σ_data ∝ √G_ii) is the correct extension of EDM-style schedules once parameterization anisotropy is strong."],"supporting_citations":[{"why":"the variational generative QST work cited as the source of the field's default Cholesky parameterization","marker":"[1]"},{"why":"supplies the diffusion QST architecture that this study inherits and challenges for its default Cholesky parameterization","marker":"[2]"},{"why":"the alternative diffusion QST baseline, also Cholesky-based, that the design-space comparison extends","marker":"[3]"},{"why":"documents the fix-trace parameterization that the paper recommends as the best conditioning–constraint tradeoff","marker":"[6]"},{"why":"provides the EDM diffusion framework, noise schedule, and training setup used in all end-to-end experiments","marker":"[7]"},{"why":"the classical maximum-likelihood QST baseline against which diffusion reconstructions are compared","marker":"[9]"},{"why":"defines the generalized Bloch/Gell-Mann representation for n-level systems, the parameterization that wins at 3-qubit scale","marker":"[25]"}],"fun_headline_variants":["Isotropy doesn't predict QST performance at 3 qubits","Bounded beats isotropic in 3-qubit quantum tomography","Worse-conditioned wins: QST ranking flips at 3 qubits","Projection loss reverses QST ranking at 3 qubits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The end-to-end fidelity rankings presuppose a specific measurement protocol—POVM, shot allocation, and the computation of the measurement-consistency loss—that the paper never specifies, so the observed ordering could be partly a property of that protocol rather than pure parameterization geometry.","fun_headline_variants_meta":{"raw":{"variants":["Isotropy doesn't predict QST performance at 3 qubits","Bounded beats isotropic in 3-qubit quantum tomography","Worse-conditioned wins: QST ranking flips at 3 qubits","Projection loss reverses QST ranking at 3 qubits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000537,"raw_usage":{"total_tokens":2638,"prompt_tokens":1067,"completion_tokens":1571,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":683,"completion_tokens_details":{"reasoning_tokens":1495}},"tokens_in":683,"tokens_out":1571,"duration_ms":10811,"temperature":1.0,"reasoning_tokens":1495,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:14:31.850803+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the three 3-qubit models under two or more explicitly specified informationally complete measurement schemes (e.g., Pauli 6-basis measurements with different shot allocations, and a SIC-POVM) with identical training protocols; if Hermitian direct outperforms Bloch at 300 shots under any such scheme, the claimed geometric explanation is incomplete. A cheaper check: compute the correlation between each model's raw-output projection cost F_proj − F_raw and its end-to-end fidelity across shot levels; the paper's mechanism predicts the projection cost ordering fully explains the fidelity ordering.","supporting_citations":[{"cited_title":"The diagonal of the multiplihedra and the tensor product of A-infinity morphisms","cited_arxiv_id":"2206.05566","evidence_quote":"the variational generative QST work cited as the source of the field's default Cholesky parameterization"},{"cited_title":"Denoising diffusion models for quantum state tomography","cited_arxiv_id":null,"evidence_quote":"supplies the diffusion QST architecture that this study inherits and challenges for its default Cholesky parameterization"},{"cited_title":"Quan- tum state tomography with regularized lin- ear estimation","cited_arxiv_id":null,"evidence_quote":"documents the fix-trace parameterization that the paper recommends as the best conditioning–constraint tradeoff"},{"cited_title":"Elucidating the design space of diffusion-based generative models (EDM).Advances in Neural Information Processing Systems (NeurIPS), 35, 2022","cited_arxiv_id":null,"evidence_quote":"provides the EDM diffusion framework, noise schedule, and training setup used in all end-to-end experiments"},{"cited_title":"Efficient method for comput- ing the maximum likelihood quantum state from measurements with additive Gaus- sian noise","cited_arxiv_id":null,"evidence_quote":"the classical maximum-likelihood QST baseline against which diffusion reconstructions are compared"},{"cited_title":"Bloch vector forn-level systems","cited_arxiv_id":null,"evidence_quote":"defines the generalized Bloch/Gell-Mann representation for n-level systems, the parameterization that wins at 3-qubit scale"}],"review_version":2}