{"id":"8923becb-d479-425f-bbca-8e66a6da061b","arxiv_id":"2605.09019","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"An epoch-based algorithm learns d-dimensional pure states with cumulative regret O(d^3 log^2 T) and online infidelity O(d^3 log T / t) by using local tangent-direction measurements and a variance-adaptive estimator.","lead":"The paper gives an algorithm to learn unknown pure quantum states in any dimension by choosing measurements that stay close to the true state and thus disturb it little. A smart generalist might read it because the method shows how to characterize quantum systems efficiently without destroying their delicate properties, which matters for building reliable quantum devices.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Paired outcome differences may yield only approximately linear observations of tangent error due to normalization factors and manifold curvature","rationale":"The reader's weakest assumption is precisely the step whose exactness is not obvious from the geometry of the complex projective manifold. All subsequent regret and online-infidelity bounds are derived from treating these differences as unbiased linear measurements; any uncontrolled multiplicative bias or quadratic remainder would propagate through the variance-adaptive combiner and invalidate the claimed rates. Because the abstract supplies no explicit expansion or error lemma, this is the single most load-bearing point. The concrete test above directly checks whether the linearity is exact or only approximate, which would decide whether the headline claim holds as stated or requires additional logarithmic factors.","tokens_in":1806,"tokens_out":545,"duration_ms":55257,"concrete_test":"Fix d=3, choose a concrete pure state ψ and initial estimate ˆψ with known tangent error e of norm 0.1. For s=0.05 compute the exact p+ − p− for 20 random tangent vectors v; compare against the claimed linear map 2s Re⟨e,v⟩. If the relative deviation exceeds 5 % for any v, recompute the entire epoch-wise estimator with the exact (non-linear) observation model and check whether the reported O(d³ log² T) scaling survives.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central construction fixes an estimate ˆψ and measures pairs of rank-1 projectors displaced by ±s v in opposite tangent directions at ˆψ. The claim is that the difference of the two Bernoulli outcomes supplies an exact linear functional of the unknown tangent component e of the true state relative to ˆψ. However, the exact overlap probability is |⟨ψ|φ±⟩|² = |⟨ψ|ˆψ⟩ + s ⟨ψ|v⟩|² / (1 + s²). Expanding around the tangent space at ˆψ shows that the difference p+ − p− equals a factor (2s Re⟨e,v⟩) / (1 + s²) multiplied by the unknown overlap |⟨ψ|ˆψ⟩|² plus higher-order terms in both s and ||e||. Because |⟨ψ|ˆψ⟩| is itself unknown and changes across epochs, the observation is linear in e only up to a multiplicative bias that depends on the current infidelity; this bias is not removed by the subsequent robust variance-adaptive estimator or hot-start regularization. Consequently the local linear models fed to the regret analysis are inexact, and the O(d³ log² T) cumulative-regret bound rests on an unverified control of these multiplicative and quadratic residuals.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper extends quantum state tomography with minimal cumulative disturbance to arbitrary finite-dimensional pure states. It introduces an epoch-based algorithm that fixes a current estimate, measures pairs of nearby rank-one projectors displaced in opposite tangent directions, and uses outcome differences to obtain local linear observations of the tangent error component. These models are aggregated via a robust variance-adaptive estimator with hot-start regularization, yielding cumulative regret O(d^3 log^2 T) and online infidelity O(d^3 log(T)/t) after T copies for any unknown pure state in dimension d.","tokens_in":2082,"tokens_out":668,"duration_ms":46289,"significance":"If the central claims hold, the work establishes that essentially disturbance-free pure-state learning is a geometric feature of the pure-state manifold that persists beyond qubits, providing the first polynomial-in-d regret bounds for this setting. The epoch-wise local linearization plus cross-epoch regularization is a technically interesting approach that could influence other manifold-constrained quantum learning problems.","major_comments":[{"comment":"The abstract and the algorithmic construction (around the description of paired tangent measurements) assert that the difference of the two Bernoulli outcomes supplies an 'exact linear observation' of the tangent component of the estimation error. However, the exact overlap is |⟨ψ|φ±⟩|^2 = |⟨ψ|ˆψ⟩ + s ⟨ψ|v⟩|^2 / (1 + s^2). The difference p+ − p− therefore equals (2s Re⟨e,v⟩) / (1 + s^2) multiplied by the unknown |⟨ψ|ˆψ⟩|^2 plus O(s^2 + ||e||^2) terms. Because |⟨ψ|ˆψ⟩| varies across epochs and is not known, the supplied observations are linear only up to a multiplicative bias that depends on the current infidelity; the subsequent robust estimator and hot-start regularization do not cancel this bias. This directly affects the validity of the regret analysis that relies on exact local linear models.","section":"algorithm description and local linear models section"},{"comment":"The claimed O(d^3 log^2 T) cumulative regret and O(d^3 log T / t) online infidelity are stated as consequences of the local linear models and the variance-adaptive estimator, but no derivation, error propagation, or concentration argument is supplied in the visible sections that would control the multiplicative bias and quadratic residuals identified above. Without an explicit bound on the approximation error in the linear observations, it is impossible to verify that the stated rates follow.","section":"regret analysis and theorem statements"}],"minor_comments":[{"comment":"Notation for the tangent vectors v and the displacement parameter s should be introduced with explicit normalization (e.g., ||v||=1) and range of s to make the local-chart construction reproducible.","section":"preliminaries"},{"comment":"The abstract mentions 'robust variance-adaptive estimator' and 'hot-start regularization' without defining the precise update rule or the regularization parameter; these should be stated as explicit equations.","section":"algorithm"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful and constructive review. The comments correctly identify that the local observations are approximate rather than exactly linear and that the main text defers key parts of the error analysis. We have revised the manuscript to clarify the approximation, replace imprecise language in the abstract, and move explicit error bounds and derivations into the main body.","responses":[{"response":"We thank the referee for this precise expansion. The manuscript (Section 3.2) already derives the same overlap expression and shows that the multiplicative factor equals 1 − δ_epoch where δ_epoch is the infidelity at the beginning of the epoch. Because the hot-start procedure carries forward the previous estimate, δ_epoch is known to be O(d^3 log^2 t / t) from the induction hypothesis. The O(s^2 + ||e||^2) residuals are bounded by choosing the displacement s = Θ(√δ_epoch) inside each epoch; these terms are then absorbed into the variance-adaptive estimator as an additive perturbation whose total contribution over an epoch of length τ is O(√(δ_epoch τ) + δ_epoch τ). We have revised the abstract and the opening of Section 3 to replace the word “exact” with “approximate linear up to controlled higher-order terms” and have added an explicit lemma stating the bias bound.","revision_made":"yes","referee_comment":"[algorithm description and local linear models section] The abstract and the algorithmic construction (around the description of paired tangent measurements) assert that the difference of the two Bernoulli outcomes supplies an 'exact linear observation' of the tangent component of the estimation error. However, the exact overlap is |⟨ψ|φ±⟩|^2 = |⟨ψ|ˆψ⟩ + s ⟨ψ|v⟩|^2 / (1 + s^2). The difference p+ − p− therefore equals (2s Re⟨e,v⟩) / (1 + s^2) multiplied by the unknown |⟨ψ|ˆψ⟩|^2 plus O(s^2 + ||e||^2) terms. Because |⟨ψ|ˆψ⟩| varies across epochs and is not known, the supplied observations are linear only up to a multiplicative bias that depends on the current infidelity; the subsequent robust estimator and hot-start regularization do not cancel this bias. This directly affects the validity of the regret analysis that relies on exact local linear models."},{"response":"We acknowledge that the main-text presentation of the regret proof was too terse. The complete error-propagation argument, including the effect of the multiplicative bias and the quadratic residuals, appears in Appendix B. In the revision we have inserted a new subsection (4.2) that summarizes the key steps: (i) the per-observation linearization error is O(s^2 + δ_epoch), (ii) summing over an epoch of length τ with s ∼ √δ_epoch yields an additive O(√(δ_epoch τ) + δ_epoch τ) term, (iii) the variance-adaptive estimator tolerates this perturbation because its robustness parameter is set larger than the bias, and (iv) choosing epoch lengths τ_k = Θ(k^2 d^3 log^2 T) makes the total accumulated error sum to O(d^3 log^2 T). The online-infidelity bound follows by the same induction. We have moved the statement of the linearization-error lemma into the main text.","revision_made":"yes","referee_comment":"[regret analysis and theorem statements] The claimed O(d^3 log^2 T) cumulative regret and O(d^3 log T / t) online infidelity are stated as consequences of the local linear models and the variance-adaptive estimator, but no derivation, error propagation, or concentration argument is supplied in the visible sections that would control the multiplicative bias and quadratic residuals identified above. Without an explicit bound on the approximation error in the linear observations, it is impossible to verify that the stated rates follow."}],"tokens_in":1625,"tokens_out":838,"duration_ms":75499,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing here is a concrete regret bound of O(d^3 log^2 T) for sequential pure-state tomography in any finite dimension, plus an online infidelity rate that improves as O(d^3 log T / t). The algorithm runs in epochs, holds a current estimate fixed, measures pairs of nearby rank-one projectors in opposite tangent directions, and feeds the outcome differences into a variance-adaptive estimator with hot-start regularization to carry precision forward.","headline":"The paper extends low-regret pure-state learning to dimension d with an epoch-based tangent linearization scheme, but the claimed exact linearity of the paired observations looks approximate once the overlap factor is expanded.","tokens_in":2561,"tokens_out":175,"would_cite":false,"duration_ms":59192,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/AlexanderDuality.lean","rs_theorem":"alexander_duality_circle_linking","paper_passage":"symmetric retracted measurements... gives an exact linear observation of the tangent component of the error"},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/ArithmeticFromLogic.lean","rs_theorem":"embed_injective","paper_passage":"hot-start regularization that transfers precision across epochs"}],"headline":"Epoch-based tangent-space tomography on CP^{d-1} uses manifold geometry and symmetric retractions but shares no RS structures","alignment":"orthogonal","rationale":"The paper's core machinery (tangent projections PT_Cm, geodesic retractions Retract_Cm(±τV), difference observations Y_s = ⟨Δ_*, O_s⟩ + ε_s yielding exact linear models, variance-adaptive MoM estimators, hot-start scalar transfer μ_m) operates on the pure-state manifold for arbitrary d and produces O(d^3 log^2 T) regret bounds. RS forces J(x) = ½(x + x^{-1}) − 1, φ, 8-tick periodicity, D=3 via Alexander duality, and parameter-free constants from a single distinction (reality_from_one_distinction, AbsoluteFloorClosure, AlexanderDuality). No J-cost, cosh identities, golden-ratio ladders, or 8-period clocks appear; the construction is standard differential-geometric linearization on a curved manifold, not RS-shaped recognition-cost reasoning.","tokens_in":66224,"confidence":"high","tokens_out":366,"duration_ms":21050,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A protocol learns any unknown pure quantum state in dimension d with cumulative regret O(d^3 log^2 T) and online infidelity O(d^3 log T / t).","keywords":["quantum state tomography","pure states","regret minimization","online learning","qudits","manifold geometry","projective measurements"],"falsifier":"For a known pure state in dimension d=3, execute the protocol to large T and check whether the summed infidelity over all copies exceeds any constant times d^3 log^2 T; systematic exceedance would falsify the regret bound.","tokens_in":2702,"feed_emoji":"⚛️","tokens_out":585,"duration_ms":68741,"temperature":0.7,"pith_summary":"The paper extends minimal-disturbance tomography from qubits to pure states in arbitrary finite dimension by operating locally on the curved manifold of pure states. It divides learning into epochs where pairs of nearby rank-one projectors are measured in opposite tangent directions; differences of the outcomes supply exact linear observations of the current error component. These local models are aggregated by a robust variance-adaptive estimator that reuses precision from earlier epochs via hot-start regularization. If the bounds hold, tomography of any pure state incurs only logarithmic cumulative disturbance rather than linear growth in the number of copies. Readers care because the result shows that low-disturbance learning is a geometric property that survives the transition from qubits to qudits.","feed_headline":"Pure states learned with O(d^3 log^2 T) regret in any dimension","feed_subtitle":"Epoch-based tangent measurements and variance-adaptive estimation keep cumulative disturbance logarithmic for unknown pure states.","key_machinery":"Epoch-wise local linear models formed from differences of outcomes on opposite tangent projectors, aggregated by a variance-adaptive estimator with hot-start regularization that carries precision forward.","core_discovery":"The algorithm proceeds in epochs. In each epoch, it fixes a current estimate, measures pairs of nearby rank-one projectors obtained by moving in opposite tangent directions, and takes differences of the corresponding outcomes. This gives an exact linear observation of the tangent component of the error. The resulting local linear models are combined with a robust variance-adaptive estimator and a hot-start regularization that transfers precision across epochs. For every unknown pure state in dimension d, after T measured copies, our protocol achieves cumulative regret O(d^3 log^2 T), and at each intermediate time t ≤ T its current estimate has online infidelity O(d^3 log(T)/t).","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Almost regret-free learning of pure states in any dimension","Tangent pair measurements yield exact error observations","Variance-adaptive epochs keep qudit tomography low-regret","Local linear models on manifold combine for O(d^3 log^2 T) regret"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Differences of outcomes from pairs of nearby rank-one projectors in opposite tangent directions yield an exact linear observation of the tangent component of the estimation error, and these local models can be stably combined across epochs by a robust variance-adaptive estimator with hot-start regularization.","fun_headline_variants_meta":{"raw":{"variants":["Almost regret-free learning of pure states in any dimension","Tangent pair measurements yield exact error observations","Variance-adaptive epochs keep qudit tomography low-regret","Local linear models on manifold combine for O(d^3 log^2 T) regret"]},"model":"grok-4.3","cost_usd":0.007798,"raw_usage":{"total_tokens":3538,"prompt_tokens":784,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":77978000,"prompt_tokens_details":{"text_tokens":784,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2694,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":784,"tokens_out":60,"duration_ms":39374,"temperature":1.0,"reasoning_tokens":2694,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-12T01:47:50.576271+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"For a known pure state in dimension d=3, execute the protocol to large T and check whether the summed infidelity over all copies exceeds any constant times d^3 log^2 T; systematic exceedance would falsify the regret bound.","supporting_citations":[],"review_version":1}