{"id":"bb13004e-e672-4207-b234-263b4f6c1e87","arxiv_id":"2509.10437","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Two pure qubit states and one mixed state suffice to prove that maximally psi-epistemic ontological models cannot explain a three-outcome guessing game called Quantum Gambling.","lead":"The paper introduces a three-outcome gambling game on two qubit states and shows that no maximally psi-epistemic hidden-variable model can reproduce its quantum statistics. The result gives a minimal no-go proof using just two pure qubit states and one equal mixture, and it also implies a quantum communication advantage with two qubit messages.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The no-go proof depends on the preparation-convexity identity Eq. (15) for the equal mixture rho; if that identity is an extra assumption rather than a consequence of the ontological-models framework, the theorem's scope is narrower than stated.","rationale":"The reader identified Eq. (15) as the weakest assumption; I agree that this is the most load-bearing point, because it is the only step in the proof of Theorem 2 that goes beyond the standard axioms of ontological models. However, the paper's communication scenario explicitly defines the equal mixture as the random-mixture preparation, which would make Eq. (15) valid for that specific preparation procedure even in a preparation-contextual model. The concern therefore is not that the bound is necessarily wrong, but that the theorem's statement is broader than what the proof establishes unless the random-mixture preparation is explicitly fixed. This is a conditionality issue, not a refutation: if the authors clarify that rho is the random-mixture preparation, or if [19] provides a proof of Eq. (15) from the standard axioms, the no-go result stands. The numerical certification of Theorem 3 is also a gap for the maximality claim, but the existence of the no-go only requires a positive B_Q for a specific pair of states, which the seesaw lower bound already provides. Thus the reader's CONDITIONAL verdict remains appropriate: the paper should be revised to either prove or explicitly assume preparation convexity, and to provide reproducible numerical artifacts for the claimed maximum. I do not see a reason to move the verdict to REJECT or ACCEPT on the basis of this concern alone.","tokens_in":13895,"tokens_out":30845,"duration_ms":254486,"concrete_test":"Re-derive Lemma 1 from Eq. (8) without invoking Eq. (15), attempting to bound T2 and T3 directly using the original definition tilde mu3 = (alpha+beta)(mu(lambda|psi1)+mu(lambda|psi2)). If inequality (29) cannot be obtained without the convex-combination substitution, then Eq. (15) is load-bearing and Theorem 2 must be restated as conditional on preparation convexity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central no-go theorem (Theorem 2) is proven via Lemma 1 in the Appendix. The proof of Lemma 1 uses Eq. (15), asserting that mu(lambda|rho) = (mu(lambda|psi1)+mu(lambda|psi2))/2 for rho = (|psi1><psi1|+|psi2><psi2|)/2. This identity is used to rewrite the three-answer game's third distribution tilde mu3(lambda) = (alpha+beta)(mu(lambda|psi1)+mu(lambda|psi2)) as 2(alpha+beta) mu(lambda|rho), which is essential for bounding the terms T2 and T3 in the generalized epistemic overlap (Eq. 8) by weighted distinguishability quantities involving rho (Eqs. 31-32). In the standard ontological-models framework, preparation procedures with the same density operator can have different epistemic states if the model is preparation-contextual; the convex combination is forced only for the specific procedure 'toss a fair coin and prepare psi1 or psi2'. The paper states Eq. (15) as a general fact and cites [19] for it. If [19] assumes preparation convexity rather than deriving it from the standard axioms, then Theorem 2 holds only for models satisfying that extra condition, and the conclusion 'no maximally psi-epistemic model' is narrower than claimed. The communication scenario (main text before Eq. 17) does define x=3 as the random mixture, which would validate Eq. (15) for that specific rho preparation, but the theorem is stated for rho as a density operator and the proof applies Eq. (15) globally, so the scope of the no-go result is not made precise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a penalized three-outcome distinguishability game ('Quantum Gambling') between two quantum states, defines generalized quantum and epistemic overlaps from the optimal reward in this game, and proves (Theorem 1) that the quantum value is upper bounded by the ontic-state value. Lemma 1 and Theorem 2 provide a lower bound B_Q on the difference between the generalized quantum and epistemic overlaps, expressed in terms of quantum distinguishability quantities. Theorem 3 claims, via numerical semidefinite programming, that the maximum of B_Q over all pairs of states is approximately 0.0639, achieved by a pair of qubit states at alpha approximately 0.7124. The authors conclude that no maximally psi-epistemic model (in their generalized sense) can explain the gambling statistics, and they relate the gap to a quantum advantage in a constrained communication scenario.","tokens_in":14220,"tokens_out":21076,"duration_ms":129781,"significance":"If the results hold, the paper provides the first no-go theorem for maximally psi-epistemic models using only two pure qubit states, which is minimal both in dimension and in the number of preparations. The proof of Theorem 1 is clean and the derivation of the operational bound in Theorem 2 is elegant given the preparation-convexity condition. The numerical finding that qubits achieve the maximum gap is falsifiable and potentially testable. However, the generality of the no-go statement is limited by the unstated convexity assumption behind Eq. (15), and the maximality claim rests on a numerical hierarchy that is not independently certified in the manuscript.","major_comments":[{"comment":"The proof of Lemma 1 uses the identity mu(lambda|rho) = (mu(lambda|psi1)+mu(lambda|psi2))/2 as a general fact, citing reference [19]. This equality is guaranteed for the specific preparation procedure that generates rho as an equal random mixture, but it is not a consequence of the standard ontological models framework for an arbitrary preparation of the density operator rho. Since Theorem 2 is stated for a density operator rho without specifying its preparation, the no-go theorem is presently proven only for models satisfying this preparation-convexity assumption (or, equivalently, for the communication scenario in which x=3 is explicitly defined as the random mixture). The statement that 'no maximally psi-epistemic model can explain Quantum Gambling with qubits' is therefore narrower than the proof supports. Please restate the theorem's scope explicitly, or reformulate it directly for the communication task where Eq. (15) is operationally guaranteed by the protocol.","section":"Theorem 2, Lemma 1 (Appendix), Eq. (15)"},{"comment":"The maximality claim, namely that B_Q(alpha) is achieved by a pair of qubit states and that max_alpha B_Q(alpha) is approximately 0.0639, is justified only by the matching of seesaw SDP lower bounds and level-2 tracial noncommuting polynomial hierarchy upper bounds 'up to machine precision'. No code, certificates, or explicit convergence argument are provided, and the level-2 hierarchy is not exact in general. The lower-bound part of the seesaw computation, which evaluates B_Q for a specific qubit pair, is sufficient to establish a positive gap and hence a no-go result for those states, but the stronger statement that qubits achieve the maximum gap requires a rigorous upper bound. Please provide reproducible code and formal error bounds, or explicitly weaken the maximality statement to a conjecture.","section":"Theorem 3 and Appendix: numerical upper bounds, Eqs. (36)-(43)"}],"minor_comments":[{"comment":"The displayed objective function for the tracial noncommuting polynomial optimization has a missing '+' before the (1-alpha) term and the subscript 'I' in [Gamma_rho2]_{I,M^3_1} should be '1'; the expression would benefit from a cleaner typeset.","section":"Eq. (42)"},{"comment":"The theorem states the interval alpha in (approximately 0.49, 1] but the origin of the approximate threshold 0.49 is not explained; a precise bound or a note on how it is obtained would improve clarity.","section":"Theorem 3"},{"comment":"The paper uses the term 'maximally psi-epistemic' to refer to equality of the generalized overlaps defined through the gambling game, which differs from the standard definition based on distinguishability (alpha=beta=0). A sentence explicitly clarifying this generalization and its relation to the standard notion would avoid potential misinterpretation.","section":"Introduction and Discussion"},{"comment":"If reference [19] itself assumes preparation convexity rather than deriving it from the ontological-models axioms, this should be stated in the text so that readers understand that Eq. (15) is an additional condition.","section":"Eq. (15) and reference [19]"},{"comment":"The phrase 'experimentally robust' may overstate the case given the relatively small gap of approximately 0.0639; a comment on the required experimental precision would be appropriate.","section":"Abstract and Discussion"},{"comment":"The regions Lambda'_3, Lambda''_3, and Lambda'''_3 are used in the caption but are not defined before the figure; adding a sentence with their definitions would improve readability.","section":"Figure 1 caption"}],"recommendation":"major_revision","confidential_remarks":"The analytical structure of Theorem 2 is sound under an explicit preparation-convexity assumption, but the theorem is presently stated more broadly than the proof supports. The numerical maximality claim of Theorem 3 would be significantly strengthened by depositing code or providing certificates; without these, the 'two qubits achieve the maximum' result remains an unverified numerical observation. The paper is likely to interest the quantum foundations audience, but the authors should be asked to either prove the no-go under standard ontological-models axioms or carefully scope the claim to preparation-convex models or to the specific communication protocol."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my read. The genuinely new thing is the Quantum Gambling game and the generalized epistemic overlap built from it. For the parameter region alpha > (1-beta)/2, the overlap is not just weighted distinguishability, and the paper shows that two pure qubit states plus their equal mixture give a positive gap between the quantum and epistemic overlaps. If correct, this is the first constraint on maximally psi-epistemic models for single qubits, where the Kochen-Specker model had previously blocked progress. Theorems 1 and 2 are analytically clean given the standard ontological-models framework, and the operational lower bound B_Q is a nice construction.\n\nThe soft spots are in proportion. The main one is Eq. (15), the preparation-convexity condition. The proof of Lemma 1 uses it to rewrite the game's third distribution as 2(alpha+beta) mu(rho). That is not a consequence of the usual ontological-models axioms; it is an extra assumption about how the model represents mixed preparations. A preparation-contextual model could assign a different epistemic state to the same density operator, and then the bound B_Q need not hold. The paper states (15) as if it were automatic and cites prior work, but it should be framed as a restriction on the class of models being ruled out. The title and abstract say \"no maximally psi-epistemic model\" without that qualification. The communication scenario in the main text does define x=3 as a random mixture, so for that particular procedure Eq. (15) is just probability calculus; but the theorem itself is stated for rho as a density operator and the proof applies the identity globally. That scope mismatch is fixable, but it is real.\n\nSecond, Theorem 3 is numerical. The authors use a seesaw SDP for lower bounds and a level-2 tracial NPA hierarchy for upper bounds, and the two match to machine precision over a range of alpha. No code or certificates are included, and the appendix has typos (e.g., Eq. 34). Matching bounds is good evidence, but it is not the same standard of proof as the analytic theorems. The paper should either ship the code, or label Theorem 3 as a numerical result that would need an analytic proof or certified bounds for full rigor.\n\nThe citation pattern is fine; prior no-go results are acknowledged, and the point about ratio vs difference of overlaps is well-taken. This is a serious paper with a novel idea and mostly solid analytic core under an explicit assumption. It deserves a serious referee. I would send it to peer review, though not accept it as is. The authors need to fix the no-go scope and make the numerics reproducible.","headline":"Novel two-qubit gambling game gives a real constraint on psi-epistemic models, but the no-go claim is broader than the proof and the numerical part needs artifacts.","tokens_in":14793,"tokens_out":6579,"would_cite":true,"duration_ms":56851,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that no maximally $\\psi$-epistemic ontological model can reproduce the statistics of a three-outcome betting game played with two pure qubit states and their equal mixture, closing the qubit loophole left by existing…","keywords":["quantum foundations","psi-epistemic models","ontological models","epistemic overlap","quantum gambling","qubit","non-projective POVM","tracial noncommuting polynomial optimization"],"falsifier":"Check the next level of the tracial noncommuting polynomial hierarchy for the states $|0\\rangle$ and $\\cos(\\pi/3)|0\\rangle+\\sin(\\pi/3)|1\\rangle$ at $\\alpha=0.7124$, $\\beta=1$; if the upper bound on $B_Q$ falls below the claimed 0.0639, Theorem 3's numerical certification fails. Alternatively, construct an explicit ontological model satisfying the preparation-convexity condition whose generalized epistemic overlap equals the generalized quantum overlap at these parameters, which would refute the no-go theorem directly.","tokens_in":13667,"feed_emoji":"🎲","tokens_out":10366,"duration_ms":83809,"temperature":0.7,"pith_summary":"This paper sets out to show that the strongest form of $\\psi$-epistemic interpretation, in which the indistinguishability of two quantum states is fully explained by overlapping distributions over hidden variables, fails even for the smallest quantum system: a single qubit. It does this by designing a three-answer betting game, Quantum Gambling, in which guessing right earns a reward, guessing wrong costs a penalty, and a third abstention answer pays a fixed reward $\\alpha$. The paper generalizes the standard quantum and epistemic overlaps to this game and proves that their difference is bounded below by an operational quantity $B_Q$. It then shows numerically that two pure qubit states and their equal mixture maximize $B_Q$, reaching about $0.0639$ at $\\alpha\\approx0.7124$, $\\beta=1$. If correct, this closes the qubit loophole: no maximally $\\psi$-epistemic model that respects preparation convexity can explain the gambling statistics, and the gap is operational enough to be probed experimentally.","feed_headline":"Gambling with two qubits rules out maximally epistemic quantum models","feed_subtitle":"A three-outcome betting game exposes a gap between quantum and epistemic overlap that just two qubits can beat.","key_machinery":"The central object is the generalized epistemic overlap $\\omega_\\Lambda(\\psi_1,\\psi_2;\\alpha,\\beta)$, built from three unnormalized distributions $\\tilde{\\mu}_1=(1+\\beta)\\mu(\\lambda|\\psi_1)$, $\\tilde{\\mu}_2=(1+\\beta)\\mu(\\lambda|\\psi_2)$, and $\\tilde{\\mu}_3=(\\alpha+\\beta)(\\mu(\\lambda|\\psi_1)+\\mu(\\lambda|\\psi_2))$, combined through min-integrals $T_1,T_2,T_3,T_4$. This quantity is the best average reward available to a gambler who knows the ontic state, and the corresponding quantum value $S^{\\mathrm{Gam}}_Q$ is computed by a semidefinite program. The argument's load-bearing step is Theorem 2, which lower-bounds the quantum-epistemic gap by $B_Q$; its proof uses the preparation-convexity identity $\\mu(\\lambda|\\rho)=\\frac12(\\mu(\\lambda|\\psi_1)+\\mu(\\lambda|\\psi_2))$ for the equal mixture. The numerical part maximizes $B_Q$ by alternating semidefinite programs for dimension-dependent lower bounds and a level-2 tracial noncommuting polynomial hierarchy for dimension-independent upper bounds; wherever the two match, the value is certified.","core_discovery":"The central claim is that no maximally $\\psi$-epistemic ontological model can explain Quantum Gambling with two pure qubits. The proof works by defining generalized overlaps $\\omega_Q(\\psi_1,\\psi_2;\\alpha,\\beta)$ and $\\omega_\\Lambda(\\psi_1,\\psi_2;\\alpha,\\beta)$ from the best average reward in the game, then showing $\\omega_Q-\\omega_\\Lambda \\ge B_Q(\\psi_1,\\psi_2,\\rho;\\alpha,\\beta)$, where $\\rho$ is the equal mixture. The gap $B_Q$ is strictly positive for a range of $\\alpha$: its maximum value is approximately $0.0639$, attained for $|\\psi_1\\rangle=|0\\rangle$ and $|\\psi_2\\rangle=\\cos(\\pi/3)|0\\rangle+\\sin(\\pi/3)|1\\rangle$ at $\\alpha\\approx0.7124$ with $\\beta=1$. Because the gap is positive, the epistemic overlap cannot fully account for the quantum statistics, so any model that tries to do so must be non-maximally $\\psi$-epistemic. The paper also reads the same gap as a communication advantage: two qubit messages outperform classical messages in a constrained communication task with bounded gambling reward.","pith_inferences":["The no-go theorem is conditional on the preparation-convexity condition of Eq. (15); dropping that condition restores a possible escape route for maximally $\\psi$-epistemic models, so the result is most naturally read as ruling out convexity-respecting ontological models.","The same three-answer gambling game with tunable $(\\alpha,\\beta)$ gives a quantitative scale for 'how epistemic' a model can be: measuring the gambling success of a physical implementation would place an upper bound on the epistemic overlap actually realized.","A natural next test is to apply Quantum Gambling to higher-dimensional state pairs; if the gap grows toward the theoretical maximum $\\omega_{\\max}=1-2\\alpha-\\beta+\\min(1+\\beta,2(\\alpha+\\beta))$, the approach could produce a dimension-independent no-go theorem for all maximally $\\psi$-epistemic models."],"forward_implications":["For every $\\alpha\\in(\\approx0.49,1]$ with $\\beta=1$, the qubit pair achieves the maximum of $B_Q$, so the no-go theorem applies to an entire interval of game parameters, not just a single point.","Since two pure qubit states and their equal mixture suffice, the minimum resources for a no-go theorem of this kind are two states in dimension two, improving on earlier constructions that needed many four-dimensional states.","The gap $\\omega_Q-\\omega_\\Lambda\\ge B_Q$ is an operational lower bound, so an experiment that estimates the success probabilities of Quantum Gambling can certify the absence of maximally $\\psi$-epistemic models without needing to characterize the hidden-variable distributions.","In the communication reading, the same construction shows that two qubit messages achieve a quantum advantage over classical messages under a bounded-reward gambling constraint."],"supporting_citations":[{"why":"Defines the ontological models framework and the epistemic states and response functions used throughout.","marker":"[6]"},{"why":"Provides the maximally $\\psi$-epistemic model for qubit projective measurements that the paper's game aims to rule out.","marker":"[13]"},{"why":"Provides an earlier no-go result that required many four-dimensional states, the baseline the two-qubit result improves on.","marker":"[10]"},{"why":"Supplies the see-saw semidefinite and tracial noncommuting polynomial optimization techniques used to bound $B_Q$.","marker":"[14]"},{"why":"Supplies the level-2 tracial moment-matrix hierarchy used for the dimension-independent upper bounds.","marker":"[15]"},{"why":"Shows every two-outcome POVM can be simulated projectively, motivating the need for a three-outcome game to constrain qubit models.","marker":"[18]"},{"why":"Supplies the preparation-convexity condition that Theorem 2's proof invokes for the equal mixture.","marker":"[19]"}],"fun_headline_variants":["Two-qubit gambling rules out maximally epistemic models","Max ψ-epistemic fails two-qubit betting test","Quantum betting gap: two qubits expose epistemic limits","Two qubits beat any maximally ψ-epistemic model","Gambling with two qubits kills ψ-epistemic explanations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a mixed preparation's epistemic state is the convex mixture of the epistemic states of its preparations, $\\mu(\\lambda|\\rho)=\\frac12(\\mu(\\lambda|\\psi_1)+\\mu(\\lambda|\\psi_2))$; if a model is preparation-contextual and violates this equality, the no-go theorem's bound need not apply.","fun_headline_variants_meta":{"raw":{"variants":["Two-qubit gambling rules out maximally epistemic models","Max ψ-epistemic fails two-qubit betting test","Quantum betting gap: two qubits expose epistemic limits","Two qubits beat any maximally ψ-epistemic model","Gambling with two qubits kills ψ-epistemic explanations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000681,"raw_usage":{"total_tokens":3073,"prompt_tokens":908,"completion_tokens":2165,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":2083}},"tokens_in":524,"tokens_out":2165,"duration_ms":15267,"temperature":1.0,"reasoning_tokens":2083,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:54:41.299008+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the next level of the tracial noncommuting polynomial hierarchy for the states $|0\\rangle$ and $\\cos(\\pi/3)|0\\rangle+\\sin(\\pi/3)|1\\rangle$ at $\\alpha=0.7124$, $\\beta=1$; if the upper bound on $B_Q$ falls below the claimed 0.0639, Theorem 3's numerical certification fails. Alternatively, construct an explicit ontological model satisfying the preparation-convexity condition whose generalized epistemic overlap equals the generalized quantum overlap at these parameters, which would refute the no-go theorem directly.","supporting_citations":[{"cited_title":"Harrigan and R","cited_arxiv_id":null,"evidence_quote":"Defines the ontological models framework and the epistemic states and response functions used throughout."},{"cited_title":"Kochen and E","cited_arxiv_id":null,"evidence_quote":"Provides the maximally $\\psi$-epistemic model for qubit projective measurements that the paper's game aims to rule out."},{"cited_title":"Oszmaniec, L","cited_arxiv_id":null,"evidence_quote":"Shows every two-outcome POVM can be simulated projectively, motivating the need for a three-outcome game to constrain qubit models."}],"review_version":1}