{"id":"f64bc206-143a-4e1a-9c0b-49c052aa77ca","arxiv_id":"2412.16806","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using a designed linguistic schema and BERT probabilities, the paper reports 77,118 sheaf-contextual and 36.9 million CbD-contextual instances from Simple English Wikipedia, with Euclidean distance as the best statistical predictor.","lead":"The authors built a sentence template about two nouns with three adjective pairs, ran it through 52 million Wikipedia-derived examples using BERT, and found many that match the mathematical pattern of quantum contextuality. The result suggests that quantum-style reasoning tools might eventually be useful for language tasks, though the connection to real quantum advantage is speculative.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The CbD-contextual count is close to tautological for the PR-prism template and the sheaf-contextual count has no null baseline, so the large-scale evidence for natural-language contextuality is unestablished.","rationale":"The paper's bookkeeping is careful: the PR-prism support is fixed by the template, the CF/SF formulas are derived for that support, and the counts follow deterministically from the recorded BERT outputs. I would not object to the arithmetic. The problem is interpretive: the two headline numbers are presented as evidence that natural language is contextual, but neither number is benchmarked. The CbD number is very close to a mathematical consequence of the template and of BERT's calibration: whenever the three masked predictions are consistently biased toward one noun with probability below 1, Δ = 2|ε| < 2 and the instance is counted as CbD-contextual. That is the modal case in Figure 5. The sheaf number is more meaningful but has no null comparison; without a randomised pairing baseline we cannot tell whether 0.148% is above the rate expected from BERT's softmax geometry alone. The similarity analysis is further weakened by post-hoc selection and tiny R2. None of this proves the authors wrong, but it does mean the central claim as stated outruns the evidence. A permutation null is the cheapest decisive check; it should be run before accepting the 'natural language at scale' conclusion. The reader's conditional verdict already anticipates most of this, so I do not propose changing it.","tokens_in":21580,"tokens_out":16141,"duration_ms":150158,"concrete_test":"Run a permutation null on a random subsample of the dataset: for each selected noun pair, pair its three adjective contexts with a randomly chosen noun pair from the same frequency and shared-adjective strata, re-run BERT, and recompute the SF<1/6 and Δ<2 rates. If the sheaf-contextual rate is not significantly above the permuted baseline, the 77,118 count is a softmax/template artifact rather than evidence about the specific noun-adjective semantics. The CbD rate should also be compared to a synthetic null drawn from the empirical ε distribution.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing gap is the absence of any null model against which the two headline counts are interpreted. On the CbD side, the count is close to a tautology: the template fixes PR-prism support, and in the common case where the three BERT probabilities favour the same noun with the same non-extreme value ε, the CbD quantity is Δ = 2|ε| (Section 4), so Δ < 2 holds for every ε in (−1,1). BERT almost never assigns probability exactly 1, which explains the 71.1% CbD-contextual rate without any distinctive natural-language phenomenon. The sheaf-contextual count (77,118; 0.148%) is the only nontrivial evidence, but it is never compared to a null distribution over randomised noun-adjective pairings; 0.148% may simply be the rate at which BERT is within 1/6 of uniform on three arbitrary contexts (Section 6(b), criterion SF < 1/6). The later similarity analysis is also confounded: the similar-noun subset is selected after seeing the data and the regression R2 is 0.006 full and 0.08 subset, so the 'best predictor' claim is weak. Even accepting the renormalization in Section 5(a) as valid, the counts need a baseline before they can support 'contextuality in natural language at scale'.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constructs a PR-prism-inspired anaphora schema, instantiates it with 51,966,480 instances from Simple English Wikipedia, and uses BERT's [MASK] predictions (normalized to the two candidate nouns) to define empirical models. It reports 77,118 sheaf-contextual instances (SF < 1/6) and 36,938,948 CbD-contextual instances (Delta < 2), derives an algebraic relation between BERT logit differences and the PR-like epsilon parameter (Prop. 7.1), and uses correlation/regression analyses to argue that Euclidean distance between BERT noun embeddings is the best statistical predictor of contextuality. The paper also provides analytic formulas for the contextual and signalling fractions of PR-like models.","tokens_in":21777,"tokens_out":3865,"duration_ms":31296,"significance":"If the headline counts were properly benchmarked, this would be the first large-scale evidence for quantum-like contextuality in natural language, with potential implications for using contextuality as a resource in language models. The paper has genuine strengths: the analytic derivation of CF and SF for PR-like models (Props. 3.3 and 3.4), the large automatically constructed dataset, the clean algebraic identity in Prop. 7.1, and the release of code and data in a public repository. However, the significance is currently conditional because the CbD count is close to a tautology of the template, the sheaf count lacks a null baseline, the similarity claim rests on an unverified isotropy assumption, and the reported correlations are very weak.","major_comments":[{"comment":"The claim that 36,938,948 CbD-contextual instances (71.1%) constitute evidence of natural-language contextuality is not supported, because for the PR-prism template the CbD criterion Delta < 2 is satisfied by construction for almost any non-extreme BERT predictions: when the three epsilon_i values are equal to a common non-zero value, Delta = 2|epsilon| < 2, and since BERT's softmax outputs are almost never exactly 1 or 0, the majority of instances automatically fall in this region. A null model that randomizes the noun pairs or replaces BERT predictions with a plausible random distribution should be used to show that the observed rate exceeds chance; without this, the count is a property of the template and softmax geometry, not of natural language.","section":"Section 6(b)"},{"comment":"The sheaf-contextual rate of 0.148% (77,118 instances) is the main non-trivial evidence, but it is not compared to any null distribution. The criterion SF < 1/6 with three contexts and normalized two-outcome probabilities can be satisfied when all three contexts produce near-uniform predictions (epsilon_i all near 0); the observed 0.148% may simply be the base rate at which BERT is within 1/6 of uniform on three arbitrary contexts. A permutation or randomized-pair baseline is needed to establish that this rate is significantly elevated.","section":"Section 6(b)"},{"comment":"Equation (7.3) is an algebraic identity that follows from the softmax definition (7.2) and the PR-like parametrization; it is correct but does not by itself provide empirical evidence about language. The subsequent conclusion that contextual instances come from semantically similar words relies on the 'isotropic distribution of prediction vectors' assumption, which is stated in Section 7(a) without verification, and the regression results (Tables 4 and 6, R^2 approximately 0.006-0.009 full dataset, approximately 0.08 subset) show that Euclidean distance explains only a small fraction of variance. The strength of the 'best predictor' and 'came from' claims should be scaled to these effect sizes, or the isotropy assumption should be tested directly.","section":"Section 7(a), Prop. 7.1"},{"comment":"The similar-noun subset is constructed by selecting noun pairs with the highest cosine similarity after seeing the full data, and the contextual rates are then recomputed on this selected subset. Without a matched control (e.g., random subsets of equal size or a prespecified hypothesis), the observed increase from 0.148% to 0.50% sheaf-contextual and from 71.1% to 81.83% CbD-contextual could be a selection effect; the claim that similarity causes contextuality is not established by this procedure.","section":"Section 6(c)"},{"comment":"The interpretation of normalized BERT [MASK] probabilities as empirical probability distributions over referents is assumed but not validated. If these probabilities reflect word co-occurrence, template artifacts, or BERT's training biases rather than referent likelihood, then the reported 'contextuality' may be a property of BERT's softmax geometry rather than of natural language. The paper should either validate this interpretation or explicitly frame the results as properties of BERT's predictions.","section":"Section 5(a)"}],"minor_comments":[{"comment":"The word 'quantity' should be 'quantify' in the sentence 'One can try to define a signalling fraction (SF), in the same way CF is defined, to quantity the degree of signalling.'","section":"Section 3(a)"},{"comment":"There are several typos: 'contexutaliy' should be 'contextuality' in the text above Table 3, 'descirbing' should be 'describing' in Section 7(a), and 'polynomail' should be 'polynomial' in the caption of Figure 10.","section":"Section 7(b)"},{"comment":"Two different tables are both labelled 'Table 2' (the random sample in Section 6(a) and the similar-noun sample in Section 6(c)), which makes cross-referencing confusing.","section":"Section 6"},{"comment":"The caption says the contextual models are 'Highlighted', but the sheaf-contextual bars are too small to be visible in the histogram; consider using a log-scale inset or a different visualization.","section":"Section 6(b), Figure 5"},{"comment":"The definition of s_odd uses sigma dot x, but the notation 'p(sigma) = product_i sigma_i' combined with the maximization over sigma is not fully explained; please spell out the parity condition explicitly.","section":"Eq. (4.2)"},{"comment":"The R^2 values are reported to four decimal places, but the caption and text inconsistently use 'R2'; please standardize the notation as R² throughout.","section":"Tables 4 and 6"}],"recommendation":"major_revision","confidential_remarks":"The paper's analytic contributions and reproducibility are solid, but the empirical headline claims need to be recalibrated with null baselines and a test of the isotropy assumption. The CbD count in particular is close to a tautology of the template, and the sheaf count needs a randomized baseline. The semantic-similarity claim should be downgraded given the low R^2 values and the post hoc subset selection. These are fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The new thing is scale: 51.9M PR-anaphora instances mined from Simple English Wikipedia, with BERT probabilities, plus a clean algebraic relation between BERT logits and the epsilon parameters (Prop 7.1). The CF/SF bookkeeping for PR-like models is correct, and the code and data are public. That part is solid.\n\nThe soft spots are exactly where the stress-test puts them. The 71.1% CbD-contextual rate is close to tautological: the template forces PR-prism support, and when BERT is confident in the same noun across the three contexts (which is most of the time), Δ = 2|ε|, so any ε < 1 gives Δ < 2. That count doesn't tell you much about natural language. The 0.148% sheaf-contextual rate is the real signal, but there's no null baseline. Without comparing against randomized noun-adjective pairings, we don't know whether that's just BERT landing within 1/6 of uniform on three contexts by chance. The claim that contextual instances 'came from' semantically similar words is an overclaim: Prop 7.1 is an identity from softmax, and the leap to Euclidean distance as driver depends on an unverified isotropy assumption. The regression R2 is 0.006 full and 0.08 in the post hoc similar-noun subset, which is weak, even if statistically significant on 52M samples. The subset selection after seeing the data makes the improvement hard to trust.\n\nTo be fair: the bookkeeping is careful, the formal parts check out, and the authors are transparent about their methods. The weakness is interpretive, not computational, but the abstract and conclusion go beyond the evidence.\n\nWho is this for? Researchers in quantum-like cognition, compositional semantics, and anyone interested in whether LLM probability tables exhibit formal contextuality. It deserves a serious referee because it's a substantial empirical study with a formal core, but it needs major revision: add null baselines, separate the tautological CbD count from the informative one, temper the causal claims, and report the regression results with proper effect sizes.\n\nMy recommendation: send it to peer review, but expect heavy revision. The current version doesn't establish 'contextuality in natural language at scale.'","headline":"A serious large-scale study of contextuality in LLM probability tables, but the headline counts need null baselines and the semantic-similarity claim is weaker than presented.","tokens_in":22401,"tokens_out":2308,"would_cite":false,"duration_ms":18761,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Quantum-like contextuality occurs in natural language at scale: out of 51,966,480 pronoun-reference instances built from Simple English Wikipedia and BERT, 77,118 are sheaf-contextual and 36,938,948 are Contextuality-by-Default contextual.","keywords":["quantum contextuality","sheaf theory","large language models","BERT","anaphora resolution","contextuality-by-default","semantic similarity","natural language"],"falsifier":"Re-probe the same 77,118 sheaf-contextual templates with human plausibility judgments or with a different masked language model; if the contextual instances largely disappear, or if randomly permuting the adjectives across contexts preserves the same fraction of contextual instances, the effect is an artifact of the probe rather than a property of natural language.","tokens_in":21261,"feed_emoji":"⚛️","tokens_out":12472,"duration_ms":92847,"temperature":0.7,"pith_summary":"This paper seeks to establish that quantum-style contextuality—the failure of a set of local observations to admit one global joint explanation—occurs in ordinary natural language, and at a scale far beyond hand-built examples. The authors build a pronoun-reference schema whose measurement structure is the minimal contextual quantum scenario, instantiate it automatically from Simple English Wikipedia, and treat BERT's masked-word predictions as probability distributions over the two candidate referents. Of 51,966,480 generated instances, 77,118 meet the signalling-corrected sheaf criterion and 36,938,948 meet the Contextuality-by-Default criterion. If the claim is right, a statistical structure known to be a resource for quantum advantage is present in a core language task, which would make quantum-inspired methods a plausible target for coreference resolution and related language engines.","feed_headline":"51 million language patterns show quantum-style contextuality","feed_subtitle":"A BERT probe of pronoun reference in Simple Wikipedia finds quantum-style contextuality in 37 million cases.","key_machinery":"The load-bearing object is the PR-anaphora schema: a three-context measurement scenario whose observables are anaphoric modifiers and whose outcomes are two candidate noun referents, instantiated as sentences such as 'There is an apple and a strawberry. One of them is red and the very same one is round.' Its support is the PR prism, the strongly contextual model of the minimal 3-cyclic scenario, so contextuality is decided by the specialised inequalities $SF < 1/6$ (signalling-corrected sheaf theory) and $\\Delta < 2$ (Contextuality-by-Default). The analytical engine is Proposition 7.1, which connects BERT's masked-token logits to the empirical table: $\\varepsilon = \\tanh\\bigl(\\tfrac12 (p\\cdot \\Delta x + \\Delta b)\\bigr)$. Under the paper's isotropy assumption, this identity turns contextuality into geometry: the farther apart the two noun embeddings are, the more extreme BERT's probabilities become and the less contextual the instance, which is why Euclidean distance, not entropy or bias difference, emerges as the dominant statistical predictor.","core_discovery":"The central discovery is that a systematically constructed linguistic scenario—the PR-anaphora schema, a variant of coreference resolution with two candidate noun phrases and three adjectival modifiers—produces contextual empirical models when the probabilities come from a large language model. An instance is sheaf-contextual when its signalling fraction satisfies $SF < 1/6$ and CbD-contextual when its direct influence satisfies $\\Delta < 2$; the first condition is the signalling-corrected sheaf criterion specialised to three contexts, and the second is the Contextuality-by-Default criterion for cyclic systems. Of the 51,966,480 instances built from 866,108 noun pairs, 77,118 are sheaf-contextual and 36,938,948 are CbD-contextual, and restricting to the 1% most semantically similar noun pairs raises these fractions to 0.50% and 81.83%. The paper also proves that the probability imbalance in each context is governed by the identity $\\varepsilon = \\tanh\\bigl(\\tfrac12 (p\\cdot \\Delta x + \\Delta b)\\bigr)$, which ties contextuality to the geometry of BERT's embedding space through the Euclidean distance between the two candidate nouns; regression over several candidate features confirms that Euclidean distance is the best statistical predictor of contextuality.","pith_inferences":["Editorial extension: collecting human plausibility judgments for the same PR-anaphora templates would test whether the inequalities survive without BERT in the loop; survival would move the phenomenon from language-model geometry to natural language itself.","Editorial extension: the same 51 million instances could be probed with a decoder-only model prompted to fill the anaphor, checking whether contextuality is specific to bidirectional masked language modelling or generic to neural language-model probability geometry.","Editorial extension: because the sheaf-contextual region is a strict subset of the CbD-contextual region, the huge gap between 0.148% and 71.1% means claims about 'contextuality in language' must specify which framework's resource is meant; the two frameworks are not interchangeable measures.","Editorial extension: a control with the adjectives randomly permuted across the three contexts would show what fraction of contextual instances is carried by lexical co-occurrence statistics rather than by the semantic relations the schema is designed to probe."],"forward_implications":["Coreference resolution becomes a candidate domain for quantum-inspired algorithms, since its ambiguity structure can host contextual empirical models at scale.","The PR-anaphora schema offers an automatically instantiable generalisation of the Winograd Schema Challenge, moving beyond hand-crafted pronoun puzzles.","Because Euclidean distance between embeddings predicts contextuality, corpus searches can be seeded by semantic similarity rather than by enumerating all adjective triples.","The earlier anecdotal evidence from eleven hand-picked noun pairs is replaced by corpus-level evidence from 866,108 noun pairs, making the phenomenon reproducible and measurable.","If quantum-like contextuality behaves like the quantum resource, contextual instances mark language tasks where quantum methods could in principle outperform classical ones."],"supporting_citations":[{"why":"Defines the sheaf-theoretic empirical model and the gluing-failure criterion for contextuality that the paper builds on.","marker":"[8]"},{"why":"Supplies the signalling-corrected inequality that certifies contextuality in the presence of noise or signalling.","marker":"[9]"},{"why":"Establishes CbD contextuality tests on behavioural data, the precedent for applying contextuality outside physics.","marker":"[21]"},{"why":"Previous small-scale anaphora schemas that the present paper scales up to hundreds of thousands of noun pairs.","marker":"[24,25]"},{"why":"The Simple English Wikipedia snapshot is the corpus from which all adjective-noun phrases are extracted.","marker":"[26]"},{"why":"Defines the contextual fraction used to measure the degree of contextuality.","marker":"[28]"},{"why":"Defines CbD cyclic systems and the CNT1 criterion generalising Bell-CHSH to signalling systems.","marker":"[29–31]"},{"why":"BERT is the masked language model whose masked-token predictions provide the empirical probability distributions.","marker":"[32]"}],"fun_headline_variants":["37M language patterns show quantum-style contextuality","BERT reveals quantum contextuality in 37 million cases","Quantum-like contextuality found in language via BERT","Coreference task in LLMs displays quantum contextuality","37M instances of quantum contextuality from BERT on Wikipedia"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that BERT's normalized [MASK] probabilities for the two candidate nouns are faithful empirical probability distributions over referents; if those numbers largely reflect word co-occurrence or template artifacts, the contextuality belongs to BERT's softmax geometry rather than to natural language.","fun_headline_variants_meta":{"raw":{"variants":["37M language patterns show quantum-style contextuality","BERT reveals quantum contextuality in 37 million cases","Quantum-like contextuality found in language via BERT","Coreference task in LLMs displays quantum contextuality","37M instances of quantum contextuality from BERT on Wikipedia"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000634,"raw_usage":{"total_tokens":2977,"prompt_tokens":1047,"completion_tokens":1930,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":663,"completion_tokens_details":{"reasoning_tokens":1853}},"tokens_in":663,"tokens_out":1930,"duration_ms":13810,"temperature":1.0,"reasoning_tokens":1853,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:15:36.759398+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-probe the same 77,118 sheaf-contextual templates with human plausibility judgments or with a different masked language model; if the contextual instances largely disappear, or if randomly permuting the adjectives across contexts preserves the same fraction of contextual instances, the effect is an artifact of the probe rather than a property of natural language.","supporting_citations":[{"cited_title":"2024 Wikimedia Downloads","cited_arxiv_id":null,"evidence_quote":"The Simple English Wikipedia snapshot is the corpus from which all adjective-noun phrases are extracted."}],"review_version":1}