REVIEW 5 major objections 6 minor 39 references
Quantum-Like Contextuality in Large Language Models
T0 review · 5 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Quantum-like contextuality occurs in natural language at scale: out of 51,966,480 pronoun-reference instances built from Simple English Wikipedia and BERT, 77,118 are sheaf-contextual and 36,938,948 are Contextuality-by-Default contextual.
desk verdict A serious large-scale study of contextuality in LLM probability tables, but the headline counts need null baselines and the semantic-similarity claim is weaker than presented. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the PR-anaphora schema: a three-context measurement scenario whose observables are anaphoric modifiers and whose outcomes are two candidate noun referents, instantiated as sentences such as 'There is an apple and a strawberry. One of them is red and the very same one is round.' Its support is the PR prism, the strongly contextual model of the minimal 3-cyclic scenario, so contextuality is decided by the specialised inequalities $SF < 1/6$ (signalling-corrected sheaf theory) and $\Delta < 2$ (Contextuality-by-Default). The analytical engine is Proposition 7.1, which connects BERT's masked-token logits to the empirical table: $\varepsilon = \tanh\bigl(\tfrac12 (p\cdot \Delta x + \Delta b)\bigr)$. Under the paper's isotropy assumption, this identity turns contextuality into geometry: the farther apart the two noun embeddings are, the more extreme BERT's probabilities become and the less contextual the instance, which is why Euclidean distance, not entropy or bias difference, emerges as the dominant statistical predictor.
What would settle it
Re-probe the same 77,118 sheaf-contextual templates with human plausibility judgments or with a different masked language model; if the contextual instances largely disappear, or if randomly permuting the adjectives across contexts preserves the same fraction of contextual instances, the effect is an artifact of the probe rather than a property of natural language.
Extended reading notes
Core claim
The central discovery is that a systematically constructed linguistic scenario—the PR-anaphora schema, a variant of coreference resolution with two candidate noun phrases and three adjectival modifiers—produces contextual empirical models when the probabilities come from a large language model. An instance is sheaf-contextual when its signalling fraction satisfies $SF < 1/6$ and CbD-contextual when its direct influence satisfies $\Delta < 2$; the first condition is the signalling-corrected sheaf criterion specialised to three contexts, and the second is the Contextuality-by-Default criterion for cyclic systems. Of the 51,966,480 instances built from 866,108 noun pairs, 77,118 are sheaf-contextual and 36,938,948 are CbD-contextual, and restricting to the 1% most semantically similar noun pairs raises these fractions to 0.50% and 81.83%. The paper also proves that the probability imbalance in each context is governed by the identity $\varepsilon = \tanh\bigl(\tfrac12 (p\cdot \Delta x + \Delta b)\bigr)$, which ties contextuality to the geometry of BERT's embedding space through the Euclidean distance between the two candidate nouns; regression over several candidate features confirms that Euclidean distance is the best statistical predictor of contextuality.
Load-bearing premise
The load-bearing premise is that BERT's normalized [MASK] probabilities for the two candidate nouns are faithful empirical probability distributions over referents; if those numbers largely reflect word co-occurrence or template artifacts, the contextuality belongs to BERT's softmax geometry rather than to natural language.
Editorial extensions
If this is right
- Coreference resolution becomes a candidate domain for quantum-inspired algorithms, since its ambiguity structure can host contextual empirical models at scale.
- The PR-anaphora schema offers an automatically instantiable generalisation of the Winograd Schema Challenge, moving beyond hand-crafted pronoun puzzles.
- Because Euclidean distance between embeddings predicts contextuality, corpus searches can be seeded by semantic similarity rather than by enumerating all adjective triples.
- The earlier anecdotal evidence from eleven hand-picked noun pairs is replaced by corpus-level evidence from 866,108 noun pairs, making the phenomenon reproducible and measurable.
- If quantum-like contextuality behaves like the quantum resource, contextual instances mark language tasks where quantum methods could in principle outperform classical ones.
Reading between the lines
- Editorial extension: collecting human plausibility judgments for the same PR-anaphora templates would test whether the inequalities survive without BERT in the loop; survival would move the phenomenon from language-model geometry to natural language itself.
- Editorial extension: the same 51 million instances could be probed with a decoder-only model prompted to fill the anaphor, checking whether contextuality is specific to bidirectional masked language modelling or generic to neural language-model probability geometry.
- Editorial extension: because the sheaf-contextual region is a strict subset of the CbD-contextual region, the huge gap between 0.148% and 71.1% means claims about 'contextuality in language' must specify which framework's resource is meant; the two frameworks are not interchangeable measures.
- Editorial extension: a control with the adjectives randomly permuted across the three contexts would show what fraction of contextual instances is carried by lexical co-occurrence statistics rather than by the semantic relations the schema is designed to probe.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper constructs a PR-prism-inspired anaphora schema, instantiates it with 51,966,480 instances from Simple English Wikipedia, and uses BERT's [MASK] predictions (normalized to the two candidate nouns) to define empirical models. It reports 77,118 sheaf-contextual instances (SF < 1/6) and 36,938,948 CbD-contextual instances (Delta < 2), derives an algebraic relation between BERT logit differences and the PR-like epsilon parameter (Prop. 7.1), and uses correlation/regression analyses to argue that Euclidean distance between BERT noun embeddings is the best statistical predictor of contextuality. The paper also provides analytic formulas for the contextual and signalling fractions of PR-like models.
Significance. If the headline counts were properly benchmarked, this would be the first large-scale evidence for quantum-like contextuality in natural language, with potential implications for using contextuality as a resource in language models. The paper has genuine strengths: the analytic derivation of CF and SF for PR-like models (Props. 3.3 and 3.4), the large automatically constructed dataset, the clean algebraic identity in Prop. 7.1, and the release of code and data in a public repository. However, the significance is currently conditional because the CbD count is close to a tautology of the template, the sheaf count lacks a null baseline, the similarity claim rests on an unverified isotropy assumption, and the reported correlations are very weak.
major comments (5)
- [Section 6(b)] The claim that 36,938,948 CbD-contextual instances (71.1%) constitute evidence of natural-language contextuality is not supported, because for the PR-prism template the CbD criterion Delta < 2 is satisfied by construction for almost any non-extreme BERT predictions: when the three epsilon_i values are equal to a common non-zero value, Delta = 2|epsilon| < 2, and since BERT's softmax outputs are almost never exactly 1 or 0, the majority of instances automatically fall in this region. A null model that randomizes the noun pairs or replaces BERT predictions with a plausible random distribution should be used to show that the observed rate exceeds chance; without this, the count is a property of the template and softmax geometry, not of natural language.
- [Section 6(b)] The sheaf-contextual rate of 0.148% (77,118 instances) is the main non-trivial evidence, but it is not compared to any null distribution. The criterion SF < 1/6 with three contexts and normalized two-outcome probabilities can be satisfied when all three contexts produce near-uniform predictions (epsilon_i all near 0); the observed 0.148% may simply be the base rate at which BERT is within 1/6 of uniform on three arbitrary contexts. A permutation or randomized-pair baseline is needed to establish that this rate is significantly elevated.
- [Section 7(a), Prop. 7.1] Equation (7.3) is an algebraic identity that follows from the softmax definition (7.2) and the PR-like parametrization; it is correct but does not by itself provide empirical evidence about language. The subsequent conclusion that contextual instances come from semantically similar words relies on the 'isotropic distribution of prediction vectors' assumption, which is stated in Section 7(a) without verification, and the regression results (Tables 4 and 6, R^2 approximately 0.006-0.009 full dataset, approximately 0.08 subset) show that Euclidean distance explains only a small fraction of variance. The strength of the 'best predictor' and 'came from' claims should be scaled to these effect sizes, or the isotropy assumption should be tested directly.
- [Section 6(c)] The similar-noun subset is constructed by selecting noun pairs with the highest cosine similarity after seeing the full data, and the contextual rates are then recomputed on this selected subset. Without a matched control (e.g., random subsets of equal size or a prespecified hypothesis), the observed increase from 0.148% to 0.50% sheaf-contextual and from 71.1% to 81.83% CbD-contextual could be a selection effect; the claim that similarity causes contextuality is not established by this procedure.
- [Section 5(a)] The interpretation of normalized BERT [MASK] probabilities as empirical probability distributions over referents is assumed but not validated. If these probabilities reflect word co-occurrence, template artifacts, or BERT's training biases rather than referent likelihood, then the reported 'contextuality' may be a property of BERT's softmax geometry rather than of natural language. The paper should either validate this interpretation or explicitly frame the results as properties of BERT's predictions.
minor comments (6)
- [Section 3(a)] The word 'quantity' should be 'quantify' in the sentence 'One can try to define a signalling fraction (SF), in the same way CF is defined, to quantity the degree of signalling.'
- [Section 7(b)] There are several typos: 'contexutaliy' should be 'contextuality' in the text above Table 3, 'descirbing' should be 'describing' in Section 7(a), and 'polynomail' should be 'polynomial' in the caption of Figure 10.
- [Section 6] Two different tables are both labelled 'Table 2' (the random sample in Section 6(a) and the similar-noun sample in Section 6(c)), which makes cross-referencing confusing.
- [Section 6(b), Figure 5] The caption says the contextual models are 'Highlighted', but the sheaf-contextual bars are too small to be visible in the histogram; consider using a log-scale inset or a different visualization.
- [Eq. (4.2)] The definition of s_odd uses sigma dot x, but the notation 'p(sigma) = product_i sigma_i' combined with the maximization over sigma is not fully explained; please spell out the parity condition explicitly.
- [Tables 4 and 6] The R^2 values are reported to four decimal places, but the caption and text inconsistently use 'R2'; please standardize the notation as R² throughout.
Circularity Check
One claimed 'proof' (Eq. 7.3) is an algebraic identity from the softmax and epsilon definitions; the headline contextuality counts are empirical and not circular.
-
self definitional
[Section 7(a), Proposition 7.1, Eq. (7.3)]
"Using equations 7.1 and 7.2, below in Prop. 7.1 we prove a result which connects the BERT logit scores to the empirical table of the PR-like model descirbing the PR-anaphora schema: ... Proposition 7.1. The logit scores of the masked token given by BERT relates to the ϵ parametrisation of the PR-like model as follows: ϵ = tanh(1/2 (p · ∆x + ∆b)). (7.3)"
The parameter ϵ is defined by the PR-like table entries (1±ϵ)/2, and those entries are taken to be the renormalized BERT softmax probabilities of Eq. (7.2). Substituting P_O1 = e^{l_O1}/(e^{l_O1}+e^{l_O2}) and P_O2 = e^{l_O2}/(e^{l_O1}+e^{l_O2}) into P_O1=(1+ϵ)/2 immediately gives log((1+ϵ)/(1-ϵ)) = l_O1-l_O2, i.e. Eq. (7.3). The proposition is therefore an algebraic identity that unpacks definitions, not an independent first-principles result. The abstract's claim that this 'proved that the contextual instances came from semantically similar words' goes beyond the identity: it additionally requires the untested isotropic-distribution assumption on p and treats the tautological link as empirical support.
full rationale
The two headline counts are empirical measurements: 51,966,480 empirical tables are constructed from BERT's renormalized [MASK] probabilities and evaluated with the external sheaf-theoretic criterion SF < 1/6 and the CbD criterion Delta < 2. These counts do not reduce by construction to a fitted parameter, and no parameter is fitted to a subset and then renamed as a prediction. The sheaf criterion is imported from ref. [9], which includes an author of the present paper, but it is a published external inequality and is not derived from the present target result, so it is not load-bearing circularity. The CbD count's closeness to a tautology for the PR-prism template (for all-equal epsilon, Delta = 2|epsilon|) is a real null-baseline and interpretation concern, but it is a validity issue rather than a circular derivation. The regression analysis is empirical, and the 'Euclidean distance best predictor' claim is not fully forced by Eq. (7.3) because the logit difference also contains the bias difference and the isotropy assumption is additional. The main circular element is Proposition 7.1: it is presented as a proof connecting contextuality to Euclidean distance, but it is an analytic identity following from the definitions of epsilon and the BERT softmax. This warrants a score of 3 rather than 0, while the overall empirical contribution remains substantially independent.
Assumptions & free parameters
free parameters (3)
- top_common_adjectives_count =
5
- similar_noun_percentile =
1%
- polynomial_degree_cutoff =
3 (cubic)
assumptions (5)
- domain assumption BERT's [MASK] probabilities, renormalized to the two candidate nouns, represent empirical probability distributions over the referents.
- domain assumption The PR-anaphora schema's linguistic constraints (same one/other one) force the empirical table to have PR-prism support.
- domain assumption The signalling-corrected criterion CF > 2|M| SF from Ref [9] applies to natural-language empirical models.
- ad hoc to paper BERT prediction vectors are isotropically distributed, so Delta_l is proportional to ||Delta x||.
- standard math Standard linear algebra and convex decomposition properties of empirical models.
Cite this review
Pith. "Pith review of Quantum-Like Contextuality in Large Language Models." pith.science (2026). https://pith.science/paper/FTSWJDHF
@misc{pith2026241216806,
author = {Pith},
title = {Pith review of: Quantum-Like Contextuality in Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/FTSWJDHF}},
note = {Machine review of arXiv:2412.16806}
}
read the original abstract
Contextuality is a distinguishing feature of quantum mechanics and there is growing evidence that it is a necessary condition for quantum advantage. In order to make use of it, researchers have been asking whether similar phenomena arise in other domains. The answer has been yes, e.g. in behavioural sciences. However, one has to move to frameworks that take some degree of signalling into account. Two such frameworks exist: (1) a signalling-corrected sheaf theoretic model, and (2) the Contextuality-by-Default (CbD) framework. This paper provides the first large scale experimental evidence for a yes answer in natural language. We construct a linguistic schema modelled over a contextual quantum scenario, instantiate it in the Simple English Wikipedia and extract probability distributions for the instances using the large language model BERT. This led to the discovery of 77,118 sheaf-contextual and 36,938,948 CbD contextual instances. We proved that the contextual instances came from semantically similar words, by deriving an equation between degrees of contextuality and Euclidean distances of BERT's embedding vectors. A regression model further reveals that Euclidean distance is indeed the best statistical predictor of contextuality. Our linguistic schema is a variant of the co-reference resolution challenge. These results are an indication that quantum methods may be advantageous in language tasks.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
1935 Can Quantum-Mechanical Description of Physical Reality Be Considered Complete?
Einstein A, Podolsky B, Rosen N. 1935 Can Quantum-Mechanical Description of Physical Reality Be Considered Complete?. Phys. Rev. 47, 777–780. (10.1103/PhysRev.47.777)
-
[2]
1966 On the Problem of Hidden Variables in Quantum Mechanics
Bell JS. 1966 On the Problem of Hidden Variables in Quantum Mechanics. Reviews of Modern Physics 38, 447–452. (10.1103/RevModPhys.38.447)
-
[3]
1964 On the Einstein Podolsky Rosen paradox
Bell JS. 1964 On the Einstein Podolsky Rosen paradox. Physics Physique Fizika 1, 195–200. (10.1103/PhysicsPhysiqueFizika.1.195)
-
[4]
2009 Computational Power of Correlations
Anders J, Browne DE. 2009 Computational Power of Correlations. Physical Review Letter 102, 050502. (10.1103/PhysRevLett.102.050502)
-
[5]
2013 Contextuality in measurement-based quantum computation
Raussendorf R. 2013 Contextuality in measurement-based quantum computation. Physical Review A - Atomic, Molecular, and Optical Physics 88, 1–7. (10.1103/PhysRevA.88.022322)
-
[6]
2018 Quantum Advantage from Sequential-Transformation Contextuality
Mansfield S, Kashefi E. 2018 Quantum Advantage from Sequential-Transformation Contextuality. Physical Review Letters 121, 1–8. (10.1103/PhysRevLett.121.230401)
-
[7]
2014 Contextuality supplies the ’magic’ for quantum computation
Howard M, Wallman J, Veitch V , Emerson J. 2014 Contextuality supplies the ’magic’ for quantum computation. Nature 510, 351–355. (10.1038/nature13460)
-
[8]
2011 The sheaf-theoretic structure of non-locality and contextuality
Abramsky S, Brandenburger A. 2011 The sheaf-theoretic structure of non-locality and contextuality. New Journal of Physics 13, 113036. (10.1088/1367-2630/13/11/113036)
Show all 39 references
-
[9]
2024 Corrected Bell and non-contextuality inequalities for realistic experiments
Vallée K, Emeriau PE, Bourdoncle B, Sohbi A, Mansfield S, Markham D. 2024 Corrected Bell and non-contextuality inequalities for realistic experiments. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 382, 20230011. (10.1098/rst...
2024
-
[10]
2017 Sheaves are the canonical data structure for sensor integration
Robinson M. 2017 Sheaves are the canonical data structure for sensor integration. Information Fusion 36, 208–224. (10.1016/j.inffus.2016.12.002)
2017 doi
-
[11]
2021 Sheaves as a Framework for Understanding and Interpreting Model Fit
Kvinge H, Jefferson B, Joslyn C, Purvine E. 2021 Sheaves as a Framework for Understanding and Interpreting Model Fit. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) pp. 4205–4213. IEEE. (10.1109/ICCVW54120.2021.00469)
2021
-
[12]
2021 On the Design of Social Robots Using Sheaf Theory and Smart Contracts
Murimi R. 2021 On the Design of Social Robots Using Sheaf Theory and Smart Contracts. Frontiers in Robotics and AI 8. (10.3389/frobt.2021.559380)
2021
-
[13]
2014 Sheaves, Cosheaves and Applications
Curry J. 2014 Sheaves, Cosheaves and Applications. PhD thesis, University of Pennsylvania
2014
-
[14]
2019 Toward a spectral theory of cellular sheaves
Hansen J, Ghrist R. 2019 Toward a spectral theory of cellular sheaves. Journal of Applied and Computational Topology 3, 315–358. (10.1007/s41468-019-00038-7)
2019 doi
-
[15]
2024 Prospects for inconsistency detection using large language models and sheaves
Huntsman S, Robinson M, Huntsman L. 2024 Prospects for inconsistency detection using large language models and sheaves. arXiv preprint arXiv:2401.16713
2024 arXiv
-
[16]
2023 Knowledge Sheaves: A Sheaf-Theoretic Framework for Knowledge Graph Embedding
Gebhart T, Hansen J, Schrater P . 2023 Knowledge Sheaves: A Sheaf-Theoretic Framework for Knowledge Graph Embedding. In Proceedings of The 26th International Conference on Artificial Intelligence and Statistics vol. 206 pp. 9094–9116. PMLR
2023
-
[17]
2022 Neural sheaf diffusion: a topological perspective on heterophily and oversmoothing in GNNs
Bodnar C, Giovanni FD, Chamberlain B, Lio P , Bronstein M. 2022 Neural sheaf diffusion: a topological perspective on heterophily and oversmoothing in GNNs. In Proceedings of the Thirty-Sixth Conference on Neural Information Processing Systems vol. 35
2022
-
[18]
Abramsky S, Sadrzadeh M. 2014 pp. 1–13. InSemantic Unification, pp. 1–13. Berlin, Heidelberg: Springer Berlin Heidelberg. (10.1007/978-3-642-54789-8_1)
2014 doi
-
[19]
2019 A Universal Construction for Semantic Compositionality
Philips S. 2019 A Universal Construction for Semantic Compositionality. Phil. Trans. R. Soc. B375
2019
-
[20]
2022 An Enriched Category Theory of Language: From Syntax to Semantics
Bradley TD, Terilla J, Vlassopoulos Y. 2022 An Enriched Category Theory of Language: From Syntax to Semantics. La Matematica 1, 551—-580
2022
-
[21]
2016 On contextuality in behavioural data
Dzhafarov EN, Kujala JV , Cervantes VH, Zhang R, Jones M. 2016 On contextuality in behavioural data. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 374, 20150234. (10.1098/rsta.2015.0234)
2016
-
[22]
2021a On the Quantum-like Contextuality of Ambiguous Phrases
Wang D, Sadrzadeh M, Abramsky S, Cervantes V . 2021a On the Quantum-like Contextuality of Ambiguous Phrases. In Proceedings of the 2021 Workshop on Semantic Spaces at the Intersection 23royalsocietypublishing.org/journal/rspa Proc R Soc A 0000000. . . . . . . . . . . . . . . ....
-
[23]
2021b Analysing Ambiguous Nouns and Verbs with Quantum Contextuality Tools
Wang D, Sadrzadeh M, Abramsky S, Cervantes VH. 2021b Analysing Ambiguous Nouns and Verbs with Quantum Contextuality Tools. Journal of Cognitive Science 22, 391–420. (10.17791/jcs.2021.22.3.391)
2021 doi
-
[24]
2022 A Model of Anaphoric Ambiguities using Sheaf Theoretic Quantum-like Contextuality and BERT.Electronic Proceedings in Theoretical Computer Science 366, 23–34
Lo KI, Sadrzadeh M, Mansfield S. 2022 A Model of Anaphoric Ambiguities using Sheaf Theoretic Quantum-like Contextuality and BERT.Electronic Proceedings in Theoretical Computer Science 366, 23–34. (10.4204/eptcs.366.5)
2022 doi
-
[25]
2023 Developments in Sheaf-Theoretic Models of Natural Language Ambiguities
Lo KI, Sadrzadeh M, Mansfield S. 2023 Developments in Sheaf-Theoretic Models of Natural Language Ambiguities. In Proceedings of the 13th International Workshop on Developments in Computational Models, a satellite event of FSCD/CADE, Rome . To appear in EPTCS (10.48550/arXiv.24...
-
[26]
2024 Wikimedia Downloads
Wikimedia Foundation. 2024 Wikimedia Downloads
2024
-
[27]
2014 On monogamy of non-locality and macroscopic averages: examples and preliminary results
Soares Barbosa R. 2014 On monogamy of non-locality and macroscopic averages: examples and preliminary results. Electronic Proceedings in Theoretical Computer Science 172, 36–55. (10.4204/EPTCS.172.4)
2014 doi
-
[28]
2017 Contextual Fraction as a Measure of Contextuality
Abramsky S, Barbosa RS, Mansfield S. 2017 Contextual Fraction as a Measure of Contextuality. Physical Review Letter 119, 050504. (10.1103/PhysRevLett.119.050504)
2017 doi
-
[29]
2013 All-Possible-Couplings Approach to Measuring Probabilistic Context
Dzhafarov EN, Kujala JV . 2013 All-Possible-Couplings Approach to Measuring Probabilistic Context. PLoS ONE 8, e61712. (10.1371/journal.pone.0061712)
2013 doi
-
[30]
2016 Contextuality-by-Default 2.0: Systems with Binary Random Variables
Dzhafarov EN, Kujala JV . 2016 Contextuality-by-Default 2.0: Systems with Binary Random Variables. arXiv preprint arXiv:1604.04799
2016 arXiv
-
[31]
2015 Contextuality-by-Default: A Brief Overview of Ideas, Concepts, and Terminology
Dzhafarov EN, Kujala JV , Cervantes VH. 2015 Contextuality-by-Default: A Brief Overview of Ideas, Concepts, and Terminology. Lecture Notes in Computer Science 9535, 12-23, 2016 . (10.1007/978-3-319-28675-4-2)
2015 doi
-
[32]
2019 BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin J, Chang MW, Lee K, Toutanova K. 2019 BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol...
2019 doi
-
[33]
2019 Linguistic matrix theory
Kartsaklis D, Ramgoolam S, Sadrzadeh M. 2019 Linguistic matrix theory. Ann. Inst. Henri Poincare Comb. Phys. Interact 6, 385–426
2019
-
[34]
O’Reilly Media, Inc
Bird S, Klein E, Loper E. 2009 Natural language processing with Python: analyzing text with the natural language toolkit. " O’Reilly Media, Inc."
2009
-
[35]
2017 Attention is All you Need
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Lu, Polosukhin I. 2017 Attention is All you Need. In Guyon I, Luxburg UV , Bengio S, Wallach H, Fergus R, Vishwanathan S, Garnett R, editors, Advances in Neural Information Processing Systems vol. 30. Curra...
2017
-
[36]
2013 Efficient Estimation of Word Representations in Vector Space
Mikolov T, Chen K, Corrado G, Dean J. 2013 Efficient Estimation of Word Representations in Vector Space. arXiv preprint arXiv:1301.3781
2013 arXiv
-
[37]
2020 Transformers: State-of-the-Art Natural Language Processing
Wolf T, Debut L, Sanh V , Chaumond J, Delangue C, Moi A, Cistac P , Rault T, Louf R, Funtowicz M, Davison J, Shleifer S, von Platen P , Ma C, Jernite Y, Plu J, Xu C, Le Scao T, Gugger S, Drame M, Lhoest Q, Rush A. 2020 Transformers: State-of-the-Art Natural Language Processing...
2020 doi
-
[38]
2012 The Winograd Schema Challenge
Levesque HJ, Davis E, Morgenstern L. 2012 The Winograd Schema Challenge. In Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning KR’12 pp. 552–561. AAAI Press. (10.5555/3031843.3031909)
2012
-
[39]
2023 Generalised Winograd Schema and its Contextuality
Lo KI, Sadrzadeh M, Mansfield S. 2023 Generalised Winograd Schema and its Contextuality. Electronic Proceedings in Theoretical Computer Science 384, 187–202. (10.4204/eptcs.384.11)
2023 doi
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.