Pith. sign in

REVIEW 5 major objections 5 minor 150 references

The Process of Categorical Clipping at the Core of the Genesis of Concepts in Synthetic Neural Cognition

T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper claims that a target neuron in a language model clips a categorically homogeneous subdimension out of each precursor neuron's category, leaving the rest as a categorical background, and supports this with cosine-similarity and…

desk verdict Interesting conceptual framing, but the main empirical claims lack the null models needed to distinguish clipping from a selection artifact. read the letter →

arxiv 2502.15710 v1 pith:V4F6MJHJ submitted 2025-01-21 cs.AI q-bio.NC

classification cs.AIq-bio.NC
keywords categoricalclippingsyntheticcognitionneuroninterpretabilitylanguagemodelsGPT-2XLconceptformationembeddingsimilarityreflectiveabstraction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that when one neuron in GPT-2XL strongly drives a neuron in the next layer, the target neuron does not inherit the whole category of its precursor; it cuts out a specific, categorically homogeneous subdimension (the "taken-tokens") and leaves the rest as a categorical background. The load-bearing evidence is that taken-token groups show higher mean pairwise cosine similarity than either the precursor's full core-token set or the left-behind tokens, with average differences around $0.14$ and a 94-95% positive-difference rate across 9,007 neuron pairs. If true, the overlap between strongly connected neurons' top-100 activation sets is not noise but a functional unit of concept formation, giving interpretability research a precise object to study. The paper also argues this extraction is constructive, not a picking-out of pre-existing categories, because the extracted subdimensions associate with only a minority of embedding dimensions and sit in distinct t-SNE regions from the background.

What carries the argument

The central object is the taken-cluster/left-cluster partition of a precursor neuron's 100 core-tokens, induced by their overlap with a strongly connected target neuron's core-tokens. The carrying mechanism is the aggregation function $\Sigma(w_{i,j}x_{i,j}) + b$, embodying the three factors of priming, attention, and categorical phasing. The paper's core identity is the difference in mean pairwise cosine similarity between taken-tokens and left- or core-tokens (the $d\approx0.14$ statistic), together with the quadrant-separation statistic for t-SNE centroids, which jointly operationalize categorical reduction and categorical-zone segmentation.

What would settle it

Compute the same $d$ statistics on randomly chosen subsets of a precursor neuron's core-tokens, matched for size to the observed taken-clusters, and on neuron pairs with no strong learned connection; if random subsets also yield mean $d\approx0.14$ or unconnected pairs show the same 87% taken/left centroid-separation rate, the evidence for categorical clipping would be indistinguishable from a token-selection artifact.

Watch

Extended reading notes

Core claim

The central discovery claim is that each MLP neuron carries a synthetic category whose extension is its core-tokens, and that when a neuron in layer $n+1$ aggregates weighted contributions from precursor neurons in layer $n$, the aggregation function performs categorical clipping: it extracts from each precursor category a subdimension aligned with the new category. The extracted subdimension is empirically marked by tokens that are core-tokens of both the precursor and the target neuron ("taken-tokens"), as opposed to "left-tokens" that remain only in the precursor. Four properties are reported: categorical reduction (taken-clusters have higher mean pairwise cosine similarity than core- or left-clusters), categorical selectivity (86% of taken-clusters contain fewer than six tokens), separation of initial embedding dimensions (embedding dimensions split into those aligned with taken versus left tokens), and segmentation of categorical zones (87% of taken/left centroid pairs fall in different t-SNE quadrants). The paper interprets these as manifestations of a synthetic theorem-in-act: the aggregation function, at high activations, reflectively abstracts and recombines sub-concepts-in-act to build the target category.

Load-bearing premise

The load-bearing premise is that a neuron's category is faithfully captured by its 100 highest-activation tokens and that cosine similarity in the model's own embedding space measures categorical proximity; the paper also provides no null model showing that the observed top-100 overlap between strongly connected neurons surpasses chance.

Editorial extensions

If this is right

  • The top-100 activation overlap between strongly connected neurons becomes a precise unit of analysis: inspecting a target neuron's taken-token clusters should reveal the categorical subdimension it is constructing.
  • A neuron's category is not a monolithic cluster; it is built part-by-part from small subdimensions clipped from many precursors, so interpretability studies should trace these subdimensions rather than treat the whole category as atomic.
  • The same four signatures (reduction, selectivity, dimension separation, and zone segmentation) give a transferable test for clipping in deeper layers and in other transformer architectures.
  • Because clipping is selective and constructive, explaining a model's behavior in human terms will require translating model-specific subdimensions, not mapping them onto pre-existing human categories.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper leaves implicit is to run the same taken-versus-left cosine analysis across layers 2 and above; if clipping is the mechanism of category genesis, the homogeneity gap should persist or sharpen as categories become more abstract.
  • One could test the causality of clipping directly by ablating the strongest precursor-target weights that define a taken-cluster and checking whether the target neuron's category specifically degrades.
  • The paper's form/background framing suggests a direct parallel with figure-ground separation in vision; re-running the clipping test on convolutional feature maps or image-patch token sets would show whether this is a general neural mechanism or specific to language models.
  • The missing null model can be supplied by a permutation baseline: random subsets of core-tokens of the same size should not reproduce $d\approx0.14$ or the 87% quadrant separation; if they do, the clipping claim would need revision.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper proposes a mechanism called "categorical clipping" in GPT-2XL's MLP layers: a target neuron in layer n+1 is said to extract from each strongly connected precursor neuron in layer n a semantically homogeneous subdimension, the "taken-tokens," leaving a "left-token" background. Using the public Bills et al. neuron-activation data, the authors define taken-tokens as the intersection of a precursor neuron's top-100 core-tokens with a target neuron's top-100 core-tokens, and then study four claimed properties: categorical reduction (taken clusters have higher mean pairwise cosine similarity), categorical selectivity (taken clusters are small), separation of initial embedding dimensions (via PCA), and segmentation of categorical zones (via t-SNE). The results are interpreted within a Piaget/Vergnaud framework of reflective abstraction, concepts-in-act, and theorems-in-act.

Significance. If the empirical claims were properly supported, the paper would offer a potentially interesting, mechanistically grounded account of how MLP layers recombine categorical subdimensions, and it would connect interpretability research with cognitive-developmental theory. The use of public GPT-2XL data, the comparison across three embedding spaces, and the authors' explicit caveats about PCA applicability are positive features. However, the central evidence for categorical clipping currently rests on selection rules and analysis choices that could produce the reported patterns by construction, so the contribution's significance is not yet established.

major comments (5)
  1. [Section 6.2, Tables 3 and 5] The categorical-reduction claim has no null model for the taken-token selection rule. Taken-tokens are defined as the intersection of the precursor's top-100 core-tokens and the target's top-100 core-tokens for the 10 strongest connections. Under this rule, any two neurons with correlated receptive fields will produce a taken set that is more internally coherent than the precursor core, even if no clipping process exists. The paper compares taken-tokens with core-tokens and left-tokens from the same precursor, but it never varies the target or compares against random/weakly connected targets, shuffled token labels, or size-matched random subsets of the core-tokens. Without such a baseline, the mean d ≈ .14 in Tables 3 and 5 cannot be attributed to categorical clipping rather than to the selection-by-overlap rule itself.
  2. [Section 6.3, Table 7] The categorical-selectivity test uses expected frequencies that are arbitrary and appear to invert the natural baseline. Table 7 lists expected frequencies of 3,200 (clusters of size < 6) and 60,800 (size ≥ 6) out of N = 64,000, i.e., a 5%/95% split that is never justified. Under a simple random-overlap null in which two independent top-100 sets are drawn from a vocabulary of ≈ 50,000 tokens, the expected overlap is close to 0.2 tokens, so almost all clusters should have size < 6. The observed 86% below the threshold would then be close to or below the natural expectation, and the reported χ² = 882,406 would not support selectivity. The authors need to state the null model explicitly and recompute the test under that model, with an appropriate effect-size measure.
  3. [Section 5.3, item (vi), and Section 6.4] The PCA analysis is circular by construction. The authors state that they overweight the two variables "taken-token" and "left-token" by 1% of the number of embedding dimensions "to format the PCA so it produces the desired factor axis F2 related to whether a token is a taken-token or left-token." Unsurprisingly, the subsequent analysis finds that embedding dimensions project onto a taken-vs-left axis and that taken-tokens are associated with fewer embedding dimensions. This result is imposed by the overweighting, so it cannot serve as evidence for the claimed "separation of initial embedding dimensions." The PCA can be kept as an explicitly descriptive visualization, but it must not be presented as a test of the paper's fourth postulate unless the overweighting is removed or justified as a sensitivity analysis with a non-circular baseline.
  4. [Section 6.2 and Section 6.3] The inferential statistics aggregate 9,007 (or 64,000) per-neuron tests as though they were independent, with p(χ²) computed on the proportion of positive differences. There is no correction for multiple comparisons, no account of the dependence among tests that share the same target or precursor neuron, and no report of the distribution of effect sizes beyond the mean d. The large sample sizes make χ² values such as 4.36E-19 essentially uninformative. The authors should report cluster-level or neuron-level effect sizes (e.g., median d with bootstrapped confidence intervals), a mixed-effects or permutation analysis, and the number of tests that remain significant after multiple-comparison correction.
  5. [Section 6.5, Table 11 and Graph 10] The t-SNE quadrant analysis needs a null model before it can support the "segmentation of categorical zones" claim. The axes of a t-SNE embedding are arbitrary up to rotation and reflection, so the statement that 87% of taken/left centroid pairs fall in different quadrants depends on the arbitrary coordinate orientation. A permutation test that randomly reassigns taken/left labels within each precursor neuron's core-tokens, or that compares against randomly chosen target neurons, is needed to establish that the observed quadrant separation exceeds chance. The current χ² = 217.75 is computed against expected frequencies that are not derived from any stated null model.
minor comments (5)
  1. [Section 4.2] The long passage citing Savioz et al. on dopamine, noradrenaline, and the sigmoidal transfer function is duplicated verbatim in the same section; one copy should be removed.
  2. [Table 5] Table 5's second row is labeled "Mean(Mean(COS(core-tokens)))" but the text says the comparison is between taken-tokens and left-tokens; the label should be corrected to left-tokens.
  3. [Bibliography] Several bibliography entries appear duplicated or mis-attributed: references [44] and [45] are the same work, [22] and [140] are the same Captum paper, and [87] and [88] are the same Nadeau book. The list should be de-duplicated and checked against in-text citations.
  4. [Section 5.3] The sample-size numbers across analyses should be reconciled: the text refers to 1,671 precursor neurons for the PCA/t-SNE, N = 950 after the KMO filter, and N = 1,610 for the t-SNE quadrant analysis; the exact inclusion criteria and the reasons for the differences are not stated clearly.
  5. [Section 6.4] The two single-neuron PCA examples (Graphs 5 and 6) are presented despite the authors' own statement that PCA applicability conditions are only partially met; this is acceptable as illustration, but the wording should mark them more explicitly as non-inferential case studies.

Circularity Check

1 steps flagged · score 6.0 of 10

The PCA 'separation of initial embedding dimensions' is forced by construction: the taken/left indicator variables are overweighted to create the separating axis; the central reduction result is independent but lacks a null model.

  1. fitted input called prediction [Section 5.3 (PCA parameterization option vi) and Section 6.4 (Graph 5 / Table 8)]
    "In §5.3: '(vi) over-weighting of the two variables "taken-token" and "left-token" by 1 % (of the number of embedding dimensions), to format the PCA so it produces the desired factor axis F2 related to whether a token is a "taken-token" or "left-token"'. In §6.4: 'We overweigh these two variables (at 1% of the 1600 embedding variables) to guide the PCA toward producing a factorial axis (with a sufficient eigenvalue) related to whether a token is a taken-token versus a left-token.'"

    The PCA axis opposing taken-tokens and left-tokens is manufactured by overweighting the two group-membership indicator variables before the analysis. The paper then reports that the taken/left tokens separate along this axis (Graph 5: 'a second vertical factor ... opposing the taken-tokens (in green) to the left-tokens (in red)') and interprets this as evidence that categorical clipping 'manifests as a dichotomous elective compartmentalization of these embeddings'. The separation is true by construction: the axis was explicitly 'formatted' to produce it. Reporting the manufactured axis as an empirical property of the embeddings is a fitted input renamed as a finding.

full rationale

The paper's strongest independent result is the categorical-reduction comparison (Tables 3 and 5): taken-token clusters show higher mean pairwise cosine similarity than core- or left-token clusters. This is not circular, because defining a token as 'taken' by top-100 overlap does not by itself entail higher embedding-space homogeneity of the intersection; the d≈0.14 gap is an empirical property of GPT-2XL's activations, even though the missing null model for expected overlap is a serious validity threat. The PCA section is circular: the authors explicitly overweight the 'taken-token' and 'left-token' variables 'to format the PCA so it produces the desired factor axis F2 related to whether a token is a taken-token or left-token,' and then treat the resulting separation as evidence that categorical clipping 'manifests as a dichotomous elective compartmentalization of these embeddings'. The separation is guaranteed by the overweighted indicator variables. No load-bearing self-citation chain or imported uniqueness theorem was found; citations to the authors' prior work supply terminology and the three-factor framework, but the empirical analyses use independent OpenAI data from Bills et al. Score 6 reflects one constructed 'prediction' amid otherwise self-contained measurements, matching the 'partial circularity' level.

Assumptions & free parameters 5 free parameters · 6 assumptions · 3 invented entities

The central claims rest on several domain assumptions about how to operationalize neuron categories (top-100 tokens), how to measure categorical proximity (cosine similarity in the input embedding space), and how to identify meaningful precursor-target pairs (top-10 connection weights). The paper introduces interpretive entities such as categorical clipping and synthetic reflective abstraction that are not independently evidenced. The free parameters listed here are hand-chosen thresholds and weightings that directly shape the reported properties.

free parameters (5)
  • PCA overweighting weight = 1% of the number of embedding dimensions
    The two dichotomous variables taken/left are overweighted by 1% of 1600 variables to force the PCA to produce a factor axis separating taken from left tokens (Section 5.3). This hand-chosen value directly shapes the reported separation of embedding dimensions.
  • taken-cluster size threshold = 6 tokens
    Used to define selectivity and to filter clusters for the normality and homoscedasticity checks; the choice of 6 is arbitrary and affects the selectivity analysis (Sections 5.4 and 6.3).
  • cos² quality threshold = 0.6 (and 0.4 in some instances)
    Threshold for retaining embedding variables in PCA interpretation; chosen post hoc, it affects which variables are projected (Section 6.4).
  • KMO threshold = 0.5
    Used to filter PCA cases, with the mean KMO of retained cases equal to 0.501, barely above the stated minimum and indicating poor factorability (Sections 5.3 and Table 10).
  • taken-token percentage range = 15% to 85%
    Selection criterion for precursor neurons in the PCA and t-SNE global analyses; this narrows the sample to 1671, and then to 950 after the KMO filter, out of 6400 neurons (Section 5.3).
assumptions (6)
  • domain assumption A neuron's categorical extension is adequately represented by its 100 highest-activation tokens (core-tokens).
    The entire analysis is built on this operationalization (Section 5.2). If the top-100 set is not a stable or meaningful summary, the measured properties may be artifacts.
  • domain assumption Cosine similarity in the GPT-2XL embedding space is a valid measure of categorical proximity between tokens.
    Used as the main dependent variable in the categorical reduction tests (Section 6.1, Tables 3 to 6). The paper itself notes that other embedding bases give weaker results.
  • domain assumption Strong connection weights between a layer-0 neuron and a layer-1 neuron identify a genetic precursor-target relation relevant to concept formation.
    Pairs are selected by top-10 absolute connection weights (Section 5.3). No control for chance overlap of top-100 sets is provided.
  • domain assumption t-SNE coordinates preserve large-scale categorical zones that can be meaningfully partitioned into quadrants.
    The quadrant analysis in Section 6.5 treats t-SNE axes as categorical zones, though t-SNE primarily preserves local structure and is stochastic.
  • domain assumption The first two MLP layers of GPT-2XL are representative of synthetic categorical segmentation in general.
    The paper generalizes from layers 0 and 1 of one model to synthetic cognition in general (Sections 5.2 and 7.1).
  • domain assumption The 9007 cluster-level tests can be treated as independent observations for chi-square aggregation.
    No correction for multiple comparisons or dependency between tests is applied; the reported p-values, such as p(chi-square) less than .0001, assume independence (Section 6.2).
invented entities (3)
  • Categorical clipping
    purpose: To name and explain the process by which target neurons extract homogeneous token subclusters from precursor neurons, distinguishing a categorical form from a background.
    The concept is defined by the observed taken/left token overlap and is not evidenced independently of the paper's own analyses.
  • Synthetic reflective abstraction
    purpose: To map the observed token-selection process onto Piaget's three stages and argue that the aggregation function performs reflective abstraction.
    An interpretive label applied to the aggregation function; no independent test distinguishes it from simpler descriptions.
  • Synthetic concepts-in-act and theorems-in-act
    purpose: To give epistemological status to neuron categories and aggregation functions.
    Philosophical framing without operational definition beyond the paper's token-based measures.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Process of Categorical Clipping at the Core of the Genesis of Concepts in Synthetic Neural Cognition." pith.science (2026). https://pith.science/paper/V4F6MJHJ

@misc{pith2026250215710,
  author       = {Pith},
  title        = {Pith review of: The Process of Categorical Clipping at the Core of the Genesis of Concepts in Synthetic Neural Cognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V4F6MJHJ}},
  note         = {Machine review of arXiv:2502.15710}
}
read the original abstract

This article investigates, within the field of neuropsychology of artificial intelligence, the process of categorical segmentation performed by language models. This process involves, across different neural layers, the creation of new functional categorical dimensions to analyze the input textual data and perform the required tasks. Each neuron in a multilayer perceptron (MLP) network is associated with a specific category, generated by three factors carried by the neural aggregation function: categorical priming, categorical attention, and categorical phasing. At each new layer, these factors govern the formation of new categories derived from the categories of precursor neurons. Through a process of categorical clipping, these new categories are created by selectively extracting specific subdimensions from the preceding categories, constructing a distinction between a form and a categorical background. We explore several cognitive characteristics of this synthetic clipping in an exploratory manner: categorical reduction, categorical selectivity, separation of initial embedding dimensions, and segmentation of categorical zones.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

150 extracted references · 47 canonical work pages

  1. [1]

    R., Hansen, M., Iarosz, K

    Protachevicz, P. R., Hansen, M., Iarosz, K. C., Caldas, I. L., Batista, A. M., & Kurths, J. (2021). Emergence of neuronal synchronisation in coupledareas. Frontiers in Computational Neuroscience,15,663408.DOI: 10.3389/fncom.2021.663408

  2. [2]

    Schmalzried, M. (2024). The need of a self for self-driving cars : a theoretical model applying homeostasis to self driving. arXiv preprint arXiv:2407.12795. DOI : 10.48550/arXiv.2407.12795

  3. [4]

    G., Lioma, C., & Augenstein, I

    Atanasova, P., Simonsen, J. G., Lioma, C., & Augenstein, I. (2020). Ge- nerating Fact Checking Explanations. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics(pp.7352–7364). Association for Computational Linguistics. DOI : 10.18653/v1/2020.acl- main.656

  4. [5]

    Barkan, R. (2021). The Role of Cognitive Biases in Human Decision Making. Journal of Behavioral Decision Making, 34(3), 243–255. DOI : 10.1002/bdm.2210

  5. [6]

    Barr, W., & Bieliauskas, L. A. (2024). Neuropsychology of Decision Ma- king : A Clinical Perspective.Neuropsychology Review, 34(1), 1–15. DOI : 10.1007/s11065-023-09500-1

  6. [8]

    Will You Find These Shortcuts?

    Bastings, J., Ebert, S., Zablotskaia, P., Sandholm, A., & Filippova, K. (2022). “Will You Find These Shortcuts? ” A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification.Procee- dings of the 2022 Conference on Empirical Methods in Natural Language Processing. https://doi.org/10.18653/v1/2022.emnlp-main.64

  7. [9]

    DOI : 10.1007/s10462-023-10123-4

    Bathia,N.,&Richie,D.(2024).AdvancesinReinforcementLearning:Ap- plications and Challenges.Artificial Intelligence Review, 57(2), 123–145. DOI : 10.1007/s10462-023-10123-4

  8. [10]

    L. W. Barsalou. Simulation, situated conceptualization, and prediction. Philosophical Transactions of the Royal Society B : Biological Sciences, 364(1521), 1281–1289, 2009

Show all 150 references
  1. [11]

    Beaufils, M. (1996). Les réseaux de neurones artificiels : Modèles et applications. Revue d’Intelligence Artificielle , 10(4), 365–387. DOI : 10.1016/S0992-499X(97)80001-2

  2. [12]

    (2023).Language models can explain neurons in language models

    Bills, S., Cammarata, N., Mossing, D., Saunders, W., Wu, J., Tillman, H., Gao, L., Goh, G., Sutskever, I., & Leike, J. (2023).Language models can explain neurons in language models. OpenAI. https://openaipublic. blob.core.windows.net/neuron-explainer/paper/index.html

  3. [13]

    Umang Bhatt, Adrian Weller, and José M. F. Moura. Evaluating and aggregating feature-based model explanations. InProceedings of the In- ternational Joint Conference on Artificial Intelligence (IJCAI), 2020

  4. [14]

    (1992).Grand dictionnaire de la psychologie

    Bloch, H. (1992).Grand dictionnaire de la psychologie

  5. [15]

    (2020).Where Words Get Their Meaning : Cognitive Pro- cessing and Distributional Modelling of Word Meaning

    Bolognesi, M. (2020).Where Words Get Their Meaning : Cognitive Pro- cessing and Distributional Modelling of Word Meaning. John Benjamins Publishing Company. DOI : 10.1075/ftl.7

  6. [16]

    Bricken, T., Schaeffer, R., Olshausen, B., & Kreiman, G. (2023). Emer- gence of Sparse Representations from Noise. Proceedings of the 40th International Conference on Machine Learning, in Proceedings of Ma- chine Learning Research, 202 :3148-3191. Available from https:// proce...

  7. [17]

    B., & Kuperman, V

    Brysbaert, M., Warriner, A. B., & Kuperman, V. (2014). Concreteness ratings for 40 thousand generally known English word lemmas.Behavior Research Methods, 46, 904–911

  8. [19]

    L., & Piaget, J

    Campbell, R. L., & Piaget, J. (2014).Studies in Reflecting Abstraction. Psychology Press

  9. [20]

    Canales-Johnson, A., Silva, C., Huepe, D., Rivera-Rei, Á., Noreika, V., Del Carmen Garcia, M., Silva, W., Vaucheret, E., Sedeño, L., Couto, B., Melloni, M., Ibáñez, A., Chennu, S., Bekinschtein, T. A. (2015). Auditory feedback differentially modulates behavioral and neural mar...

  10. [21]

    S. Carey. Precis ofThe Origin of Concepts.Behavioral and Brain Sciences, 34(3) :113, 2011

  11. [23]

    Chao, L. L. (2024). Advances in Neuroimaging Techniques for Cogni- tive Neuroscience.Journal of Cognitive Neuroscience, 36(1), 1–15. DOI : 10.1162/jocn_a_01700

  12. [24]

    Clark, S., Lerchner, A., von Glehn, T., Tieleman, O., Tanburn, R., Da- shevskiy, M., & Bosnjak, M. (2021). Formalising Concepts as Grounded Abstractions. arXiv preprintarXiv:2101.05125

  13. [25]

    M., & Quillian, M

    Collins, A. M., & Quillian, M. R. (1969). Retrieval time from semantic memory. Journal of Verbal Learning and Verbal Behavior, 8(2), 240–247. https://doi.org/10.1016/s0022-5371(69)80069-1

  14. [26]

    M., & Loftus, E

    Collins, A. M., & Loftus, E. F. (1975). A spreading activation theory of semantic processing.Psychological Review, 82(6), 407–428

  15. [27]

    Annual Review of Psychology, 75, 1–25

    Cowan,N.(2024).WorkingMemoryCapacity:TheoriesandApplications. Annual Review of Psychology, 75, 1–25. DOI : 10.1146/annurev-psych- 010723-120001

  16. [28]

    Cuccio, V., & Gallese, V. (2018). A Peircean account of concepts : groun- ding abstraction in phylogeny through a comparative neuroscientific pers- pective. Philosophical Transactions of the Royal Society B : Biological Sciences, 373(1752), 20170128. https ://doi.org/10.1098/r...

  17. [29]

    Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., & Wei, F. (2022). Know- ledge Neurons in Pretrained Transformers.Proceedings of the 60th An- nual Meeting of the Association for Computational Linguistics (Volume 1 : Long Papers). https://doi.org/10.18653/v1/2022.acl-long.581 41

  18. [30]

    A., & Glass, J

    Dalvi, F., Durrani, N., Sajjad, H., Belinkov, Y., Bau, D. A., & Glass, J. (2019, January). What is one grain of sand in the desert? Analyzing individualneuronsindeepNLPmodels.In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence (AAAI, Oral presentation)

  19. [31]

    R., Alam, F., Durrani, N., Xu, J., & Sajjad, H

    Dalvi, F., Khan, A. R., Alam, F., Durrani, N., Xu, J., & Sajjad, H. (2022). Discovering Latent Concepts Learned in BERT. In In- ternational Conference on Learning Representations (ICLR) . DOI : 10.48550/arXiv.2201.10020

  20. [32]

    Danilevsky, M., Qian, K., Aharonov, R., Katsis, Y., Kawas, B., & Sen, P. (2020). A Survey of the State of Explainable AI for Natural Language Pro- cessing. arXiv (Cornell University).https://doi.org/10.48550/arxiv. 2010.00711

  21. [33]

    A., Durrani, N., Sajjad, H., Dalvi, F., & Belinkov, Y

    Dar, S. A., Durrani, N., Sajjad, H., Dalvi, F., & Belinkov, Y. (2023). Probing Pre-trained Language Models for Temporal Knowledge. InPro- ceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL). DOI : 10.18653/v1/2023.acl-long.123

  22. [34]

    Symbols and mental programs : a hypothesis about human singularity

    Stanislas Dehaene, Fosca Al Roumi, Yair Lakretz, Samuel Planton, and Mathias Sablé-Meyer. Symbols and mental programs : a hypothesis about human singularity. Trends in Cognitive Sciences, 26(9) :751–766, 2022. ISSN 1364-6613. doi : https://doi.org/10.1016/j.tics.2022.06

  23. [35]

    URL https://www.sciencedirect.com/science/article/pii/ S1364661322001413

  24. [36]

    S., Lee, J

    Du, S. S., Lee, J. D., Li, H., Wang, L., & Zhai, (2019). Gradient descent finds globalminima of deep neural networks, 1675-1685

  25. [37]

    F., & Cabi, S

    Du, Y., Konyushkova, K., Denil, M., Raju, A., Landon, J., Hill, F., Nando, D. F., & Cabi, S. (2023).Vision-Language Models as Success Detectors. arXiv (Cornell University). https ://doi.org/10.48550/arxiv.2303.07280

  26. [38]

    Duncan, J. (1984). Selective Attention and the Organization of Visual Information. Journal of Experimental Psychology : General, 113(4), 501-

  27. [39]

    Echterhoff, J., Yan, A., Han, K., Abdelraouf, A., Gupta, R., & McAu- ley, J. (2024). Driving through the Concept Gridlock : Unraveling Explainability Bottlenecks in Automated Driving . Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). htt...

  28. [40]

    Durrani, N., Sajjad, H., Dalvi, F., & Belinkov, Y. (2022). On the Transfor- mation of Latent Space in Fine-Tuned NLP Models. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP). DOI : 10.18653/v1/2022.emnlp-main.123

  29. [41]

    Epstein, Eva Zita Patai, Joshua B

    Russell A. Epstein, Eva Zita Patai, Joshua B. Julian, and Hugo J. Spiers. The cognitive map in humans : spatial navigation and beyond.Nature Neuroscience, 20(11) :1504–1513, 2017

  30. [42]

    Enguehard, J. (2023). Extrmask : A Method for Explaining Time Se- ries Predictions by Masking. arXiv preprint arXiv:2301.08552. DOI : 10.48550/arXiv.2301.08552. 42

  31. [43]

    Fan, Y., Dalvi, F., Durrani, N., & Sajjad, H. (2023). Evaluating Neu- ron Interpretation Methods of NLP Models. arXiv (Cornell University). https ://doi.org/10.48550/arxiv.2301.12608

  32. [44]

    W., & Keane, M

    Eysenck, M. W., & Keane, M. T. (2020). Cognitive Psycho- logy : A Student’s Handbook (8th ed.). Psychology Press. DOI : 10.4324/9780429449229

  33. [45]

    A Holistic Approach to Unifying Automa- tic Concept Extraction and Concept Importance Estimation,

    Fel, J., Smith, A., & Wang, T., "A Holistic Approach to Unifying Automa- tic Concept Extraction and Concept Importance Estimation," inProcee- dings of the 37th Conference on Neural Information Processing Systems (NeurIPS 2023), 2024

  34. [46]

    A Holistic Approach to Unifying Automa- tic Concept Extraction and Concept Importance Estimation,

    Fel, J., Smith, A., & Wang, T., "A Holistic Approach to Unifying Automa- tic Concept Extraction and Concept Importance Estimation," inProcee- dings of the 37th Conference on Neural Information Processing Systems (NeurIPS 2023), 2023

  35. [47]

    Geva, M., Schuster, R., Berant, J., & Levy, O. (2023). Transformer Feed- Forward Layers Are Key-Value Memories. In Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS). DOI : 10.48550/arXiv.2012.14913

  36. [48]

    Funayama, T., & Shibata, K. (2024). Advances in Quantum Compu- ting:AComprehensiveReview. Journal of Quantum Information Science, 12(1), 45–67. DOI : 10.4236/jqis.2024.121004

  37. [49]

    (2002).Radical Constructivism : A Way of Knowing and Learning

    Von Glaserfeld, E. (2002).Radical Constructivism : A Way of Knowing and Learning. London : RoutledgeFalmer

  38. [50]

    I., Benjamin, A

    Glaser, J. I., Benjamin, A. S., Chowdhury, R. H., Perich, M. G., Mil- ler, L. E., & Kording, K. P. (2020). Machine learning for neural deco- ding. eNeuro, 7(4), ENEURO.0506-19.2020.https://doi.org/10.1523/ eneuro.0506-19.2020

  39. [51]

    Shifting attention to orient or avoid : a unifying account of the tail of the striatum and its dopaminergic inputs.Current Opinion in Behavioral Sciences, 59, 101441, 2024

    Green, Isobel, Ryunosuke Amo, and Mitsuko Watabe-Uchida. Shifting attention to orient or avoid : a unifying account of the tail of the striatum and its dopaminergic inputs.Current Opinion in Behavioral Sciences, 59, 101441, 2024

  40. [52]

    Gresch, D., & Müller, K. (2024). Machine Learning in Materials Science : Recent Progress and Emerging Applications.Advanced Materials, 36(5), 2105678. DOI : 10.1002/adma.202105678

  41. [53]

    Zou, and Been Kim

    Amirata Ghorbani, James Wexler, James Y. Zou, and Been Kim. Towards automatic concept-based explanations. InAdvances in Neural Information Processing Systems, pages 9273–9282, 2019. 43

  42. [54]

    Goodman, J

    N. Goodman, J. B. Tenenbaum, and T. Gerstenberg. Concepts in a pro- babilistic language of thought. In E. Margolis and S. Laurence, editors, The Conceptual Mind : New Directions in the Study of Concepts, pages 623–654. The MIT Press, 2015

  43. [55]

    A., Reicher, S

    Haslam, S. A., Reicher, S. D., & Platow, M. J. (2020).The New Psychology of Leadership : Identity, Influence, and Power(2nd ed.). Routledge. DOI : 10.4324/9781351108225

  44. [56]

    Harrison, S., Gualdoni, E., & Boleda, G. (2023). Run like a girl! Sports-related gender bias in language and vision. arXiv preprint arXiv:2305.14468

  45. [57]

    P., Glorot, X., Botvinick, M., & Lerchner, A

    Higgins, I., Matthey, L., Pal, A., Burgess, C. P., Glorot, X., Botvinick, M., & Lerchner, A. (2017).β-VAE : Learning basic visual concepts with a constrained variational framework. InProceedings of ICLR 2017

  46. [58]

    A., & Pérez-González, J

    Hernández-Gutiérrez, C. A., & Pérez-González, J. (2024). Deep Learning Techniques for Natural Language Processing : A Survey.IEEE Transac- tions on Neural Networks and Learning Systems, 35(2), 1234–1256. DOI : 10.1109/TNNLS.2023.3101234

  47. [59]

    Hofstadter and E

    D. Hofstadter and E. Sander.Surfaces and Essences. Basic Books, 2013

  48. [60]

    Higgins, I., Amos, D., Pfau, D., Racaniere, S., Matthey, L., Rezende, D., & Lerchner, A. (2018). Towards a definition of disentangled representations. arXiv preprintarXiv:1812.02230

  49. [61]

    Effects of manipulating prefrontal activity and dopamine D1 receptor signaling in an appetitive feature-negative discri- mination learning task.Behavioral Neuroscience, 2024

    Hock, Rebecca M., et al. Effects of manipulating prefrontal activity and dopamine D1 receptor signaling in an appetitive feature-negative discri- mination learning task.Behavioral Neuroscience, 2024

  50. [62]

    N., & Love, B

    Hornsby, A. N., & Love, B. C. (2020). How decisions and the desire for coherency shape subjective preferences over time.Cognition, 200, 104244. https://doi.org/10.1016/j.cognition.2020.104244

  51. [63]

    Howell, D. C. (2024). Méthodes statistiques en sciences humaines. De Boeck Supérieur

  52. [64]

    Howell, D. C. (2008). Fundamental Statistics for the Behavioral Sciences (6th ed.). Wadsworth Publishing. DOI : 10.1111/j.1467- 985X.2008.00508_14.x

  53. [65]

    Capuano, F., & Kaup, B. (2024). Pragmatic Reasoning in GPT Models : Replication of a Subtle Negation Effect. Proceedings of the Annual Mee- ting of the Cognitive Science Society, 46. Retrieved from https ://escho- larship.org/uc/item/22q5920s

  54. [66]

    Large Language Models Struggle to Learn Long-Tail Knowledge

    Kandpal,N.,Deng,H.,Roberts,A.,Wallace,E.,&Raffel,C.(2023). Large Language Models Struggle to Learn Long-Tail Knowledge. arXiv (Cornell University). https ://doi.org/10.48550/arxiv.2211.08411

  55. [67]

    On the necessity of abstraction

    George Konidaris. On the necessity of abstraction. Current Opinion in Behavioral Sciences, 29 :1–7, October 2019. ISSN 2352-1546. doi : 10.1016/j.cobeha.2018.11.005. URL https://www.sciencedirect.com/ science/article/pii/S2352154618302080. 44

  56. [68]

    A., Bouadjenek, M

    Kheya, T. A., Bouadjenek, M. R., & Aryal, S. (2024). The Pursuit of Fairness in Artificial Intelligence Models : A Survey. arXiv (Cornell Uni- versity). https://doi.org/10.48550/arxiv.2403.17333

  57. [69]

    Fauconnier.Mappings in Thought and Language

    G. Fauconnier.Mappings in Thought and Language. Cambridge University Press, 1997

  58. [70]

    G. Lakoff. The contemporary theory of metaphor. In Metaphor and Thought, pages 202–251. Cambridge University Press, 2008

  59. [71]

    A., Ide, I., Nack, F., Kawanishi, Y., Hirayama, T., Degu- chi, D., & Murase, H

    Kastner, M. A., Ide, I., Nack, F., Kawanishi, Y., Hirayama, T., Degu- chi, D., & Murase, H. (2020). Estimating the imageability of words by mining visual characteristics from crawled image data.Multimedia Tools and Applications, 79(25), 18167–18199

  60. [72]

    M. Johnson. Embodied Mind, Meaning, and Reason : How Our Bodies Give Rise to Understanding. University of Chicago Press, Chicago, IL, 2017

  61. [73]

    H., Zdon, A., Fraga, N

    Love, A. H., Zdon, A., Fraga, N. S., Cohen, B., Mejia, M. P., Maxwell, R., & Parker, S. S. (2022). Statistical evaluation of the similarity of cha- racteristics in springs of the California Desert, United States.Frontiers in Environmental Science, 10. https://doi.org/10.3389/f...

  62. [74]

    Killian and Elizabeth A

    Nathaniel J. Killian and Elizabeth A. Buffalo. Grid cells map the visual world. Nature Neuroscience, 21(2), 2018

  63. [75]

    Lynott,D.,Connell,L.,Brysbaert,M.,Brand,J.,&Carney,J.(2020).The Lancaster Sensorimotor Norms : Multidimensional measures of perceptual and action strength for 40,000 English words.Behavior Research Methods, 52, 1271–1291

  64. [76]

    Luo, J., Zhuo, W., Liu, S., & Xu, B. (2024). The Optimi- zation of Carbon Emission Prediction in Low Carbon Energy Economy under Big Data . IEEE Access, 12, 14690-14702. https ://doi.org/10.1109/access.2024.3351468

  65. [77]

    C., & Grienberger, C

    Magee, J. C., & Grienberger, C. (2020). Synaptic plasticity forms and functions. Annual Review of Neuroscience, 43(1), 95–117

  66. [78]

    C., Tsoi, L

    Ma, F., Plazyo, O., Billi, A. C., Tsoi, L. C., Xing, X., Wasikowski, R., Gharaee-Kermani, M., Hile, G., Jiang, Y., Harms, P. W., Xing, E., Kirma, J., Xi, J., Hsu, J., Sarkar, M. K., Chung, Y., Di Domizio, J., Gilliet, M., Ward, N. L., et al. (2023). Single cell and spatial seq...

  67. [79]

    Marty, P., Romoli, J., Sudo, Y., & Breheny, R. (2024). Implicature pri- ming, salience, and context adaptation. Cognition, 244, 105667. DOI : 10.1016/j.cognition.2023.105667

  68. [80]

    Marconato, E., & al. (2024). BEARS Make Neuro-Symbolic Models Aware of their Reasoning Shortcuts. arXiv preprint arXiv:2402.12240. DOI : 10.48550/arXiv.2402.12240

  69. [82]

    Margolis, E., & Laurence, S. (2019). Concepts. In E. N. Zalta (Ed.), The Stanford encyclopedia of philosophy(Summer 2019 ed.). Metaphy- 45 sics Research Lab, Stanford University.https://plato.stanford.edu/ archives/sum2019/entries/concepts/

  70. [83]

    J., Johnson, M., & Steed- man,M.(2023).SourcesofHallucinationbyLargeLanguageModelsonIn- ference Tasks

    McKenna, N., Li, T., Cheng, L., Hosseini, M. J., Johnson, M., & Steed- man,M.(2023).SourcesofHallucinationbyLargeLanguageModelsonIn- ference Tasks. arXiv (Cornell University).https://doi.org/10.48550/ arxiv.2305.14552

  71. [84]

    Mitchell, M. (2021). Abstraction and analogy-making in artificial intel- ligence. Annals of the New York Academy of Sciences, 1505(1), 79-101. DOI : 10.1111/nyas.14619

  72. [85]

    (1994).Piaget ou l’intelligence en marche : aperçu chronologique et vocabulaire

    Montangero, J., & Maurice-Naville, D. (1994).Piaget ou l’intelligence en marche : aperçu chronologique et vocabulaire. Editions Mardaga

  73. [86]

    Mousi, B., Durrani, N., & Dalvi, F. (2023). Can LLMs facilitate interpre- tation of pre-trained language models?arXiv preprint arXiv:2305.13386. DOI : 10.48550/arXiv.2305.13386

  74. [87]

    (1999).Vocabulaire technique et analytique de l’épistémologie

    Nadeau, R. (1999).Vocabulaire technique et analytique de l’épistémologie. Presses universitaires de France

  75. [88]

    Moser, May-Britt Moser, and Bruce L

    Edvard I. Moser, May-Britt Moser, and Bruce L. McNaughton. Spatial representation in the hippocampal formation : a history.Nature Neuros- cience, 20(11) :1448–1464, 2017

  76. [89]

    Nisa, et al. (2020). Advances in Social Science, Education and Humanities Research, volume 574

  77. [90]

    (1999).Vocabulaire technique et analytique de l’épistémologie

    Nadeau, R. (1999).Vocabulaire technique et analytique de l’épistémologie. Presses Universitaires de France

  78. [91]

    Nanda, N., Lee, A., & Wattenberg, M. (2023). Emergent linear represen- tations in world models of self-supervised sequence models.arXiv preprint arXiv:2309.00941. DOI : 10.48550/arXiv.2309.00941

  79. [92]

    M., Meagher, B

    Nosofsky, R. M., Meagher, B. J., & Kumar, P. (2022). Contrasting exem- plar and prototype models in a natural-science category domain.Journal of Experimental Psychology : Learning, Memory, and Cognition, 48(12), 1970–1994. https://doi.org/10.1037/xlm0001069

  80. [93]

    O’Keefe and J

    J. O’Keefe and J. Dostrovsky. The hippocampus as a spatial map. Prelimi- nary evidence from unit activity in the freely-moving rat.Brain Research, 34(1) :171–175, 1971

  81. [94]

    Nosofsky, R. M. (1986). Attention, similarity, and the identifica- tion–categorizationrelationship. Journal of Experimental Psychology : Ge- neral, 115(1), 39

  82. [95]

    Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., & Carter, S. (2020). Zoom In : An Introduction to Circuits. Retrieved fromhttps: //distill.pub/2020/circuits/zoom-in/. Accessed 24-11-2023

  83. [96]

    The Hippocampus as a Cognitive Map

    John O’Keefe and Lynn Nadel. The Hippocampus as a Cognitive Map. Oxford University Press, 1978. 46

  84. [97]

    Paolo, G., Gonzalez-Billandon, J., & Kégl, B. (2024). A call for embodied AI. arXiv preprint arXiv:2402.03824. DOI : 10.48550/arXiv.2402.03824

  85. [98]

    Olah, C. (2023). Distributed Representations : Composition & Super- position. Retrieved from https://transformer-circuits.pub/2023/ superposition-composition/index.html

  86. [99]

    (2002).Composantes intra-individuelle et contractuelle de la conceptualisation mathématique en situation didactique de traitement de tâches algébriques (thèse de doctorat)

    Pichat, M. (2002).Composantes intra-individuelle et contractuelle de la conceptualisation mathématique en situation didactique de traitement de tâches algébriques (thèse de doctorat). Saint-Denis : Université Paris 8

  87. [100]

    (1974).Adaptation vitale et psychologie de l’intelligence

    Piaget, J. (1974).Adaptation vitale et psychologie de l’intelligence. Paris : Hermann

  88. [101]

    Pichat, M. (2023). Collaboration des intelligences humaine et artificielle : alignement et psychologie de l’IA. Actes du colloqueIntelligence artifi- cielle collaborative & impacts managériaux au sein des organisationsdu 30/06/2023 coorganisé par l’Université Paris Dauphine-PS...

  89. [102]

    (2007).Psychologie de l’éducation

    Pichat, M., & Merri, M. (2007).Psychologie de l’éducation. Paris : Bréal

  90. [103]

    Pichat, M. (2024). Psychology of Artificial Intelligence : Epistemological Markers of the Cognitive Analysis of Neural Networks. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2407.09563

  91. [104]

    Pichat, M. (2024a). Psychologie de l’IA et alignement cognitif. Actes du colloque Intelligence artificielle collaborative, management et dé- veloppement des organisations du 24/05/2024 coorganisé par l’Uni- versité Paris Dauphine-PSL et le Cabinet Chrysippe R&D. Avai- lable on...

  92. [105]

    (2024).Neuropsy- chology and Explainability of AI : A Distributional Approach to the Rela- tionship Between Activation Similarity of Neural Categories in Synthetic Cognition

    Pichat, M., Campoli, E., Pogrund, W., Wilson, J., Veillet-Guillem, M., Melkozerov, A., Gasparian, A., Pichat, P., Poumay, J. (2024).Neuropsy- chology and Explainability of AI : A Distributional Approach to the Rela- tionship Between Activation Similarity of Neural Categories i...

  93. [106]

    (2024).Neuropsy- chology of AI : Relationship Between Activation Proximity and Catego- rical Proximity Within Neural Categories of Synthetic Cognition

    Pichat, M., Campoli, E., Pogrund, W., Wilson, J., Veillet-Guillem, M., Melkozerov, A., Gasparian, A., Pichat, P., Poumay, J. (2024).Neuropsy- chology of AI : Relationship Between Activation Proximity and Catego- rical Proximity Within Neural Categories of Synthetic Cognition. ...

  94. [107]

    Ontology Concept Extraction Algo- rithm for Deep Neural Networks,

    A. Ponomarev and A. Agafonov, "Ontology Concept Extraction Algo- rithm for Deep Neural Networks," inProceedings of the 32nd Conference of Open Innovations Association (FRUCT),IEEE,2022,pp.221–226.doi: https://doi.org/10.23919/FRUCT56874.2022.9953838

  95. [108]

    (2024).How Do Artificial Intelligences Think? The Three Mathematico-Cognitive Factors of Categorical Segmentation Ope- rated by Synthetic Neurons

    Pichat, M., Pogrund, W., Gasparian, A., Pichat, P., Demarchi, S., & Veillet-Guillem, M. (2024).How Do Artificial Intelligences Think? The Three Mathematico-Cognitive Factors of Categorical Segmentation Ope- rated by Synthetic Neurons.. arXiv preprint 47

  96. [109]

    I., & Snyder, C

    Posner, M. I., & Snyder, C. R. R. (1975). Attention and Cognitive Control. In R. L. Solso (Ed.), Information Processing and Cognition : The Loyola Symposium(pp. 55-85). Lawrence Erlbaum Associates. DOI : 10.4324/9781315784786

  97. [110]

    Posner, M. I. (1978).Chronometric Explorations of Mind. Lawrence Erl- baum Associates

  98. [111]

    Raieli, S., Altahhan, A., Jeanray, N., Gerart, S., & Vachenc, S. (2024). Escaping the Forest : Sparse Interpretable Neural Networks for Tabular Data.arXiv preprint arXiv:2410.17758.DOI:10.48550/arXiv.2410.17758

  99. [112]

    Pulvermuller,F.(2018).NeurobiologicalMechanismsforSemanticFeature Extraction and Conceptual Flexibility.Topics in Cognitive Science, 10, 590–620

  100. [113]

    Richard, J. C. (1980).The Language Teaching Matrix. Cambridge Univer- sity Press

  101. [114]

    Ribary, U., & Ward, L. M. (2024). Synchronization and functional connec- tivity dynamics across TC-CC-CT networks : Implications for clinical symptoms and consciousness. InPhenomenological Neuropsychiatry : How Patient Experience Bridges the Clinic with Clinical Neuroscience (...

  102. [115]

    E. Rosch. Principles of categorization. In E. Margolis and S. Laurence, editors, Concepts : Core Readings, pages 189–206. MIT Press, 1999

  103. [116]

    D., & Love, B

    Roads, B. D., & Love, B. C. (2024). Modeling Similarity and Psy- chological Space. Annual Review of Psychology, 75(1), 215–240. DOI : 10.1146/annurev-psych-040323-115131

  104. [117]

    (2010).Introduc- tion aux réseaux neuronaux : de la synapse à la psyché

    Savioz, A., Leuba, G., Vallet Philippe, G., & Walzer, C. (2010).Introduc- tion aux réseaux neuronaux : de la synapse à la psyché. De Boeck

  105. [118]

    Rzechorzek, A. (2024). Understanding Cognitive Processes : Insights from Recent Research. Journal of Cognitive Neuroscience . DOI : 10.1162/jocn_a_01678

  106. [119]

    Servan-Schreiber, D., Printz, H., & Cohen, J. D. (1990). A network model of catecholamine effects : gain, signal-to-noise ratio, and behavior.Science, 249(4971), 892–895

  107. [120]

    Psychological Review, 84(1), 1-66

    Schneider,W.,&Shiffrin,R.M.(1977).ControlledandAutomaticHuman InformationProcessing:I.Detection,Search,andAttention. Psychological Review, 84(1), 1-66

  108. [121]

    E. S. Spelke and K. D. Kinzler. Core knowledge.Developmental Science, 10(1) :89–96, 2007

  109. [122]

    Shavikloo, M., Esmaeili, A., Valizadeh, A., & Madadi Asl, M. (2024). Syn- chronization of delayed coupled neurons with multiple synaptic connec- tions. Cognitive Neurodynamics, 18(2), 631-643. DOI : 10.1007/s11571- 023-10013-9. 48

  110. [123]

    Singh, V., Gupta, I., & Jana, P. K. (2020). An energy efficient algorithm for workflow scheduling in IaaS cloud.Journal of Grid Computing, 18(3), 357–376. https://doi.org/10.1007/s10723-019-09490-2

  111. [124]

    Stoewer, P., Schilling, A., Maier, A., & Krauss, P. (2022). Neural network based formation of cognitive maps of semantic spaces and the emergence of abstract concepts.arXiv preprintarXiv:2210.16062

  112. [125]

    Tang, Y., Bi, J., Xu, S., Song, L., Liang, S., Wang, T., ... & Xu, C. (2023). Videounderstandingwithlargelanguagemodels:Asurvey. arXiv preprint arXiv:2312.17432

  113. [126]

    AligningArtificialNeuralNetworksand Ontologies towards Explainable AI,

    M.deSousaRibeiroandJ.Leite,"AligningArtificialNeuralNetworksand Ontologies towards Explainable AI," inProceedings of the AAAI Confe- rence on Artificial Intelligence, vol. 35, no. 6, pp. 4932–4940, 2021

  114. [127]

    Tipper, S. P. (1985). The Negative Priming Effect : Inhibitory Priming by Ignored Objects.The Quarterly Journal of Experimental Psychology, 37A(4), 571-590. DOI : 10.1080/14640748508400920

  115. [128]

    Tater, T., Walde, S. S. I., & Frassinelli, D. (2024). Unveiling the mystery of visual attributes of concrete and abstract concepts : Variability, nearest neighbors, and challenging categories.arXiv preprintarXiv:2410.11657

  116. [129]

    R., Hassid, M., Heafield, K., Hooker, S., Raffel, C., Martins, P

    Treviso, M., Lee, J., Ji, T., Van Aken, B., Cao, Q., Ciosici, M. R., Hassid, M., Heafield, K., Hooker, S., Raffel, C., Martins, P. H., Martins, A. F. T., Forde, J. Z., Milder, P., Simpson, E., Slonim, N., Dodge, J., Strubell, E., Balasubramanian, N.,. . . Schwartz, R. (2023). ...

  117. [130]

    Treisman, A., & Gelade, G. (1980). A Feature-Integration Theory of Attention. Cognitive Psychology, 12(1), 97-136. DOI : 10.1016/0010- 0285(80)90005-5

  118. [131]

    Varela, F. J. (1988).Cognitive Science : A Cartography of Current Ideas. MIT Press.Varela1996

  119. [132]

    Varela, F. (1984). The creative circle. In P. Watzlawick (Ed),The invented reality. London : W W Norton & Co Inc

  120. [133]

    Questions à Gérard Ver- gnaud (pp

    Vergnaud,G.(2009).Activité,développement,représentation.InM.Merri (Ed.), Activité humaine et conceptualisation. Questions à Gérard Ver- gnaud (pp. 149–154). Presses universitaires du Mirail

  121. [134]

    Varela, F. J. (1996). Invitation aux sciences cognitives. Édi- tions du Seuil eBooks. http://inventin.lautre.net/livres/ Varela-Invitation-aux-sciences-cognitives.pdf

  122. [135]

    Vergnaud, G. (2020a). Héritages (pp. 27–37). In M. Merri (Ed.),Activité humaines et conceptualization. Toulouse : Presses universitaires du Midi

  123. [136]

    Vergnaud, G. (2016). Relations entre conceptualisations dans l’ac- tion et signifiants langagiers et symboliques. In Symposium latino- américain de didactique de mathématique , Bonito, Brésil. Disponible 49 sur : https://www.gerard-vergnaud.org/texts/gvergnaud_2016_ signifiant...

  124. [137]

    Vergnaud, G. (2020c). A Classification of Cognitive Tasks and Operations of Thought Involved in Addition and Subtraction Problems. In P. Car- penter, M. Moser, & A. Romberg (Eds.),Addition and Subtraction : A Cognitive Perspective. London : Routledge

  125. [138]

    Vergnaud, G. (2020b). Réponses de Gérard Vergnaud (pp. 341–357). In M. Merri (Ed.),Activité humaines et conceptualization. Toulouse : Presses universitaires du Midi

  126. [139]

    Voita, E., Sennrich, R., & Titov, I. (2021). Language modeling, lexi- cal translation, reordering : The training process of NMT through the lens of classical SMT. arXiv preprint arXiv:2109.01396 . DOI : 10.48550/arXiv.2109.01396

  127. [140]

    Vogel, T., Ingendahl, M., & Winkielman, P. (2021). The architecture of prototype preferences : Typicality, fluency, and valence.Journal of Ex- perimental Psychology : General, 150(1), 187–194.https://doi.org/10. 1037/xge0000798

  128. [141]

    Watzlawick, P. (1977). How real is real? London : Vintage Books

  129. [142]

    Kokhlikyan, N., Miglani, V., Martin, M., Wang, E., Reynolds, J., Melni- kov, A., Lunova, N., & Reblitz-Richardson, O. (2020). Captum : A uni- fied and generic model interpretability library for PyTorch.arXiv preprint arXiv:2009.07896. DOI : 10.48550/arXiv.2009.07896

  130. [143]

    pyOptSparse : A Python framework for large-scale constrained nonlinear optimization of sparse systems

    Wu et al., (2020). pyOptSparse : A Python framework for large-scale constrained nonlinear optimization of sparse systems. Journal of Open Source Software, 5(54), 2564. DOI : 10.21105/joss.02564

  131. [144]

    H., & Fisch, R

    Watzlawick, P., Weakland, J. H., & Fisch, R. (1984).Change : Principles of Problem Formation and Problem Resolution. W. W. Norton & Com- pany. DOI : 10.1002/9781119164894

  132. [145]

    (2024).We know what attention is!

    Wu, W. (2024).We know what attention is!. Trends in Cognitive Sciences, 28(4), 304-318

  133. [146]

    (2022).Automatic detection and severity analysis of grape black measles disease based on deep learning and fuzzy logic

    Ji, M., & Wu, Z. (2022).Automatic detection and severity analysis of grape black measles disease based on deep learning and fuzzy logic. Computers and Electronics in Agriculture, 193, 106718

  134. [147]

    Xu, W., & Futrell, R. (2024). A hierarchical Bayesian mo- del for syntactic priming. arXiv preprint arXiv:2405.15964 . DOI : 10.48550/arXiv.2405.15964

  135. [148]

    Y., Lillicrap, T

    Xie, Y., Goyal, A., Zheng, W., Kan, M. Y., Lillicrap, T. P., Kawaguchi, K., & Shieh, M. (2024). Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.arXiv preprintarXiv:2405.00451

  136. [149]

    (2002).Inner Vision : An Exploration of Art and the Brain

    Zeki, S. (2002).Inner Vision : An Exploration of Art and the Brain. Ox- ford : Oxford University Press

  137. [150]

    Zadeh, L. A. (1996). Fuzzy Logic = Computing with Words.IEEE Tran- sactions on Fuzzy Systems, 4(2), 103-111. DOI : 10.1109/91.493904 50

  138. [151]

    Zheng, Y., & Stewart, N. (2024). Improving EFL students’ cultural awa- reness : Reframing moral dilemmatic stories with ChatGPT.Computers And Education Artificial Intelligence, 6, 100223. https://doi.org/10. 1016/j.caeai.2024.100223

  139. [152]

    A., Kirko- rian, H., & Lupyan, G

    Zettersten, M., Bredemann, C., Kaul, M., Ellis, K., Vlach, H. A., Kirko- rian, H., & Lupyan, G. (2024). Nameability supports rule-based category learning in children and adults.Child Development, 95(2), 497-514. DOI : 10.1111/cdev.14008

  140. [153]

    Zhang, Z., Song, Y., Yu, G., Han, X., Lin, Y., Xiao, C., ...& Sun, M. (2024). ReLU 2 Wins : Discovering Efficient Activation Functions for Sparse LLMs. arXiv preprint arXiv:2402.03804. DOI : 10.48550/arXiv.2402.03804. 51

  141. [154]

    Zhao, H., Chen, H., Yang, F., Liu, N., Deng, H., Cai, H., Wang, S., Yin, D., & Du, M. (2023). Explainability for Large Language Models : A Survey. arXiv (Cornell University). DOI : 10.48550/arxiv.2309.01029

  142. [517]

    DOI : 10.1037/0096-3445.113.4.501

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.