Pith. sign in

REVIEW 1 major objections 2 minor 2 cited by

CNNs that model the visual system well do not necessarily align better with human texture perception.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-05-21 09:22 UTC pith:D4KKCPYQ

load-bearing objection The null correlation between Brain-Score and human texture alignment is the main result, but it may rest on Gram matrices missing spatial or higher-order structure that the networks actually use. the 1 major comments →

arxiv 2604.01341 v2 pith:D4KKCPYQ submitted 2026-04-01 cs.CV q-bio.NC

Perceptual misalignment of texture representations in convolutional neural networks

classification cs.CV q-bio.NC
keywords texture perceptionconvolutional neural networksGram matricesBrain-Scorevisual systemperceptual alignmentfeature correlationsJulesz textures
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tests whether convolutional neural networks that are considered good models of the visual system also have texture representations that match human perception. It uses Gram matrices of features from different CNNs to measure how well they capture perceptual texture content and compares this to Brain-Score rankings. The authors find no connection between these, suggesting that texture perception uses mechanisms distinct from those in standard object recognition models. A reader would care because it questions the completeness of current CNN-based models for all aspects of vision.

Core claim

There is no connection between conventional measures of CNN quality as a model of the visual system and its alignment with human texture perception. The study quantifies the perceptual content captured by feature correlations in a diverse pool of CNNs and finds it does not correlate with Brain-Score.

What carries the argument

Gram matrices compiling linear correlations between nonlinear CNN features as texture representations.

Load-bearing premise

Gram matrices of CNN features are a sufficient and unbiased summary of the texture information the network uses for perceptual comparisons.

What would settle it

Finding a CNN with high Brain-Score that also shows strong alignment with human texture judgments, or a low Brain-Score one with high alignment, would test the no-connection claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Texture perception likely involves integration of contextual information not captured by standard CNN feature correlations.
  • Approaches based on CNNs trained on object recognition may miss important aspects of human texture processing.
  • New modeling strategies beyond object recognition training are needed for accurate texture perception models.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Texture perception might rely on different neural pathways or computations than those optimized for object recognition.
  • Testing CNNs on texture-specific tasks could reveal whether alignment improves independently of Brain-Score.
  • These results highlight potential gaps in using Brain-Score as a sole indicator of perceptual fidelity across all visual tasks.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 2 minor

Summary. The manuscript claims that CNNs regarded as better models of the visual system (per Brain-Score) do not exhibit correspondingly better alignment with human texture perception when texture representations are quantified via Gram matrices of CNN features. Across a diverse pool of CNNs, the authors report no correlation between these two quantities and conclude that texture perception depends on mechanisms distinct from those captured by standard object-recognition-trained CNNs, possibly requiring contextual integration.

Significance. If the dissociation holds, the result would indicate that conventional CNN-based models of the visual system, optimized for object recognition, capture only a subset of human perceptual capabilities and that texture judgments may rely on separate computations. The work explicitly leverages the independence of the external Brain-Score benchmark and human texture judgments, yielding a clear, falsifiable prediction about model alignment that can be tested with future architectures.

major comments (1)
  1. [Results (texture alignment quantification)] The central null-correlation result rests on texture alignment being measured exclusively via Gram-matrix summaries of CNN features. No ablation is reported that replaces or augments this statistic with alternatives (feature means, higher-order moments, or spatially-aware descriptors) and re-computes the correlation with Brain-Score. This is load-bearing for the headline claim: if Gram matrices omit phase or higher-order spatial structure that the network actually uses for texture, the observed lack of correlation could be an artifact of the chosen summary rather than evidence of a genuine mechanistic dissociation.
minor comments (2)
  1. [Abstract] The abstract and introduction would benefit from an explicit statement of the exact number of CNN architectures tested, the precise human texture judgment dataset, and the statistical test used for the reported null correlation.
  2. [Methods] Notation for the Gram-matrix construction and the subsequent correlation with human judgments should be defined once in a dedicated methods subsection rather than introduced piecemeal.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for their constructive comments. We address the major comment on the texture alignment quantification below.

read point-by-point responses
  1. Referee: [Results (texture alignment quantification)] The central null-correlation result rests on texture alignment being measured exclusively via Gram-matrix summaries of CNN features. No ablation is reported that replaces or augments this statistic with alternatives (feature means, higher-order moments, or spatially-aware descriptors) and re-computes the correlation with Brain-Score. This is load-bearing for the headline claim: if Gram matrices omit phase or higher-order spatial structure that the network actually uses for texture, the observed lack of correlation could be an artifact of the chosen summary rather than evidence of a genuine mechanistic dissociation.

    Authors: Gram matrices were selected because they constitute the canonical summary of second-order feature correlations in the CNN texture literature, directly extending Julesz's conjecture on local feature statistics and the successful texture synthesis framework of Gatys et al. The study was scoped to test whether better Brain-Score alignment predicts better perceptual alignment under this established representation. We acknowledge that the absence of ablations with alternative descriptors (means, higher-order moments, or spatial descriptors) leaves open the possibility that the null result is specific to Gram matrices. In the revised manuscript we will add an ablation that recomputes texture alignment using first-order moments and spatially-aware statistics and re-evaluates the correlation with Brain-Score. revision: yes

Circularity Check

0 steps flagged

No circularity: empirical null result uses independent external benchmarks

full rationale

The paper's central claim is an observed lack of correlation between Brain-Score (an external benchmark for CNN alignment with mammalian visual system data) and a texture alignment score computed from Gram matrices of CNN features versus human perceptual judgments. Neither quantity is defined in terms of the other, and the comparison is a direct empirical measurement rather than a derivation or prediction that reduces to fitted inputs or self-referential equations by construction. The Gram matrix approach follows standard prior literature applied to a pool of models, with no self-citation load-bearing the null result and no ansatz or uniqueness theorem imported from the authors' own prior work. The derivation chain remains self-contained against these external benchmarks.

Axiom & Free-Parameter Ledger

0 free parameters · 2 axioms · 0 invented entities

The work relies on the standard assumption that Gram matrices capture the relevant texture statistics and that Brain-Score is a valid proxy for visual-system alignment; no new free parameters or invented entities are introduced in the abstract.

axioms (2)
  • domain assumption Feature correlations compiled into Gram matrices provide a sufficient representation of texture content for perceptual comparison.
    Invoked when the authors quantify perceptual content captured by CNN feature correlations.
  • domain assumption Brain-Score is an appropriate measure of a model's alignment with the mammalian visual system.
    Used to compare CNN quality against texture-perception alignment.

pith-pipeline@v0.9.0 · 5739 in / 1250 out tokens · 29071 ms · 2026-05-21T09:22:36.503090+00:00 · methodology

0 comments
read the original abstract

Mathematical modeling of visual textures traces back to Julesz's intuition that texture perception in humans is based on local correlations between image features. An influential approach for texture analysis and generation generalizes this notion to linear correlations between the nonlinear features computed by convolutional neural networks (CNNs), compiled into Gram matrices. Given that CNNs are often used as models for the visual system, it is natural to ask whether such "texture representations" spontaneously align with the textures' perceptual content, and in particular whether those CNNs that are regarded as better models for the visual system also possess more human-like texture representations. Here we quantify the perceptual content captured by feature correlations computed for a diverse pool of CNNs, and we compare it to the models' perceptual alignment with the mammalian visual system as measured by Brain-Score. Surprisingly, we find that there is no connection between conventional measures of CNN quality as a model of the visual system and its alignment with human texture perception. We conclude that texture perception involves mechanisms that are distinct from those that are commonly modeled using approaches based on CNNs trained on object recognition, possibly depending on the integration of contextual information.

Figures

Figures reproduced from arXiv: 2604.01341 by Alessio Ansuini, Eugenio Piasini, Fabio Anselmi, Ludovica de Paolis.

Figure 1
Figure 1. Figure 1: Representational Dissimilarity Matrices (5640x5640) obtained with Representa [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: MI values across CNN layers identified by index (1-5). Each line corresponds to [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Plots of the correlations with Pearson’s [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: It is not simply due for instance to all models having roughly the same texture [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 4
Figure 4. Figure 4: Textures generated with the Gatys algorithm applied to a subsample of CNNs [PITH_FULL_IMAGE:figures/full_fig_p015_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Lean theorems connected to this paper

Citations machine-checked in the Pith Canon. Every link opens the source theorem in the public Lean library.

What do these tags mean?
matches
The paper's claim is directly supported by a theorem in the formal canon.
supports
The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
extends
The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
uses
The paper appears to rely on the theorem as machinery.
contradicts
The paper's claim conflicts with a theorem or certificate in the canon.
unclear
Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Texture Representations in Deep Vision Models: Comparing CNNs, Vision Transformers, and Human Perception

    cs.CV 2026-07 conditional novelty 6.0

    ViT texture representations align with each other and with human psychophysics better than VGG-19 representations, suggesting architecture drives texture coding more than training objective.

  2. A Fourier perspective on the learning dynamics of neural networks: from sample complexities to mechanistic insights

    stat.ML 2026-05 conditional novelty 6.0

    Neural networks prioritize amplitude over phase in Fourier space during training on translation-invariant data; power-law spectra accelerate phase learning despite not aiding classification.

Reference graph

Works this paper leans on

69 extracted references · 69 canonical work pages · cited by 2 Pith papers · 3 internal anchors

  1. [1]

    Textons, the Elements of Texture Perception, and Their Interactions

    Bela Julesz. “Textons, the Elements of Texture Perception, and Their Interactions”. In:Nature290.5802(Mar.1981),pp.91–97.issn:1476-4687.doi:10.1038/290091a0. url:https://www.nature.com/articles/290091a0(visited on 12/10/2025). 16

  2. [2]

    Texture Synthesis Using Convolutional Neural Networks

    L. A. Gatys et al. “Texture Synthesis Using Convolutional Neural Networks”. In: Advances in Neural Information Processing Systems 28 (NIPS 2015)(2015)

  3. [3]

    Brain-Score: Which Artificial Neural Network for Object Recog- nition is most Brain-Like?

    M. Schrimpf et al. “Brain-Score: Which Artificial Neural Network for Object Recog- nition is most Brain-Like?” In:bioRxiv(2018)

  4. [5]

    Visual Pattern Discrimination

    B. Julesz. “Visual Pattern Discrimination”. In:IRE Transactions on Information Theory8.2 (Feb. 1962), pp. 84–92.issn: 2168-2712.doi:10 . 1109 / TIT . 1962 . 1057698

  5. [6]

    On Perceptual Analyzers Underlying Visual Texture Dis- crimination: Part I

    T. Caelli and B. Julesz. “On Perceptual Analyzers Underlying Visual Texture Dis- crimination: Part I”. In:Biological Cybernetics28.3 (Sept. 1, 1978), pp. 167–175. issn: 1432-0770.doi:10 . 1007 / BF00337138.url:https : / / doi . org / 10 . 1007 / BF00337138(visited on 03/01/2022)

  6. [7]

    Textons, the elements of texture perception, and their interactions

    B. Julesz. “Textons, the elements of texture perception, and their interactions”. In: Nature290 (1981), pp. 91–97

  7. [8]

    ExploringTextureEnsemblesbyEfficientMarkov Chain Monte Carlo-Toward a

    S.C.Zhu,X.W.Liu,andY.N.Wu.“ExploringTextureEnsemblesbyEfficientMarkov Chain Monte Carlo-Toward a "Trichromacy" Theory of Texture”. In:IEEE Transac- tions on Pattern Analysis and Machine Intelligence22.6 (June 2000), pp. 554–569. issn: 1939-3539.doi:10.1109/34.862195.url:https://ieeexplore.ieee.org/ abstract/document/862195(visited on 12/08/2025)

  8. [9]

    From BoW to CNN: Two Decades of Texture Representation for Texture Classification

    Li Liu et al. “From BoW to CNN: Two Decades of Texture Representation for Texture Classification”. In:International Journal of Computer Vision127.1 (Jan. 1, 2019), pp. 74–109.issn: 1573-1405.doi:10 . 1007 / s11263 - 018 - 1125 - z.url:https : //doi.org/10.1007/s11263-018-1125-z(visited on 12/08/2025)

  9. [10]

    A parametric texture model based on joint statistics of complex wavelet coefficients

    J. Portilla E. P. Simoncelli. “A parametric texture model based on joint statistics of complex wavelet coefficients”. In:International Journal of Computer Vision40 (2000), pp. 49–70. 17

  10. [11]

    Local Statistics in Natural Scenes Predict the Saliency of Synthetic Textures

    G. Tkačik et al. “Local Statistics in Natural Scenes Predict the Saliency of Synthetic Textures”. In:Proceedings of the National Academy of Sciences107.42 (Oct. 5, 2010), pp. 18149–18154.doi:10.1073/pnas.0914916107. PMID:20923876.url:http: //dx.doi.org/10.1073/pnas.0914916107

  11. [12]

    Local Image Statistics: Maximum-Entropy Constructions and Perceptual Salience

    Jonathan D. Victor and Mary M. Conte. “Local Image Statistics: Maximum-Entropy Constructions and Perceptual Salience”. In:JOSA A29.7 (July 1, 2012), pp. 1313– 1345.issn: 1520-8532.doi:10.1364/JOSAA.29.001313.url:https://opg.optica. org/josaa/abstract.cfm?uri=josaa-29-7-1313(visited on 02/20/2022)

  12. [13]

    Variance Predicts Salience in Central Sensory Pro- cessing

    Ann M Hermundstad et al. “Variance Predicts Salience in Central Sensory Pro- cessing”. In:eLife3 (Nov. 14, 2014), e03722.issn: 2050-084X.doi:10 . 7554 / eLife . 03722.url:https : / / elifesciences . org / articles / 03722(visited on 01/10/2022)

  13. [14]

    Efficient Coding of Natural Scene Statistics Predicts Dis- crimination Thresholds for Grayscale Textures

    Tiberiu Tesileanu et al. “Efficient Coding of Natural Scene Statistics Predicts Dis- crimination Thresholds for Grayscale Textures”. In:eLife9 (Aug. 3, 2020). Ed. by Stephanie Palmer and Timothy E Behrens, e54347.issn: 2050-084X.doi:10 . 7554/eLife.54347.url:https://doi.org/10.7554/eLife.54347(visited on 05/08/2023)

  14. [15]

    Rat Sensitivity to Multipoint Statistics Is Predicted by Efficient Coding of Natural Scenes

    Riccardo Caramellino et al. “Rat Sensitivity to Multipoint Statistics Is Predicted by Efficient Coding of Natural Scenes”. In:eLife10 (Dec. 7, 2021), e72081.issn: 2050- 084X.doi:10.7554/eLife.72081.url:https://elifesciences.org/articles/ 72081(visited on 01/08/2022)

  15. [16]

    Plata, Clary B

    Mirko Zanon et al.Predisposed and Learned Preferences for Multipoint Visual Statis- tics in Visually Naïve Newly Hatched Chicks. June 17, 2025.doi:10.1101/2025. 06.16.659971.url:https://www.biorxiv.org/content/10.1101/2025.06.16. 659971v1(visited on 12/10/2025). Pre-published

  16. [17]

    Image Style Transfer Using Convolutional Neural Networks

    Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge. “Image Style Transfer Using Convolutional Neural Networks”. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016, pp. 2414–2423.url:https : / / openaccess . thecvf . com / content _ cvpr _ 2016 / html / Gatys _ Image _ Style _ Transfer_CVPR_2016_paper.html(visited o...

  17. [18]

    Diversified Texture Synthesis With Feed-Forward Networks

    Yijun Li et al. “Diversified Texture Synthesis With Feed-Forward Networks”. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017, pp. 3920–3928.url:https://openaccess.thecvf.com/content_cvpr_2017/ html/Li_Diversified_Texture_Synthesis_CVPR_2017_paper.html(visited on 06/10/2024)

  18. [19]

    Improved Texture Net- works: Maximizing Quality and Diversity in Feed-forward Stylization and Texture Synthesis

    Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. “Improved Texture Net- works: Maximizing Quality and Diversity in Feed-forward Stylization and Texture Synthesis”. In:Computer Vision and Pattern Recognition(2017)

  19. [20]

    Perceptual losses for real-time style transfer and super-resolution

    Justin Johnson, Alexandre Alahi, and Li Fei-Fei. “Perceptual Losses for Real-Time Style Transfer and Super-Resolution”. In:Computer Vision – ECCV 2016. Ed. by Bastian Leibe et al. Cham: Springer International Publishing, 2016, pp. 694–711. isbn: 978-3-319-46475-6.doi:10.1007/978-3-319-46475-6_43

  20. [21]

    Texture Mixer: A Network for Controllable Synthesis and Interpo- lation of Texture

    Ning Yu et al. “Texture Mixer: A Network for Controllable Synthesis and Interpo- lation of Texture”. In: Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition. 2019, pp. 12164–12173.url:https://openaccess. thecvf . com / content _ CVPR _ 2019 / html / Yu _ Texture _ Mixer _ A _ Network _ for _ Controllable_Synthesis_and_Inter...

  21. [22]

    High-Resolution Multi-Scale Neural Texture Synthesis

    Xavier Snelgrove. “High-Resolution Multi-Scale Neural Texture Synthesis”. In:SA ’17: SIGGRAPH Asia 2017 Technical Briefs. 2017, pp. 1–4

  22. [23]

    Deep Correlations for Texture Synthesist

    Omry Sednik and Daniel Cohen-Or. “Deep Correlations for Texture Synthesist”. In: ACM Transactions on Graphics36 (2017), pp. 161–176

  23. [24]

    Incorporating long-range consistency in CNN-based texture generation

    G. Berger and R. Memisevic. “Incorporating Long-Range Consistency in CNN-based Texture Generation”. In:International Conference on Learning Representations. In- ternational Conference on Learning Representations. arXiv, 2017.doi:10.48550/ arXiv.1606.01286. arXiv:1606.01286 [cs].url:http://arxiv.org/abs/1606. 01286(visited on 12/08/2025)

  24. [25]

    A Sliced Wasserstein Loss for Neural Texture Synthesis

    Eric Heitz et al. “A Sliced Wasserstein Loss for Neural Texture Synthesis”. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 9407–9415

  25. [26]

    Texture Interpolation for Probing Visual Perception

    J. Vacher et al. “Texture Interpolation for Probing Visual Perception”. In:NeurIPS 2020(2020). 19

  26. [27]

    Performance-optimized hierarchical models predict neural re- sponses in higher visual cortex

    D. K. L Yamins et al. “Performance-optimized hierarchical models predict neural re- sponses in higher visual cortex”. In:Proceedings of the National Academy of Sciences 111.23 (2014), pp. 8619–8624

  27. [28]

    Deep Neural Networks Rival the Representation of Primate IT Cortical Neurons

    Charles F. Cadieu et al. “Deep Neural Networks Rival the Representation of Primate IT Cortical Neurons”. In:Proceedings of the National Academy of Sciences111.23 (2014), pp. 8519–8524

  28. [29]

    Deep Supervised, But Not Unsupervised, Models May Explain IT Cortical Representation

    Seyed-Mahdi Khaligh-Razavi and Nikolaus Kriegeskorte. “Deep Supervised, But Not Unsupervised, Models May Explain IT Cortical Representation”. In:PNAS111.23 (2014), pp. 8619–8624

  29. [30]

    Cadena, George H

    SantiagoA.Cadenaetal.“Deepconvolutionalmodelsimprovepredictionsofmacaque V1 responses to natural images”. In:PLOS Computational Biology15.4 (2019), e1006897.doi:10.1371/journal.pcbi.1006897

  30. [31]

    Neural Population Control via Deep Image Synthesis

    Pouya Bashivan, Kohitij Kar, and James J. DiCarlo. “Neural Population Control via Deep Image Synthesis”. In:Science364.6439 (May 3, 2019), eaav9436.issn: 0036- 8075, 1095-9203.doi:10.1126/science.aav9436.url:https://www.science. org/doi/10.1126/science.aav9436(visited on 01/17/2022)

  31. [32]

    Diverse Deep Neural Networks All Predict Human Inferior Temporal Cortex Well, After Training and Fitting

    Katherine R. Storrs et al. “Diverse Deep Neural Networks All Predict Human Inferior Temporal Cortex Well, After Training and Fitting”. In:Journal of Cognitive Neuro- science33.10 (Sept. 1, 2021), pp. 2044–2064.issn: 0898-929X.doi:10.1162/jocn_ a_01755.url:https://doi.org/10.1162/jocn_a_01755(visited on 12/08/2025)

  32. [33]

    Deep Neural Networks Reveal a Gradient in the Complexity of Neural Representations across the Ventral Stream

    Umut Güçlü and Marcel A. J. van Gerven. “Deep Neural Networks Reveal a Gradient in the Complexity of Neural Representations across the Ventral Stream”. In:Journal of Neuroscience35.27 (July 8, 2015), pp. 10005–10014.issn: 0270-6474, 1529-2401. doi:10.1523/JNEUROSCI.5023- 14.2015. PMID:26157000.url:https://www. jneurosci.org/content/35/27/10005(visited on ...

  33. [34]

    How well do deep neural networks trained on object recog- nition characterize the mouse visual system?

    Santiago A. Cadena et al. “How well do deep neural networks trained on object recog- nition characterize the mouse visual system?” In:Real Neurons & Hidden Units: Fu- ture directions at the intersection of neuroscience and AI, NeurIPS 2019 Workshop. 2019

  34. [35]

    Mouse Visual Cortex as a Limited Resource System That Self- Learns an Ecologically-General Representation

    Aran Nayebi et al. “Mouse Visual Cortex as a Limited Resource System That Self- Learns an Ecologically-General Representation”. In:PLOS Computational Biology 19.10 (Oct. 2, 2023), e1011506.issn: 1553-7358.doi:10 . 1371 / journal . pcbi . 20 1011506.url:https : / / journals . plos . org / ploscompbiol / article ? id = 10 . 1371/journal.pcbi.1011506(visited...

  35. [36]

    Prune and distill: similar reformatting of image information along rat visual cortex and deep neural networks

    Paolo Muratore et al. “Prune and distill: similar reformatting of image information along rat visual cortex and deep neural networks”. In:Advances in Neural Informa- tion Processing Systems. Vol. 35. 2022, pp. 30206–30218

  36. [37]

    Unraveling the complexity of rat object vision requires a full convolutional network and beyond

    Paolo Muratore, Alireza Alemi, and Davide Zoccolan. “Unraveling the complexity of rat object vision requires a full convolutional network and beyond”. In:Patterns6.2 (2025). Patterns (N Y), p. 101149

  37. [38]

    Model Metamers Reveal Divergent Invariances between Bio- logical and Artificial Neural Networks

    Jenelle Feather et al. “Model Metamers Reveal Divergent Invariances between Bio- logical and Artificial Neural Networks”. In:Nature Neuroscience26.11 (Nov. 2023), pp. 2017–2034.issn: 1546-1726.doi:10.1038/s41593-023-01442-0.url:https: //www.nature.com/articles/s41593-023-01442-0(visited on 05/19/2025)

  38. [39]

    Partial Success in Closing the Gap between Human and Ma- chine Vision

    Robert Geirhos et al. “Partial Success in Closing the Gap between Human and Ma- chine Vision”. In:Advances in Neural Information Processing Systems. Vol. 34. Cur- ran Associates, Inc., 2021, pp. 23885–23899.url:https://proceedings.neurips. cc/paper/2021/hash/c8877cff22082a16395a57e97232bb6f-Abstract.html(vis- ited on 12/08/2025)

  39. [40]

    2026 , month =

    Felix A. Wichmann and Robert Geirhos. “Are Deep Neural Networks Adequate Be- havioral Models of Human Visual Perception?” In:Annual Review of Vision Science 9 (Volume 9, 2023 Sept. 15, 2023), pp. 501–524.issn: 2374-4642, 2374-4650.doi: 10.1146/annurev- vision- 120522- 031739.url:https://www.annualreviews. org/content/journals/10.1146/annurev- vision- 1205...

  40. [41]

    Large-Scale, High-Resolution Comparison of the Core Vi- sual Object Recognition Behavior of Humans, Monkeys, and State-of-the-Art Deep Artificial Neural Networks

    Rishi Rajalingham et al. “Large-Scale, High-Resolution Comparison of the Core Vi- sual Object Recognition Behavior of Humans, Monkeys, and State-of-the-Art Deep Artificial Neural Networks”. In:Journal of Neuroscience38.33 (Aug. 15, 2018), pp. 7255–7269.issn: 0270-6474, 1529-2401.doi:10 . 1523 / JNEUROSCI . 0388 - 18

  41. [42]

    PMID:30006365.url:https://www.jneurosci.org/content/38/33/7255 (visited on 12/08/2025)

  42. [43]

    Incorporating Intrinsic Suppression in Deep Neural Networks Captures Dynamics of Adaptation in Neurophysiology and Percep- tion

    K. Vinken, X. Boix, and G. Kreiman. “Incorporating Intrinsic Suppression in Deep Neural Networks Captures Dynamics of Adaptation in Neurophysiology and Percep- tion”. In:Science Advances6.42 (Oct. 16, 2020), eabd4205.issn: 2375-2548.doi: 21 10 . 1126 / sciadv . abd4205.url:https : / / www . science . org / doi / 10 . 1126 / sciadv.abd4205(visited on 01/31/2022)

  43. [44]

    Texture Discrimination by Gabor Functions

    M. R. Turner. “Texture Discrimination by Gabor Functions”. In:Biological Cyber- netics55.2 (Nov. 1, 1986), pp. 71–82.issn: 1432-0770.doi:10.1007/BF00341922. url:https://doi.org/10.1007/BF00341922(visited on 12/08/2025)

  44. [45]

    Preattentive Texture Discrimination with Early Vision Mechanisms

    Jitendra Malik and Pietro Perona. “Preattentive Texture Discrimination with Early Vision Mechanisms”. In:JOSA A7.5 (May 1, 1990), pp. 923–932.issn: 1520-8532. doi:10.1364/JOSAA.7.000923.url:https://opg.optica.org/josaa/abstract. cfm?uri=josaa-7-5-923(visited on 12/08/2025)

  45. [46]

    Metamers of the Ventral Stream

    Jeremy Freeman and Eero P. Simoncelli. “Metamers of the Ventral Stream”. In: Nature Neuroscience14.9 (9 Sept. 2011), pp. 1195–1201.issn: 1546-1726.doi:10. 1038/nn.2889.url:https://www.nature.com/articles/nn.2889(visited on 02/16/2022)

  46. [47]

    A Functional and Perceptual Signature of the Second Visual Area in Primates

    Jeremy Freeman et al. “A Functional and Perceptual Signature of the Second Visual Area in Primates”. In:Nature Neuroscience16.7 (July 2013), pp. 974–981.issn: 1546- 1726.doi:10.1038/nn.3402.url:https://www.nature.com/articles/nn.3402 (visited on 12/08/2025)

  47. [48]

    Visual processing of informa- tive multipoint correlations arises primarily in V2

    Yaping Yu, Ann M. Schmid, and Jonathan D. Victor. “Visual processing of informa- tive multipoint correlations arises primarily in V2”. In:eLife4 (2015), e06604

  48. [49]

    Rodrigue, Kristen M

    Gouki Okazawa, Satohiro Tajima, and Hidehiko Komatsu. “Gradual Development of VisualTexture-SelectivePropertiesBetweenMacaqueAreasV2andV4”.In:Cerebral Cortex27.10 (Oct. 1, 2017), pp. 4867–4880.issn: 1047-3211.doi:10.1093/cercor/ bhw282.url:https://doi.org/10.1093/cercor/bhw282(visited on 12/08/2025)

  49. [50]

    Selectivity and tolerance for visual texture in macaque V2

    Corey M. Ziemba et al. “Selectivity and tolerance for visual texture in macaque V2”. In:Proceedings of the National Academy of Sciences113.22 (2016)

  50. [51]

    A Texture Statistics En- coding Model Reveals Hierarchical Feature Selectivity across Human Visual Cortex

    Margaret M. Henderson, Michael J. Tarr, and Leila Wehbe. “A Texture Statistics En- coding Model Reveals Hierarchical Feature Selectivity across Human Visual Cortex”. In:The Journal of Neuroscience43 (2023), pp. 4144–4161

  51. [52]

    Unsupervised learning of mid-level visual representations

    Giorgio Matteucci, Eugenio Piasini, and Davide Zoccolan. “Unsupervised learning of mid-level visual representations”. In:Current Opinion in Neurobiology84 (2024), p. 102834. 22

  52. [53]

    Neuronal and behavioral responses to naturalistic texture images in macaque monkeys

    Corey M. Ziemba et al. “Neuronal and behavioral responses to naturalistic texture images in macaque monkeys”. In:Journal of Neuroscience44.42 (2024)

  53. [54]

    Describing Textures in the Wild

    M. Cimpoi et al. “Describing Textures in the Wild”. In:CVPR 2014(2014)

  54. [55]

    Very Deep Convolutional Networks for Large Scale Image Recognition

    K. Simonyan A. Zisserman. “Very Deep Convolutional Networks for Large Scale Image Recognition”. In:ICLR 2015(2015)

  55. [56]

    Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton.The CIFAR-10 Dataset. 1994. url:https://www.cs.toronto.edu/~kriz/cifar.html

  56. [57]

    Densely Connected Convolutional Networks

    G. Huang et al. “Densely Connected Convolutional Networks”. In:Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2017, pp. 4700–4708

  57. [58]

    MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications

    A. G. Howard et al. “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications”. In:arXiv preprint arXiv:1704.04861(2017)

  58. [59]

    Rethinking the Inception Architecture for Computer Vision

    C. Szegedy et al. “Rethinking the Inception Architecture for Computer Vision”. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2016, pp. 2818–2826

  59. [60]

    Deep Residual Learning for Image Recognition

    K. He et al. “Deep Residual Learning for Image Recognition”. In:Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2016, pp. 770–778

  60. [61]

    The Texture Lexicon: Understanding the Categorization of Visual Texture Terms and Their Relationship to Texture Images

    N. Bhushan et al. “The Texture Lexicon: Understanding the Categorization of Visual Texture Terms and Their Relationship to Texture Images”. In:Cognitive Science, 21 (1997), pp. 219–246

  61. [62]

    Representational similarity analysis - connecting the branches of systems neuroscience

    Nikolaus Kriegeskorte, Marieke Mur, and Peter A. Bandettini. “Representational Similarity Analysis – Connecting the Branches of Systems Neuroscience”. In:Fron- tiers in Systems Neuroscience2 (2008), p. 4.doi:10.3389/neuro.06.004.2008. url:https://doi.org/10.3389/neuro.06.004.2008

  62. [63]

    A Tutorial for Information Theory in Neuroscience

    Nicholas M. Timme and Christopher Lapish. “A Tutorial for Information Theory in Neuroscience”. In:eNeuro5.3 (2018)

  63. [64]

    Entropy and Inference, Revis- ited

    Ilya Nemenman, Fariel Shafee, and William Bialek. “Entropy and Inference, Revis- ited”. In:Advances in Neural Information Processing Systems 14. 2002, pp. 471–478

  64. [65]

    Simone Marsili.ndd.https://github.com/simomarsili/ndd. 2021. 23

  65. [66]

    Supplementary materials: Representational Dissimilarity Matrix (RDM) plots for texture representations in CNNs

    Ludovica de Paolis. “Supplementary materials: Representational Dissimilarity Matrix (RDM) plots for texture representations in CNNs.” In: (Dec. 2025).doi:10.5281/ zenodo.17856135.url:https://doi.org/10.5281/zenodo.17856135

  66. [67]

    A vailable: https://www.physiology.org/doi/full/10.1152/jn

    Kalathupiriyan A. Zhivago and Sripati P. Arun. “Texture Discriminability in Monkey Inferotemporal Cortex Predicts Human Texture Perception”. In:Journal of Neuro- physiology112.11 (Dec. 2014), pp. 2745–2755.issn: 0022-3077.doi:10.1152/jn. 00532.2014.url:https://journals.physiology.org/doi/full/10.1152/jn. 00532.2014(visited on 12/08/2025)

  67. [68]

    Learning Transferable Visual Models From Natural Language Supervision

    Alec Radford et al. “Learning Transferable Visual Models From Natural Language Supervision”.In:Proceedings of the 38th International Conference on Machine Learn- ing. 2021. arXiv:2103.00020

  68. [69]

    Attention Is All You Need

    Ashish Vaswani et al. “Attention Is All You Need”. In:Advances in Neural Informa- tion Processing Systems. Vol. 30. 2017

  69. [70]

    U-Attention to Textures: Hierarchical Hourglass Vision Trans- former for Universal Texture Synthesis

    Shouchang Guo et al. “U-Attention to Textures: Hierarchical Hourglass Vision Trans- former for Universal Texture Synthesis”. In:Proceedings of the 19th ACM SIG- GRAPH European Conference on Visual Media Production. 2022, pp. 1–10. 24 9 Appendix 9.1 Loss normalization In this section we derive the weighting rule we used to combine the Gram losses from mult...