Pith. sign in

REVIEW 2 major objections 4 minor 60 references

Where visual information lands on the retina shapes what cortex learns: fovea-only and periphery-only training on natural egocentric video produces models whose task performance and fMRI predictivity mirror the cortical eccentricity bias, w

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 12:45 UTC pith:ZQKJJVGA

load-bearing objection Solid, honest empirical study of eccentricity-constrained pretraining on egocentric video; the central fMRI claim hinges on a single training seed per condition. the 2 major comments →

arxiv 2607.19316 v1 pith:ZQKJJVGA submitted 2026-07-21 q-bio.NC

Eccentricity-Constrained CNN Training Reveals Adaptive Information Coding Around the Visual Field

classification q-bio.NC
keywords eccentricity biasegocentric visiongaze-contingent processingcontrastive learningfMRI encoding modelsscene-selective cortexvisual cortex organizationfoveal vs peripheral vision
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks whether the center-periphery organization of visual cortex can arise from natural experience rather than being hard-wired. It trains self-supervised vision models on first-person video frames that are either gaze-centered fovea-only crops, periphery-only crops, or periphery-only crops with a texture-like NeuroFovea transform, then tests these models on classification tasks and on predicting human fMRI responses. The central finding is that task performance and neural predictivity systematically split along the fovea/periphery axis: fovea-trained models are better at classifying fine-grained and social content, while periphery-trained models explain more unique variance in scene-selective cortex (PPA, RSC) and less in primary visual cortex. The authors interpret this as evidence that naturalistic egocentric experience provides an organizing constraint, adaptively aligning cortical information processing with the statistics of where information lands on the retina.

Core claim

The paper claims that eccentricity-constrained self-supervised training on egocentric video produces representations whose differential performance and fMRI predictivity match the cortical eccentricity bias. Specifically, periphery-trained models were found to have higher predictive accuracy in scene-selective cortical regions (PPA, RSC) compared to fovea-trained models, while fovea-trained models were favored in V1, and these differences were small but consistent across all eight participants. The authors also show that models pretrained on egocentric frames achieve neural predictivity comparable to models trained on mid-sized curated datasets (e.g., ImageNet-100), supporting the idea that

What carries the argument

The central machinery is the set of gaze-contingent input transformations applied to egocentric video frames before contrastive learning: Fovea-Gaze (a fixed 112×112 crop centered on the per-frame gaze position, upsampled and masked), Periph (the complementary peripheral region), and Periph-NF (peripheral region with NeuroFovea metamerization approximating peripheral texture pooling). These conditions are compared through SimCLR pretraining of a ResNet-18 backbone, followed by linear-probe classification and voxelwise fMRI encoding models with variance partitioning to separate unique contributions of each model in each cortical ROI.

Load-bearing premise

The load-bearing premise is that the Fovea-Gaze, Periph, and Periph-NF transforms actually isolate central vs. peripheral visual information in a way that matches cortical eccentricity maps; the paper itself notes that Fovea-Gaze is a gaze-centered central-only input, not a biological fovea, and there is no center-crop control to separate gaze alignment from central-image advantage.

What would settle it

Run the same training and fMRI encoding pipeline with a non-gaze center-crop control (a fixed 112×112 crop centered on the image rather than on the gaze point). If that model produces the same pattern — periphery-trained better in PPA/RSC and fovea-trained better in V1 — then the effect is driven by central vs. peripheral image statistics, not by gaze-contingent experience, undercutting the paper's egocentric-experience interpretation.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • If the central claim is correct, scene-selective cortical regions (PPA, RSC) may be adapted to the statistics of peripheral vision, potentially supporting navigation and layout processing.
  • Fovea-only training captures the most diagnostic features for fine-grained and social tasks, suggesting that central-vision statistics are especially informative for object- and face-related behavior — though the paper found no advantage in face-selective cortex.
  • Natural egocentric experience, despite being temporally correlated and less diverse than curated image datasets, is sufficient to produce representations that align with human visual cortex at a level comparable to mid-sized curated datasets.
  • The NeuroFovea transform, which simulates peripheral texture pooling, provided a small robustness benefit for out-of-domain transfer, suggesting that peripheral metamerization can act as a useful augmentation.
  • The differences between fovea- and periphery-trained models are statistically reliable in PPA, RSC, and V1 but modest in magnitude, indicating that further modeling work is needed to capture the full structure of eccentricity-dependent cortical organization.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural extension is to test whether the foveal advantage in face- and word-selective cortex would appear with a smaller foveal crop (e.g., 10×10 pixels) rather than the 112×112 crop used here; the paper's own Discussion suggests the large crop may explain the null result.
  • The absence of a non-gaze center-crop control leaves open whether the observed differences are driven by central vs. peripheral image statistics or by gaze alignment per se; adding a fixed-center crop condition would disentangle these.
  • The small PPA/RSC advantage for periphery-trained models may reflect sensitivity to layout and spatial-navigational features rather than semantic scene content; testing on tasks that isolate geometry (e.g., depth estimation or spatial layout classification) would sharpen the interpretation.
  • Because the paper used static frames, training on video with temporal structure could produce stronger eccentricity-dependent effects, since real visual experience is inherently dynamic and gaze-contingent over time.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper trains ResNet-18 SimCLR models on egocentric video frames from the Visual Experience Dataset under four input conditions: Baseline (full frame), Fovea-Gaze (a 112x112 gaze-centered crop upsampled to 224x224), Periph (peripheral-only with a central scotoma), and Periph-NF (peripheral-only with a NeuroFovea texture-pooling transform). The models are evaluated with frozen-backbone linear probes on in-domain VEDB category classification, VGGFace2 face recognition, and Places365 scene recognition, and with voxelwise encoding models on the Natural Scenes Dataset fMRI data. The paper reports that VEDB-pretrained models achieve neural predictivity comparable to models trained on mid-sized non-egocentric datasets, that Fovea-Gaze models outperform peripheral models on both downstream face and scene classification, and that periphery-trained models explain more variance in scene-selective PPA/RSC while fovea-trained models explain more variance in V1. The authors interpret these results as evidence that naturalistic egocentric experience provides an organizing constraint that adaptively shapes cortical information coding across visual field eccentricities.

Significance. If the central fMRI result is robust, the paper makes a valuable contribution by connecting natural egocentric image statistics to the known center-periphery organization of high-level visual cortex. The experimental design is non-circular: models are trained self-supervised on VEDB without using brain data, and neural predictivity is evaluated on held-out NSD test images. The paper also has several methodological strengths, including session-level splits to prevent temporal leakage, permutation-based repeated-measures statistics, variance partitioning, in-domain bootstrap confidence intervals, and comparison against non-egocentric pretraining baselines. The main caveat is that each condition was trained once, so the key PPA/RSC and V1 differences rest on a single encoder per condition; participant-level permutation tests do not protect against training stochasticity. This is the load-bearing weakness that drives my recommendation.

major comments (2)
  1. [Methods — SimCLR Pre-Training; Results — fMRI encoding model performance; Supp. Table 3] The central neuroimaging claim rests on a single trained model per condition. The reported PPA/RSC advantage of Periph over Fovea-Gaze and the V1 advantage of Fovea-Gaze over Periph are computed from one randomly initialized encoder per condition. The permutation tests in Supp. Methods shuffle condition labels across the 8 NSD participants, but all participants share the same encoder, so participant-level consistency does not control for seed-induced variability in representation learning. Because no code or checkpoints are released, this dependence cannot be checked externally. The Limitations section acknowledges that some downstream performance differences are modest and should be interpreted as suggestive, but the PPA/RSC result is described as 'small but consistent' and is the main fMRI support for the adaptive-coding claim. I request either (a) retraining each condition with multip
  2. [Results — fMRI encoding model performance; Supp. Table 3] The statistical evidence for condition differences is based on many paired tests across 11 ROIs and several condition pairs, with a p<0.01 threshold and no multiple-comparison correction. With this number of tests, chance is expected to produce some 'significant' differences. The PPA/RSC comparisons are directional and a priori, which mitigates but does not eliminate the concern, and the consistency across all 8 participants is reassuring. Nevertheless, the paper should report FDR-adjusted p-values or explicitly justify the uncorrected threshold, and it should include effect sizes or confidence intervals for the key R2 differences (e.g., Periph vs. Fovea-Gaze in PPA, RSC, and V1).
minor comments (4)
  1. [Methods — Fovea-Gaze condition] The 112x112 crop is much larger than a biological fovea, and the authors correctly state that the condition 'should be interpreted as gaze-centered central-only input, not a simulation of biological fovea.' Adding a matched center-crop control without gaze would help disentangle gaze alignment from a general central-image advantage; this is not blocking for the broad central-vs-peripheral comparison but would strengthen the interpretation.
  2. [Supp. Methods — Statistical testing] The description of the ANOVA p-value calculation is confusing: 'proportion of iterations where the true F-statistic was less than or equal to the shuffled F-statistic' should be reworded as the proportion of permutations in which the shuffled statistic equals or exceeds the observed statistic.
  3. [Discussion] Minor typo: 'face recogniton' should be 'face recognition'.
  4. [Discussion — Limitations] The Limitations section explicitly says that a full subsampling-based robustness analysis would better assess stability and that some results should be interpreted as suggestive. This is appropriately candid, but it should be reconciled with the strength of the claims made for the fMRI PPA/RSC effect in the Results and Abstract.

Circularity Check

0 steps flagged

No significant circularity; the central fMRI and transfer predictions are tested against independent external data.

full rationale

The paper's derivation chain is not circular. Models are pretrained self-supervised on VEDB frames modified to isolate foveal or peripheral input, then evaluated on held-out tasks and on NSD fMRI data. Crucially, the models are not fit to brain data: features are extracted from frozen pretrained backbones, and encoding models are fit separately to held-out voxel responses. The paper explicitly states: 'condition-specific transforms were applied only during SimCLR pretraining and were not reapplied at the encoding stage.' Thus the reported PPA/RSC advantage for Periph over Fovea-Gaze, and the V1 advantage for Fovea-Gaze, are empirical outcomes rather than built-in consequences of the training scheme. The transfer results (VGGFace2, Places365) are likewise external benchmarks. The paper's caveat that Fovea-Gaze 'should be interpreted as gaze-centered central-only input, not a simulation of biological fovea' further prevents a definitional collapse between the input transform and the cortical eccentricity claim. Self-citations (Henderson et al., 2022, 2023, 2025; Jinsi et al., 2023) are used to motivate hypotheses or justify permutation-testing conventions, but none is load-bearing: the central claims do not reduce to those citations. The single-seed training concern raised by the skeptic is a statistical robustness limitation, not circularity, and the paper itself labels several downstream differences as 'suggestive rather than definitive.' Overall, the core comparisons are independent of the conclusions, so the circularity score is low.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 0 invented entities

No new physical or theoretical entities are introduced. The paper's central comparisons rest on hand-chosen preprocessing transforms and standard domain assumptions about what self-supervised features and fMRI encoding models measure.

free parameters (5)
  • Fovea crop size = 112×112 px
    Defines the Fovea-Gaze condition; hand-chosen. The claimed fovea/periphery differences depend on this size.
  • Gaze confidence threshold = 0.30
    Filters raw gaze samples used to compute per-frame fixation centroids; hand-chosen.
  • NeuroFovea scale = 0.4
    Controls peripheral texture pooling strength in the Periph-NF condition; hand-chosen within the model's standard regime.
  • ROI inclusion thresholds = localizer t>2; noise ceiling 0.10
    Voxel selection thresholds for category-selective ROIs and noise-ceiling filtering; influence the ROI-level R2 comparisons.
  • PCA components retained = 200
    Dimensionality reduction for encoding model features; standard but hand-chosen.
axioms (6)
  • domain assumption VEDB egocentric video with synchronized gaze is a representative sample of natural visual experience.
    Used as the training data for all central claims; if unrepresentative, the connection to natural experience fails.
  • domain assumption A fixed 112×112 gaze-centered crop and a gray scotoma mask isolate central vs peripheral visual information.
    This is the key operationalization of eccentricity; the paper acknowledges it is not a biological fovea simulation.
  • domain assumption SimCLR contrastive learning on static frames learns visual representations relevant to biological cortex.
    The paper relies on self-supervised features transferring to fMRI encoding models.
  • domain assumption Voxelwise linear encoding models with PCA features from a ResNet-18 measure cortical alignment.
    The fMRI conclusions are based on R2 from linear regressions on model features.
  • domain assumption NSD ROI definitions from localizers are valid.
    ROI-level comparisons assume the localizer-based functional regions accurately reflect scene- and face-selectivity.
  • domain assumption Non-VEDB baselines trained with slightly different pipelines are comparable to VEDB models.
    Comparisons to ImageNet-100, STL-10, and ImageNet-1K assume differences in training pipelines do not dominate the comparisons.

pith-pipeline@v1.3.0-alltime-deepseek · 20242 in / 12632 out tokens · 117496 ms · 2026-08-01T12:45:40.960062+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Eccentricity-Constrained CNN Training Reveals Adaptive Information Coding Around the Visual Field." pith.science (2026). https://pith.science/paper/ZQKJJVGA

@misc{pith2026260719316,
  author       = {Pith},
  title        = {Pith review of: Eccentricity-Constrained CNN Training Reveals Adaptive Information Coding Around the Visual Field},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZQKJJVGA}},
  note         = {Machine review of arXiv:2607.19316}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In the primate visual system, center-preferring cortical populations have higher spatial resolution and overlap face- and word-selective regions while periphery-preferring populations have lower spatial resolution and overlap scene-selective regions. This "eccentricity bias" may reflect differential task-relevance: central vision may better support fine-grained tasks like face recognition and reading, while peripheral vision may better support scene understanding. To test whether eccentricity-dependent coding can emerge from natural experience, we used egocentric video and eye-tracking data from the Visual Experience Dataset (VEDB). We trained ResNet-18 models using contrastive learning (SimCLR) on frames modified to isolate different eccentricities (gaze-contingent fovea-only crops, periphery-only crops, and periphery-only crops with a NeuroFovea transform applied). We evaluated downstream task performance and model alignment with human fMRI data (Natural Scenes Dataset; encoding models). In-domain VEDB frame classification showed systematic differences between fovea- and periphery-only models across categories, indicating differential informativeness across tasks. On downstream classification, VEDB-pretrained models generalized better to scene categorization (Places365) than face recognition (VGGFace2), with fovea-only models stronger on both. Across visual cortex, VEDB-pretrained models matched neural predictivity of models trained on mid-sized non-egocentric datasets (ImageNet-100), suggesting egocentric data supports emergence of cortically-aligned representations. In scene-selective cortex (PPA, RSC), periphery-only models held a small but consistent advantage in explained variance over fovea-only models, suggesting these regions are aligned with peripheral statistics. Together, these results suggest egocentric experience may adaptively constrain cortical information processing.

Figures

Figures reproduced from arXiv: 2607.19316 by Dylan M. Diaz, Margaret M. Henderson.

Figure 1
Figure 1. Figure 1: Method overview. VEDB frames and gaze metadata are processed to create eccentricity-constrained inputs [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Performance of VEDB models during pretraining and linear probing. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Multidimensional scaling (MDS) was performed [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: fMRI encoding model performance. (A) Voxel￾wise prediction accuracy (cross-validated R 2 ) for Base￾line VEDB and non-VEDB comparison models, averaged across voxels in functional ROIs for each participant. (B) As in (A), for the four VEDB conditions. See Supp. Fig￾ure 4 for comparison of R 2 vs. voxelwise noise ceiling, and see Supp [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

60 extracted references · 6 canonical work pages

  1. [1]

    Chalnick and D

    A. Chalnick and D. Billman , title =. Proceedings of the Tenth Annual Conference of the Cognitive Science Society , pages =

  2. [2]

    E. A. Feigenbaum , title =. Computers and thought , publisher =

  3. [3]

    J. A. C. Hill , title =. Cognition and Brain Theory , year = 1983, volume = 6, pages =

  4. [4]

    Ohlsson and P

    S. Ohlsson and P. Langley , title =

  5. [5]

    Teenie Matlock , title =

  6. [6]

    Newell and H

    A. Newell and H. A. Simon , title =

  7. [7]

    Computational models of scientific discovery and theory formation , publisher =

  8. [8]

    2022 , eprint=

    Ego4D: Around the World in 3,000 Hours of Egocentric Video , author=. 2022 , eprint=

  9. [9]

    Eccentricity bias as an organizing principle for human high-order object areas , volume =

    Uri Hasson and Ifat Levy and Marlene Behrmann and Talma Hendler and Rafael Malach , doi =. Eccentricity bias as an organizing principle for human high-order object areas , volume =. Neuron , month =

  10. [10]

    Center-periphery organization of human object areas , volume =

    Ifat Levy and Uri Hasson and Galia Avidan and Talma Hendler and Rafael Malach , doi =. Center-periphery organization of human object areas , volume =. Nature Neuroscience , keywords =

  11. [11]

    Pourlxadian and Roger B.H

    Xiaomin Yue and Irene S. Pourlxadian and Roger B.H. Tootell and Leslie G. Ungerleider , doi =. Curvature-processing network in macaque visual cortex , volume =. Proceedings of the National Academy of Sciences of the United States of America , keywords =

  12. [12]

    Ungerleider , doi =

    Xiaomin Yue and Sophia Robert and Leslie G. Ungerleider , doi =. Curvature processing in human visual cortical areas , volume =. NeuroImage , keywords =

  13. [13]

    Ponce and Till S

    Carlos R. Ponce and Till S. Hartmann and Margaret S. Livingstone , doi =. End-stopping predicts curvature tuning along the ventral stream , volume =. Journal of Neuroscience , keywords =

  14. [14]

    Vincent and Margaret S

    Krishna Srihasam and Justin L. Vincent and Margaret S. Livingstone , doi =. Novel domain formation reveals proto-architecture in inferotemporal cortex , volume =. Nature Neuroscience , keywords =

  15. [15]

    eLife , year =

    Arcaro, Michael J and Livingstone, Margaret S , title =. eLife , year =. doi:10.7554/eLife.26196 , url =

  16. [16]

    Journal of Vision , volume=

    The visual experience dataset: Over 200 recorded hours of integrated eye movement, odometry, and egocentric video , author=. Journal of Vision , volume=. 2024 , publisher=

  17. [17]

    , title =

    Deza, Arturo and Jonnalagadda, Aditya and Eckstein, Miguel P. , title =. International Conference on Learning Representations (ICLR) , year =

  18. [18]

    Henderson and Michael J

    Margaret M. Henderson and Michael J. Tarr and Leila Wehbe , doi =. Low-level tuning biases in higher visual cortex reflect the semantic informativeness of visual features , volume =. Journal of Vision , keywords =

  19. [19]

    A Simple Framework for Contrastive Learning of Visual Representations , volume =

    Ting Chen and Simon Kornblith and Mohammad Norouzi and Geoffrey Hinton , isbn =. A Simple Framework for Contrastive Learning of Visual Representations , volume =. 37th International Conference on Machine Learning, ICML 2020 , month =

  20. [20]

    Allen and Ghislain St-Yves and Yihan Wu and Jesse L

    Emily J. Allen and Ghislain St-Yves and Yihan Wu and Jesse L. Breedlove and Jacob S. Prince and Logan T. Dowdle and Matthias Nau and Brad Caron and Franco Pestilli and Ian Charest and J. Benjamin Hutchinson and Thomas Naselaris and Kendrick Kay , doi =. A massive 7T fMRI dataset to bridge cognitive neuroscience and artificial intelligence , volume =. Natu...

  21. [21]

    Machine Learning and Knowledge Discovery in Databases (ECML PKDD 2011) , pages=

    On the Stratification of Multi-label Data , author=. Machine Learning and Knowledge Discovery in Databases (ECML PKDD 2011) , pages=. 2011 , series=

  22. [22]

    Simoncelli , doi =

    Jeremy Freeman and Eero P. Simoncelli , doi =. Metamers of the ventral stream , volume =. Nature Neuroscience , month =

  23. [23]

    Kay , doi =

    Thomas Naselaris and Kendrick N. Kay , doi =. Resolving Ambiguities of MVPA Using Explicit Models of Representation , volume =. Trends in Cognitive Sciences , keywords =

  24. [24]

    Prince and Ian Charest and Jan W

    Jacob S. Prince and Ian Charest and Jan W. Kurzawski and John A. Pyles and Michael J. Tarr and Kendrick N. Kay , doi =. Improving the accuracy of single-trial fMRI response estimates using GLMsingle , volume =. eLife , pmid =

  25. [25]

    Journal of Vision , volume=

    Central and peripheral vision for scene recognition: A neurocomputational modeling exploration , author=. Journal of Vision , volume=. 2017 , publisher=. doi:10.1167/17.4.9 , url=

  26. [26]

    IEEE Transactions on Multimedia , volume=

    A Gated Peripheral-Foveal Convolutional Neural Network for Unified Image Aesthetic Prediction , author=. IEEE Transactions on Multimedia , volume=. 2019 , publisher=

  27. [27]

    3rd Workshop on Shared Visual Representations in Human and Machine Intelligence (SVRHM) at NeurIPS , year=

    On the use of Cortical Magnification and Saccades as Biological Proxies for Data Augmentation , author=. 3rd Workshop on Shared Visual Representations in Human and Machine Intelligence (SVRHM) at NeurIPS , year=

  28. [28]

    Psychological Science , volume=

    The Role of Fixation Position in Detecting Scene Changes Across Saccades , author=. Psychological Science , volume=. 1999 , publisher=

  29. [29]

    Journal of Experimental Psychology: Human Learning and Memory , volume=

    The functional visual field during picture viewing , author=. Journal of Experimental Psychology: Human Learning and Memory , volume=. 1980 , publisher=

  30. [30]

    Nature Neuroscience , volume=

    The uncrowded window of object recognition , author=. Nature Neuroscience , volume=. 2008 , publisher=

  31. [31]

    Journal of Vision , year =

    Rosenholtz, Ruth , title =. Journal of Vision , year =

  32. [32]

    Visual Cognition , volume=

    The use of visual information in natural scenes , author=. Visual Cognition , volume=. 2005 , publisher=

  33. [33]

    Annual Review of Vision Science , volume=

    Capabilities and Limitations of Peripheral Vision , author=. Annual Review of Vision Science , volume=. 2016 , publisher=

  34. [34]

    Journal of Vision , volume=

    The contributions of central versus peripheral vision to scene gist recognition , author=. Journal of Vision , volume=. 2009 , publisher=

  35. [35]

    2021 , month = feb, publisher =

    Silva, Thalles and Marcolini, Alessia and Yuhao, Dan , title =. 2021 , month = feb, publisher =. doi:10.5281/zenodo.4486327 , url =

  36. [36]

    Journal of Vision , volume=

    A summary statistic representation in peripheral vision explains visual search , author=. Journal of Vision , volume=. 2012 , publisher=

  37. [37]

    Progress in Brain Research , volume=

    Building the gist of a scene: the role of global image features in recognition , author=. Progress in Brain Research , volume=. 2006 , publisher=. doi:10.1016/S0079-6123(06)55002-2 , pmid=

  38. [38]

    Nature Reviews Neuroscience , volume=

    The functional architecture of the ventral temporal cortex and its role in categorization , author=. Nature Reviews Neuroscience , volume=. 2014 , publisher=

  39. [39]

    arXiv preprint arXiv:2006.07991 , year=

    Emergent properties of foveated perceptual systems , author=. arXiv preprint arXiv:2006.07991 , year=

  40. [40]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Peripheral Vision Transformer , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  41. [41]

    Annual Review of Vision Science , volume=

    Scene perception in the human brain , author=. Annual Review of Vision Science , volume=. 2019 , publisher=. doi:10.1146/annurev-vision-091718-014809 , pmcid=

  42. [42]

    Trends in Cognitive Sciences , volume=

    Parahippocampal and retrosplenial contributions to human spatial navigation , author=. Trends in Cognitive Sciences , volume=. 2008 , publisher=. doi:10.1016/j.tics.2008.07.004 , pmid=

  43. [43]

    Nature Reviews Neuroscience , volume=

    What does the retrosplenial cortex do? , author=. Nature Reviews Neuroscience , volume=. 2009 , publisher=

  44. [44]

    Nature Communications , volume=

    Improved modeling of human vision by incorporating robustness to blur in convolutional neural networks , author=. Nature Communications , volume=. 2024 , publisher=

  45. [45]

    PLOS ONE , volume=

    Early experience with low-pass filtered images facilitates visual category learning in a neural network model , author=. PLOS ONE , volume=. 2023 , publisher=

  46. [46]

    Advances in Neural Information Processing Systems (NeurIPS) , volume=

    Learning From Brains How to Regularize Machines , author=. Advances in Neural Information Processing Systems (NeurIPS) , volume=

  47. [47]

    Advances in Neural Information Processing Systems (NeurIPS) , volume=

    Simulating a Primary Visual Cortex at the Front of CNNs Improves Robustness to Image Perturbations , author=. Advances in Neural Information Processing Systems (NeurIPS) , volume=

  48. [48]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Explicitly Modeling Subcortical Vision with a Neuro-Inspired Front-End Improves CNN Robustness , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  49. [49]

    Sprague, T. C. and Serences, J. T. , title =. Nature Neuroscience , volume =

  50. [50]

    Nature Communications , volume=

    Dynamic categorization rules alter representations in human visual cortex , author=. Nature Communications , volume=. 2025 , publisher=

  51. [51]

    Henderson and Rosanne L

    Margaret M. Henderson and Rosanne L. Rademaker and John T. Serences , doi =. Flexible utilization of spatial-and motor-based codes for the storage of visuo-spatial information , volume =. eLife , pmid =

  52. [52]

    Rademaker, R. L. and Chunharas, C. and Serences, J. T. , title =. Nature Neuroscience , volume =

  53. [53]

    Parkhi and Andrew Zisserman , title =

    Qiong Cao and Li Shen and Weidi Xie and Omkar M. Parkhi and Andrew Zisserman , title =. CoRR , volume =. 2017 , url =. 1710.08092 , timestamp =

  54. [54]

    IEEE Transactions on Pattern Analysis and Machine Intelligence , year=

    Places: A 10 million Image Database for Scene Recognition , author=. IEEE Transactions on Pattern Analysis and Machine Intelligence , year=

  55. [55]

    , number =

    Freeman, Jeremy and Simoncelli, Eero P. , number =. 2011 , journal =. doi:10.1038/nn.2889 , issn =

  56. [56]

    2024 , journal =

    Spatial sampling of deep neural network features improves encoding models of foveal and peripheral visual processing in humans , author =. 2024 , journal =

  57. [57]

    2025 , journal =

    Retinotopic scaffolding of high-level vision , author =. 2025 , journal =

  58. [58]

    Berg and Li Fei-Fei , Title =

    Olga Russakovsky and Jia Deng and Hao Su and Jonathan Krause and Sanjeev Satheesh and Sean Ma and Zhiheng Huang and Andrej Karpathy and Aditya Khosla and Michael Bernstein and Alexander C. Berg and Li Fei-Fei , Title =. 2015 , journal =. doi:10.1007/s11263-015-0816-y , volume=

  59. [59]

    Ng , title =

    Adam Coates and Honglak Lee and Andrew Y. Ng , title =. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (AISTATS) , year =

  60. [60]

    CoRR , volume =

    Yonglong Tian and Dilip Krishnan and Phillip Isola , title =. CoRR , volume =. 2019 , url =. 1906.05849 , timestamp =