Pith. sign in

REVIEW 25 references

Neural encoding with affine feature response transforms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.03741 v1 pith:TSW3W4Z5 submitted 2025-01-07 q-bio.NC

classification q-bio.NC
keywords encodingresponseaffineafrtfeaturemodelsbrainneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current linearizing encoding models that predict neural responses to sensory input typically neglect neuroscience-inspired constraints that could enhance model efficiency and interpretability. To address this, we propose a new method called affine feature response transform (AFRT), which exploits the brain's retinotopic organization. Applying AFRT to encode multi-unit activity in areas V1, V4, and IT of the macaque brain, we demonstrate that AFRT reduces redundant computations and enhances the performance of current linearizing encoding models by segmenting each neuron's receptive field into an affine retinal transform, followed by a localized feature response. Remarkably, by factorizing receptive fields into a sequential affine component with three interpretable parameters (for shifting and scaling) and response components with a small number of feature weights per response, AFRT achieves encoding with orders of magnitude fewer parameters compared to unstructured models. We show that the retinal transform of each neuron's encoding agrees well with the brain's receptive field. Together, these findings suggest that this new subset within spatial transformer network can be instrumental in neural encoding models of naturalistic stimuli.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

25 extracted references · 20 canonical work pages

  1. [1]

    A primer on encoding models in sensory neuroscience

    Marcel AJ van Gerven. A primer on encoding models in sensory neuroscience. Journal of Mathematical Psychology, 76:172–183, 2017

  2. [2]

    Performance-optimized hierarchical models predict neural responses in higher visual cortex

    Daniel LK Yamins, Ha Hong, Charles F Cadieu, Ethan A Solomon, Darren Seibert, and James J DiCarlo. Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proceedings of the National Academy of Sciences, 111(23):8619–8624, 2014

  3. [3]

    Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream

    Umut Güçlü and Marcel AJ van Gerven. Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream. Journal of Neuroscience, 35(27):10005–10014, 2015

  4. [4]

    The feature-weighted receptive field: an interpretable encoding model for complex feature spaces

    Ghislain St-Yves and Thomas Naselaris. The feature-weighted receptive field: an interpretable encoding model for complex feature spaces. NeuroImage, 180:188–202, 2018

  5. [5]

    Characterizing the ventral visual stream with response-optimized neural encoding models

    Meenakshi Khosla, Keith Jamison, Amy Kuceyeski, and Mert Sabuncu. Characterizing the ventral visual stream with response-optimized neural encoding models. Advances in Neural Information Processing Systems, 35:9389–9402, 2022

  6. [6]

    A task-optimized neural network replicates human auditory behavior, predicts brain responses, and reveals a cortical processing hierarchy

    Alexander JE Kell, Daniel LK Yamins, Erica N Shook, Sam V Norman-Haignere, and Josh H McDermott. A task-optimized neural network replicates human auditory behavior, predicts brain responses, and reveals a cortical processing hierarchy. Neuron, 98(3):630–644, 2018

  7. [7]

    Identifying natural images from human brain activity

    Kendrick N Kay, Thomas Naselaris, Ryan J Prenger, and Jack L Gallant. Identifying natural images from human brain activity. Nature, 452(7185):352–355, 2008. 9

  8. [8]

    Recurrence is required to capture the representational dynamics of the human visual system

    Tim C Kietzmann, Courtney J Spoerer, Lynn KA Sörensen, Radoslaw M Cichy, Olaf Hauk, and Nikolaus Kriegeskorte. Recurrence is required to capture the representational dynamics of the human visual system. Proceedings of the National Academy of Sciences, 116(43):21854–21863, 2019

Show all 25 references
  1. [9]

    Vector-based navigation using grid-like representations in artificial agents

    Andrea Banino, Caswell Barry, Benigno Uria, Charles Blundell, Timothy Lillicrap, Piotr Mirowski, Alexander Pritzel, Martin J Chadwick, Thomas Degris, Joseph Modayil, et al. Vector-based navigation using grid-like representations in artificial agents. Nature, 557(7705):429–433, 2018

  2. [10]

    Neural encoding with visual attention

    Meenakshi Khosla, Gia Ngo, Keith Jamison, Amy Kuceyeski, and Mert Sabuncu. Neural encoding with visual attention. Advances in Neural Information Processing Systems, 33:15942–15953, 2020

  3. [11]

    what” and “where

    Haibao Wang, Lijie Huang, Changde Du, Dan Li, Bo Wang, and Huiguang He. Neural encoding for human visual cortex with deep neural networks learning “what” and “where”. IEEE Transactions on Cognitive and Developmental Systems, 13(4):827–840, 2020

  4. [12]

    Deep convolutional models improve predictions of macaque V1 responses to natural images

    Santiago A Cadena, George H Denfield, Edgar Y Walker, Leon A Gatys, Andreas S Tolias, Matthias Bethge, and Alexander S Ecker. Deep convolutional models improve predictions of macaque V1 responses to natural images. PLoS Computational Biology, 15(4):e1006897, 2019

  5. [13]

    Receptive fields, binocular interaction and functional architecture in the cat’s visual cortex

    David H Hubel and Torsten N Wiesel. Receptive fields, binocular interaction and functional architecture in the cat’s visual cortex. The Journal of physiology, 160(1):106, 1962

  6. [14]

    Spatial transformer networks

    Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al. Spatial transformer networks. Advances in Neural Information Processing Systems, 28, 2015

  7. [15]

    A method for stochastic optimization

    D Kinga, Jimmy Ba Adam, et al. A method for stochastic optimization. In International Conference on Learning Representations (ICLR), volume 5, page 6. San Diego, California;, 2015

  8. [16]

    Very deep convolutional networks for large-scale image recogni- tion

    Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recogni- tion. arXiv preprint arXiv:1409.1556, 2014

  9. [17]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016

  10. [18]

    Top-down influences on visual processing

    Charles D Gilbert and Wu Li. Top-down influences on visual processing. Nature Reviews Neuroscience, 14(5):350–363, 2013

  11. [19]

    Things: A database of 1,854 object concepts and more than 26,000 naturalistic object images

    Martin N Hebart, Adam H Dickter, Alexis Kidder, Wan Y Kwok, Anna Corriveau, Caitlin Van Wicklin, and Chris I Baker. Things: A database of 1,854 object concepts and more than 26,000 naturalistic object images. PloS ONE, 14(10):e0223792, 2019

  12. [20]

    Shape perception via a high-channel- count neuroprosthesis in monkey visual cortex

    Xing Chen, Feng Wang, Eduardo Fernandez, and Pieter R Roelfsema. Shape perception via a high-channel- count neuroprosthesis in monkey visual cortex. Science, 370(6521):1191–1196, 2020

  13. [21]

    1024-channel electrophysiological recordings in macaque V1 and V4 during resting state

    Xing Chen, Aitor Morales-Gregorio, Julia Sprenger, Alexander Kleinjohann, Shashwat Sridhar, Sacha J Van Albada, Sonja Grün, and Pieter R Roelfsema. 1024-channel electrophysiological recordings in macaque V1 and V4 during resting state. Scientific Data, 9(1):77, 2022

  14. [22]

    Comparisons of the dynamics of local field potential and multiunit activity signals in macaque visual cortex

    Samuel P Burns, Dajun Xing, and Robert M Shapley. Comparisons of the dynamics of local field potential and multiunit activity signals in macaque visual cortex. Journal of Neuroscience, 30(41):13739–13749, 2010

  15. [23]

    Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems

    Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems. arXiv preprint arXiv:1512.01274, 2015

  16. [24]

    A voxel- wise encoding model for early visual areas decodes mental images of remembered scenes

    Thomas Naselaris, Cheryl A Olman, Dustin E Stansbury, Kamil Ugurbil, and Jack L Gallant. A voxel- wise encoding model for early visual areas decodes mental images of remembered scenes. Neuroimage, 105:215–228, 2015

  17. [25]

    End-to-end optimization of prosthetic vision

    Jaap de Ruyter van Steveninck, Umut Güçlü, Richard van Wezel, and Marcel van Gerven. End-to-end optimization of prosthetic vision. Journal of Vision, 22(2):20–20, 2022. A Feature shapes example In Fig. 5 we show the dimensional structure of feature spaces that are then transfo...

Pith tools