REVIEW 3 major objections 5 minor 91 references
Machine Learning Modeling for Multi-order Human Visual Motion Processing
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A dual-pathway neural network trained to estimate the motion of glossy, transparent objects acquires human-like second-order motion perception without ever seeing explicit second-order training stimuli.
desk verdict A genuinely new ecological hypothesis for second-order motion perception, tested with the right controlled experiment, but the current evidence is single-seed, partly in-distribution, and built on a second-order benchmark the authors themselves admit is not purely second-order. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a dual-channel V1-MT-style network. Stage I's first-order channel holds 256 trainable quadrature Gabor filters (spatiotemporal motion-energy sensors) arranged in a multiscale pyramid, with preferred speeds and directions learned during training. A second channel stacks five 3D convolutional layers with residual ReLU connections before the same energy computation, realizing the filter-rectify-filter preprocessing thought to underlie second-order motion. Stage II turns every spatial location into a node of a fully connected graph whose adjacency is a cosine-similarity (self-attention) matrix; a convolutional gated recurrent unit repeatedly mixes motion signals across the graph to integrate global motion, and normalized cuts on that graph give object segmentation with no extra training. The datasets that trigger the emergence are generated by a physics-based rendering pipeline, with diffuse (matte) and non-diffuse (specular, glossy, transparent, anisotropic) versions of the same scenes, so the only controlled difference is material.
What would settle it
One decisive check is to run a standard intensity-conservation optical-flow algorithm on the seven modulation movies and measure whether the ground-truth motion direction is recoverable above chance from luminance alone; if any modulation passes, re-score the model on a strictly Fourier-balanced subset—the emergence claim would collapse if the dual-channel model's advantage over the first-order channel disappears once contaminating modulations are removed.
Extended reading notes
Core claim
The discovery is that a task not obviously related to second-order motion—estimating the object-level motion of non-Lambertian materials—is sufficient to produce a human-like second-order motion system. On a benchmark with seven contrast- and texture-modulation types (drift-balanced motion, Gaussian blur, water waves, swirls, noise, and Fourier and pixel shuffles), human observers matched the physical ground truth with mean correlation 0.983; the dual-channel model trained on non-diffuse rendered materials averaged 0.902 with human responses, whereas a representative state-of-the-art optical-flow model scored 0.102. The higher-order channel's units become direction-tuned to drifting second-order gratings, and this tuning is sharpened by non-diffuse training, while the first-order channel remains tuned to luminance motion. The same model also yields training-free object segmentation from the motion graph and reproduces known component- and pattern-cell distributions across the V1-to-MT stages.
Load-bearing premise
The benchmark's seven second-order modulations are assumed to contain no usable first-order luminance cues, but only drift-balanced motion is explicitly designed to be Fourier-balanced; water waves, swirls, noise, blur, and shuffles are called near-indiscernible rather than verified cue-free, so part of the model's measured second-order advantage could be better first-order feature extraction.
Editorial extensions
If this is right
- Second-order motion perception in the model is not a fixed architectural property: it appears only after training on non-Lambertian motion, so the learning signal is the material-driven optical turbulence rather than the presence of explicit second-order labels.
- A single trained model covers both major human motion phenomena: first-order flow estimation that correlates with human perception beyond ground truth, and second-order perception that reaches an average correlation of 0.902 with human judgments.
- The motion graph encodes object structure implicitly, so object segmentation from motion—including segments defined only by drift-balanced motion—falls out of the same trained network without task-specific supervision.
- The dual-channel design stabilizes flow estimation in naturally noisy scenes such as transparent containers with moving liquid, where luminance-based flow models become unstable.
- The model replicates the physiological split between component cells and pattern cells across stages, suggesting the V1-MT pathway and the graph-integration mechanism are functionally equivalent at the level of population responses.
- A testable extension follows from the paper's logic: controlled-rearing experiments with animals in environments rich in specular surfaces (water, glossy foliage) should produce stronger second-order motion sensitivity than matte-only environments, matching the model's learning trajectory.
- The result predicts that adding other sources of optical turbulence during training—caustics, subsurface scattering, heat shimmer—should further improve generalization of the higher-order channel; the paper does not run these ablations.
- The same learning principle could be transferred to engineering: inserting an end-to-end nonlinear preprocessing stream into intensity-conservation optical-flow models might make them robust to non-Lambertian scenes, but the paper leaves that engineering transfer implicit.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a two-stage, biologically inspired model of human visual motion processing. Stage I contains a trainable motion-energy sensor bank (first-order channel) plus a 3D CNN preprocessing channel intended to extract higher-order features, and Stage II uses a self-attention-based motion graph with recurrent GRU integration to produce dense optical flow and, via graph cuts, motion segmentation. The authors train the model on naturalistic datasets and on newly rendered datasets containing diffuse or non-diffuse (glossy, transparent, metallic) objects, and they claim that training on non-Lambertian materials endows the model with human-comparable second-order motion perception. The paper reports first-order motion benchmarks against human perceived flow on Sintel, a seven-modulation second-order benchmark with human psychophysical data, comparisons with several optical-flow models, and qualitative motion-segmentation results.
Significance. If the central claim were fully supported, the paper would make an important interdisciplinary contribution: it would provide a computational account of how second-order motion perception could arise from natural statistics, and it would offer a human-aligned flow model that also handles non-Lambertian optical turbulence. The manuscript has notable strengths: it ships code and data-generation pipelines, it uses human perceived flow with partial correlations to control for ground-truth confounds, and it attempts in-silico neurophysiological validation with component/pattern cell analyses. However, the load-bearing evidence for the evolutionary hypothesis is weakened by two methodological problems: the second-order benchmark does not isolate second-order motion for six of its seven modulations, and the model is trained and evaluated on overlapping Sintel data. The paper's significance is therefore conditional on resolving these issues.
major comments (3)
- [Section 4.3.2 and Fig. 5-C] The second-order benchmark is contaminated by first-order luminance cues. The text states that water wave, swirl, and random flow field modulations 'warp pixels using specific flow fields'; pixel warping creates ordinary luminance displacement, which a first-order motion-energy sensor can lock onto. Only drift-balanced motion is explicitly designed to be Fourier-balanced, and the claim that the other modulations are 'near-indiscernible in Fourier space' is not supported by any spectral analysis or control condition. The high average correlation with human responses (r = 0.902) and the non-diffuse training advantage could therefore reflect better first-order feature tracking rather than a genuinely emergent second-order mechanism. The authors should either restrict the claim to the drift-balanced condition or provide quantitative evidence that each of the seven modulations contains no usable first-order motion information; a simple test is to show that a first-order-only baseline cannot recover the modulation direction above chance.
- [Section 4.2.1 vs. Section 2.2 and Table 1] The model is trained on MPI-Sintel and Sintel-Slow as part of Dataset A (Section 4.2.1) and then validated on the Sintel slow benchmark with human-perceived flows (Section 2.2, Table 1), with no reported train/validation split. This creates a risk of circularity: the high partial correlations with human responses on first-order natural scenes could be inflated by direct supervision on the same benchmark's ground truth. The authors should clarify whether the Sintel-Slow evaluation sequences were excluded from training, or retrain and evaluate on non-overlapping splits.
- [Section 4.2.1 and Fig. 5-C] The central diffuse-versus-non-diffuse ablation is presented without training-seed variance or inferential statistics. The text says the results 'indicate that both the dataset material properties and the model architecture significantly influence' second-order motion perception, but no significance test is reported, and the per-modulation comparisons in Fig. 5-C appear to be single runs. The authors should report multiple training runs, error bars across seeds, and a statistical comparison focused on the drift-balanced condition, which is the only modulation that unambiguously requires second-order processing.
minor comments (5)
- [Section 4.3.2] For the 'random noise' and 'Gaussian blur' modulations, it is unclear whether the carrier itself moves or whether the modulation pattern moves over a static carrier; please clarify the generation procedure.
- [Section 2.3] The phrase 'near-indiscernible in Fourier space' is vague; please report a quantitative measure, such as the ratio of first-order motion energy in the balanced and unbalanced conditions.
- [Equation (2)] There is a typographical error in the definition of Lon: the term 'S * Im[G]) * Im[T]' has an unbalanced parenthesis.
- [Section 4.2.1] The labels 'Type-I' and 'Type-II' for diffuse and non-diffuse training are not explicitly mapped to Datasets D and E; please make this mapping explicit.
- [Code Availability] The repository and project links contain the placeholder 'anoymized'; please update them to the final publicly accessible URLs.
Circularity Check
No significant circularity: the second-order benchmark is held out from training, the diffuse/non-diffuse comparison controls for architecture, and the paper's central claim rests on new human data rather than on a self-citation chain.
full rationale
The claimed derivation — train a dual-channel V1-MT-like model to estimate object motion from non-Lambertian material renderings, then evaluate it on a separate seven-modulation second-order benchmark — does not reduce to its own inputs by construction. The second-order benchmark is not part of the training loss or training dataset; the model's improvement is measured against a diffuse-material-trained control with the same architecture, so the non-Lambertian training signal carries genuine causal content. The higher-order 3D CNN pathway is admittedly inspired by filter-rectify-filter models, but that is an architectural hypothesis, not a fitted parameter, and the paper tests it empirically rather than defining second-order perception as the output of that pathway. The paper does admit that water-wave, swirl, and random-flow modulations are "not pure second-order motion" and are generated by pixel warps, which is a legitimate threat to the construct validity of the benchmark: first-order luminance cues may contaminate several conditions. That is a measurement/confounding concern about whether the task isolates second-order motion, not a case where a prediction is identical to its training target by construction; drift-balanced motion remains an independent condition, and the human-response correlation is a separate external criterion. Self-citations [7, 8, 15, 25] supply prior data, an early model version, and background mechanisms, but the central comparison here does not rest on an unverified self-citation — the human data for the second-order benchmark are newly collected in this paper, and the diffuse/non-diffuse manipulation is controlled. Accordingly, no circular step meeting the "specific reduction" standard is present.
Assumptions & free parameters
free parameters (5)
- Trainable Gabor filter parameters (fs, ft, θ, σ, γ, τ) for 256 ME units =
unknown
- 3D CNN weights (higher-order channel)
- Self-attention/GRU and flow decoder weights (Stage II)
- Learnable scalar s in adjacency exponentiation
- Normalization constants K1, σ1, K2, σ2, spontaneous firing rate α1
assumptions (6)
- domain assumption Gabor quadrature pairs (Eqs. 1-3) capture V1 motion energy as in the Adelson-Bergen model.
- domain assumption The 3D CNN with ReLU nonlinearity implements filter-rectify-filter preprocessing for second-order motion.
- domain assumption The second-order benchmark stimuli contain no usable first-order motion cues.
- ad hoc to paper Non-Lambertian optical turbulence in rendered videos is a natural source of second-order motion signals.
- domain assumption Human perceptual data from three observers reliably represent human second-order motion perception.
- standard math Normalized cuts spectral clustering yields meaningful object segmentation from the motion graph.
Cite this review
Pith. "Pith review of Machine Learning Modeling for Multi-order Human Visual Motion Processing." pith.science (2026). https://pith.science/paper/Y3BR65BD
@misc{pith2026250112810,
author = {Pith},
title = {Pith review of: Machine Learning Modeling for Multi-order Human Visual Motion Processing},
year = {2026},
howpublished = {\url{https://pith.science/paper/Y3BR65BD}},
note = {Machine review of arXiv:2501.12810}
}
read the original abstract
Our research aims to develop machines that learn to perceive visual motion as do humans. While recent advances in computer vision (CV) have enabled DNN-based models to accurately estimate optical flow in naturalistic images, a significant disparity remains between CV models and the biological visual system in both architecture and behavior. This disparity includes humans' ability to perceive the motion of higher-order image features (second-order motion), which many CV models fail to capture because of their reliance on the intensity conservation law. Our model architecture mimics the cortical V1-MT motion processing pathway, utilizing a trainable motion energy sensor bank and a recurrent graph network. Supervised learning employing diverse naturalistic videos allows the model to replicate psychophysical and physiological findings about first-order (luminance-based) motion perception. For second-order motion, inspired by neuroscientific findings, the model includes an additional sensing pathway with nonlinear preprocessing before motion energy sensing, implemented using a simple multilayer 3D CNN block. When exploring how the brain acquired the ability to perceive second-order motion in natural environments, in which pure second-order signals are rare, we hypothesized that second-order mechanisms were critical when estimating robust object motion amidst optical fluctuations, such as highlights on glossy surfaces. We trained our dual-pathway model on novel motion datasets with varying material properties of moving objects. We found that training to estimate object motion from non-Lambertian materials naturally endowed the model with the capacity to perceive second-order motion, as can humans. The resulting model effectively aligns with biological systems while generalizing to both first- and second-order motion phenomena in natural scenes.
Reference graph
Works this paper leans on
-
[1]
Yamins, D. L. & DiCarlo, J. J. Using goal- driven deep learning models to understand sensory cortex. Nature neuroscience 19, 356–365 (2016)
2016
-
[2]
Deep neural networks: a new framework for modeling biological vision and brain information processing
Kriegeskorte, N. Deep neural networks: a new framework for modeling biological vision and brain information processing. Annual review of vision science 1, 417–446 (2015)
2015
-
[3]
Wichmann, F. A. & Geirhos, R. Are deep neural networks adequate behavioral models of human visual perception? Annual Review of Vision Science 9, 501–524 (2023)
2023
-
[4]
Deng, J. et al. Imagenet: A large-scale hierar- chical image database , 248–255 (Ieee, 2009)
2009
-
[5]
& Darrell, T
Long, J., Shelhamer, E. & Darrell, T. Fully convolutional networks for semantic segmen- tation, 3431–3440 (2015)
2015
-
[6]
Dosovitskiy, A. et al. Flownet: Learning opti- cal flow with convolutional networks , 2758– 2766 (2015)
2015
-
[7]
& Nishida, S
Yang, Y.-H., Fukiage, T., Sun, Z. & Nishida, S. Psychophysical measurement of perceived 18 motion flow of naturalistic scenes. iScience in press (2023)
2023
-
[8]
& Nishida, S
Sun, Z., Chen, Y.-J., Yang, Y.-H. & Nishida, S. Comparative analysis of visual motion perception: Computer vision models versus human vision (Oxford, UK, 2023)
2023
Show all 91 references
-
[9]
& Black, M
Ranjan, A., Janai, J., Geiger, A. & Black, M. J. Attacking optical flow , 2404–2413 (2019)
2019
-
[10]
& Welchman, A
Rideaux, R. & Welchman, A. E. But still it moves: static image statistics underlie how we see motion. Journal of Neuroscience 40, 2538–2552 (2020)
2020
-
[11]
& Pack, C
Mineault, P., Bakhtiari, S., Richards, B. & Pack, C. Your head is there to move you around: Goal-driven models of the primate dorsal pathway. Advances in Neural Infor- mation Processing Systems 34, 28757–28771 (2021)
2021
-
[12]
Simoncelli, E. P. & Heeger, D. J. A model of neuronal responses in visual area mt. Vision research 38, 743–761 (1998)
1998
-
[13]
& Gallant, J
Nishimoto, S. & Gallant, J. L. A three- dimensional spatiotemporal receptive field model explains responses of area mt neurons to naturalistic movies. Journal of Neuro- science 31, 14551–14564 (2011)
2011
-
[14]
& Malik, J
Shi, J. & Malik, J. Normalized cuts and image segmentation. IEEE Transactions on pattern analysis and machine intelligence 22, 888– 905 (2000)
2000
-
[15]
& Nishida, S
Sun, Z., Chen, Y.-J., Yang, Y.-H. & Nishida, S. Modeling human visual motion processing with trainable motion energy sensing and a self-attention network (2023). URL https:// openreview.net/forum?id=tRKimbAk5D
2023
-
[16]
& Sperling, G
Chubb, C. & Sperling, G. Drift-balanced random stimuli: a general basis for studying non-fourier motion perception. JOSA A 5, 1986–2007 (1988)
1988
-
[17]
& Mather, G
Cavanagh, P. & Mather, G. Motion: the long and short of it. Spatial vision 4, 103–129 (1989)
1989
-
[18]
O’KEEFE, L. P. & Movshon, J. A. Process- ing of first-and second-order motion signals by neurons in area mt of the macaque mon- key. Visual neuroscience 15, 305–317 (1998)
1998
-
[19]
C., Duistermars, B
Theobald, J. C., Duistermars, B. J., Ringach, D. L. & Frye, M. A. Flies see second-order motion. Current Biology 18, R464–R465 (2008)
2008
-
[20]
Baker Jr, C. L. Central neural mechanisms for detecting second-order motion. Current opinion in neurobiology 9, 461–466 (1999)
1999
-
[21]
& Weiss, Y
Fleet, D. & Weiss, Y. in Optical flow estimation 237–257 (Springer, 2006)
2006
-
[22]
W., Freedman, J
Clifford, C. W., Freedman, J. N. & Vaina, L. M. First-and second-order motion per- ception in gabor micropattern stimuli: psy- chophysics and computational modelling. Cognitive brain research 6, 263–271 (1998)
1998
-
[23]
& Smith, A
Ledgeway, T. & Smith, A. T. Evidence for separate motion-detecting mechanisms for first-and second-order motion in human vision. Vision research 34, 2727–2740 (1994)
1994
-
[24]
T., Greenlee, M
Smith, A. T., Greenlee, M. W., Singh, K. D., Kraemer, F. M. & Hennig, J. The pro- cessing of first-and second-order motion in human visual cortex assessed by functional magnetic resonance imaging (fmri). Journal of Neuroscience 18, 3816–3830 (1998)
1998
-
[25]
& Ashida, H
Nishida, S. & Ashida, H. A hierarchical struc- ture of motion system revealed by interocular transfer of flicker motion aftereffects. Vision Research 40, 265–278 (2000)
2000
-
[26]
Prins, N., Hayes, A. et al. Mechanism inde- pendence for texture-modulation detection is consistent with a filter-rectify-filter mecha- nism. Visual neuroscience 20, 65–76 (2003)
2003
-
[27]
& New- some, W
Movshon, J., Adelson, E., Gizzi, M. & New- some, W. T. in The analysis of moving visual patterns (MIT Press, 1992). 19
1992
-
[28]
Amano, K., Edwards, M., Badcock, D. R. & Nishida, S. Adaptive pooling of visual motion signals by the human visual system revealed with a novel multi-element stimulus. Journal of vision 9, 4–4 (2009)
2009
-
[29]
& Adelson, E
McDermott, J., Weiss, Y. & Adelson, E. H. Beyond junctions: nonlocal form constraints on motion interpretation. Perception 30, 905–923 (2001)
2001
-
[30]
J., Wulff, J., Stanley, G
Butler, D. J., Wulff, J., Stanley, G. B. & Black, M. J. A. Fitzgibbon et al. (Eds.) (ed.) A naturalistic open source movie for optical flow evaluation . (ed.A. Fitzgibbon et al. (Eds.)) European Conf. on Computer Vision (ECCV) , Part IV, LNCS 7577, 611– 625 (Springer-Verlag, 2012)
2012
-
[31]
Fennema, C. L. & Thompson, W. B. Veloc- ity determination in scenes containing sev- eral moving objects. Computer graphics and image processing 9, 301–315 (1979)
1979
-
[32]
Two-frame motion estima- tion based on polynomial expansion , 363–370 (Springer, 2003)
Farneb¨ ack, G. Two-frame motion estima- tion based on polynomial expansion , 363–370 (Springer, 2003)
2003
-
[33]
& Deng, J
Teed, Z. & Deng, J. Raft: Recurrent all-pairs field transforms for optical flow , 402–419 (Springer, 2020)
2020
-
[34]
Luo, A. et al. Learning optical flow with adaptive graph reasoning, Vol. 36, 1890–1898 (2022)
2022
-
[35]
& Tao, D
Xu, H., Zhang, J., Cai, J., Rezatofighi, H. & Tao, D. Gmflow: Learning optical flow via global matching, 8121–8130 (2022)
2022
-
[36]
Huang, Z. et al. Flowformer: A trans- former architecture for optical flow , 668–685 (Springer, 2022)
2022
-
[37]
Solari, F., Chessa, M., Medathati, N. K. & Kornprobst, P. What can we expect from a v1-mt feedforward architecture for optical flow estimation? Signal Processing: Image Communication 39, 342–354 (2015)
2015
-
[38]
& Mikami, A
Handa, T. & Mikami, A. Neuronal corre- lates of motion-defined shape perception in primate dorsal and ventral streams. Euro- pean journal of Neuroscience 48, 3171–3185 (2018)
2018
-
[39]
Ilg, E. et al. Flownet 2.0: Evolution of optical flow estimation with deep networks , 2462–2470 (2017)
2017
-
[40]
& Van Hooser, S
Mazurek, M., Kager, M. & Van Hooser, S. D. Robust quantification of orientation selectiv- ity and direction selectivity. Frontiers in neural circuits 8, 92 (2014)
2014
-
[41]
Shi, X. et al. Videoflow: Exploiting temporal cues for multi-frame optical flow estimation. arXiv preprint arXiv:2303.08340 (2023)
2023 arXiv
-
[42]
D., Kanade, T
Lucas, B. D., Kanade, T. et al. An iterative image registration technique with an applica- tion to stereo vision (Vancouver, 1981)
1981
-
[43]
Jaegle, A. et al. Perceiver io: A general architecture for structured inputs & outputs
-
[44]
W., Lee, J.-Y
Heo, M., Hwang, S., Oh, S. W., Lee, J.-Y. & Kim, S. J. Vita: Video instance segmenta- tion via object token association. Advances in Neural Information Processing Systems 35, 23109–23120 (2022)
2022
-
[45]
Perazzi, F. et al. A benchmark dataset and evaluation methodology for video object seg- mentation (2016)
2016
-
[46]
G., Kir- illov, A
Cheng, B., Misra, I., Schwing, A. G., Kir- illov, A. & Girdhar, R. Masked-attention mask transformer for universal image seg- mentation, 1290–1299 (2022)
2022
-
[47]
& Welchman, A
Rideaux, R. & Welchman, A. E. Explor- ing and explaining properties of motion pro- cessing in biological brains using a neural network. Journal of Vision 21, 11–11 (2021)
2021
-
[48]
& Gomi, H
Nakamura, D. & Gomi, H. Decoding self- motion from visual image sequence pre- dicts distinctive features of reflexive motor responses to visual motion. Neural Networks 162, 516–530 (2023)
2023
-
[49]
& Fleming, R
Storrs, K., Kampman, O., Rideaux, R., Maiello, G. & Fleming, R. Properties of v1 20 and mt motion tuning emerge from unsuper- vised predictive learning. Journal of Vision 22, 4415–4415 (2022)
2022
-
[50]
& Tsotsos, J
Mehrani, P. & Tsotsos, J. K. Self-attention in vision transformers performs perceptual grouping, not attention. arXiv preprint arXiv:2303.01542 (2023)
2023 arXiv
-
[51]
Tsotsos, J. K. et al. Attending to visual motion. Computer Vision and Image Under- standing 100, 3–40 (2005)
2005
-
[52]
Daugman, J. G. & Downing, C. J. Demod- ulation, predictive coding, and spatial vision. JOSA A 12, 641–660 (1995)
1995
-
[53]
Schofield, A. J. What does second-order vision see in an image? Perception 29, 1071–1086 (2000)
2000
-
[54]
& Bruhn, A
Schmalfuss, J., Mehl, L. & Bruhn, A. Distracting downpour: Adversarial weather attacks for motion estimation (IEEE/CVF, 2023)
2023
-
[55]
& Fukiage, T
Nishida, S., Kawabe, T., Sawayama, M. & Fukiage, T. Motion perception: From detec- tion to interpretation. Annual review of vision science 4, 501–523 (2018)
2018
-
[56]
Angelaki, D. E. & Hess, B. J. Self-motion- induced eye movements: effects on visual acuity and navigation. Nature Reviews Neu- roscience 6, 966–976 (2005)
2005
-
[57]
E., Klieger, S
Fencsik, D. E., Klieger, S. B. & Horowitz, T. S. The role of location and motion information in the tracking and recovery of moving objects. Perception & psychophysics 69, 567–577 (2007)
2007
-
[58]
Land, M. F. & Lee, D. N. Where we look when we steer. Nature 369, 742–744 (1994)
1994
-
[59]
J., Tenenbaum, J
Gershman, S. J., Tenenbaum, J. B. & J¨ akel, F. Discovering hierarchical motion structure. Vision research 126, 232–241 (2016)
2016
-
[60]
Bill, J., Gershman, S. J. & Drugowitsch, J. Visual motion perception as online hierarchi- cal inference. Nature communications 13, 7403 (2022)
2022
-
[61]
P., Stepnoski, A
Jones, J. P., Stepnoski, A. & Palmer, L. A. The two-dimensional spectral structure of simple receptive fields in cat striate cortex. Journal of Neurophysiology 58, 1212–1232 (1987)
1987
-
[62]
Jones, J. P. & Palmer, L. A. An evaluation of the two-dimensional gabor filter model of simple receptive fields in cat striate cortex. Journal of neurophysiology 58, 1233–1258 (1987)
1987
-
[63]
Adelson, E. H. & Bergen, J. R. Spatiotem- poral energy models for the perception of motion. Josa a 2, 284–299 (1985)
1985
-
[64]
& Zanker, J
Castet, E. & Zanker, J. Long-range inter- actions in the spatial integration of motion signals. Spatial Vision 12, 287–307 (1999)
1999
-
[65]
Heeger, D. J. Modeling simple-cell direction selectivity with normalized, half-squared, lin- ear operators. Journal of neurophysiology 70, 1885–1898 (1993)
1993
-
[66]
& Heeger, D
Carandini, M. & Heeger, D. J. Summation and division by neurons in primate visual cortex. Science 264, 1333–1336 (1994)
1994
-
[67]
Pack, C. C. & Born, R. T. Temporal dynam- ics of a neural solution to the aperture prob- lem in visual area mt of macaque brain. Nature 409, 1040–1042 (2001)
2001
-
[68]
Visual motion serves but is not under the purview of the dorsal pathway
Gilaie-Dotan, S. Visual motion serves but is not under the purview of the dorsal pathway. Neuropsychologia 89, 378–392 (2016)
2016
-
[69]
& Van Den Berg, A
Noest, A. & Van Den Berg, A. The role of early mechanisms in motion transparency and coherence. Spatial Vision 7, 125–147 (1993)
1993
-
[70]
Wei, Y. et al. Revisiting dilated convolu- tion: A simple approach for weakly-and semi- supervised semantic segmentation, 7268–7277 (2018)
2018
-
[71]
Vaswani, A. et al. Attention is all you need. Advances in neural information processing 21 systems 30 (2017)
2017
-
[72]
Wang, X., Girshick, R., Gupta, A. & He, K. Non-local neural networks, 7794–7803 (2018)
2018
-
[73]
Dosovitskiy, A. et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)
2020 arXiv
-
[74]
Deep learning: the good, the bad, and the ugly
Serre, T. Deep learning: the good, the bad, and the ugly. Annual review of vision science 5, 399–426 (2019)
2019
-
[75]
Kipf, T. N. & Welling, M. Semi-supervised classification with graph convolutional net- works. arXiv preprint arXiv:1609.02907 (2016)
2016 arXiv
-
[76]
& Bengio, Y
Chung, J., Gulcehre, C., Cho, K. & Bengio, Y. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555 (2014)
2014 arXiv
-
[77]
& Kautz, J
Sun, D., Yang, X., Liu, M.-Y. & Kautz, J. Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume , 8934–8943 (2018)
2018
-
[78]
Liu, L. et al. Learning by analogy: Reliable supervision from transformations for unsu- pervised optical flow estimation , 6489–6498 (2020)
2020
-
[79]
& Yuille, A
Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K. & Yuille, A. L. Deeplab: Seman- tic image segmentation with deep convo- lutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pat- tern analysis and machine intelligence 40, 834–848 (2017)
2017
-
[80]
Wang, X., Girdhar, R., Yu, S. X. & Misra, I. Cut and learn for unsupervised object detec- tion and instance segmentation , 3124–3134 (2023)
2023
-
[81]
Weiss, Y., Simoncelli, E. P. & Adelson, E. H. Motion illusions as optimal percepts. Nature neuroscience 5, 598–604 (2002)
2002
-
[82]
Greff, K. et al. Kubric: a scalable dataset generator (2022)
2022
-
[83]
& Bai, Y
Coumans, E. & Bai, Y. Pybullet, a python module for physics simulation for games, robotics and machine learning (2016)
2016
-
[84]
Blender - a 3d modelling and rendering package (2021)
Blender Online Community. Blender - a 3d modelling and rendering package (2021)
2021
-
[85]
Zaal, G. et al. Polyhaven: A curated pub- lic asset library for visual effects artists and game designers (2021)
2021
-
[86]
Peirce, J. et al. Psychopy2: Experiments in behavior made easy. Behavior Research Methods (2019). URL https://dx.doi.org/10. 3758/s13428-018-01193-y
2019
-
[87]
J., Lisberger, S
Priebe, N. J., Lisberger, S. G. & Movshon, J. A. Tuning for spatiotemporal frequency and speed in directionally selective neurons of macaque striate cortex. Journal of Neuro- science 26, 2941–2950 (2006)
2006
-
[88]
& Crowder, N
LeDue, E., Zou, M. & Crowder, N. Spa- tiotemporal tuning in mouse primary visual cortex. Neuroscience letters 528, 165–169 (2012)
2012
-
[89]
J., Cassanello, C
Priebe, N. J., Cassanello, C. R. & Lisberger, S. G. The neural representation of speed in macaque area mt/v5. Journal of Neuro- science 23, 5650–5661 (2003)
2003
-
[90]
Perrone, J. A. & Thiele, A. Speed skills: mea- suring the visual speed analyzing properties of primate mt neurons. Nature neuroscience 4, 526–532 (2001)
2001
-
[91]
J., Cassanello, C
Priebe, N. J., Cassanello, C. R. & Lisberger, S. G. The neural representation of speed in macaque area mt/v5. Journal of Neuro- science 23, 5650–5661 (2003). 22
2003
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.