V1 Saliency Hypothesis
V1 Saliency Hypothesis: V1SH (pronounced as ‘vish’)
V1 Saliency Hypothesis, V1SH, is a theory[1] [2] about V1, the primary visual cortex. It is the only theory so far to not only endow V1 a very important cognitive function, but also to have provided multiple non-trivial theoretical predictions that have been experimentally confirmed subsequently [2][3]. According to V1SH, V1 creates a saliency map from retinal inputs to guide visual attention or gaze shifts[1]. Anatomically, V1 is the gate for retinal visual inputs to enter neocortex, and is also the largest cortical area devoted to vision. In the 1960's, Hubel and Wiesel discovered that V1 neurons are activated by tiny image patches that are large enough to depict a small bar [4] but not a discernible face. This work led to a Nobel prize , and V1 has since been seen as merely serving a back-office function (of image processing) for the subsequent cognitive processing in the brain beyond V1. However, Hubel and Wiesel commented half a century later that little progress has since been made to understand the subsequent visual processing[5]. Outside the box of the traditional views, V1SH is catalyzing a change of framework[6] to enable fresh progresses on understanding vision.
See

for where primary visual cortex is in the brain and relative to the eyes.
V1SH states that V1 transforms the visual inputs into a saliency map of the visual field to guide visual attention or direction of gaze[2][1]. Humans are essentially blind to visual inputs outside their window of attention. Therefore, attention gates visual perception and awareness, and theories of visual attention are cornerstones of theories of visual functions in the brain.
A saliency map is by definition computed from, or caused by, the external visual input rather than from internal factors such as animal’s expectations or goals (e.g., to read a book). Therefore, a saliency map is said to guide attention exogenously rather than endogenously. Accordingly, this saliency map is also called the bottom-up saliency map to guide reflexive or involuntary shifts of attention. For example, it guides our gaze shifts towards an insect flying in our peripheral visual field when we are reading a book.
In this saliency map of the visual field, each visual location has a saliency value. This value is defined as the strength of this location to attract attention exogenously[2]. So if location A has a higher saliency value than location B, then location A is more likely to attract visual attention or gaze shifts towards it than location B. In V1, each neuron can be activated only by visual inputs in a small region of the visual field. This region is called the receptive field of this neuron, and typically covers no more than the size of a coin at an arm’s length[7]. Neighbouring V1 neurons have neighbouring and overlapping receptive fields[7]. Hence, each visual location can simultaneously activate many V1 neurons. According to V1SH, the most activated neuron among these neurons signals the saliency value at this location by its neural activity[1][2]. A V1 neuron’s response to visual inputs within its receptive field is also influenced by visual inputs outside the receptive field[8]. Hence saliency value at each location depends on visual input context[1][2]. This is as it should be since saliency depends on context. For example, a vertical bar is salient in an image in which all the other visual items surrounding it are horiozntal bars, but this same vertical bar is not salient if these other items are all vertical bars instead.
Neural mechanisms in V1 to generate the saliency map

The figure above gives a schematics of the neural mechanisms in V1 to generate the saliency map. In this example, the retinal image has many purple bars, all uniformly oriented (right-tilted) except for one bar that is oriented uniquely (left-tilted). This orientation singleton is the most salient in this image, so it attracts attention or gaze, as observed in psychological experiments[9]. In V1, many neurons have their preferred orientations for visual inputs[7]. For example, a neuron's response to a bar in its receptive field is higher when this bar is oriented in its preferred orientation. Analogously, many V1 neurons have their preferred colours[7]. In this schematic, each input bar to the retina activates two (groups of) V1 neurons, one preferring its orientation and the other preferring its colour. The responses from neurons activated by their preferred orientations in their receptive fields are visualized in the schematics by the black dots in the plane representing the V1 neural responses. Similarly, responses from neurons activated by their preferred colours in their receptive fields are visualized by the purple dots. The sizes of the dots visualize the strengths of the V1 neural responses. In this example, the largest response comes from the neurons preferring and responding to the uniquely oriented bar. This is because of iso-orientation suppression: when two V1 neurons are near each other and have the same or similar preferred orientations, they tend to suppress each other’s activities[8][10]. Therefore, among the group of neurons that prefer and respond to the uniformly oriented background bars, each neuron receives iso-orientation suppression from other neurons of this group[1][8]. Meanwhile, the neuron responding to the orientation singleton does not belong to this group and thus escapes this suppression[1], hence its response is higher than the other neural responses. Iso-colour suppression[11] is analogous to iso-orientation suppression, so all neurons preferring and responding to the purple colours of the input bars are under the iso-colour suppression. According to V1SH, the maximum response at each bar’s location represents the saliency value at each bar’s location[1][2]. This saliency value is thus highest at the location of the orientation singleton, and is represented by the response from neurons preferring and responding to the orientation of this singleton. These saliency values are sent to the superior colliculus[12], a midbrain area, to execute gaze shifts to the receptive field of the most activated neuron responding to visual input space[12]. Hence, for this input image in the figure above, the orientation singleton, which evokes the highest V1 response to this image, attracts visual attention or gaze.
V1SH explains behavioral data on visual search/segmentation
V1SH can explain data on visual search, such as the short response times to find a uniquely red item among green items, or a uniquely vertical bar among horizontal bars, or an item uniquely moving to the right among items moving to the left. These kind of visual searches are called feature searches, when the search target is unique in a basic feature value like orientation, color, or motion direction[9][13]. The shortness of the search response time manifests a higher saliency value at the location of the search target to attract attention. V1SH also explains why it takes longer to find a unique red-vertical bar among red-horizontal bars and green-vertical bars. This is an example of conjunction searches when the search target is unique only by the conjunction of two features, each of which is present in the visual scene[9]. Furthermore, V1SH explains data that are difficult to be explained by alternative frameworks[9][14]. For example (see Fig. 1 of [15] ), two neighboring textures, one made of uniformly left-tilted bar and another of uniformly right-tilted bar, are very easy to be segmented from each other by human vision. This is because the border between the two textures is the most salient location to attract our attention. However, the segmentation becomes much more difficult if another texture made of horizontal and vertical bars are superposed on the original image, and the superposing texture is such that, among any two neighboring original bars, one is superposed by a horizontal bar and the other by a vertical bar[15].
Controversy
V1SH was proposed in late 1990's[16][17] by Li Zhaoping. It was uninfluential initially since for decades it has been believed that attentional guidance is essentially or only controlled by higher-level brain areas. These higher-level brain areas include the frontal eye field and parietal cortical areas[18] in the frontal and more anterior part of the brain, and they are believed to be intelligent for attentional and executive control. In addition, the primary visual cortex, V1, located in occipital lobe in the back or posterior part of the brain, has traditionally been thought of as a low-level visual area that plays mainly a supporting role to other brain areas for their more important visual functions[7]. Opinions started to change by a surprising piece of behavioral data: an item uniquely shown to one eye among similarly appearing items shown to the other eye (using e.g. a pair of glasses for watching 3D movies) can attract gaze or attention. For example, a horizontal bar shown to the left eye in a background of horizontal bars shown to the right eye attracts gaze automatically, even though all the horizontal bars appear identical to viewers. In addition, this unique, left-eye, bar attracts gaze or attention more strongly than a unique and distintive non-horizontal bar among the horizontal bars[19] in the right eye (see an illustration in[20]). This observation was counter-intuitive[20], was easily reproduced by other vision researchers, and was uniquely predicted by V1SH. Since V1 is the only visual cortical area with neurons tuned to eye of origin of visual inputs[4], this observation strongly supports V1's role in guiding attention. More experiments followed to further investigate V1SH[2], and supporting data emerged from functional brain imaging[21] , visual psychophysics[22], and from monkey electrophysiology[3][23]. V1SH has since become more popular in invited and keynote speeches in various prestigious international research conferences[24][25]. V1 is now seen as one of the corner stones in the brain's network of attentional mechanisms[26][27], and its functional role in guiding visual attention is appearing in handbooks[28][29] and textbooks[30][31]. However, some monkey electrophysiological data[3] supporting V1SH are contradicted by a previous piece of monkey electrophysiological data[32].
Zhaoping argues that If V1SH is correct, the ideas[33][34]about how visual system works, and consequently questions to ask for future vision research, should be fundamentally changed[6].
References
- ↑ 1.0 1.1 1.2 1.3 1.4 1.5 1.6 1.7 Li, Zhaoping (2002-01-01). "A saliency map in primary visual cortex". Trends in Cognitive Sciences. 6 (1): 9–16. doi:10.1016/S1364-6613(00)01817-9. ISSN 1364-6613. PMID 11849610.
- ↑ 2.0 2.1 2.2 2.3 2.4 2.5 2.6 2.7 Zhaoping, Li. The V1 hypothesis—creating a bottom-up saliency map for preattentive selection and segmentation. Oxford University Press. doi:10.1093/acprof:oso/9780199564668.001.0001/acprof-9780199564668-chapter-5. ISBN 978-0-19-177250-4. Search this book on
- ↑ 3.0 3.1 3.2 Yan, Yin; Zhaoping, Li; Li, Wu (2018-10-09). "Bottom-up saliency and top-down learning in the primary visual cortex of monkeys". Proceedings of the National Academy of Sciences. 115 (41): 10499–10504. doi:10.1073/pnas.1803854115. ISSN 0027-8424. PMID 30254154.
- ↑ 4.0 4.1 Hubel, D. H.; Wiesel, T. N. (Jan 1962). "Receptive fields, binocular interaction and functional architecture in the cat's visual cortex". The Journal of Physiology. 160 (1): 106–154.2. ISSN 0022-3751. PMC 1359523. PMID 14449617.
- ↑ "David Hubel and Torsten Wiesel". Neuron. 75 (2): 182–184. 2012-07-26. doi:10.1016/j.neuron.2012.07.002. ISSN 0896-6273.
- ↑ 6.0 6.1 Zhaoping, Li (2019-10-01). "A new framework for understanding vision from the perspective of the primary visual cortex". Current Opinion in Neurobiology. Computational Neuroscience. 58: 1–10. doi:10.1016/j.conb.2019.06.001. ISSN 0959-4388.
- ↑ 7.0 7.1 7.2 7.3 7.4 "The Primary Visual Cortex by Matthew Schmolesky – Webvision". Retrieved 2020-07-05.
- ↑ 8.0 8.1 8.2 Knierim, J. J.; van Essen, D. C. (April 1992). "Neuronal responses to static texture patterns in area V1 of the alert macaque monkey". Journal of Neurophysiology. 67 (4): 961–980. doi:10.1152/jn.1992.67.4.961. ISSN 0022-3077. PMID 1588394.
- ↑ 9.0 9.1 9.2 9.3 Treisman, Anne M.; Gelade, Garry (1980-01-01). "A feature-integration theory of attention". Cognitive Psychology. 12 (1): 97–136. doi:10.1016/0010-0285(80)90005-5. ISSN 0010-0285.
- ↑ Allman, J.; Miezin, F.; McGuinness, E. (1985). "Stimulus specific responses from beyond the classical receptive field: neurophysiological mechanisms for local-global comparisons in visual neurons". Annual Review of Neuroscience. 8: 407–430. doi:10.1146/annurev.ne.08.030185.002203. ISSN 0147-006X. PMID 3885829.
- ↑ Wachtler, Thomas; Sejnowski, Terrence J.; Albright, Thomas D. (2003-02-20). "Representation of color stimuli in awake macaque primary visual cortex". Neuron. 37 (4): 681–691. doi:10.1016/s0896-6273(03)00035-7. ISSN 0896-6273. PMC 2948212. PMID 12597864.
- ↑ 12.0 12.1 Schiller, Peter H. (1988), Held, Richard, ed., "Colliculus, Superior", Sensory System I: Vision and Visual Systems, Readings from the Encyclopedia of Neuroscience, Boston, MA: Birkhäuser, pp. 9–9, doi:10.1007/978-1-4899-6647-6_6, ISBN 978-1-4899-6647-6, retrieved 2020-07-05
- ↑ Wolfe, Jeremy. "Visual Search". psycnet.apa.org. Retrieved 2020-07-11. Unknown parameter
|url-status=ignored (help) - ↑ Itti, L.; Koch, C. (Mar 2001). "Computational modelling of visual attention". Nature Reviews. Neuroscience. 2 (3): 194–203. doi:10.1038/35058500. ISSN 1471-003X. PMID 11256080.
- ↑ 15.0 15.1 Zhaoping, Li; May, Keith A. (2007-04-06). "Psychophysical Tests of the Hypothesis of a Bottom-Up Saliency Map in Primary Visual Cortex". PLOS Computational Biology. 3 (4): e62. doi:10.1371/journal.pcbi.0030062. ISSN 1553-7358. PMC 1847698. PMID 17411335.
- ↑ Li, Zhaoping (1999-08-31). "Contextual influences in V1 as a basis for pop out and asymmetry in visual search". Proceedings of the National Academy of Sciences. 96 (18): 10530–10535. doi:10.1073/pnas.96.18.10530. ISSN 0027-8424. PMID 10468643.
- ↑ Li, Zhaoping (1998). "Primary cortical dynamics for visual grouping " as a book chapter in "Theoretical Aspects of Neural Computation", Eds K.M. Wong, I. King, and D.Y. Yeung. Springer-verlag. pp. 155–164. Search this book on
- ↑ Desimone, Robert; Duncan, John (Mar 1995). "Neural Mechanisms of Selective Visual Attention". Annual Review of Neuroscience. 18 (1): 193–222. doi:10.1146/annurev.ne.18.030195.001205. ISSN 0147-006X.
- ↑ Zhaoping, Li (2008-05-01). "Attention capture by eye of origin singletons even without awareness—A hallmark of a bottom-up saliency map in the primary visual cortex". Journal of Vision. 8 (5): 1–1. doi:10.1167/8.5.1. ISSN 1534-7362.
- ↑ 20.0 20.1 Zhaoping, Li (2014-08-21). "Are we too "smart" to understand how we see?". OUPblog. Retrieved 2020-07-11. Unknown parameter
|url-status=ignored (help) - ↑ Zhang, Xilin; Zhaoping, Li; Zhou, Tiangang; Fang, Fang (2012-01-12). "Neural Activities in V1 Create a Bottom-Up Saliency Map". Neuron. 73 (1): 183–192. doi:10.1016/j.neuron.2011.10.035. ISSN 0896-6273.
- ↑ Kennett, Matthew J.; Wallis, Guy (2019-07-01). "The face-in-the-crowd effect: Threat detection versus iso-feature suppression and collinear facilitation". Journal of Vision. 19 (7): 6–6. doi:10.1167/19.7.6. ISSN 1534-7362.
- ↑ Wagatsuma, Nobuhiko; Hidaka, Akinori; Tamura, Hiroshi (2021-01-12). "Correspondence between Monkey Visual Cortices and Layers of a Saliency Map Model Based on a Deep Convolutional Neural Network for Representations of Natural Images". eNeuro. 8 (1). doi:10.1523/ENEURO.0200-20.2020. ISSN 2373-2822. PMC 7890521 Check
|pmc=value (help). PMID 33234544 Check|pmid=value (help). - ↑ "Visual Perception meets Computational Neuroscience | www.ecvp.uni-bremen.de". www.ecvp.uni-bremen.de. Retrieved 2020-07-11.
- ↑ "CNS 2020". www.cnsorg.org. Retrieved 2021-06-27.
- ↑ Bisley, James W.; Goldberg, Michael E. (June 2010). "Attention, Intention, and Priority in the Parietal Lobe". Annual Review of Neuroscience. 33 (1): 1–21. doi:10.1146/annurev-neuro-060909-152823. ISSN 0147-006X. PMC 3683564. PMID 20192813.
- ↑ "The brain circuitry of attention". Trends in Cognitive Sciences. 8 (5): 223–230. 2004-05-01. doi:10.1016/j.tics.2004.03.004. ISSN 1364-6613.
- ↑ The Oxford Handbook of Attention. Oxford University Press. 2014-01-01. doi:10.1093/oxfordhb/9780199675111.001.0001/oxfordhb-9780199675111. ISBN 978-0-19-175301-5. Search this book on
- ↑ Wolfe, Jeremy M. (2018), "Visual Search", Stevens' Handbook of Experimental Psychology and Cognitive Neuroscience, American Cancer Society, pp. 1–55, doi:10.1002/9781119170174.epcn213, ISBN 978-1-119-17017-4, retrieved 2021-06-24
- ↑ "Visual Attention and Consciousness". Routledge & CRC Press. Retrieved 2021-06-24.
- ↑ Zhaoping, Li (2014-05-08). Understanding Vision: Theory, Models, and Data. Oxford, New York: Oxford University Press. ISBN 978-0-19-956466-8. Search this book on
- ↑ White, Brian J.; Kan, Janis Y.; Levy, Ron; Itti, Laurent; Munoz, Douglas P. (2017-08-29). "Superior colliculus encodes visual saliency before the primary visual cortex". Proceedings of the National Academy of Sciences. 114 (35): 9451–9456. doi:10.1073/pnas.1701003114. ISSN 0027-8424. PMID 28808026.
- ↑ "Webvision – The Organization of the Retina and Visual System". Retrieved 2020-07-05.
- ↑ Stone, James. "Vision and Brain | The MIT Press". mitpress.mit.edu. Retrieved 2020-07-05. Unknown parameter
|url-status=ignored (help)
This article "V1 Saliency Hypothesis" is from Wikipedia. The list of its authors can be seen in its historical and/or the page Edithistory:V1 Saliency Hypothesis. Articles copied from Draft Namespace on Wikipedia could be seen on the Draft Namespace of Wikipedia and not main one.
| This page exists already on Wikipedia. |
