How the Brain Processes Visual Information
A cornerstone overview of how light becomes sight - from the retina through the optic nerve to the visual cortex, and why that pipeline leaves room for illusions.
Seeing feels effortless: you open your eyes and the world is simply there, already sorted into objects, distances, colors, and motion. That effortlessness is an illusion in its own right. Behind it sits one of the most computation-heavy processes your body performs, involving roughly a third of the human cortex and several distinct relay stations, each of which transforms raw light into something closer to knowledge.
From photons to signals
Vision starts as physics, not psychology. Light bounces off a surface, enters the eye through the cornea and lens, and lands on the retina - a thin sheet of tissue at the back of the eyeball packed with photoreceptor cells called rods and cones. Cones handle color and fine detail in bright light; rods handle low-light, black-and-white vision. Neither of these cells "sees" in any meaningful sense. They simply convert photons into electrochemical signals, a process called phototransduction.
What happens next is where vision starts to become interesting. The retina isn't a passive relay - it's already doing computation before a signal ever leaves the eye. Layers of bipolar, horizontal, amacrine, and ganglion cells compare each photoreceptor's output to its neighbors', a process built around center-surround receptive fields: a cell fires strongest when light hits its center but is dimmed at its edges (or vice versa). This wiring produces lateral inhibition, where active neurons suppress their neighbors, sharpening edges and boosting contrast before the signal even reaches the brain. It's an early, elegant example of the visual system prioritizing change and contrast over raw brightness - and it's part of why illusions like the Hermann grid illusion exist: the ghostly gray blobs you see at grid intersections are a side effect of this contrast-sharpening machinery misfiring on a repeating pattern.
The relay through the thalamus
Roughly a million axons per eye bundle into the optic nerve and carry these processed signals toward the brain. They don't go straight to the visual cortex. Most first pass through the lateral geniculate nucleus (LGN), a way station in the thalamus that organizes incoming signals by eye of origin, spatial location, and stimulus type, and applies its own filtering - boosting some signals, suppressing others - informed by feedback from higher brain areas. Even at this early stage, vision is not a one-way camera feed; it's a conversation between incoming data and the brain's existing expectations, a theme that becomes central once you get to predictive processing.
Arrival in the cortex - and the split into "what" and "where"
From the LGN, signals arrive at the primary visual cortex (V1) at the back of the brain, where cells respond to simple features: edges at particular orientations, spots of contrast, basic motion direction. For a fuller look at how this region and its downstream partners work, see the role of the visual cortex. From V1, processing splits into two broad streams. The ventral stream runs toward the temporal lobe and specializes in identifying what something is - object recognition, face recognition, reading. The dorsal stream runs toward the parietal lobe and handles where something is and how to act on it - spatial location, motion, guiding a reaching hand. This division, first proposed as a "what" versus "where" distinction by neuroscientists Leslie Ungerleider and Mortimer Mishkin in the early 1980s and later reframed by David Milner and Melvyn Goodale as a "what" versus "how" distinction, explains why certain brain injuries can leave someone able to describe an object's shape but unable to reach for it accurately, or vice versa.
Why none of this is a camera
The natural assumption is that this whole pipeline works like a camera piping a feed to a screen in your head. It doesn't. At every stage - retina, thalamus, cortex - the system is compressing, filtering, and interpreting, not recording. Your retina has a blind spot where the optic nerve exits, yet you never see a hole in your visual field, because the brain fills it in using surrounding information. Your eyes make several small movements a second, yet the world doesn't appear to jump around. Depth, which doesn't exist on your flat retina at all, gets reconstructed from cues like relative size, overlap, and perspective lines, which is exactly what illusions like the Ponzo illusion and Müller-Lyer illusion exploit - they feed the depth-reconstruction system cues that suggest three dimensions even though the image is flat.
Why this matters for illusions
Every optical illusion on this site works by targeting a specific stage of this pipeline. Some exploit retinal-level contrast processing. Some exploit cue-combination rules the cortex uses to infer depth or lightness, as in the checker shadow illusion. Some exploit the fact that the brain organizes raw shapes into meaningful wholes, a process explored in Gestalt principles of perception. None of them are "tricks" in the sense of fooling a broken system. They're demonstrations of a system built for speed and useful inference under real-world conditions, doing exactly what it evolved to do - just under conditions engineered to expose the shortcuts. Understanding that pipeline, roughly, is the first step to understanding why optical illusions happen at all.