Skip to content
Auditory

McGurk Effect (Audio-Visual Illusion)

Dub the audio 'ba' over video of a mouth silently saying 'ga,' and most listeners hear neither - they hear 'da,' a third sound their brain invents by fusing mismatched ears and eyes.

This is an audio-visual illusion, not a still image - it only works as video with sound. The description below covers exactly what you'd see and hear.
McGurk Effect (Audio-Visual Illusion)

What you're seeing (and hearing)

In the classic setup, a video shows a person's face repeatedly mouthing the syllable "ga," but the audio track dubbed over the video actually plays the syllable "ba." Watching and listening at the same time, most people report clearly hearing a third sound: "da" - a syllable that matches neither the audio nor the visible mouth movement on its own. Close your eyes, or look away from the screen, and the illusion collapses instantly: with only the audio, the sound is unambiguously "ba."

This is the McGurk effect, and it's one of the clearest demonstrations available that hearing speech is not a purely auditory process. Your brain doesn't treat sound and sight as separate channels it can consult independently and then decide which to trust - for speech specifically, it fuses them into a single percept, and the fusion can produce a sound that was never actually spoken by either channel.

Why it happens

Every consonant sound in speech is associated with a particular configuration of the lips, tongue, and jaw needed to produce it - its articulatory gesture. The sound "ba" requires a full lip closure (a bilabial stop); "ga" is produced further back in the mouth with the lips staying relatively open (a velar stop), and a viewer can often tell the two apart from lip movement alone, well before any sound arrives. Human speech perception evolved, and develops through infancy, in a world where audio and visual speech cues are almost always in agreement, arriving together and reinforcing one another. Because of that, the brain doesn't just use vision as a backup source of information when audio is unclear - it actively integrates visual articulatory cues into the speech-recognition process by default, weighting both sources continuously rather than picking one.

When the two channels are placed in conflict, as in the McGurk-effect setup, the brain doesn't discard one source in favor of the other. It searches for the interpretation that best reconciles both. "Da" is, acoustically and visually, roughly a midpoint between "ba" and "ga" - its place of articulation sits between the full-lip-closure of "ba" and the further-back "ga," making it a plausible compromise the brain can settle on that isn't grossly inconsistent with either the sound it heard or the mouth shape it saw. The result is a genuinely new perceptual experience: a phoneme constructed from the fusion process itself, not simply picked from the two candidates on offer.

This kind of forced multisensory integration isn't unique to speech - the brain regularly combines information across senses to resolve ambiguity, weighting each sense according to how reliable it tends to be for the task at hand. Speech perception is simply a domain where visual and auditory cues are unusually tightly coupled from a neurological standpoint, given how much everyday communication depends on being able to lip-read at least a little, consciously or not, especially in noisy environments.

What makes it a genuine illusion, not just a distraction

It's tempting to assume this is just a case of visual information distracting you from correctly identifying an audio signal you could otherwise hear clearly - but that's not quite right. Listeners aren't guessing or being fooled by inattention; when asked directly, most people confidently and immediately report hearing "da," with no sense that anything is amiss, and no awareness that the audio and video don't match until it's explicitly pointed out or they close their eyes. The mismatch is resolved below the level of conscious awareness, producing a single, seamless, and mistaken percept - the hallmark of a genuine perceptual illusion rather than a simple failure of attention.

A little history

The effect is named for Harry McGurk and his research assistant John MacDonald, who published the discovery in the journal Nature in 1976 in a paper memorably titled "Hearing Lips and Seeing Voices." According to McGurk's own account, the mismatched audio-visual pairing was originally created by accident, during research into how infants perceive speech at different developmental stages, and the researchers noticed the surprising perceptual fusion themselves before realizing its broader significance. The finding has since become a cornerstone of research into multisensory integration, cited across linguistics, cognitive science, and neuroscience, and it remains one of the few illusions that reliably works on the vast majority of people who try it, though the strength of the effect varies somewhat with language background and, notably, tends to be weaker in perceivers who grew up needing to rely less on visual speech cues.

Related reading

The Shepard tone explores a very different auditory illusion - one built from pitch and timing rather than cross-sensory conflict - while inattentional blindness explores another case where perception depends heavily on what the brain is set up to expect.

Discovered / popularized by
Harry McGurk and John MacDonald
Year
1976
Category
Sound Illusions

Read the science behind why this happens →