The short version
The McGurk effect is what happens when the sound you hear and the mouth movements you see disagree, and your brain quietly picks a compromise instead of flagging the mismatch. Play an audio recording of someone saying "ba" under video of a different person's mouth shaping "ga," and most viewers report hearing neither. They hear "da," a syllable that exists in neither the soundtrack nor the video.
It only works with matched audio and video playing together. Reading the word "da" on a page does nothing on its own; the illusion needs your eyes and ears actively disagreeing about the same instant of speech before your brain steps in to referee it.
A dubbing job that wasn't supposed to prove anything
Harry McGurk and his research assistant John MacDonald weren't studying illusions when they found this. They were investigating how infants perceive language at different stages of development, and they needed a technician to dub video footage with mismatched audio for an unrelated part of that project.
When the dubbed tape came back and played, both researchers, adults with no reason to expect anything strange, heard a syllable that wasn't on either track. They assumed the technician had made a dubbing error at first. It hadn't: the mismatch was intentional, and the sound they were hearing was a fusion their own brains had manufactured. McGurk and MacDonald published the finding in Nature on 23 December 1976, in a paper titled "Hearing Lips and Seeing Voices." The pairing that started it, audio "ba" over visual "ga," heard as "da," is now called a fusion. Run the pairing in reverse, audio "ga" over visual "ba," and most people report hearing both sounds blended together as "bga" instead, a combination rather than a fusion.
What your brain is doing that you can't opt out of
There's a phonetic reason the blend lands on "da" specifically, and it comes down to where each sound is made in the mouth. Spoken "ba" is bilabial, made by pressing both lips together, which is exactly the lip movement your eyes are watching for when you lip-read. Spoken "ga" is velar, made further back in the throat, with lip movements that look nothing like "ba." "Da" sits phonetically between the two, made with the tongue against the ridge behind the upper teeth. Confronted with an ear that says bilabial and a set of lips that clearly aren't doing a bilabial sound, the brain doesn't discard either signal. It settles on the intermediate consonant that least contradicts both.
Brain-imaging research points to a specific region responsible for carrying out that settlement: the superior temporal sulcus (STS), which combines auditory and visual speech signals into a single percept before it reaches conscious awareness. That's a correlation on its own, the kind of finding that shows up whenever a brain scanner is running during the illusion, but it doesn't prove the region causes the effect rather than just lighting up alongside it.
A 2010 study by Michael Beauchamp, Audrey Nath, and Siavash Pasalar closed that gap. They used fMRI to locate each participant's own STS, then applied transcranial magnetic stimulation directly to that spot, temporarily disrupting its normal activity. The McGurk illusion weakened measurably when the STS was disrupted, while the same participants' perception of ordinary, matched speech was untouched. That's a causal result, not just an association: knock out this one region and the specific illusion drops, while everything else about how you hear speech keeps working normally.
It doesn't land the same way for everyone
The McGurk effect shows up reliably when researchers average results across a group of test subjects, but individual results underneath that average are far messier than the group number suggests. A 2018 study by Violet Brown and colleagues, published in PLOS ONE, found that the rate at which individual participants report the fused sound ranges from 0% to 100% person to person across published studies. Some people almost never get fused with a given stimulus; others get it almost every time. The researchers tested whether that gap tracked with attention, working memory, or processing speed and found it mostly didn't; the one factor that showed any real connection was a person's skill at extracting information from lip movements alone, and even that explained only a small slice of the variation.
Language and cultural background move the needle too. Kaoru Sekiyama and Yoh'ichi Tohkura's 1991 study in the Journal of the Acoustical Society of America found the effect was measurably weaker in Japanese listeners hearing Japanese syllables than in English listeners hearing English syllables, particularly when the audio itself was easy to make out on its own. Whatever is happening in the STS isn't a fixed, universal wiring diagram installed the same way in every brain. It interacts with how clearly you can hear the specific sound and with a lifetime of listening to one language's specific sound categories.
Understanding it doesn't make it stop
Most optical illusions lose their pull once you see through the trick. The McGurk effect isn't like that. Wikipedia's own overview of the research, citing a 2005 audiovisual-speech study by Colin, Radeau, and Deltenre in the European Journal of Cognitive Psychology among its sources, notes that some people, including researchers who have spent more than twenty years studying the illusion, still perceive it even while fully aware of what is happening and why. Knowing the mechanism in detail doesn't give you the ability to switch it off on command.
That's the same pattern this site keeps running into under different names. It's the reason the uncanny valley stopped working as a reliable warning sign for AI-generated faces right around the time synthetic faces got convincing: a gut sense that something's off isn't built to catch a mismatch your senses are actively smoothing over rather than flagging. It shows up again in automated deepfake detection research: Matyas Bohacek and Hany Farid's 2024 lip-sync detection method doesn't try to make the mismatch visible to a human viewer at all. It transcribes the audio and the video's mouth movements separately with software and compares the two transcripts, because the mismatch itself is exactly the kind of signal a brain built to fuse audio and video into one clean percept isn't wired to surface on its own.
What this is actually useful for
The takeaway isn't "don't trust your senses," which is a fairly useless instruction to live by. It's narrower: your confidence that you would notice if a video's audio and mouth movements didn't quite line up isn't backed by how perception actually works. The McGurk effect is a demonstration, not a metaphor, that a mismatch between what you hear and what you see can get resolved into something that feels like a single, uncomplicated fact before you ever get a chance to question it.
That's a reason to lean on something other than gut feeling when the stakes are real: reverse-searching a clip, checking whether an original unedited version exists somewhere, or just noticing when a video is doing a lot of work to make you feel certain about something you only encountered secondhand. None of that requires distrusting every video you watch. It just means treating a video sounding right to you as the weak evidence it actually is.