The short version
The uncanny valley is the queasy feeling you get from something that looks almost, but not quite, human: a realistic doll, a CGI character, an android. The term comes from a 1970 hypothesis, not a controlled experiment, and for most of its life it functioned more as a shared cultural shorthand than a measured phenomenon.
That shorthand mattered for a practical reason beyond horror movies and robotics conferences: a lot of people, without ever reading Mori's essay, learned to trust the "something feels off" reaction as an informal detector for synthetic faces. Recent research on AI-generated images and deepfakes has started measuring whether that reflex still works, and the answer depends heavily on whether you're looking at a still image, a video, or an audio clip.
Where the idea actually comes from
Masahiro Mori, a robotics professor at the Tokyo Institute of Technology, published a short essay in 1970 titled "Bukimi no Tani" ("The Uncanny Valley") in the Japanese journal Energy. His argument: as a robot or humanoid figure is made more lifelike, people's emotional response to it becomes warmer, up to a point. Just before the figure becomes fully convincing, that warmth collapses into revulsion, and Mori argued movement makes the drop steeper still, before recovering once the figure is indistinguishable from an actual human.
Mori's original essay was never formally translated into English until 2012, when Karl MacDorman and Norri Kageki published an authorized translation in IEEE Robotics & Automation Magazine, the version IEEE Spectrum republished as "the original essay." Before that, the concept spread mostly through a looser 1978 translation by art critic Jasia Reichardt, which is how "uncanny valley" entered English usage well before most readers had access to what Mori actually argued.
Mori himself was explicit that the hypothesis was a suggestion for engineers to think about, not a finding backed by data. He recommended designers aim for the moderate, clearly non-human end of the curve rather than chase full realism, precisely because the valley is hardest to escape once you're near the bottom of it.
The theory, in one graph
The shape of Mori's hypothesis is usually shown as a curve: affinity rising with human likeness, a sharp drop just before "fully human," and a partial recovery. Two things are easy to miss if you've only heard the term secondhand rather than seen the graph. The dip isn't at the low end of realism (a toy robot, a stick figure) and it isn't at the high end (an actual person). It sits specifically at "almost passes."
That detail is exactly why the uncanny valley became a useful shorthand for AI-generated content. A cartoon avatar or an obviously synthetic voice was never in the valley; it reads as artificial and nobody expects otherwise. The valley is the zone modern generative models are aiming directly at, and increasingly hitting.
For still images, the valley has already been crossed
The first hard test of whether the uncanny valley still functions as a detector came from Sophie Nightingale (Lancaster University) and Hany Farid (UC Berkeley), published in the Proceedings of the National Academy of Sciences in February 2022. Across a set of 800 faces, half real and half generated by a GAN (a type of AI image generator), each participant classified a subset as real or fake. Accuracy came in close to chance, around the coin-flip line, meaning the eerie-feeling early-warning system mostly wasn't firing at all. Giving participants brief training with feedback on common synthetic-face artifacts (irregular teeth, odd earrings, background warping) raised accuracy only to about 59%, still far short of reliable.
The more striking finding was about trust, not just accuracy. Participants rated the AI-generated faces as, on average, 7.7% more trustworthy than the real ones. Nightingale and Farid's own framing was blunt: synthesis engines have crossed the uncanny valley for still faces, and the faces on the other side aren't just passable, they're rated more favorably than the real thing, a pattern researchers link to synthetic faces trending toward an average, symmetric appearance that people tend to rate as more trustworthy regardless of whether it's genuine.
For a phenomenon that spent fifty years as a byword for "you'll know it when you see it," a controlled study finding coin-flip accuracy on exactly that task is a specific, falsifiable result rather than a change in mood around AI images.
Video and audio tell a more complicated story
A study from the University of Florida, published in Cognitive Research: Principles and Implications in January 2026 (led by Didem Pehlivanoglu, with Mengdi Zhu and senior author Natalie Ebner among the co-authors), tested both a CNN-based detection algorithm and thousands of human participants on the same deepfake faces, across both still images and video. The pattern split in a way that complicates any simple "AI wins" narrative: on still images, the algorithm was accurate up to 97% of the time while human participants performed no better than chance, mirroring the PNAS result. On video, the pattern reversed. The algorithm's accuracy dropped to chance level, while human participants correctly classified real versus fake video about two-thirds of the time. The researchers also found that participants in a better mood performed worse at spotting fake video, suggesting mood affects skepticism as much as attention does.
A separate 2024 study from MIT Media Lab, published in Nature Communications, ran five pre-registered experiments with 2,215 participants specifically on political speech deepfakes, comparing performance across transcripts, audio, and video. Two findings stand out: people are consistently better at spotting a fake when they can see and hear it than when they only read a transcript of what was supposedly said, and deepfakes voiced with modern text-to-speech synthesis are harder for people to catch than the identical fabricated script performed by a human voice actor. Put together with the University of Florida result, the emerging picture is that the format doing the least to trigger any "something feels off" reaction, synthetic audio, is also the one humans are worst at catching.
That people do better with both sound and picture together than with either alone fits a much older finding about how those two channels combine: the McGurk effect shows mismatched audio and video don't just sit side by side in perception, one can actively override what the other reports, which is a reminder that a convincing deepfake only has to win the channel a viewer is actually paying attention to.
What one photo did to a community of AI-spotters
That split between theory and the current state of AI images is playing out in public in real time, well outside any journal. In late 2026, a poster in the Reddit community r/Aiorfake, which exists specifically for people to test each other on real-versus-AI images, shared a close-up portrait with the claim that it was fully synthetic and that the uncanny valley was, in their words, officially over. The comment section split roughly in half. Several commenters said the image passed a casual glance entirely and only gave itself away under scrutiny of the eye reflections or a repeating texture pattern in the skin. Others pushed back hard, insisting the picture still read as obviously artificial and unsettling to them on a first look, not a delayed one.
Neither side of that argument is wrong, exactly, which is the point worth taking from it. The uncanny valley was never a universal, fixed threshold. It's a perceptual reaction that varies by viewer, by image, and increasingly by how recently that viewer has been exposed to what today's generative models actually produce. A debate that would have been settled instantly a decade ago ("obviously fake") is now a genuine coin flip inside a community of people who look at this content daily for fun.
Why this matters more than a robotics trivia fact
None of this means people have gotten worse at noticing fakes. It means the specific, informal detector a lot of people were quietly relying on (a gut-level "this looks a little wrong") has stopped being reliable for exactly the content where getting it wrong matters most: a scam video call, a fabricated political clip, a cloned voice on the phone asking for money. Concern about what's real online has been rising faster than any measured improvement in people's actual ability to tell the difference, and the uncanny valley research helps explain why that gap exists rather than closing on its own.
The practical upshot mirrors what the research on both still images and video actually rewards: people did better when they had more channels of information to check against each other, not when they leaned harder on a single gut reaction to any one channel. Checking a claim against its source, the way the Five Key Questions of media literacy frame it, does more work than waiting for a feeling of wrongness that, per Nightingale and Farid's data, is now about as reliable as a coin flip for a still photo. That gap between how convincing something feels and whether it holds up under actual scrutiny is the same one the Mandela effect runs into with memory: confidence and accuracy are not the same measurement, and mistaking one for the other is exactly the failure mode synthetic media is now built to exploit.