MEDIA LITERACY

The Uncanny Valley: Why Your Gut Feeling No Longer Catches AI Fakes

A 1970 hypothesis about eerie robots became the go-to explanation for why AI faces feel "off." Recent research shows it stopped working for still images years ago.

LAST UPDATED 2026-08-23

Roboticist Hiroshi Ishiguro seated beside Geminoid HI-1, the android built as his physical double, at Ars Electronica Festival 2009.
Photo: Shervinafshar, Hiroshi Ishiguro and Geminoid HI-1 at Ars Electronica Festival 2009 — CC BY-SA 3.0

CORE SUMMARY

The uncanny valley is a 1970 hypothesis by Japanese roboticist Masahiro Mori: as something artificial becomes more human-like, our affinity for it rises, then plunges sharply just short of realistic, before recovering at true human likeness. For decades, that dip functioned as an informal early-warning system for spotting artificial faces. A 2022 study in PNAS by Sophie Nightingale and Hany Farid found that no longer holds for still images: participants classified AI-synthesized faces correctly at close to chance, and rated the synthetic faces as more trustworthy than real ones. A 2026 University of Florida study found the same pattern (AI-detection software hit 97% accuracy on fake face images while humans were at chance) but the reverse for video, where humans still outperformed the algorithm. A 2024 MIT Media Lab study of 2,215 participants found audio deepfakes made with modern text-to-speech are the hardest format of all for humans to catch. The intuitive "something feels off" reflex the uncanny valley describes has not disappeared. It has just stopped firing reliably for exactly the content it would be most useful for.

The short version

The uncanny valley is the queasy feeling you get from something that looks almost, but not quite, human: a realistic doll, a CGI character, an android. The term comes from a 1970 hypothesis, not a controlled experiment, and for most of its life it functioned more as a shared cultural shorthand than a measured phenomenon.

That shorthand mattered for a practical reason beyond horror movies and robotics conferences: a lot of people, without ever reading Mori's essay, learned to trust the "something feels off" reaction as an informal detector for synthetic faces. Recent research on AI-generated images and deepfakes has started measuring whether that reflex still works, and the answer depends heavily on whether you're looking at a still image, a video, or an audio clip.

Where the idea actually comes from

Masahiro Mori, a robotics professor at the Tokyo Institute of Technology, published a short essay in 1970 titled "Bukimi no Tani" ("The Uncanny Valley") in the Japanese journal Energy. His argument: as a robot or humanoid figure is made more lifelike, people's emotional response to it becomes warmer, up to a point. Just before the figure becomes fully convincing, that warmth collapses into revulsion, and Mori argued movement makes the drop steeper still, before recovering once the figure is indistinguishable from an actual human.

Mori's original essay was never formally translated into English until 2012, when Karl MacDorman and Norri Kageki published an authorized translation in IEEE Robotics & Automation Magazine, the version IEEE Spectrum republished as "the original essay." Before that, the concept spread mostly through a looser 1978 translation by art critic Jasia Reichardt, which is how "uncanny valley" entered English usage well before most readers had access to what Mori actually argued.

Mori himself was explicit that the hypothesis was a suggestion for engineers to think about, not a finding backed by data. He recommended designers aim for the moderate, clearly non-human end of the curve rather than chase full realism, precisely because the valley is hardest to escape once you're near the bottom of it.

The theory, in one graph

The shape of Mori's hypothesis is usually shown as a curve: affinity rising with human likeness, a sharp drop just before "fully human," and a partial recovery. Two things are easy to miss if you've only heard the term secondhand rather than seen the graph. The dip isn't at the low end of realism (a toy robot, a stick figure) and it isn't at the high end (an actual person). It sits specifically at "almost passes."

That detail is exactly why the uncanny valley became a useful shorthand for AI-generated content. A cartoon avatar or an obviously synthetic voice was never in the valley; it reads as artificial and nobody expects otherwise. The valley is the zone modern generative models are aiming directly at, and increasingly hitting.

Graph of Masahiro Mori's uncanny valley hypothesis, showing affinity (familiarity) rising with human likeness, then dropping sharply near-but-not-quite-human before recovering at full human likeness, with a steeper dip for moving figures than still ones
Diagram: Smurrayinchester, based on Mori and MacDorman's original chart — CC BY-SA 3.0

For still images, the valley has already been crossed

The first hard test of whether the uncanny valley still functions as a detector came from Sophie Nightingale (Lancaster University) and Hany Farid (UC Berkeley), published in the Proceedings of the National Academy of Sciences in February 2022. Across a set of 800 faces, half real and half generated by a GAN (a type of AI image generator), each participant classified a subset as real or fake. Accuracy came in close to chance, around the coin-flip line, meaning the eerie-feeling early-warning system mostly wasn't firing at all. Giving participants brief training with feedback on common synthetic-face artifacts (irregular teeth, odd earrings, background warping) raised accuracy only to about 59%, still far short of reliable.

The more striking finding was about trust, not just accuracy. Participants rated the AI-generated faces as, on average, 7.7% more trustworthy than the real ones. Nightingale and Farid's own framing was blunt: synthesis engines have crossed the uncanny valley for still faces, and the faces on the other side aren't just passable, they're rated more favorably than the real thing, a pattern researchers link to synthetic faces trending toward an average, symmetric appearance that people tend to rate as more trustworthy regardless of whether it's genuine.

For a phenomenon that spent fifty years as a byword for "you'll know it when you see it," a controlled study finding coin-flip accuracy on exactly that task is a specific, falsifiable result rather than a change in mood around AI images.

Video and audio tell a more complicated story

A study from the University of Florida, published in Cognitive Research: Principles and Implications in January 2026 (led by Didem Pehlivanoglu, with Mengdi Zhu and senior author Natalie Ebner among the co-authors), tested both a CNN-based detection algorithm and thousands of human participants on the same deepfake faces, across both still images and video. The pattern split in a way that complicates any simple "AI wins" narrative: on still images, the algorithm was accurate up to 97% of the time while human participants performed no better than chance, mirroring the PNAS result. On video, the pattern reversed. The algorithm's accuracy dropped to chance level, while human participants correctly classified real versus fake video about two-thirds of the time. The researchers also found that participants in a better mood performed worse at spotting fake video, suggesting mood affects skepticism as much as attention does.

A separate 2024 study from MIT Media Lab, published in Nature Communications, ran five pre-registered experiments with 2,215 participants specifically on political speech deepfakes, comparing performance across transcripts, audio, and video. Two findings stand out: people are consistently better at spotting a fake when they can see and hear it than when they only read a transcript of what was supposedly said, and deepfakes voiced with modern text-to-speech synthesis are harder for people to catch than the identical fabricated script performed by a human voice actor. Put together with the University of Florida result, the emerging picture is that the format doing the least to trigger any "something feels off" reaction, synthetic audio, is also the one humans are worst at catching.

That people do better with both sound and picture together than with either alone fits a much older finding about how those two channels combine: the McGurk effect shows mismatched audio and video don't just sit side by side in perception, one can actively override what the other reports, which is a reminder that a convincing deepfake only has to win the channel a viewer is actually paying attention to.

What one photo did to a community of AI-spotters

That split between theory and the current state of AI images is playing out in public in real time, well outside any journal. In late 2026, a poster in the Reddit community r/Aiorfake, which exists specifically for people to test each other on real-versus-AI images, shared a close-up portrait with the claim that it was fully synthetic and that the uncanny valley was, in their words, officially over. The comment section split roughly in half. Several commenters said the image passed a casual glance entirely and only gave itself away under scrutiny of the eye reflections or a repeating texture pattern in the skin. Others pushed back hard, insisting the picture still read as obviously artificial and unsettling to them on a first look, not a delayed one.

Neither side of that argument is wrong, exactly, which is the point worth taking from it. The uncanny valley was never a universal, fixed threshold. It's a perceptual reaction that varies by viewer, by image, and increasingly by how recently that viewer has been exposed to what today's generative models actually produce. A debate that would have been settled instantly a decade ago ("obviously fake") is now a genuine coin flip inside a community of people who look at this content daily for fun.

Why this matters more than a robotics trivia fact

None of this means people have gotten worse at noticing fakes. It means the specific, informal detector a lot of people were quietly relying on (a gut-level "this looks a little wrong") has stopped being reliable for exactly the content where getting it wrong matters most: a scam video call, a fabricated political clip, a cloned voice on the phone asking for money. Concern about what's real online has been rising faster than any measured improvement in people's actual ability to tell the difference, and the uncanny valley research helps explain why that gap exists rather than closing on its own.

The practical upshot mirrors what the research on both still images and video actually rewards: people did better when they had more channels of information to check against each other, not when they leaned harder on a single gut reaction to any one channel. Checking a claim against its source, the way the Five Key Questions of media literacy frame it, does more work than waiting for a feeling of wrongness that, per Nightingale and Farid's data, is now about as reliable as a coin flip for a still photo. That gap between how convincing something feels and whether it holds up under actual scrutiny is the same one the Mandela effect runs into with memory: confidence and accuracy are not the same measurement, and mistaking one for the other is exactly the failure mode synthetic media is now built to exploit.

Frequently asked questions

What is the uncanny valley?

A hypothesis proposed by Japanese roboticist Masahiro Mori in a 1970 essay, describing how human emotional affinity for a humanlike figure rises with realism, then drops sharply into revulsion just short of fully convincing, before recovering once the figure is indistinguishable from an actual human. Movement, Mori argued, makes both the rise and the drop steeper.

Who coined the term and when?

Masahiro Mori, in a 1970 essay titled "Bukimi no Tani" published in the Japanese journal Energy. An informal English translation by art critic Jasia Reichardt spread the phrase "uncanny valley" starting in 1978; the first translation authorized by Mori himself, by Karl MacDorman and Norri Kageki, was published in IEEE Robotics & Automation Magazine in 2012.

Does the uncanny valley still work as a way to spot AI-generated content?

For still images, largely no. A 2022 PNAS study by Sophie Nightingale and Hany Farid found people classify real versus AI-generated faces at close to chance accuracy, and rate the AI faces as more trustworthy on average. A 2026 University of Florida study found the same near-chance result for human judgment of fake face images, though humans still outperformed AI detection software on video, correctly classifying about two-thirds of fake videos.

Why are audio deepfakes harder to catch than video ones?

A 2024 MIT Media Lab study of 2,215 participants, published in Nature Communications, found that people rely more on how something is said (voice, expression, delivery) than on the words themselves, and that deepfakes voiced with modern text-to-speech synthesis are harder to detect than the same script performed by a real voice actor. Audio strips away the visual cues (facial micro-expressions, lighting inconsistencies) that still give away some fake video.

READER VERDICT

Did this entry hold up?

Written and edited by the Hollowvane Editorial Team