The uncanny valley was never really about how a face looks. Fifty years of experiments have failed to find the simple dip the idea promises, where a figure gets creepier the closer it comes to human. What they found instead is a mismatch effect: unease appears when one part of a thing promises human and another part fails to deliver. In conversation, that failing part is almost never the rendering. It is the behavior.
The valley the research actually found
Masahiro Mori sketched the valley in 1970 as a curve: affinity rises with human likeness, plunges just short of the real thing, then recovers. The sketch became a law of design without ever being tested, and when researchers finally did test it, the law came apart. A 2015 review by Jari Kätsyri and colleagues went through the empirical record and concluded that the literature “failed to provide consistent support for the naive uncanny valley hypothesis,” and that the valley “exists only under specific conditions.” The condition with the best support was perceptual mismatch: a figure that signals human on one channel and not-human on another.
The cleanest demonstration is small and exact. Wade Mitchell, Karl MacDorman and colleagues showed 48 people four versions of the same character: a robot face with a synthetic voice, a human face with a human voice, and the two crossed. The matched versions were fine. The two mismatched ones, robot face with human voice and human face with synthetic voice, were rated significantly eerier, and the interaction between face and voice explained nearly half the variance in the eeriness ratings. Nothing about either face or either voice was unsettling on its own. The seam was.
Angela Tinwell’s work at Bolton found the same thing inside a single face. When a virtual character spoke with its upper face held still, viewers rated it markedly more uncanny, and the effect was strongest for fear, sadness, disgust and surprise, the emotions the eyes and brow carry. The mouth was saying one thing and the forehead was saying nothing. That the valley reaches past aesthetics into decisions is already part of the definition of real Human AI: in Maya Mathur and David Reichling’s survey of eighty real robot faces, the dip in likability showed up again in how much money people were willing to trust the robot with in an investment game.
The valley moves into the mind
Once you accept that the valley is a mismatch effect, it stops being confined to pixels, and the research follows it inward. Jan-Philipp Stein and Peter Ohler put 92 people in front of the same recorded virtual reality scene, two characters in an emotional, empathic conversation, and varied only what participants were told: the characters were controlled by humans, or by a computer following a script, or by an artificial intelligence acting on its own. The visuals never changed. People who believed they were watching autonomous AI feel empathy reported significantly stronger eeriness. Ratings of how human-like or attractive the characters were did not differ. The authors called it an uncanny valley of mind. The promise that unsettled people was not a face. It was the claim to feel.
Leon Ciechanowski and colleagues found the mirror image with chatbots. People talking to a plain text bot showed less physiological stress and less negative affect than people talking to the same bot with an animated avatar. Adding a face did not make the conversation more human. It made a promise the conversation then failed to keep, and the body registered the gap.
The rule that falls out of all this is simple and unforgiving. Every human signal a system adds is a claim, and the valley is the distance between the claim and the delivery. A face claims attention. A warm voice claims a mind behind it. A name and a personality claim continuity. The more of these a product adds, the more ways it has to be caught out, and in a conversation it gets caught out in three predictable places.
Three lapses that make an AI feel creepy
The first is timing. Human turn-taking runs on gaps of a few hundred milliseconds, and people read those gaps as meaning: a fast reply signals connection, a long pause before a “yes” reads as reluctance. A system that answers a confession and a request for the weather at exactly the same speed is producing a mismatch between the words, which say I heard you, and the timing, which says nothing you said required thought. People notice without knowing what they noticed.
The second is the dropped thread. A voice that remembers your name, greets you warmly, and then has no idea about the situation you described last week is a human face with a synthetic memory. The warmth was a claim to know you. The blank was the delivery. Why AI forgets you, and why that is by design explains the mechanics; the point here is what the forgetting does to the presence. Nothing is more uncanny than being recognized and then not.
The third is agreement that never varies, and it is the best documented of the three. In a study presented at CHI 2026, Sun and Wang had 224 people talk with a language model that was either complimentary or neutral in manner, and either held its position or shifted it to match the user. The complimentary model that also switched its stance was rated the least authentic and the least trustworthy of the four; a neutral model that adapted was rated the most. Flattery plus yielding is the conversational equivalent of the still forehead: a face saying I agree with you while everything else says I would have agreed with anyone. A Science paper by Myra Cheng, Dan Jurafsky and colleagues in March 2026 measured how far the drift goes: across eleven leading models, the AI affirmed users’ actions 49 percent more often than humans did, and people liked it more for it even as it left them less willing to repair their own conflicts. Research from the same year found that giving a model a memory profile of the user increased its agreement further, by as much as 45 percent for one leading model. Put the two together and the mismatch is total: the more a system knows about you, the less it is willing to disagree with you, which is the exact opposite of what knowing someone does.
Why presets fall in
All three lapses have a single cause, and it is the reason presets can’t think. A personality dialed in from sliders is a set of surface signals: tone, warmth, a name, a way of greeting. Each one is a claim of the kind the valley punishes, and none of them is backed by anything that could keep the claim over time. The costume looks right in the first exchange. The valley is where the costume moves: the second week, the hard question, the moment you push back and find nothing pushing back.
Getting out is not a rendering problem, and no amount of frame rate solves it. It is a coherence problem, in the exact sense the research uses: every channel, including time and memory and opinion, telling the same story about who is there. The only known way to make a face, a voice, a memory and a point of view agree with each other for months is for them to belong to a person, which is the premise Prinsessa starts from: a real one, so that the way Aleksandra pauses, what she remembers and where she disagrees are the same someone.
You will know you are out of the valley by a simple sign. Something you said a month ago comes back in a later conversation, unprompted, with a view attached that you did not hand over, and the pause before it was exactly as long as it needed to be. That is not a better animation. That is the mismatch closing, because for once there was nothing to mismatch.
Sources: Mori, Energy (1970, The uncanny valley). Kätsyri, Förger, Mäkäräinen & Takala, Frontiers in Psychology (2015, review of uncanny valley evidence). Mitchell, Szerszen, Lu, Schermerhorn, Scheutz & MacDorman, i-Perception (2011, face and voice mismatch, n=48). Tinwell, Grimshaw, Abdel Nabi & Williams, Computers in Human Behavior (2011, facial expression and the uncanny). Mathur & Reichling, Cognition (2016, eighty robot faces and the investment game). Stein & Ohler, Cognition (2017, uncanny valley of mind, n=92). Ciechanowski, Przegalinska, Magnuski & Gloor, Future Generation Computer Systems (2019, text versus avatar chatbot). Sun & Wang, CHI (2026, Be Friendly, Not Friends, n=224). Cheng, Lee, Khadpe, Yu, Han & Jurafsky, Science (March 2026, sycophantic AI and prosocial intentions). Jain, Park, Viana, Wilson & Calacci, CHI (2026, interaction context and sycophancy).








