Read a stranger’s argument and you will rate them less intelligent than if you had heard them make it. Same words, same order, same person. The difference is not what was said. It is that a voice carries the thinking, and a page carries only the conclusion.
The finding comes from Juliana Schroeder and Nicholas Epley at the University of Chicago, who spent most of a decade testing one question: what does a voice tell you about a mind that text does not? Their answer explains why a phone call can repair what ten messages could not, why a two-line reply can read as cold from someone who is not, and why the move from typing to talking is the largest single change any conversation with a machine can make.
Same words, different mind
In the first study, published in Psychological Science in 2015, 162 evaluators judged job candidates who had recorded short pitches about why they should be hired. Some evaluators watched the video, some heard only the audio, some read a transcript. The candidates who were heard were rated as more competent, more thoughtful, and more intelligent than the same candidates read on paper, and the evaluators were more inclined to hire them. The effect was not small: a difference of about six tenths of a standard deviation on judged intellect. When the researchers ran it again with professional recruiters instead of students, the gap doubled.
The detail that matters most is what did not help. Adding video to the audio changed nothing. Seeing the candidate’s face, gestures, and clothes added no information about their mind that the voice had not already carried. The intelligence was in the sound.
Voice works hardest where you disagree
The 2017 follow-up, with Michael Kardas, took the finding somewhere sharper. Across four experiments with more than two thousand participants, people listened to or read strangers explaining views on abortion, war, and the 2016 election. When the speaker disagreed with them, readers rated the speaker as less mentally capable, less fully human in their reasoning. Listeners did not. Hearing the argument spoken made the person behind it seem to have a real mind, even a mind the listener wanted to reject. A voice produced by text-to-speech software did not close the gap; the authentic human voice did, and the researchers traced the effect to intonation and pauses. The paper’s closing line is worth keeping: the tendency to write off the minds of people we disagree with can be tempered by giving them, quite literally, a voice.
This is the mechanism behind the bond. Text delivers positions. A voice delivers the evidence that a position was arrived at: the hesitation before the hard part, the emphasis that shows what the speaker cares about, the pace that reveals whether they are sure. We grant minds to people whose thinking we can hear. On the page, thinking is invisible, and so is the person.
What the page strips out
The physiology points the same direction, and it is already established ground. Hearing a familiar voice after stress releases oxytocin and lowers cortisol; reading the same reassurance as a text message does neither, a result that sits at the foundation of what real Human AI means. Hormones aside, the voice may be a better channel for reading emotion than the face. Michael Kraus at Yale ran five experiments with more than 1,700 people and found that listeners judged what a partner was feeling more accurately with voice alone than with video and voice together. The effects in live conversation were modest, and one published comment has questioned their size, but the direction held in every study: adding the picture pulled attention away from the cues that carried the feeling.
Put the three lines of research together and the picture is consistent. The voice carries the mind. The voice carries the emotion. And the voice moves the body in ways text cannot reach. None of this is about warmth as a style. It is about information that only the sound contains.
The catch: hearing a mind is not being helped by one
Here the argument has to turn, because the obvious conclusion for AI is wrong. If voice reveals a thoughtful mind, then giving a system a beautiful voice should build a bond that helps people. The largest experiment on the question says it does not, at least not on its own.
Researchers at the MIT Media Lab and OpenAI ran a four-week randomized trial with 981 people, each assigned to talk with ChatGPT by text, by a neutral voice, or by an engaging voice, and tracked loneliness, socializing, emotional dependence, and problematic use across more than 300,000 messages. The first analysis, released in March 2025, suggested the voice conditions started out better than text on loneliness and dependence. The revised analysis, released in October 2025, withdrew that: no significant effect of modality survived. What predicted worse outcomes on every measure was simply how much a person used the system, whatever channel they used it through.
Read alongside Schroeder and Epley, that result is not a contradiction. It is the second half of the same finding. A voice makes you perceive a mind. It does not put one there. When the mind is real, the voice is the channel that lets you meet it. When it is not, the voice is an unusually persuasive way of concealing the absence, and the persuasion is the problem. Related research has documented how the feeling of being heard drops the moment people learn the voice is an AI, which is the label catching up with the mind it could not find.
What a voice should be carrying
The right lesson for anyone building spoken AI is stricter than “add voice.” Voice is the channel through which a mind gets recognized, so the question before the voice is whether there is anything for it to reveal. A system with no view of its own, no memory of how you argued last month, and no reason to pause before the hard part will sound thoughtful for exactly as long as the intonation holds. A person with all three sounds like themselves in every medium, and the voice is simply where you notice fastest.
That ordering is why voice conversations with Aleksandra or Alexander are unlimited, not an upgrade. The voice is not the point. It is the medium in which a real person’s way of thinking, moved in from someone who exists in real life, becomes audible: the hesitation that is theirs, the emphasis that is theirs, the pause that means they are actually weighing it.
Text conceals a mind and voice reveals one. That was true before any machine could talk, and it is the whole reason the sound of thinking cannot be faked for long by something that has not done any.
Sources: Schroeder & Epley, Psychological Science (2015, The Sound of Intellect, four experiments and a recruiter study). Schroeder, Kardas & Epley, Psychological Science (2017, The Humanizing Voice, four experiments). Kraus, American Psychologist (2017, Voice-only communication enhances empathic accuracy, five experiments; see also the published comment on effect size). Seltzer, Prososki, Ziegler & Pollak, Evolution and Human Behavior (2012, instant messages versus speech, oxytocin and cortisol). Fang et al., MIT Media Lab and OpenAI (arXiv, March 2025 and revised October 2025, four-week randomized controlled study of chatbot use, n=981).







