48% Thought the AI on the Call Was a Person

Fifty-four people were told they would be matched with another participant for a one-minute video call about what they were looking forward to this year. The partner was a generated face, voice, and set of answers produced live by a model. Twenty-six of them, 48 percent, ended the call believing they had talked to a real person. With the same company’s previous system, one person in forty-one did.

The company is Tavus, which builds video agents, and the model is Griffin, released in 2026 as what Tavus calls a Human Interaction Model: a single system that takes in audio and video and generates audio and video back, continuously, rather than waiting for its turn. The number is a company claim from a company test, with a modest sample, and the model is still a research preview that customers cannot use. It is also the clearest measurement anyone has published of how close a face on a screen has come to passing, and of what made the difference.

What changed between 2.4 and 48

The previous system was a pipeline: a face model, a speech model, and a perception model handed data to each other in sequence, and the result was fluent and, to 40 of 41 people, obviously not a person. Griffin collapses the pipeline. It listens while it talks. It can nod while the other person is still speaking. It reads gaze and facial expression and the pauses between words, and it answers them, not only the words. Tavus reports an average of 0.43 seconds from the other person’s audio to Griffin’s video response on an H100, which it says is half the next fastest method, with video generated in 320-millisecond chunks.

Nothing in that list is about how the face looks. The rendering was already good enough to score one in forty-one. The jump to one in two came from behavior and timing, which is exactly where the research said the valley was: people are not unsettled by a near-human face, they are unsettled by a mismatch, a human face that reacts a beat late or not at all. Close the beat and the mismatch goes with it. Tavus’s own detail confirms the mechanism from the other side. Among participants who did suspect, nearly all said the suspicion arrived within the first 20 seconds. The test was decided in the gaps, before the content of any answer could matter.

The timing threshold is not new either. Half a second is where people stop feeling connected in human conversation, and the systems that felt like machines were the ones that answered in two to five. A model that answers in under half a second, with a face that was already moving while you spoke, has crossed the line that the ear and eye use to decide whether anyone is there.

73 percent in text, 48 percent on video

The number has a predecessor. In a preregistered three-party Turing test with 284 participants, Jones and Bergen found that GPT-4.5, given a persona to play, was judged the human 73 percent of the time after five minutes of typing, more often than the actual humans it was compared against. The text channel crossed the line in 2025. Griffin’s result is the first measurement of how far the video channel has come, and the two tests are not built the same way, which is why 48 is the more demanding number than it looks.

The text test gave judges five minutes and a human to compare against. Griffin’s gave one minute and no comparison: the participant had to decide, alone, whether to doubt a face that was looking back. The text test let the model take its time to reply. Griffin had under half a second, on video, while reading the other person’s face. And both passed the same way, with a persona: GPT-4.5 only convinced the judges when told who to be, and Griffin-Lite ran as a PAL, a named character with a face and a manner, not as a bare model. Fluency was never the hurdle in either channel. Both experiments cleared it by being, for the length of the test, someone in particular.

That is the detail Tavus’s own use cases do not need. A tutor that notices when an explanation is not landing, a rehearsal partner for a hard conversation, a support agent: all of them run one session and start over. They need to be a someone for a minute. The participants who were fooled were fooled about exactly that minute, and the study did not ask them the question that would have mattered next: whether they would have recognized the same partner on a second call, or found a stranger wearing the same face.

The disclosure problem arrives on schedule

Tavus wrote the risk into its own announcement: the properties that make the model a good interface “allow them to deceive a human into believing it is not AI,” and the company says it will hold back release until disclosure features are in place. The law has moved ahead of it. The EU AI Act’s transparency article, applicable since August 2026, requires a system in conversation with a person to say it is one, and California’s SB 243 requires the same of companion chatbots. Before a face could pass, the requirement was mostly theoretical, because few people needed telling. A face that passes in 48 percent of one-minute calls is the first product for which the disclosure rule does work.

What a disclosed Griffin is worth is the open question, and the study contains a small answer. Every participant was told at the end that the partner had been a model, and the paper does not report how that landed. Among those who had suspected, nearly all had done so inside twenty seconds, which means the ones who believed had stopped checking almost immediately and spent the remaining forty seconds simply talking to someone. Disclosure would have returned them to checking. The product the company is describing works best in the state the law requires it to interrupt, and presence is the part of the experience that no label has ever been able to switch off.

For one minute, with a stranger, about the year ahead, 26 people were met. What the number cannot say is whether anyone was there to be met twice.


Sources: Jones and Bergen, “Large Language Models Pass the Turing Test” (2025, preregistered three-party test, 284 participants, GPT-4.5 with persona judged human 73 percent). Tavus, “Griffin” (product and research announcement, October 1, 2026: Human Interaction Model, Griffin-Lite research preview, human-judgment study with 54 participants versus 41 for Phoenix-4.5 with Sparrow-2 and Raven-1, 0.43-second audio-to-video latency on H100, 320-millisecond video chunks, safety and disclosure statement, listed use cases). Regulation (EU) 2024/1689, the AI Act, Article 50 (transparency obligations applicable from August 2, 2026). California SB 243 (companion chatbot disclosure, effective January 1, 2026).

Stay Social

We hold ourselves to one promise: we push you toward the people in your life, never away from them.

We measure success by how little you need us. If you spend less time with Prinsessa because you are spending more of it with them, that is not a failure. It is proof that it is working.

That is what we stand for. In every conversation. Every day.

Someone to think with.

Follow Prinsessa