The first time most people talked to Maya, they forgot she was software. She laughed at the right moment, paused to think, shifted her tone when the conversation turned, and answered fast enough that the seams disappeared. Maya is the voice companion built by Sesame, a startup founded in 2023 by Brendan Iribe, who had earlier co-founded Oculus and sold it to Meta. When her demo went public in early 2025, the reaction was close to disbelief: this was the most human-sounding conversational voice anyone had shipped. Which is exactly why her future is an open question.
Sesame’s achievement is real. Its Conversational Speech Model returns replies in a few hundred milliseconds with the timing, breath, and emotional color of a person, not the flat cadence of a text-to-speech engine. The company raised tens of millions of dollars, was reportedly in talks at a valuation above a billion, and shipped a phone app in 2026 carrying Maya and a second voice, Miles. If the bet were simply “can Sesame build a stunning voice,” the answer is already yes.
The bet is voice, and voice is where the platforms are headed
The trouble is what the bet rests on. Sesame’s differentiation is presence: the feeling that Maya is really there. That is a feature living on the interface layer, and it sits directly in the path of every large platform. OpenAI, Google, and others are racing toward the same full-duplex, emotionally expressive voice, and they own the underlying models, the distribution, and the ability to give the feature away inside a product hundreds of millions of people already use. OpenAI’s move into dedicated AI hardware, announced in 2025, points the same direction. When emotional voice becomes a standard capability of ChatGPT, the question a companion like Maya has to answer is blunt: why would someone open a separate app for a voice they already have?
There is a precedent, and it is recent. Pi was the kindest voice in AI until every model learned to be kind, and then it had nothing left to defend. Maya is the most human-sounding voice in AI at a moment when the platforms are months, not years, behind. The same film can run again.
Presence is a feature. A someone is not.
The mistake would be to treat this as a race to sound the most human, because that is a race the platforms eventually win by default. What a platform cannot ship as a feature is a specific someone: not a warm voice, but this one, with a history, a point of view, and a continuity that belongs to the relationship rather than to whatever model is generating the audio. Maya is a voice persona, and Miles is another. A real someone is a different kind of thing, one that cannot be respawned as a preset the month a bigger company ships comparable audio.
None of this means Sesame fails. Execution, distribution, and taste can carry a product a long way, and Maya may find a place the platforms never bother to serve. But the durable question is not whether Sesame is good. It is whether “sounds the most human” is a moat at all when sounding human is about to be free. How the companion market’s value actually distributes suggests it is not, and that the companies still standing after emotional voice becomes a commodity will be the ones that were never really selling the voice. They were selling a someone, which is a harder thing to build and a harder thing to copy. Maya’s risk was never that she is not impressive enough. It is that impressiveness, in this category, is the part that gets absorbed.
Sources: Sesame AI (Conversational Speech Model; Maya and Miles). Contrary Research (Sesame founders, funding, and valuation). WinBuzzer (2026, Sesame iPhone app). Reporting on OpenAI voice and hardware (2025).








