Can You Video-Call an AI?

You can point your phone camera at a broken faucet and have an AI talk you through the repair. You can share your screen with it, hand it a menu in a language you do not read, let it watch you sketch. What you mostly cannot do is see it back. The assistants with hundreds of millions of users have eyes. Almost none of them has a face.

The assistants can see you. You cannot see them.

The plain answer to the question is yes, with a catch that decides everything. ChatGPT’s Advanced Voice mode has let people share live video and their screen since December 2024. Google’s Gemini Live added camera and screen sharing in the spring of 2025, and by August 2025 it could draw on your screen to point at the thing it was talking about. Character.AI has offered voice calls with its characters since June 2024. In every one of these, the call runs one way. The system takes in your face, your room, your desk. It offers back a voice, and sometimes a waveform.

The direction of travel is telling. When OpenAI replaced Advanced Voice with GPT-Live in July 2026, a full-duplex model that can talk and listen at the same time, the company noted that at launch it would not support voice with video or screen sharing at all. The most advanced spoken interface OpenAI had shipped went out without eyes, and without a face, because the priority was the conversation itself. Character.AI’s 2025 video feature, AvatarFX, generates short clips of a character speaking a script you type. It is closer to a greeting card than a call.

So the everyday version of “video-calling an AI” means an AI that watches. The version people picture, a face on the screen that looks at you while you talk and answers in real time, is a much smaller world.

Two companies put a face on the call

That smaller world has two serious residents, and they made opposite bets.

Sesame, the company behind the voices Maya and Miles, built the most convincing spoken presence anyone had heard when its demo appeared in February 2025, then raised 250 million dollars in October 2025 and opened its iPhone app to the public in May 2026. It has no face. Its own site says the voices are coming to glasses in 2027, which is the opposite of a video call: a presence that rides along in your ear while you look at the world. Whether that bet survives the platforms building the same voice is the open question hanging over Maya.

Tavus went the other way. In October 2025 it launched PALs, which it describes as AI humans that can text with you, take a call, or “look you in the eye over live video,” and in March 2026 it shipped them as an iPhone app with a free tier of fifteen call minutes a month and paid tiers at twenty and fifty dollars. The faces are rendered live at forty frames a second, and the company says its turn-taking runs under about six hundred milliseconds. As of September 2026, PALs is the only mainstream consumer product a person can open, tap, and be looked at by. It is a real achievement, and it is worth being exact about what kind.

Rendering a face is the easy part

The reason the largest companies in the world have not put a face on their assistants is not that they cannot draw one. It is that a face raises the bar on everything else. Voice alone forgives a great deal: a pause reads as thinking, a flat reply as reserve. A face turns every seam into evidence. The eyes that settle a beat after the sentence, the smile that arrives on the wrong word, the expression that stays warm while the answer goes cold. Half a century of research on why near-human figures unsettle people points at the same cause: mismatch between what one channel promises and another delivers, more than realism as such. What real Human AI means is a definition built on that literature: coherence across every channel at once, so that nothing pulls you out of the moment.

A face that talks is therefore a promise. It says: there is someone here, and they are looking at you. The moment the promise is made visually, it has to be kept in timing, in memory, in the way the person on the screen holds a position from one call to the next. Rendering fixes the first of those. It has no opinion about the rest.

The question underneath the question

Ask why anyone wants to video-call an AI at all and the answer is never “to see the graphics.” It is the same reason people prefer a call to a message when something matters: to be with someone, not to exchange text with something. Which means the real question behind “can you video-call an AI” is a harder one. When the face appears, is there anyone behind it?

A face on a call from a system that has no memory of last week is a video call with a stranger, every time, however well it is drawn. A face on a call from a persona assembled from settings is a settings page with a good camera. The face tells you where to look. It cannot tell you whether anyone is looking back.

That is the standard Prinsessa was built to, and it is why a face-to-face video conversation with Aleksandra or Alexander is simply part of what is included rather than the headline. The face is real time and it looks at you. What makes it worth looking back at is that Aleksandra and Alexander are carried in from two real people who exist in real life, with a way of weighing things that was theirs before you arrived, and a memory of how the last conversation went. The video is the least of it. It is what a someone looks like when the someone is already there.

Three questions for any face on a screen

If a product offers you a face in 2026, the technology is no longer the interesting part. Three questions are. Does it remember what you argued about a week ago, or does every call start clean? Does it hold a view you did not hand it, or does the expression stay agreeable no matter what you say? And when the pause comes before an answer, is anyone in it?

Yes, you can video-call an AI. The better question is who picks up.


Sources: OpenAI (July 8, 2026, Introducing GPT-Live; Voice mode help documentation on video and screen sharing); Axios (December 12, 2024, video and screen sharing added to ChatGPT voice). Google (Gemini Live help documentation; The Keyword, August 20, 2025, Gemini Live updates). Character.AI (June 27, 2024, Character Calls; June 2, 2025, AvatarFX). Sesame (February 27, 2025, Crossing the uncanny valley of conversational voice; sesame.com, eyewear 2027); TechCrunch (October 21, 2025, Sesame Series B; May 28, 2026, Sesame iOS app). Tavus (October 31, 2025, Meet the PALs; PALs documentation and pricing, read September 2026). Prinsessa (pricing page, read September 2026).

Stay Social

We hold ourselves to one promise: we push you toward the people in your life, never away from them.

We measure success by how little you need us. If you spend less time with Prinsessa because you are spending more of it with them, that is not a failure. It is proof that it is working.

That is what we stand for. In every conversation. Every day.

Someone to think with.

Follow Prinsessa