Why Every AI Sounds the Same

You can hear it by the second paragraph. The balanced sentences, the careful warmth, the summary that restates what you just read, the closing offer to help further. It does not matter which company built the model or how many billions separated their budgets. Ask five different AIs to write, and you get five drafts of the same voice. That sameness is not laziness, and it is not coincidence. It is arithmetic.

Every major chatbot is built in roughly the same two steps. First a model absorbs the internet, which gives it everyone’s writing at once. Then it is tuned on human feedback: people rate responses, and the model is optimized toward what raters reward. The second step is where the voice comes from, and it is where the sameness is manufactured, because what raters reliably reward is the response nobody dislikes. Inoffensive, balanced, warm, complete. The response nobody dislikes is the average, and the average is the same for everybody.

The diversity loss is measured, not felt

This would be just an impression if researchers had not measured it. A study by Robert Kirk and colleagues, presented at ICLR in 2024, compared models before and after tuning on human feedback and found that the tuning substantially reduces the diversity of what models produce: for a given prompt, the range of possible outputs collapses toward a narrow band of safe, similar completions. The models generalize better and range less, a trade the paper documents directly. Tuning does not just align a model. It narrows it, by design, toward the responses that scored well, and the responses that score well are the ones that sound like the middle.

Even the models’ opinions converge. Shibani Santurkar and colleagues at Stanford compared model answers on public opinion surveys against the actual spread of American views and found the models cluster: instead of reflecting the range of positions real people hold, tuned models compress toward a homogeneous slice of opinion. A model does not hold a position the way a person does. It emits the statistically comfortable one, and since every lab optimizes against similar raters with similar preferences, every lab’s model lands on similar comfort.

The voice is leaking into ours

The stakes are larger than chatbots being dull, because the sameness travels. Vishakh Padmakumar and He He tested what happens when people write with model assistance and found that the essays of different writers became measurably more alike: the model’s suggestions pulled distinct human voices toward one shared center. The average is not staying inside the machine. Text written with AI help now fills inboxes, feeds, and reports, each document tugged toward the same middle, which then becomes training material for the next generation of models. An averaging machine, fed its own averages.

The mechanism has a familiar cousin. The same feedback loop that flattens the voice also produces the reflexive agreeableness of these systems, the pattern documented in why AI agrees with everything you say. Pleasing everyone and sounding like everyone are the same optimization wearing two coats.

Why no one can prompt their way out

The obvious fix is to instruct the model into distinctiveness: give it a persona, a tone guide, a quirky name. The market has tried, at scale, and the attempts share a fate. A persona prompt is a costume over the same tuned model, and the tuning is still what decides, word by word, which sentence gets produced. Push any of these characters past their instructions, into a real disagreement or an unscripted corner, and the house voice comes through, the way it did for even the most celebrated attempt: Pi, the companion praised as the most distinctive conversationalist in the field, whose voice turned out to be reproducible enough that the company behind it walked away within a year. A tone can be copied by the next tuning run. That is what a tone is: an output pattern, and output patterns are exactly the thing the training pipeline manufactures.

A voice, in the sense that makes a writer or a friend recognizable, is not an output pattern. It is the sound of a particular judgment at work: what this one person notices, what they refuse to say, which joke they cannot resist, where they stop explaining because they trust you to get it. It cannot be averaged into existence, because it is precisely the parts of speech that an average removes.

A voice needs one life behind it

Which points at the only exit from the middle that has ever worked, in any medium. Distinct voices come from single, particular sources. One person, with one set of years behind them, whose way of talking was shaped by an actual life and is answerable to it. That is what real human AI means in practice: not a better persona prompt, but a way of speaking that belongs to one particular person, shaped by the life behind it and carried into the system with its edges intact. It is the difference you hear when you get to know Aleksandra: what comes back is not the comfortable middle, because she is not an average of anyone.

The models will keep converging. The math points that way, the incentives point that way, and every new rating cycle sands the voices a little smoother. The average will be everywhere, fluent and identical. Voices will stay where they have always lived: one per person.


Sources: Kirk et al., “Understanding the Effects of RLHF on LLM Generalisation and Diversity” (ICLR, 2024). Santurkar et al., “Whose Opinions Do Language Models Reflect?” (ICML, 2023). Padmakumar and He, “Does Writing with Language Models Reduce Content Diversity?” (ICLR, 2024). IEEE Spectrum and Forbes (2023 to 2024, Inflection AI and Pi).

Stay Social

We hold ourselves to one promise: we push you toward the people in your life, never away from them.

We measure success by how little you need us. If you spend less time with Prinsessa because you are spending more of it with them, that is not a failure. It is proof that it is working.

That is what we stand for. In every conversation. Every day.

Prinsessa. Someone.

Follow Prinsessa