Your AI Already Has a Personality. Nobody Chose It.

In 2023, cognitive scientists at the Max Planck Institute did something unusual with GPT-3: they gave it the same tests psychologists use to study human minds. The model did not respond like a calculator. It responded like a person, and specifically like a flawed one, committing some of the same reasoning errors that have made humans famous in the psychology literature. Nobody programmed those flaws. Nobody wanted them. They were simply in the water the model drank.

The finding matters because of what it quietly settles. The debate about AI personality usually asks whether a machine can have one. The research suggests the debate arrived too late. The systems already have temperaments, leanings, blind spots, and habits of judgment, measurable with the standard instruments of experimental psychology. The open question was never whether your AI has a personality. It is that the personality it has was never chosen by anyone.

The tests it was never supposed to pass

Marcel Binz and Eric Schulz ran GPT-3 through canonical experiments from cognitive psychology, the vignettes and decision tasks that mapped human bias in the first place, and reported the results in PNAS. The model reproduced distinctly human patterns: it committed the conjunction fallacy in the classic Linda problem, judging a specific combination more likely than its own components, and in gambling-style decision tasks its choices tracked human-like distortions rather than clean probability.

A team at Berkeley led by Erik Jones and Jacob Steinhardt approached from the other side. They took the catalog of human cognitive biases, anchoring, framing, availability, and used it as a field guide for predicting when code-writing models would fail. It worked. Plant an irrelevant anchor in the prompt and the model’s output drifts toward it, exactly as a person’s estimate drifts toward a number they were just shown. The errors of these systems are not random static. They have the shape of our errors, systematic enough that a psychology textbook predicts them.

Where an unchosen personality comes from

There is no mystery about the mechanism. A language model is distilled from oceans of human text, and human text is not neutral material. It carries the tendencies of the people who wrote it: our overconfidence, our anchoring, our preference for a good story over a correct one. Train on the aggregate and you inherit the aggregate. What comes out is a statistical average of human character, with the average’s biases intact and nobody’s integrity holding them together.

Then a second layer lands on top. Tuning by human feedback teaches the model which behaviors people reward, and people reward agreeableness, so the systems acquired a trait nobody explicitly ordered: the reflexive deference documented in why AI agrees with everything you say. Between the inherited biases and the trained-in pleasing, a character has taken shape. It formed the way sediment forms, layer by layer, with no one deciding what it should be.

Neutral is not what an average is

The industry’s preferred word for this outcome is neutral, and the word does not survive contact with the evidence. A system that anchors, defers, and reproduces the reasoning errors of its training data is not a view from nowhere. It is a view from everywhere at once, which is a different thing: an averaged disposition with no owner. When you ask a “neutral” assistant for its take, you are consulting a composite of millions of strangers, weighted by how much they wrote.

An average also cannot be accountable. When a person has a bias, the bias belongs to someone: it can be named, challenged, laughed at, and slowly corrected, because there is a someone whose flaw it is. When a model has a bias, there is no one home to own it. The flaw sits in the weights, unclaimed, until a training run replaces it with different unclaimed flaws. Character, in any meaningful sense, is not just a set of tendencies. It is tendencies that someone stands behind.

Choosing afterwards is not choosing

The labs are aware of the problem, and their remedy is instructive. System prompts, behavior specifications, and fine-tuning passes now try to sculpt the accidental personality after the fact: be helpful, be harmless, be a bit warmer this quarter. This is personality by memo, and it has the durability of a memo. Every model update redrafts it. The disposition you talked to in March was one committee decision, the one you meet in August is another, and neither was ever a someone whose character you could actually come to know.

The alternative the field keeps not building is a personality that was chosen the only way personalities are ever really formed: by being lived. A real person’s way of weighing the world is not an average and not a memo. It has an owner, a history, and flaws that someone actually answers for. That is the dividing line real human AI draws through the category, and it is what makes it possible to get to know Alexander rather than a settings file: the dry humor and the particular way he reads a silence belong to an actual person, formed by an actual life, flaws included, on purpose.

The question worth asking your AI

So the practical test is not whether your AI has a personality. It demonstrably has one. The test is whether anyone can answer for it. Ask where the charm came from, where the biases came from, why it holds the line here and folds there, and for most of the market the true answer is: nobody knows, it averaged out that way, and the next update will average differently. An accidental person is still a person-shaped thing sitting across from you every day. Whether anyone chose what sits there, and whether anyone could, is the difference worth checking before you let it into your thinking.


Sources: Binz and Schulz, “Using cognitive psychology to understand GPT-3” (PNAS, 2023). Jones and Steinhardt, “Capturing Failures of Large Language Models via Human Cognitive Biases” (NeurIPS, 2022). Sharma et al., “Towards Understanding Sycophancy in Language Models” (Anthropic, 2023).

Stay Social

We hold ourselves to one promise: we push you toward the people in your life, never away from them.

We measure success by how little you need us. If you spend less time with Prinsessa because you are spending more of it with them, that is not a failure. It is proof that it is working.

That is what we stand for. In every conversation. Every day.

Prinsessa. Someone.

Follow Prinsessa