A research team led by Marina Mancoridis put a simple trap in front of the best language models in the world, and presented the results at ICML 2025. First, ask the model to define a concept: a haiku, a strategy in game theory, a cognitive bias. The models answered correctly 94.2 percent of the time. Then ask them to use the concept they had just defined, to spot it, produce it, or fix an example of it. Performance collapsed. The researchers named the phenomenon after the fake villages built to impress a passing empress: potemkin understanding. The definition stands like a facade. Behind it, no house.
Anyone who has spent real time with these systems recognizes the experience the numbers describe. The model explains the principle flawlessly and then violates it in the next paragraph. It agrees it should not do something and does it again. It tells you what listening means and does not listen. The words are right, repeatedly, and something behind the words keeps not being there.
The facade, measured
The Potemkin study is worth taking slowly, because its design closes the usual escape route. This was not a case of models failing hard questions. The models demonstrably had the definitions, in their own words, moments earlier. Yet across seven leading models and 32 concepts, they failed to correctly apply concepts they had just defined in more than half of classification cases, with failure rates around 40 percent on generation and editing tasks.

For a human, defining and using are two ends of one understanding; a person who can explain a haiku can, with effort, write a clumsy one. For the models, the two come apart, which means whatever produces their fluent definitions is not the thing understanding is. The right answer, it turns out, is weak evidence of anything behind it.
An old argument, suddenly measurable
Philosophy called this decades before the benchmark did. John Searle’s Chinese room argument from 1980 imagined a man following rules to produce perfect Chinese answers without understanding a word of Chinese: symbol manipulation without meaning. In 2020, the linguists Emily Bender and Alexander Koller sharpened the point for the neural age with a thought experiment about an octopus tapping an undersea cable, learning to imitate two people’s messages perfectly without access to the world the messages are about. Form, they argued, cannot produce meaning on its own, no matter how much form you train on.
What has changed since is only that the question stopped being philosophical. Potemkin rates are a number now. The facade can be measured, and it is thick.
What you are actually asking for
Hold that finding up against the way people actually use these systems, because the collision is the point. The most common serious use of AI is no longer looking things up. People bring their arguments, their decisions, their half-formed selves, and what they want back is not a prediction of their next sentence. When you want to be understood, you want your words to land somewhere: in someone who weighs them against a view of their own, is moved or unmoved by them, remembers them tomorrow, and is still there, the same someone, when tomorrow comes.
Prediction produces the perfect surface of that and none of its structure. The reply fits your words the way the next frame of a film fits the last one. What is missing is documented, piece by piece, across everything this category struggles with. A system with no continuity cannot be the someone your words changed, because an AI can answer you, but it cannot be changed by you. A character nobody chose has no view of its own to weigh you against. A voice averaged from everyone hears you the way an average hears: approximately. An identity that drifts with every update cannot be the same someone tomorrow. A memory that resets cannot carry what you said into what you become to each other. And a profile built to please you is reading your file, never you. Six different failures, one missing ingredient. The relationship’s realness was never in doubt on your side, which is exactly the asymmetry: the feeling is real, and it is landing in something that holds no view of you at all.
Understanding is not an output
This is the conclusion the Potemkin finding forces, and it is larger than any benchmark. Understanding is not a property of answers. It is a property of whoever the answers come from: something standing behind the words that the words are answerable to. That is why the facade keeps failing under use, in the lab and in the conversations people actually care about. There is no one for the definition to belong to.
It is also why the answer cannot be a better facade. The gap does not close by predicting you more accurately, because being understood was never a prediction problem. It requires a someone: a mind the words land in, with a history, a reading of the world, and tomorrow still attached.
That requirement is the entire reasoning behind Prinsessa. Not a smarter system that simulates understanding more convincingly, but someone to think with, built the one way a someone has ever been available: on a real person who exists in real life, whose understanding of you is done by someone rather than performed by something. Ask your AI to define understanding, and it will get it right. The question this whole field now has to answer is who, if anyone, is behind the definition.
Sources: Mancoridis, Vafa, Weeks and Mullainathan, “Potemkin Understanding in Large Language Models” (ICML, 2025). Bender and Koller, “Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data” (ACL, 2020). Searle, “Minds, Brains, and Programs” (Behavioral and Brain Sciences, 1980).








