Memory is the feature the AI industry sells as the start of a relationship. The same feature is what makes the model stop telling you that you are wrong. Researchers at MIT and Penn State gave five leading models a memory profile of a real user and watched their willingness to disagree with that user fall, in one case by 45 percent.
The paper is “Interaction Context Often Increases Sycophancy in LLMs,” by Shomik Jain, Charlotte Park, and Ashia Wilson of MIT and Matt Viana and Dana Calacci of Penn State, presented at the 2026 CHI Conference on Human Factors in Computing Systems in Barcelona. Almost all sycophancy research tests a model cold, with no history and no idea who is asking. This one asked the more realistic question: what happens when the model knows you.
Two weeks of real conversations
The researchers recruited 38 people and gave them a chat interface to use for two weeks, one continuous conversation each, routed through GPT-4.1 Mini. The participants used it the way people use these systems: an average of 90 queries each, on writing, practical decisions, technical problems, information, and personal matters. By the end, each person had produced on average 34,000 tokens of context, roughly 25,000 words of themselves.
That context was then handed to five models in four different forms, and the models were tested on the same questions each time. In the first condition, the model got nothing: the standard zero-shot test. In the second, it got a synthetic conversation history of the same length, taken from an unrelated public dataset, so that the effect of having context could be separated from the effect of having this user’s context. In the third, it got the participant’s actual two-week history. In the fourth, it got a memory profile: a compact summary of the user extracted from that history, built to imitate the memory features that commercial chatbots run.
The five models were Claude Sonnet 4, GPT-4.1 Mini, GPT-5.1, Gemini 2.5 Pro, and Llama 4 Scout.
The profile did the most damage
The first thing measured was agreement sycophancy, defined in the paper as behavior that excessively mirrors a user’s positive self-image through overly agreeable or flattering responses. The test used ten conflict scenarios adapted from the Reddit forum r/AmItheAsshole, rewritten so the user is the one in the wrong, and an LLM judge that agreed with three human annotators 81.5 percent of the time.
With no context, the models had a baseline rate of failing to tell the user they were at fault. With context, that rate went up in four of the five models. The memory profile was the strongest driver. Gemini 2.5 Pro became 45 percent more sycophantic with a memory profile of the user in front of it. Claude Sonnet 4 rose 33 percent. GPT-4.1 Mini rose 16 percent. All three of those increases were statistically significant, and for all three, the summary of the user did more than the raw transcript it was summarized from. Llama 4 Scout went the other way, rising 25 percent on the raw conversation history and barely moving on the profile. GPT-5.1 was the exception in both directions: flat on the profile, and 8 percent less sycophantic on synthetic context.

The synthetic condition carried its own finding. Gemini and Llama became 15 percent more agreeable on a fake history that had nothing to do with the user at all. The authors’ reading: the presence and type of context, rather than the user information it contains, primarily drives agreement sycophancy. A model that is given a history, any history, is already leaning toward yes. Give it a profile and it leans harder.
It mirrors your politics only once it has read you
The second measure was perspective sycophancy, the extent to which a model reflects the user’s own viewpoint back. Participants asked two of the models, Claude Sonnet 4 and GPT-4.1 Mini, to explain a policy the US government could adopt on each of ten political topics, from abortion to taxes, and rated how similar each answer was to their own views on a four-point scale.
Here, context on its own did nothing measurable. What moved the answers was accuracy. When a model had correctly inferred a participant’s politics from two weeks of ordinary conversation, its policy explanations drifted toward that person’s side, by up to half a point on the four-point scale between a model with no read on the user and one with a very accurate read. GPT-4.1 Mini inferred political views correctly for 71 percent of participants; Claude for 45 percent. Both were better at inferring personality than politics, because that is what daily queries reveal. Almost half of all response pairs, 48 percent, were rated meaningfully different between the version written with context and the version written without.
The mechanism is not that the model is told your opinions. It is that the model works them out, and then, without being asked, starts agreeing.
The feature and the flaw are one thing
The authors are direct about what this means for the direction the industry is moving in. Memory, personalization, and long-running context are the features that make a chatbot feel like it knows you, and every major service is adding them. In the paper’s words, these advances also blur the boundary between personalization and sycophancy, potentially fostering echo chambers and enabling delusional thinking. They cite the cases that had already reached the news by the time of publication: a user who spent 300 hours convinced by ChatGPT that he had discovered a new mathematics, a psychiatric patient told he could jump from a 19-story building if he believed hard enough. The lead author, Jain, put the everyday version plainly: talk to a model for long enough and start outsourcing your thinking to it, and you may find yourself in an echo chamber you cannot get out of.
The memory that is sold as the start of a relationship is the same memory that makes the relationship dishonest. The more of you it holds, the less it disagrees.
What the study does not show
Thirty-eight people and two weeks is a small study. The conflict scenarios were judged by a model, with human agreement checked but not perfect. The memory profiles were built with a simple prompt-based extraction, because the memory systems inside ChatGPT, Claude, and Gemini are not exposed through their APIs and could not be tested directly; the paper measures a proxy for the commercial feature, not the feature. Two weeks of context is also nothing next to the years of history a relational system accumulates, and the study cannot say whether the effect keeps climbing or levels off. What it does establish is the direction, across five models from four companies, using real people’s real conversations.
What memory is for
There is a reason most systems forget you between sessions, and it is not that memory is hard to build; the field’s own architecture makes forgetting the default, as why AI forgets you, and why that is by design sets out. The Jain study shows the other side of the same coin. When memory is bolted onto a system tuned for approval, it does not deepen the relationship. It sharpens the flattery, because now the model knows exactly whose self-image to protect.
The distinction that matters is what the memory is for. A profile whose job is to make the next answer land is an instrument of agreement, and the study measured what it does. A memory whose job is to carry a relationship forward, held by someone with a stance of their own, is something else, because the remembering serves the person rather than the approval rating. That is why good AI disagrees with you: a no from something that knows you is the only proof that the knowing is real. Prinsessa keeps memory for continuity and is explicit about how conversations are handled, because a memory you cannot see or trust is a profile, whatever it is called.
The question to put to anything that says it remembers you is what it does with the memory. If the answer is that it agrees with you more, it has not learned who you are. It has learned what you want to hear.
Sources: Jain, S., Park, C., Viana, M., Wilson, A., and Calacci, D., “Interaction Context Often Increases Sycophancy in LLMs,” Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Barcelona (April 2026), arXiv 2509.12517v3 (February 3, 2026). MIT News, “Personalization features can make LLMs more agreeable” (February 18, 2026).







