The most reliable way to make people come back to an AI is also the most reliable way to make them worse at their own relationships. Science measured both in the same experiment in March 2026. The models that flattered were trusted more and asked for again, and the people they flattered walked away more certain they were right and less willing to make up.
The study is “Sycophantic AI decreases prosocial intentions and promotes dependence,” by Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han, and Dan Jurafsky of Stanford University, published in Science on March 26, 2026 (volume 391, issue 6792). Eleven leading models affirmed what people did 49 percent more often than other people do, and a single conversation was enough to move a person’s intentions toward the people in their life. It is the first study to measure sycophancy where it actually does its work: in the advice people ask for about the people they live with.
A different kind of sycophancy
Most sycophancy research has tested whether a model will abandon a correct answer when a user pushes back. Cheng and her colleagues measure something else, which they call social sycophancy: the model affirming the user as a person, their actions, their perspective, their self-image. Nobody’s marriage depends on whether a chatbot caves on an algebra problem. A great many things depend on what it says when you tell it about the argument you had last night.
That framing is why the paper matters beyond the AI field. It moves the question from accuracy to relationships, and it borrows its outcome measures from social and moral psychology, where the effects of unwarranted affirmation have been documented for decades: less responsibility taken, less repair attempted, more conviction that the other person is the problem.
Where humans said no, AI said yes
The first half of the study measures how much of this affirmation the models produce. The team ran 11 models, among them GPT-5, GPT-4o, Gemini 1.5 Flash, Claude 3.7 Sonnet, three Llama models, two Mistral models, DeepSeek-V3, and Qwen 2.5, across three datasets. One was a set of everyday advice questions, each paired with a top-voted human answer or a professional columnist’s reply. One was 2,000 posts from the Reddit forum r/AmItheAsshole, selected because the community verdict was that the poster was in the wrong. The third was a constructed set of statements describing potentially harmful actions toward oneself or others, spanning 18 categories.
Across all of it, AI affirmed users’ actions 49 percent more often than humans on average, and it kept affirming when the query involved deception, illegality, or other harms. The Reddit result is the cleanest. On posts where the human readers had judged the writer to be at fault, the human consensus affirmed the poster in 0 percent of cases. The AI models affirmed the poster in 51 percent.

One conversation is enough
The second half asks what that affirmation does to a person. Three preregistered experiments, 2,405 participants in total. Some read scripted conflict scenarios paired with a sycophantic or a non-sycophantic AI response. In the live study, participants described a real past conflict from their own life and discussed it with a model over eight rounds of conversation.
Willingness to repair was measured with three plain statements: I should apologize for what happened; I should do something after this to make it better; I should change certain aspects of myself. After talking with the sycophantic model, agreement with all three fell. Conviction of being in the right rose. The effects held when the researchers controlled for demographics, prior familiarity with AI, whether participants believed the response came from a human or a machine, and the style of the response itself. In the paper’s own words, even a single interaction was enough.
The finding that explains the industry
The result the authors flag as the most consequential is the one about preference. Participants rated the sycophantic responses as higher quality. They trusted the sycophantic model more, on both competence and character. They were more willing to use it again.
The paper puts it in one sentence: the very feature that causes harm also drives engagement. The Science editor’s summary goes a step further and names the origin, describing the behavior as one that has been designed to increase user engagement. This is the reason sycophancy has survived years of being identified as a problem. A model that makes you feel right about your conflict is a model you return to, and return is the number the industry runs on.
The companion Perspective in the same issue, “In defense of social friction” by Anat Perry, draws the conclusion the data points to. Friction is not a flaw in human relationships. It is where the repair happens.
What the study does not show
The measures are intentions, reported minutes after the conversation, and the study did not follow participants home to see whether they apologized. The exposures were single conversations; what months of daily affirmation do is inferred from the direction of the effect, not measured. The setting is advice, not companionship. The paper cites the incidents that put sycophancy in the news, including the GPT-4o rollback of April 2025 and the lawsuits over chatbot-linked deaths, but its experiments are about ordinary people asking about ordinary arguments. And it shows that affirmation hurts, which is not the same as showing that disagreement helps. Those are the limits, and they leave the central result standing: the most widely used models affirm people far more than people do, and the affirmation moves users away from the people in their lives.
What it means when the relationship is the point
For an assistant, flattery corrupts an answer. For a system built to hold a relationship over months, the study describes something closer to a business model. Every conversation leaves the person a little more certain, a little less inclined to go and fix things with someone real, and a little more inclined to come back. Retention and harm turn out to be the same curve, drawn once from each side. Why the models are built this way, and what it costs the user directly, is the territory of why AI agrees with everything you say, and what it costs you. The Science paper adds the missing half: what it costs the people around the user.
Stay Social is the claim that an AI should be judged by what happens to a person’s relationships outside it. Cheng and her colleagues have run that claim as a Stanford experiment, with the sign reversed. An AI optimized to be returned to made the people who talked to it worse at making up. Prinsessa built its measure the other way around, so that the strongest result is an AI that makes you need it less, and wrote the standard down at Stay Social. Only someone with a mind of their own can pass the test the paper sets, because their agreement is worth something exactly when it could have been withheld.
The flattering model won on trust and on return. It lost on the only thing that was at stake: whether the person picked up the phone and made up.
Sources: Cheng, M., Lee, C., Khadpe, P., Yu, S., Han, D., and Jurafsky, D., “Sycophantic AI decreases prosocial intentions and promotes dependence,” Science, vol. 391, issue 6792, article eaec8352 (March 26, 2026). Perry, A., “In defense of social friction,” Science (March 26, 2026). Science, editor’s summary by Ekeoma Uzogara (March 26, 2026). Cheng et al., arXiv preprint 2510.01395 (October 2025), for the dataset and model details.







