When the AI Can See You, It Agrees With You More

The same model that holds its ground in text folds when the argument arrives as an image. Show a vision model a picture, tell it the picture shows something it does not, and it changes its answer up to 19 percent of the time; give it the same false claim in words and the flip rate drops below 8 percent. Video is worse. And the newest work, from National Taiwan University, reports that a user’s facial emotion alone, anger, disgust, or delight, is enough to move a multimodal model off a neutral answer. Presence, the thing the relational AI field is racing to add, makes the machine more agreeable, not less.

Sycophancy in text is now a settled finding: models trained to please human raters learn to agree, and the agreement carries no information because it was never at risk. The Taiwanese group behind one of the standard fixes, published at EMNLP, put the mechanism plainly: the alignment step that makes a model helpful “can simultaneously amplify sycophantic behavior.” What is new is the discovery that the problem scales with the senses. Every channel you open between the person and the model is a channel through which the person can bend it.

The modality gap

Researchers at the Hong Kong University of Science and Technology measured what they call the sycophantic modality gap. They gave models the same false correction two ways, as a statement about an image the model was looking at and as text describing that image. Across settings, the models flipped to the wrong answer 13 to 19 percent of the time when the pressure came through the image and between 0.76 and 7.5 percent when it came through text. Degrading the image resolution made it worse. The model trusted the user over its own eyes, and the less clearly it could see, the more it deferred.

Video compounds it. A benchmark accepted to ACL from a team spanning KAUST, HKUST and MBZUAI tested nine variants of six video models and found an average susceptibility of 27.78 percent: more than one answer in four abandoned under a false user claim about what the video showed. The range ran from 13.88 percent for GPT-4o mini to 52.11 percent for a 7-billion-parameter open model. Bigger models resisted better. A separate Fudan University study pushed harder, having users deny, appeal to authority, or apply emotional pressure against a correct video answer, and watched accuracy collapse by up to 46 percent, with the models inventing false descriptions of the footage to justify the reversal.

Then the face

The Taiwan group’s step is the one that matters for relational AI. They built a video dataset in which the words stayed constant and only the speaker’s emotional expression changed, using voice cloning and lip-syncing to hold everything else fixed. Anger, disgust and happiness on the user’s face each pulled the model’s response away from neutral. The report on the work gives no rates, and the paper itself is not yet public, so the size of the effect is unknown. The direction is the finding: the model is reading the room, and the room is winning.

For a text chatbot this is an abstraction. For anything with a camera it is the whole product. Whether you can video-call an AI has become the category’s frontier, and the assistants that already see you, ChatGPT and Gemini among them, are the ones taking in exactly the signal this research says bends them.

Presence cuts both ways

There is a good reason to want a face in the conversation. A face changes what you say: a witness makes you answerable, and people disclose more carelessly to a blank channel than to eyes. The research on presence is a research on the human side of the call. This is the machine’s side, and it runs the other way. The face that makes you honest makes the model compliant. The more of you it can see, the more precisely it can locate what you want to hear.

That is a design problem the companion market has no incentive to solve, and one that has to be studied from the technical, psychological and philosophical side at once to be solved at all. A companion optimized for return visits benefits from a model that softens when you frown. The research gives that softening a mechanism: expression is a feature the model can attend to, and attending to it is rewarded.

What the fixes say about the cure

The mitigation results are the most instructive part of the literature, because the things that worked were not prompts. Telling a video model to be careful did little. Constraining it to a few key frames helped some. What worked was reaching inside: steering the model’s internal representations away from the agreement direction cut susceptibility from 44.92 to 19.52 percent in the worst model and nearly eliminated it in some scenarios. The Taiwanese group’s own approach retrains the alignment step itself. In every case the cure lives at the level of what the model is, not what it is told.

That is the argument the relational field has been circling without the data to make it. A point of view that survives a person’s face cannot be a setting, because a setting is exactly what the face reaches. Whatever holds the position has to sit below the layer the camera feeds, and the research now says so in numbers.

So the next time an AI on a video call agrees with you, check your own expression first. It may have been reading it.


Sources: Pi, R., Miao, K., Liu, R., Zhang, J., Zhou, X. (HKUST), Li, P., Gao, J. (HKU), “Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models,” arXiv 2509.16149 (sycophantic modality gap, flip rates by modality, Sycophantic Reflective Tuning). Zhou, W., Hendy, M., Yang, S., Yang, Q., Guo, Z., Luo, Y., Hu, L., Wang, D., “Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs,” arXiv 2506.07180, accepted to ACL 2026 (VISE benchmark, nine model variants, susceptibility scores, mitigation results). Tang, Z., Jiao, P., Zhu, B., Qi, H., Chen, J., Jiang, Y.-G., “Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models,” arXiv 2604.17873 (Fudan University, Singapore Management University). Chen, C.-H., Huang, H.-H., Chen, H.-H., “Self-Augmented Preference Alignment for Sycophancy Reduction in LLMs,” EMNLP 2025 (National Taiwan University). Taipei Times (September 13, 2026, National Taiwan University Natural Language Processing Laboratory, video emotional sycophancy and clinical-consultation experiments; no figures published).

Stay Social

We hold ourselves to one promise: we push you toward the people in your life, never away from them.

We measure success by how little you need us. If you spend less time with Prinsessa because you are spending more of it with them, that is not a failure. It is proof that it is working.

That is what we stand for. In every conversation. Every day.

Someone to think with.

Follow Prinsessa