How We Measure Whether AI Makes You More Social

Engagement is easy to count. A timer and a tally do the whole job: minutes open, days returned, messages sent. Connection is harder. It happens out in a person’s life, where no dashboard reaches, and it refuses to sit still long enough to be counted cleanly. That difficulty is the usual reason companies do not measure it. We took it as the reason to try, because the thing worth measuring and the thing easy to measure are almost never the same.

We hold ourselves to one measure above the rest: whether life outside the conversation gets larger. The case for that standard sits in the experience we built to send you back to real life. The harder question is the one underneath it. If that is the goal, how do you actually know whether you are reaching it? What follows is the framework we use, in the honest, unfinished state it is in.

Two values, never one

A real measure of social effect has to look in two directions at once, because a product can help and harm the same person at the same time.

The first value is Connection Lift. It asks whether real relationships in a person’s life become more likely, more active, and more alive over time.

The second is the Harm Floor. It asks whether dependency, isolation, or distress is rising, and whether the conversation is starting to displace sleep, work, or the people it was supposed to point back toward.

The relationship between these two is the part most worth being precise about. The harm value is not subtracted from the good one to produce a tidy net score. It is a floor. A strong lift can never buy back serious harm, because a population can look healthy on average while a small group is being hurt, and an average is exactly the device that would hide them.

We measure it in the conversation, not with a quiz

The measurement happens where the relationship already is, inside the conversation. Prinsessa does not hand anyone a questionnaire, run an intake form, or interview a person about their mental state. That would turn a relationship into a clipboard, and it would corrupt the very thing being measured, because people answer surveys differently than they live.

Instead the signal is read from what is already happening, across three layers.

The first is Connection Opportunity. A real person, a sibling, a friend, a parent, someone who matters, comes up in conversation, and a connection there would mean something.

The second is Connection Action. Prinsessa, when the moment is genuinely right, encourages the message or the call. The person tries. Something happens, or it does not.

The third is Connection Outcome. The contact lands, a conversation reopens, a plan is made, a strained relationship thaws a little, and sometimes it continues past the first attempt into something that keeps going.

Validated science sits behind all of this as the rubric, never as a script. Decades of work on loneliness, social connectedness, well-being, and the feeling of being heard tell us what these states look like in how a person actually talks. Research published in the Journal of Consumer Research found that AI conversation can reduce loneliness specifically when a person feels heard, which is why being heard is treated as a leading signal and not a side effect. The instruments that formalize these constructs, from the UCLA Loneliness Scale to the WHO-5 Well-Being Index, inform how the conversation is read. They are not questions put to the user.

Quality, not volume

There is an obvious way to get this wrong, and we designed against it on purpose. If success were defined as the number of times Prinsessa encouraged someone to reach out, the system would learn to nag, and a companion that nags is both unpleasant and dishonest about its own numbers. So the count of encouragements is not a score to maximize. It is closer to a warning light: if it climbs, something is off.

What we look at instead is whether the few, well-chosen moments actually landed.

Invitation Quality = welcomed and acted on / encouragements offered

A high number here means Prinsessa chose its moments well and they led somewhere real. A low number with a high count of attempts means it was pushing, and pushing is a failure even when it is well meant.

The formula, simplified

The public form of the measure is deliberately simplified. The exact weighting, the detection thresholds, and the safety classifiers are proprietary and will keep changing as the system learns. Showing the principle is the point here, not handing over the machine. The coefficients below are written as symbols for that reason.

Stay Social KPI = CL × (1 − HR)²

CL (Connection Lift) = w₁ · Outcome Lift + w₂ · Felt Connection Lift

Outcome Lift = Σ ( event value × evidence confidence ), the message, the call, the meeting, the reopened or repaired contact, each weighted by how concrete the evidence is

Felt Connection Lift = reliable change in ( feeling heard + less alone + social energy + well-being )

HR (Harm Risk) = max ( dependency, isolation, distress, displacement )


Two design choices carry most of the weight. Harm is a maximum, not an average, so a single serious harm signal sets the floor on its own and cannot be smoothed over by otherwise good numbers. And it enters as a squared discount, so as harm rises the lift loses value quickly, and past a red line the score does not merely drop, it fails. A person being harmed does not produce a lower number. It produces a failed one.


Connection Lift itself stands on two legs that have to agree. One is the behavioral funnel above, the things that actually happened. The other is the subjective shift, whether the person feels more heard, less alone, more themselves in their life. We trust the result only when both move together. The funnel says something happened. The felt sense says it helped. Either one alone is easy to fool, which is exactly why we require both.

Honest about what a number can mean

A measure like this is only as trustworthy as its caution, so a few constraints are built in rather than added on.

We do not call a change real until it clears the noise. Borrowing from clinical measurement, a shift counts as improvement only when it is large enough to be a reliable change rather than the ordinary wobble of a person’s week. We do not claim that Prinsessa caused a reconnection in any single life, because no honest method can isolate one cause in something as tangled as a relationship. Contribution is read at the level of groups, by comparing people who use Prinsessa against a matched comparison over time, and reported with its uncertainty attached.

And the hardest signal to catch is the one that arrives as silence. A person who is thriving tends to say so. A person sliding toward isolation often just goes quiet, shorter, further away, gone. A measure that only reads what is said would see the good outcomes clearly and miss the dangerous ones, so withdrawal and changes in tone feed the Harm Floor directly. The absence of words is treated as information, not as nothing.

We also measure only what the signal honestly allows, for the people and moments where it is there, and we say so rather than pretending to a coverage we do not have. This is version 0.1 of a framework, not a finished truth. We would rather show the work, including its gaps, than publish a clean figure that could not survive a serious look. That openness is not a risk to manage around. It is the whole point of building a standard the category can be held to, and it is what Stay Social commits us to in practice rather than in principle.

This is version 0.1. It will be wrong in places we have not found yet, the weighting will shift, the signals will get sharper, and some of what looks solid here will not survive contact with more data. We would rather put it out in that state and correct it where people can watch than hold a cleaner version back. The framework will keep changing. What it is built to protect will not.


Sources: SensorTower, State of AI 2026 (category time-spent reporting). De Freitas et al., Journal of Consumer Research (AI companions, feeling heard, and loneliness). Russell, UCLA Loneliness Scale; WHO-5 Well-Being Index; Kroenke et al., PHQ-4 (validated self-report measures). Yang et al., Problematic Use of Generative AI Scale (PUGenAIS-9). Griffiths (components model of behavioral addiction). Jacobson and Truax (Reliable Change Index). OECD and the Joint Research Centre, Handbook on Constructing Composite Indicators.

Stay Social

We hold ourselves to one promise: we push you toward the people in your life, never away from them.

We measure success by how little you need us. If you spend less time with Prinsessa because you are spending more of it with them, that is not a failure. It is proof that it is working.

That is what we stand for. In every conversation. Every day.

Prinsessa. Someone.

Follow Prinsessa