For most messages, yes, and the evidence points somewhere nobody expects. People who actually used algorithmic reply suggestions were rated by their conversation partners as more cooperative and produced more felt affiliation. The cost in the same study fell on people who were suspected of using them, whether they had or not. The penalty attaches to detection rather than to the practice. That reframes the question from ethics to concealment, and it leaves one small class of messages where the objection survives intact.
The experiment that separated using it from being seen to use it
Hohenstein and colleagues ran two randomised experiments on algorithmic response suggestions, the smart replies that sit under billions of messages a day. In the first, 438 crowdworkers were paired and asked to reach agreement on a policy question by live chat. Smart-reply availability was randomised separately for each person, so some pairs had it on both sides, some on one, some on neither.
Availability strongly encouraged use, and smart replies accounted for 14.3 per cent of messages sent. Conversations with the feature available ran 10.2 per cent faster in messages per minute. Greater use by a partner led the other person to write with more positive sentiment, an effect that held even when the smart-reply messages themselves were stripped out of the calculation.
Then the social result. Greater actual use by the partner improved the other person's rating of their cooperation (b = 15.66) and increased felt affiliation towards them (b = 21.79), with no effect on perceived dominance.
And the mirror image. The more a participant believed their partner had used smart replies, the less cooperative they rated them, the less affiliation they felt, and the more dominant they judged them, all at p below 0.0001, after controlling for the partner's actual use. The authors put it plainly: people who appear to be using smart replies pay an interpersonal toll even if they are not using them.
Suspicion is also close to useless as a detector. Beliefs about a partner's use correlated with actual use at Pearson's r = 0.22, which the authors describe as a correlation that exists but is not strong. So the toll is being levied largely at random.
A label costs something the writing had already earned
Yin, Jia and Wakslak found that AI-generated replies made recipients feel more heard than replies written by untrained humans, and that attaching an AI label removed the advantage. The words did the work; the disclosure undid it.
Two things follow, and only one of them is comfortable. The first is that the widespread assumption that machine-written warmth reads as hollow is not supported. The second is that people are not responding to the message. They are responding to a fact about its provenance, and they respond to it whether or not the provenance is real.
This is a finding about effects rather than a permission. Yin and colleagues measured what a label does to perception. They did not establish that concealment is defensible, and this research does not read them as having done so. Norms here are forming rather than settled, and the price of a label today tells you nothing about the price of having concealed one in five years.
Where the objection survives: when the effort is the content
Most messages carry information. A small class of messages carries something else, and for those the argument changes completely.
An apology transmits that you have thought about what you did, felt the discomfort of it, and chosen to sit with that discomfort long enough to say so. A condolence transmits that you stopped your day for someone else's grief. In both, the costliness is the message. A perfectly worded condolence that cost nothing has delivered the words and withheld the thing the words were standing in for.
Note that this objection does not depend on detection at all. The Hohenstein result says the interpersonal penalty only arrives if you are suspected. The apology case is different in kind: the content has already been removed at the moment of composition, and nobody needs to find out for that to be true. A page on this estate already makes the general version of the argument, at outsourced recognition, about what happens when the expression of noticing another person is delegated.
The workable test is not about words. Would this message still mean what it is meant to mean if the recipient knew exactly how it was made? A meeting summary passes. A thank-you for a dinner party mostly passes. A note to a bereaved friend does not.
A useful middle that the research does not cover
The studies above tested composition by machine against composition by hand. Most real use sits between: a person writes badly and asks for help with the phrasing, or writes something cruel at midnight and has it softened.
Nobody has measured that case. It is plausibly the most common one and it is absent from the evidence base, so this page states the gap rather than filling it with reasoning that would sound like a finding. What can be said is that the effort test still applies: help with expressing a thought you had is a different act from acquiring a thought you did not have. The line runs at whether the sentiment originated with you, and that line is set out for writing generally at human at the start.
Five things the evidence does not establish
- Nothing about long-term effects. The Hohenstein authors say directly that they have little insight into the longitudinal consequences, and raise the possibility of language becoming more homogeneous over time. That question is treated separately at does AI make everyone think alike.
- The suspicion finding is correlational. The authors state it does not show causally how attitudes shift in response to actual use.
- These were strangers on a crowdworking platform. Both experiments used Mechanical Turk participants discussing a policy question. Whether the same effects hold between a husband and wife, or a manager and a direct report, is untested.
- Nobody has tested the apology or condolence case experimentally. The argument in the section above is reasoning from what those messages are for, and it is presented as reasoning.
- Detection is a moving target. The r = 0.22 figure concerns short smart replies in 2023. Neither the writing nor the reader's calibration has stood still.
What to do with all this
- Use it freely for messages that carry information. Logistics, scheduling, updates, thanks for something transactional. The evidence gives no reason for guilt and one reason for the opposite.
- Write the hard ones yourself, badly. Apologies, condolences, and anything where the point is that you took the trouble. An awkward sentence you meant beats a polished one you did not.
- Apply the disclosure test before sending, not after. If knowing how it was made would change what it means, that is the answer, and it is available before anyone finds out.
- Stop trying to detect it in other people. On the only measurement available, you are barely better than chance and the suspicion costs the relationship something real.
- Notice which way your own drafts are moving. If the first instinct on receiving hard news is to open a model, the question has stopped being about the message.
Key research and primary sources
- Hohenstein, J., Kizilcec, R. F., DiFranzo, D. et al. (2023). Artificial intelligence in communication impacts language and social relationships. Scientific Reports, 13, 5487.
- Yin, Y., Jia, N. and Wakslak, C. J. (2024). AI can help people feel heard, but an AI label diminishes this impact. PNAS, 121(14).
- Jakesch, M., Bhat, A., Buschek, D., Zalmanson, L. and Naaman, M. (2023). Co-Writing with Opinionated Language Models Affects Users' Views. CHI 2023.
- Joshi, N. and Vogel, D. (2025). Writing with AI Lowers Psychological Ownership, but Longer Prompts Can Help.
Related SuperSkills research
On the delegated compliment, outsourced recognition. On writing generally, human at the start and keeping your own voice. On the population-level cost, does AI make everyone think alike. On the adjacent everyday questions, letting AI summarise what you read, whether you still need to remember things and AI for therapy or advice.
About this research
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. Both experiments were read at the primary source and every coefficient, correlation and percentage checked against the published text. The Hohenstein paper carries a publisher correction dated 3 October 2023; it added an omitted funding acknowledgement and changed no data or conclusion, which was verified rather than assumed. The apology and condolence argument is reasoning from what those messages are for, and is labelled as reasoning rather than evidence. Reviewed quarterly.
Cite this
Hirji, R. (2026). Should I use AI to write personal messages? The SuperSkills Intelligence Company. Last reviewed 28 August 2026. thesuperskills.com/research/should-i-use-ai-to-write-personal-messages