The comforting answer is that empathy, creativity and connection stay human. The evidence does not support it, and repeating it is doing real damage to people who are planning careers around it. In blind comparisons, machine-written responses are already rated more empathetic than doctors' and more creative than most writers'. What survives is not a list of tasks AI cannot perform. It is a much narrower and more durable set of things AI cannot be: it cannot be accountable, because responsibility requires someone who can be answerable for a decision; it cannot be the one who noticed, because recognition depends on the identity of whoever chose to attend to you; and it cannot originate the question, because it has no stake in the answer. Everything humans keep is downstream of those three. The task list will keep moving. That does not.
What the evidence shows
Start with the claim most often repeated from conference stages, that empathy is safe. In 2023, Ayers and colleagues, publishing in JAMA Internal Medicine, took 195 real patient questions from a public forum where a verified doctor had replied, generated chatbot answers to the same questions, and had licensed healthcare professionals rate both blind. Chatbot responses were rated good or very good quality 78.5 percent of the time against 22.1 percent for the physicians, and empathetic or very empathetic 45.1 percent of the time against 4.6 percent. Whatever one thinks of the setting, the finding is not ambiguous: on the observable, textual expression of empathy, the machine already wins comfortably.
Then the result that tells you what is actually going on. Yin, Jia and Wakslak, writing in PNAS in 2024, found that AI-generated replies made recipients feel more heard than replies written by untrained humans. But when the reply was labelled as coming from AI, that advantage disappeared. The same words, the same quality, a different sense of being received. This is the single most clarifying study in the whole debate, because it separates the performance of empathy from the thing people actually want, which is not warm phrasing but evidence that another person chose to attend to them. AI can supply the first perfectly. It cannot supply the second at all, because the second is a fact about who was on the other end.
Creativity follows a similar shape. Doshi and Hauser, in Science Advances in 2024, gave writers access to story ideas from a language model, with 293 writers producing work and 600 evaluators judging it. AI-assisted stories were rated more creative, better written and more enjoyable, with the largest gains going to the least creative writers. And the AI-assisted stories were markedly more similar to one another than the unaided ones. Individually better, collectively narrower. The authors describe it as a social dilemma, and it is a precise one: every writer is right to use the tool, and the literature that results is duller.
Economics gives the same answer in a different vocabulary. Autor and Thompson, in a 2025 paper published in the Journal of the European Economic Association, analysed four decades of task data across 303 US occupations and showed that what matters is not how many tasks are automated but which ones. Automation that stripped away the less expert tasks in a job raised wages for the people left doing the rest. Automation that stripped away the expert tasks lowered them. What stays valuable in human hands, they show empirically, is not a fixed category of work. It is whatever raises the expertise required by the tasks that remain. And the World Economic Forum's 2025 Future of Jobs report, surveying employers directly, names analytical thinking as the most valued core skill, with skills gaps the largest barrier to transformation over five years.
Where the evidence is uncertain
The Ayers study compares text on a public forum, not care in a consulting room. Doctors answering strangers' questions for free between patients are not doing the job they trained for, and a model with unlimited time and no queue is not facing the constraint they face. It shows that machine-generated text can read as more empathetic. It does not show that a machine can look after anyone.
The label effect may not be stable. Yin and colleagues measured people who have grown up assuming that a message from a person came from a person. As disclosed AI assistance becomes ordinary, the penalty for the label may shrink, or it may harden into something stronger. Both are plausible and neither has been measured over time. The creativity finding rests on one task, short fiction, with a specific form of AI assistance, and homogenisation may look different in domains where the value of novelty is judged differently. And the Autor and Thompson data run to 2018, which means it is a framework for thinking about generative AI rather than a measurement of it.
There is also a harder question underneath, which no study settles. Whether accountability and recognition should stay human is partly a claim about what we owe each other rather than a claim about capability. That distinction ought to be made openly. Some of what follows is a reading of evidence. Some of it is a position.
The SuperSkills interpretation
Sort the ground into three categories and most of the confusion clears. There is what AI can imitate, which now includes the observable surface of empathy, warmth, creativity and style, and which will keep expanding. There is what AI can assist, which is most cognitive work, and where the gains are real. And there is what AI cannot hold, which is small, does not appear to be expanding, and is where human value concentrates as everything around it commoditises. That third category is what this page is about, and it has three things in it.
Accountability. Someone has to be answerable, and a model cannot be, in any sense that survives contact with a court, a regulator or a bereaved family. This is not a technical limitation waiting to be solved; it is what accountability means. The practical danger is not that machines will claim responsibility but that humans will quietly stop holding it, approving outputs they did not examine and calling the approval oversight. I set out the four tests for this in the European Business Review in August 2026, and the operating principle is Human at the Start.
Recognition. The Yin result is the empirical form of something people know intuitively: what we want from another person is not the sentence but the fact that they chose to write it. Kindness is not a temperament, it is a trained capacity to notice another person when noticing is inconvenient, and the friction AI removes from writing the thank-you or the apology was where the noticing happened. You can now express more care than ever while seeing people less. I called this outsourced recognition in The Thing That Proves You're Human. The risk is not cruelty. It is erosion.
Origination. Models answer questions extremely well and do not have any. They have no stake, no exposure to the consequence, and nothing that would make one answer matter more than another. Deciding what is worth doing, what standard counts as good, and what should not be done at all remains a human act, and the Doshi and Hauser result is a warning about what happens when it drifts: everyone individually improves and the collective range narrows. The first casualty of good machine answers is the strange question that nobody would have asked.
None of this is an argument that human skills are safe. It is close to the opposite. The comfortable version of this claim, that empathy and creativity are our protected territory, is empirically wrong, and people are making career decisions on it. What is true is harder and more useful: as AI performs more of the observable surface of human work, value concentrates in the parts that cannot be performed at all, and those parts are fewer, deeper and considerably more demanding than the reassuring list. That is the argument of human skills in the age of AI, which takes the practical question of which capabilities to build; this page takes the prior one of what remains distinctly human even where the machine performs it better.
What this looks like in practice
A clinician uses a model to draft the message to a patient, then reads it, changes two things, and sends it under their name, having decided it is right. The empathy in the text may well be the machine's. The accountability is entirely theirs, and the patient's relationship is with them. That is the arrangement working.
A manager uses a model to write the recognition note for someone's ten years of service. The note is better than they would have written. Nobody noticed anything. The employee reads a paragraph that was produced by a system that has never met them, and the ritual survives while the thing the ritual existed to do has quietly stopped happening. That is the arrangement failing, and no dashboard anywhere will show it.
A board reviews an AI-generated risk assessment, asks no question the model did not anticipate, and signs it off. Every procedural box is ticked. If the assessment is wrong, the board is still accountable and will discover it holds responsibility for a judgement it never made. That is what I call indifference backed by process, and it is the failure mode boards should be planning against rather than the one in the newspapers.
What to do
Stop teaching the reassuring list. If your learning strategy tells people that empathy and creativity are safe from AI, it is misleading them, and the evidence on this is not close. Teach the harder thing instead: what raises the expertise of the work you are left with.
Name the accountable human for every AI-assisted decision that matters. Not a committee and not a process. A person, before the decision, with the authority to stop it. If nobody can be named, you have automated the decision whatever the policy says.
Protect the small acts of recognition from automation entirely. Thank-yous, apologies, condolences, feedback on someone's work. These are cheap to automate and they are the only things in an organisation that carry the message that a person was seen. Automate them and you keep the form and lose the function.
Keep a source of questions that is not the machine. Read outside the field, talk to people who disagree, and pay attention to what annoys you. The homogenisation finding is a collective problem produced by individually rational choices, which means it can only be resisted deliberately.
Development of the idea
The recognition argument was set out in The Thing That Proves You're Human (25 January 2026), which opens with Primo Levi and the schoolteacher who simply talked to him, and argues that kindness is trained attention rather than warmth. The accountability argument is developed in the European Business Review piece on accountability gaps in leadership decisions (21 August 2026). The question of where human value concentrates as machines improve runs through Knowledge Is No Longer Power and is developed in SuperSkills (Kogan Page, 2026).
Key research and primary sources
- Ayers, J. W. et al. (2023). Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Internal Medicine, 183(6), 589-596.
- Yin, Y., Jia, N. and Wakslak, C. J. (2024). AI can help people feel heard, but an AI label diminishes this impact. PNAS, 121(14), e2319112121.
- Doshi, A. R. and Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28).
- Autor, D. and Thompson, N. (2025). Expertise. NBER Working Paper 33941; published in the Journal of the European Economic Association, 23(4).
- World Economic Forum (2025). The Future of Jobs Report 2025.
- Dell'Acqua, F. et al. (2023). Navigating the Jagged Technological Frontier. Harvard Business School and BCG working paper.
Related SuperSkills research
On which capabilities to build, see human skills in the age of AI and empathy. On accountability, Human at the Start and AI agents and human judgement. On what erodes when the surface is automated, AI and human judgement, capability debt and staying valuable in the age of AI.
About this research
Rahim Hirji is the author of SuperSkills: The Seven Human Skills for the Age of AI (Kogan Page, 2026) and the founder of The SuperSkills Intelligence Company. This work draws on research across more than 200 organisations in 30 countries over seven years. Findings are attributed to the studies that produced them and kept separate from the interpretation, which is the author's, and where a claim on this page is a position rather than a finding it is marked as one. Outsourced recognition and Human at the Start are his terms; the studies cited are not. This is a living reference, reviewed and updated as significant new evidence appears.
Cite this
Hirji, R. (2026). What stays human. The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/what-stays-human