For some decisions the question is whether the advice is any good. For others the question is whether the deciding was the point. Choosing a mortgage product belongs to the first kind: you want the best outcome and a model will consider options you would not have raised. Choosing whether to forgive someone belongs to the second: hand it over and the outcome may look identical while the thing you actually wanted has gone. Most personal decisions people ask about sit somewhere between, and working out which part you care about is the whole task.
The evidence below is about the first kind, because that is where the experiments have been run. The second kind has no experiments and probably cannot have any, which does not make it the softer question.
The advice moves you, and knowing it is a machine does not help#
In December 2022, two weeks after ChatGPT was released, Krügel, Ostermaier and Uhl asked it repeatedly whether it would be right to sacrifice one life to save five. It argued sometimes for and sometimes against, on the same question in different words, resetting the conversation each time.
They then ran a preregistered experiment with 1,851 US residents, of whom 767 passed both comprehension checks and form the analysis sample as preregistered. Each read a transcript of that advice before giving their own judgement on a trolley dilemma. The advice was attributed either to ChatGPT, introduced as an AI-powered chatbot, or to a human moral advisor.
Three results.
- The advice moved them. In both versions of the dilemma, participants found the sacrifice more or less acceptable depending on which way they had been advised. In the harder version, the advice flipped the majority verdict.
- Disclosure made almost no difference. The effect was statistically indistinguishable whether the source was named as a chatbot or as a person.
- They could not see it happening. Asked whether they would have made the same judgement without advice, 80 per cent said yes. Their judgements say otherwise. Asked the same about other participants, only 67 per cent thought so, and 79 per cent rated themselves more ethical than the others.
The authors' conclusion is blunter than most papers allow themselves: ChatGPT threatens to corrupt rather than promises to improve moral judgement. Their remedy is not disclosure, which they had just tested and found wanting, but the user's own ability to notice.
Hold the shape rather than the specifics. Advice from a source with no settled view still moved the view of the person reading it, and the person could not feel it move.
You are miscalibrated in both directions at once#
The obvious defence is that you would be more sceptical than an experimental participant. Two older literatures suggest scepticism is not the variable.
Logg, Minson and Moore, across six experiments on estimates and forecasts, found that people often weight algorithmic advice more heavily than advice from another human. Domain experts were the notable exception, and weighted it less. So before anything visibly goes wrong, the ordinary tendency is to over-trust rather than under-trust, and the people who resist are the ones who know the domain.
Dietvorst, Simmons and Massey, across five experiments, found the opposite failure on the other side of a single error. Once people have seen an algorithm make a mistake, they abandon it, even when it demonstrably outperforms them and even when their own record is worse. This is algorithm aversion, and one visible error is enough to trigger it.
Put them together and the pattern is over-trust until the first mistake you happen to notice, then under-trust afterwards, with neither position tracking whether the thing is any good at the task in front of you. What you actually feel, at any moment, is confidence. That is why a rule set in advance beats an assessment made in the moment.
Some decisions are constituted by the deciding#
Now the part with no experiment behind it.
A class of personal decisions has the property that the deciding is the product. An apology works because somebody sat with what they did. Forgiveness means something because the person forgiving weighed it and chose. A commitment is worth having because it was made by someone who understood what it would cost. In each case the outcome and the process are not separable, and an identical outcome produced by delegation is a different thing wearing the same clothes.
This research has made the argument in a neighbouring case, on whether to use AI for personal messages and in outsourced recognition: where the effort is the signal being sent, removing the effort removes the signal, whether or not anybody finds out. Decisions work the same way, with one difference. In a message the other person is the one deprived. In a decision, you are.
There is a quieter cost too. Deciding is how you find out what you think. A person who has outsourced twenty small decisions has not merely saved twenty small amounts of time; they have skipped twenty occasions on which they would have had to work out what they actually valued. This is the personal-scale version of what this research calls missed reps, and nobody has measured it, because measuring it would require knowing what a person would have become.
Three questions, and the second is the one people skip#
Before handing a decision over, work through these in order.
- Could you be wrong, and would you find out? A decision with a feedback loop is safe to delegate and cheap to correct. Which flight, which supplier, which of four phrasings. A decision whose consequences arrive in five years and are unattributable is one where bad advice is indistinguishable from good advice for as long as it matters.
- Is being the one who decided part of what you want? If yes, delegating gets you an outcome and loses the reason you wanted it. This is the question people skip, because it is not about the quality of the answer and so it does not feel like a real objection.
- Would you tell the people affected that a model chose? If you find yourself saying no, that reluctance is worth reading rather than overriding. That instinct is usually right, and usually points at question two.
A fourth move, for anything that passes all three: decide the rule before you ask. Write down what would change your mind, then ask, then check whether what changed your mind was on the list. That is the only practical defence against the Krügel finding, because you cannot notice the shift from inside the moment, but you can notice a gap between two written positions. The organisational version of the same idea is the Delegation Boundary Map.
How thin this evidence actually is#
The main study is a single sitting, one dilemma type, a model version from December 2022 and a 41 per cent comprehension pass rate. It reports test statistics rather than an effect size in the text, so the size of the shift is not something this page can state. It says nothing about repeated real-world decisions, about whether the influence persists past the session, or about whether people who use these tools daily are more or less susceptible than people encountering one in an experiment.
The advice-taking literature is older than the technology and was built on numerical estimates rather than personal choices. Applying it here is reasonable and is still an extension.
And the second half of this page, on decisions constituted by the deciding, is an argument rather than a finding. It is the kind of claim that would be very hard to test and is not therefore untrue. It is offered as reasoning you can check against your own case, which is the appropriate strength for it.
Key sources
- Krügel, S., Ostermaier, A. and Uhl, M. (2023). ChatGPT's inconsistent moral advice influences users' judgment. Scientific Reports, 13, 4569. Open access.
- Logg, J. M., Minson, J. A. and Moore, D. A. (2019). Algorithm Appreciation: People Prefer Algorithmic to Human Judgment. Organizational Behavior and Human Decision Processes, 151, 90-103.
- Dietvorst, B. J., Simmons, J. P. and Massey, C. (2015). Algorithm Aversion: People Erroneously Avoid Algorithms After Seeing Them Err. Journal of Experimental Psychology: General, 144(1).
Related SuperSkills research#
On the neighbouring everyday questions, personal messages, AI for therapy or advice and whether AI should remember everything about you. On the thinking underneath, what judgement is, decision quality and how to know when AI is wrong. On the boundary, the Delegation Boundary Map and letting an agent act on your behalf. On what is lost rather than what is risked, outsourced recognition and human at the start.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The Krügel paper was read in full at the primary source and every figure quoted here is from its text. No effect size is given for the shift in judgement, because the paper reports none in prose and the proportions exist only inside its figures. This page is not therapy, legal advice or financial advice, and if a decision you are weighing is affecting your wellbeing, a person you trust is a better first call than either a model or a website.
Cite this
Hirji, R. (2026). Should I let AI make personal decisions for me? The SuperSkills Intelligence Company. Last reviewed 31 August 2026. thesuperskills.com/research/should-i-let-ai-make-personal-decisions-for-me
