By compressing the analysis and leaving intact the two things a client is actually buying: a recommendation somebody is answerable for, and the judgement to know where the analysis stops being reliable. Consulting has better evidence about this than any other profession, for an accidental reason. The two strongest field experiments on generative AI in professional work were both run on consultants, and they point in different directions.
The experiment that gave the field its most useful concept#
Dell'Acqua and colleagues ran a field experiment with 758 BCG consultants, assigning tasks deliberately placed inside and just outside GPT-4's competence. Inside, AI-assisted consultants were dramatically better and faster. Outside, they performed worse than consultants given no AI at all.
The concept the study named, the jagged frontier, is that model competence is uneven rather than smoothly graded by apparent difficulty. Tasks that look equivalent to a human sit on opposite sides of it. And the model's tone does not change when it crosses the line, so the confident output arrives identically whether it is right or wrong.
The study does not tell any consultant where the frontier runs in their own domain. That is local, it moves with each model release, and it has to be learned by being wrong. The paper's own limit is that it establishes the shape rather than the map.
One person with a model matched two people without one#
The second experiment is the one with consequences for the organisation chart. Dell'Acqua and colleagues, with Lakhani, Sadun, Mollick and others, ran a pre-registered field experiment with 776 professionals at Procter and Gamble working on real product innovation problems. Participants were randomised twice: with or without AI, and working alone or in a two-person new-product-development team. It was published as an NBER working paper in April 2025 and in Organization Science in June 2026.
Three findings.
- Individuals with AI matched the performance of teams without it. The tool replicated some of what a second professional was contributing.
- Functional silos dissolved. Without AI, research and development staff proposed technical solutions and commercial staff proposed commercial ones. With AI, both produced balanced proposals regardless of their background.
- The interface did social work. Participants using AI reported more positive emotional responses, which the authors read as the tool filling part of the motivational role a human teammate plays.
The finding about silos deserves more attention than the productivity one. A consultancy's structure exists partly to assemble people whose different training produces different proposals, and then to reconcile them. If a model flattens that variation, the reconciliation was the value being added and it has just become cheaper. Whether the flattened output is better or merely more balanced is a separate question, and the study measures the second. This research has argued elsewhere that population-level convergence is the cost that individual-level improvement conceals, at does AI make everyone think alike.
The exposed part is the pricing, not the skill#
Consulting sells leverage: a partner's judgement, delivered through a pyramid of analysts whose hours are billed. The Procter and Gamble result attacks the arithmetic of that pyramid directly. A firm charging for two people to do what one person and a model now do is running a pricing model its own clients can read the research about.
What the evidence does not support is the conclusion that the skill is obsolete. The jagged frontier result says the opposite: the consultants who did worst were the ones who trusted confident output on a task the model could not do. Detecting that requires knowing the domain well enough to feel the answer is wrong before being able to prove it. That capability was previously built by doing the analyst work.
Rahim Hirji made the commercial version of this argument in Entrepreneur UK in July 2026: AI does not create bad decisions, it exposes them faster, because a weak call that used to take weeks to surface now arrives by push notification. Applied to professional services, a firm whose recommendation quality rested on the analysis being slow and expensive finds that out quickly. A firm whose quality rested on judgement finds the judgement more valuable and considerably more visible.
The firms have now said this out loud, and prescribed the wrong remedy#
On 27 August 2026 the Financial Times reported that consulting firms are considering requiring junior staff into the office more often, because AI has raised the value of interpersonal skills. EY's UK head of consulting is quoted saying firms will have to reduce flexibility, but in order to help the human skills
, and that training in empathy, storytelling and leadership was dropped during the remote-working period while AI and technical skills were prioritised. KPMG describes reinventing in-person training. BCG is expanding office social activities. Deloitte and PwC began extra coaching for their youngest UK recruits in 2023 after finding weaker teamwork and communication than earlier cohorts.
Read carefully, that is the apprenticeship argument arriving in the trade press with named executives attached, which makes it the strongest external corroboration this research has. The consulting apprenticeship works by juniors watching seniors handle a client and then talking about it afterwards, and the firms are saying that pathway has thinned.
The diagnosis is right and the remedy does not follow from it. If juniors are weaker because AI absorbed the tasks that used to build judgement, then attendance does not repair it. A junior sitting in an office while a model still does the first draft has gained proximity and not repetitions. Presence is being prescribed for a problem of practice, and the two are easy to confuse because they were bundled together for a century.
Note also what the reporting does not contain. No measurement appears anywhere in it. The 2023 cohort effects at Deloitte and PwC are attributed to pandemic lockdowns rather than to AI, EY as a firm restated its existing flexibility policy alongside its executive's comments, and every speaker has an interest in the answer. It is strong evidence that large firms now believe this and are acting; it is not evidence of the mechanism. The mechanism is at missing rungs.
Why the feedback loop protects consulting more than it protects medicine#
Consulting scores high on the first condition of deskilling risk, because the tool substitutes for analytical judgement rather than for preparation. It scores low on the fourth, because being wrong in consulting becomes apparent fast and expensively. A recommendation that fails shows up in a client relationship within a year.
Compare a radiologist who misses a nodule and may never learn of it. Fast, painful feedback is an underrated form of protection against capability loss, and consulting has more of it than most professions. That is an argument for keeping the feedback loop rather than an argument for complacency: a firm that stops tracking which of its recommendations worked has removed its own best defence.
There is a self-inflicted risk too. Firms selling AI transformation to clients while running unmeasured pilots internally are describing capability they have not built, which is the pattern this research calls usage theatre.
What these two experiments do not settle#
- Neither measured client outcomes. Both graded task performance under experimental conditions. Whether AI-assisted consulting produces better decisions for the organisations buying it is unmeasured.
- Both are single-firm studies. BCG consultants and Procter and Gamble professionals are not a representative sample of professional services, and both firms co-operated with researchers who had access because of that relationship. Procter and Gamble provided financial support to the institute involved, which the paper discloses.
- Neither followed anyone over time. A one-day experiment cannot detect whether the consultants who leaned on the model became less able to work without it.
- Self-assessment is unreliable in the opposite direction. METR randomised 16 experienced open-source developers across 246 real tasks and found them 19 per cent slower with AI permitted, while they estimated afterwards that it had made them about 20 per cent faster. Different work, but a standing warning about any productivity claim resting on what professionals report.
- The consulting labour market has not been measured separately. Claims about analyst hiring in the sector circulate widely; this page has no verified figure and does not offer one.
Six things a professional services firm can act on#
- Map your own frontier, in writing. Which task types has the model been reliably right about in your practice, and which has it been confidently wrong about? Nobody can tell you; it has to be recorded case by case.
- Price the judgement, not the hours. The Procter and Gamble result is public. Clients can read it too.
- Keep a human view before the model's. Forming a position first is the only reliable protection against confident output on a task outside the frontier. See human at the start.
- Track which recommendations worked. The feedback loop is the profession's structural advantage and most firms let it decay.
- Decide deliberately what juniors still do by hand. If analyst work is the training ground, automating all of it is a decision about partners in 2036.
- Watch for flattening. If proposals from different functions are converging, some of what the firm sells has just been standardised, and the first person to notice should be inside the firm.
Key research and primary sources
- Dell'Acqua, F., McFowland, E., Mollick, E. et al. (2023). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Harvard Business School and BCG working paper.
- Dell'Acqua, F., Ayoubi, C., Lifshitz, H., Sadun, R., Mollick, E., Mollick, L., Han, Y., Goldman, J., Nair, H., Taub, S. and Lakhani, K. (2025). The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise. NBER Working Paper 33641; Organization Science, 2026.
- Model Evaluation and Threat Research (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity.
- Autor, D. and Thompson, N. (2025). Expertise. NBER Working Paper 33941.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful. Nature Human Behaviour, 8.
Related SuperSkills research#
On the concept, the jagged frontier and human and AI collaboration. On the method, deskilling risk by profession, and on the neighbouring cases, medicine and law. On the internal risk, usage theatre and measuring adoption properly. On what stays valuable, staying valuable and decision quality.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Both field experiments were read at the primary source and the sample sizes and findings checked against them, including the funding disclosures. The consulting labour market figures that circulate in trade press are not used here because none could be verified against a primary dataset.
Cite this
Hirji, R. (2026). How will AI change consulting? The SuperSkills Intelligence Company. Last reviewed 28 August 2026. thesuperskills.com/research/how-will-ai-change-consulting
