By compressing the analysis and leaving intact the two things a client is actually buying: a recommendation somebody is answerable for, and the judgement to know where the analysis stops being reliable. Consulting has better evidence about this than any other profession, for an accidental reason. The two strongest field experiments on generative AI in professional work were both run on consultants, and they point in different directions.
The experiment that gave the field its most useful concept
Dell'Acqua and colleagues ran a field experiment with 758 BCG consultants, assigning tasks deliberately placed inside and just outside GPT-4's competence. Inside, AI-assisted consultants were dramatically better and faster. Outside, they performed worse than consultants given no AI at all.
The concept the study named, the jagged frontier, is that model competence is uneven rather than smoothly graded by apparent difficulty. Tasks that look equivalent to a human sit on opposite sides of it. And the model's tone does not change when it crosses the line, so the confident output arrives identically whether it is right or wrong.
The study does not tell any consultant where the frontier runs in their own domain. That is local, it moves with each model release, and it has to be learned by being wrong. The paper's own limit is that it establishes the shape rather than the map.
One person with a model matched two people without one
The second experiment is the one with consequences for the organisation chart. Dell'Acqua and colleagues, with Lakhani, Sadun, Mollick and others, ran a pre-registered field experiment with 776 professionals at Procter and Gamble working on real product innovation problems. Participants were randomised twice: with or without AI, and working alone or in a two-person new-product-development team. It was published as an NBER working paper in April 2025 and in Organization Science in June 2026.
Three findings.
- Individuals with AI matched the performance of teams without it. The tool replicated some of what a second professional was contributing.
- Functional silos dissolved. Without AI, research and development staff proposed technical solutions and commercial staff proposed commercial ones. With AI, both produced balanced proposals regardless of their background.
- The interface did social work. Participants using AI reported more positive emotional responses, which the authors read as the tool filling part of the motivational role a human teammate plays.
The finding about silos deserves more attention than the productivity one. A consultancy's structure exists partly to assemble people whose different training produces different proposals, and then to reconcile them. If a model flattens that variation, the reconciliation was the value being added and it has just become cheaper. Whether the flattened output is better or merely more balanced is a separate question, and the study measures the second. This research has argued elsewhere that population-level convergence is the cost that individual-level improvement conceals, at does AI make everyone think alike.
The exposed part is the pricing, not the skill
Consulting sells leverage: a partner's judgement, delivered through a pyramid of analysts whose hours are billed. The Procter and Gamble result attacks the arithmetic of that pyramid directly. A firm charging for two people to do what one person and a model now do is running a pricing model its own clients can read the research about.
What the evidence does not support is the conclusion that the skill is obsolete. The jagged frontier result says the opposite: the consultants who did worst were the ones who trusted confident output on a task the model could not do. Detecting that requires knowing the domain well enough to feel the answer is wrong before being able to prove it. That capability was previously built by doing the analyst work.
Rahim Hirji made the commercial version of this argument in Entrepreneur UK in July 2026: AI does not create bad decisions, it exposes them faster, because a weak call that used to take weeks to surface now arrives by push notification. Applied to professional services, a firm whose recommendation quality rested on the analysis being slow and expensive finds that out quickly. A firm whose quality rested on judgement finds the judgement more valuable and considerably more visible.
Why the feedback loop protects consulting more than it protects medicine
Consulting scores high on the first condition of deskilling risk, because the tool substitutes for analytical judgement rather than for preparation. It scores low on the fourth, because being wrong in consulting becomes apparent fast and expensively. A recommendation that fails shows up in a client relationship within a year.
Compare a radiologist who misses a nodule and may never learn of it. Fast, painful feedback is an underrated form of protection against capability loss, and consulting has more of it than most professions. That is an argument for keeping the feedback loop rather than an argument for complacency: a firm that stops tracking which of its recommendations worked has removed its own best defence.
There is a self-inflicted risk too. Firms selling AI transformation to clients while running unmeasured pilots internally are describing capability they have not built, which is the pattern this research calls usage theatre.
What these two experiments do not settle
- Neither measured client outcomes. Both graded task performance under experimental conditions. Whether AI-assisted consulting produces better decisions for the organisations buying it is unmeasured.
- Both are single-firm studies. BCG consultants and Procter and Gamble professionals are not a representative sample of professional services, and both firms co-operated with researchers who had access because of that relationship. Procter and Gamble provided financial support to the institute involved, which the paper discloses.
- Neither followed anyone over time. A one-day experiment cannot detect whether the consultants who leaned on the model became less able to work without it.
- Self-assessment is unreliable in the opposite direction. METR randomised 16 experienced open-source developers across 246 real tasks and found them 19 per cent slower with AI permitted, while they estimated afterwards that it had made them about 20 per cent faster. Different work, but a standing warning about any productivity claim resting on what professionals report.
- The consulting labour market has not been measured separately. Claims about analyst hiring in the sector circulate widely; this page has no verified figure and does not offer one.
Six things a professional services firm can act on
- Map your own frontier, in writing. Which task types has the model been reliably right about in your practice, and which has it been confidently wrong about? Nobody can tell you; it has to be recorded case by case.
- Price the judgement, not the hours. The Procter and Gamble result is public. Clients can read it too.
- Keep a human view before the model's. Forming a position first is the only reliable protection against confident output on a task outside the frontier. See human at the start.
- Track which recommendations worked. The feedback loop is the profession's structural advantage and most firms let it decay.
- Decide deliberately what juniors still do by hand. If analyst work is the training ground, automating all of it is a decision about partners in 2036.
- Watch for flattening. If proposals from different functions are converging, some of what the firm sells has just been standardised, and the first person to notice should be inside the firm.
Key research and primary sources
- Dell'Acqua, F., McFowland, E., Mollick, E. et al. (2023). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Harvard Business School and BCG working paper.
- Dell'Acqua, F., Ayoubi, C., Lifshitz, H., Sadun, R., Mollick, E., Mollick, L., Han, Y., Goldman, J., Nair, H., Taub, S. and Lakhani, K. (2025). The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise. NBER Working Paper 33641; Organization Science, 2026.
- Model Evaluation and Threat Research (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity.
- Autor, D. and Thompson, N. (2025). Expertise. NBER Working Paper 33941.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful. Nature Human Behaviour, 8.
Related SuperSkills research
On the concept, the jagged frontier and human and AI collaboration. On the method, deskilling risk by profession, and on the neighbouring cases, medicine and law. On the internal risk, usage theatre and measuring adoption properly. On what stays valuable, staying valuable and decision quality.
About this research
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. Both field experiments were read at the primary source and the sample sizes and findings checked against them, including the funding disclosures. The consulting labour market figures that circulate in trade press are not used here because none could be verified against a primary dataset. Reviewed quarterly.
Cite this
Hirji, R. (2026). How will AI change consulting? The SuperSkills Intelligence Company. Last reviewed 28 August 2026. thesuperskills.com/research/how-will-ai-change-consulting