← Research
Research

AI frameworks compared

Offered as unique rather than best. On most of these questions somebody has published an answer with more testing behind it.

Last reviewed: 6 September 2026

Which framework answers which question, the peer-reviewed critique that applies to all of them including these, and the three measured findings underneath the lot.

Question this page answersAll 811 questions this research covers

Six frameworks for using AI well, each set against the established alternatives rather than presented on its own. The comparison is the point: on most of these questions somebody has already published an answer with more testing behind it, and knowing which one to reach for matters more than adopting any single set. Where the outside framework is better, these pages say so.

The answer, in one line

It depends on the decision in front of you. For the order of work, think AI think, or Dell'Acqua's centaur and cyborg patterns. For naming what you are asking a model for, the four levels, or Bloom's revised taxonomy for a curriculum.

Share as a card

Definition#

A framework, here: a named, memorable structure for making a decision about AI use. Almost none of them, including these, has been tested against an alternative or against no framework at all, which is the first thing anyone choosing one should know.

Share this definition as a card

Which one answers which question#

Think, AI, think is the spine. Write your own position first, use the model, then decide what to keep. Everything else sits underneath it. Nearest published relatives: Dell'Acqua's centaur and cyborg patterns, and Mollick's seven approaches for students.

The four levels, Extract, Explore, Examine, Extend, name what you are asking for. Nearest relatives: Bloom's revised taxonomy, which describes the learner, and SAMR, which describes the technology. Bloom is the evidenced one.

Goal, context, friction, standard is how to write the request. Nearest relatives: Lo's CLEAR, Google's TCREI, CO-STAR. CLEAR is tighter and TCREI is easier to teach; the only thing this adds is a slot for what the model should hand back to you.

Keep, share, hand over decides what gets delegated at all. Nearest relatives: Sheridan and Verplank's ten levels of automation, and Parasuraman, Sheridan and Wickens across four processing stages. Those are far better specified and are what a safety engineer should use.

The five rungs, Chat, Project, Skill, Automation, Agent, describe the tooling rather than the skill. Nearest relatives: Anthropic's workflow and agent distinction, and SAE's levels of driving automation.

The source rule handles a reference a model gave you. Nearest relatives: Caulfield's SIFT and the CRAAP test. SIFT is better and that page says so; the one amendment is that a model requires you to check a source exists before checking whether it is any good.

The standard every framework on this page fails#

In 2016 Hamilton, Rosenberg and Akcaoglu reviewed SAMR, at that point one of the most widely taught models in educational technology, and found it largely absent from the peer-reviewed literature despite heavy practitioner adoption, with thin theoretical and foundational evidence. They named three faults: the absence of context, a rigid hierarchy implying that higher is better, and an emphasis on product over process.

Every framework listed above is vulnerable to at least two of those, and none of the six has been tested against an alternative. What is evidenced is the mechanisms they are built on, which is a weaker claim and the accurate one. A memorable structure is a teaching aid, and the reason to prefer one is that it fits the decision in front of you rather than that it has been shown to work.

What the evidence underneath them does support#

Three findings recur across these pages and each is measured rather than argued.

Interface beats policy. Nearly a thousand school students split between unrestricted GPT-4, a hints-only tutor and nothing: with the tool removed, the unrestricted group scored 17 per cent below students who never had it, while the tutor group kept most of its gain. Same model, different interaction.

The gap is invisible until you measure unaided. Among 78 novice programmers, both AI groups produced working code and looked identical on every measure taken, until the tool was cut off and unrestricted users failed at 77 per cent against 39 for a scaffolded group.

It happens fast. In randomised trials with 1,222 people, the withdrawal effect appeared after roughly ten minutes and showed up as reduced persistence rather than lost knowledge.

These are older than the pages that describe them#

All six are taught material, carried forward through successive versions of a deck and most recently delivered in September 2026. None has a dated first publication, and each page says so in its own words rather than implying a coinage. They are also due a revision: the five rungs in particular track product features that did not exist as named things three years ago, and a ladder pinned to a vendor's menu will need rewriting.

They are offered as unique rather than as best. Where an established framework does the job better, the page for that framework names it and recommends it.

Key sources

The argument these operationalise, drift versus design. Applied to a student, how to use AI at university. Applied to a junior, how juniors become senior. On why a framework is not a substitute for measuring, the capability audit. On how this research grades what it cites, how this research works.

About these frameworks#

The six frameworks described here are used by Rahim Hirji in teaching and in the Mastering AI deck. No claim of first use is made for any of them except the source rule, which was written on 6 September 2026 and is dated from that page. Bloom's taxonomy, SAMR, CLEAR, TCREI, CO-STAR, SIFT, the CRAAP test and the levels of automation all belong to the authors named on the individual pages. The judgement that a framework should be chosen by which decision it fits, and that most of these have less behind them than the mechanisms they rest on, is an interpretation by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), and is marked as an interpretation and not a finding.

Reference · SS-2026-185

Cite this page

Hirji, R. (2026). AI frameworks compared. The SuperSkills evidence base, SS-2026-185. https://thesuperskills.com/research/ai-frameworks-compared. Last reviewed 6 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Which AI framework should I use?

It depends on the decision in front of you. For the order of work, think AI think, or Dell'Acqua's centaur and cyborg patterns. For naming what you are asking a model for, the four levels, or Bloom's revised taxonomy for a curriculum. For writing the request, Lo's CLEAR is tighter and Google's TCREI is easier to teach than goal-context-friction-standard, which adds only a slot for what the model should hand back to you. For what to delegate, the levels of automation literature is far better specified than a three-column sort. For checking a source, use Caulfield's SIFT.

Are AI frameworks like SAMR or CLEAR supported by evidence?

Mostly not, and that includes the six on this page. Hamilton, Rosenberg and Akcaoglu reviewed SAMR in TechTrends in 2016 and found it largely absent from the peer-reviewed literature despite heavy practitioner adoption, with thin theoretical and foundational evidence, naming three faults: absence of context, a rigid hierarchy implying higher is better, and product over process. Lo presents CLEAR as a framework proposal with no trial behind it. What is evidenced is the mechanisms these frameworks rest on, which is a weaker and more accurate claim.

What is actually proven about how to use AI well?

Three things recur and each is measured. Interface beats policy: among nearly a thousand school students, an unrestricted group scored 17 per cent below students who never had the tool once it was removed, while a hints-only tutor group kept most of its gain. The gap is invisible until you measure unaided: among 78 novice programmers both AI groups looked identical until the tool was cut off, at which point unrestricted users failed at 77 per cent against 39 for a scaffolded group. And it happens fast: in randomised trials with 1,222 people the withdrawal effect appeared after roughly ten minutes.

Are these frameworks original?

They are offered as unique rather than as best, and none of the five carried forward from teaching claims a first use, because no dated first publication exists for any of them. The source rule was written on 6 September 2026 and is dated from its own page. They are also due revision: the five rungs in particular track product features that did not exist as named things three years ago, and a ladder pinned to a vendor's menu will need rewriting.

In this hub

Frameworks for using AI well

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire