Home  ›  How to decide where AI belongs
The method

How to decide where AI belongs in a regulated process

The scoring method in full, published so you can judge it before you commission anything. Six factors, two modifiers, seven rungs.

The question the score answers

Every step in a validated process map receives an AIQ score from 0 to 100. It answers exactly one question: how amenable is this work to machine execution?

It deliberately does not answer "how risky would it be to change this?" Keeping those two apart is the single most important design decision in the method, and conflating them is the most common error in AI opportunity work.

The six factors

FactorWhat it measures
Volume and frequency: 20%Annual instances multiplied by handling time. The size of the prize, and the reason a tedious high-volume step usually beats a glamorous low-volume one.
Standardisation and rule clarity: 20%Whether the decision can be written down: rule-based, policy-based, or genuinely discretionary.
Input structure and data accessibility: 15%Structured data, semi-structured documents, free text, or information that exists only in a conversation.
Judgement intensity, inverse: 15%The degree of expert interpretation, negotiation or personal accountability the step carries.
Error, rework and exception rate: 15%Where quality is poor today the benefit is larger and, importantly, easier to evidence afterwards.
Systems accessibility: 15%API, database, portal-only, or a green screen. Feasibility, not desirability.

The two constraint modifiers

Neither changes the score. Both change what you are permitted to do with it.

Regulatory and risk sensitivity. Whether the step touches a regulated outcome, client money, personal data, a reportable obligation or a senior management responsibility. This drives the mandatory control pattern: human approval points, four-eyes requirements, sampling rates, explainability expectations, audit trail.

Change readiness. Data quality, system stability, team capacity, and whether a related programme is already in flight. This drives sequencing. It is why the highest-scoring opportunity is frequently not the one to do first.

High regulatory sensitivity does not reduce the opportunity. It changes the control pattern that has to sit around it. Folding risk into the score hides good candidates and produces a ranking nobody can interrogate.

The disposition ladder

The output is never a binary automate-or-do-not. Each step is placed on a seven-rung ladder, and the cheapest rungs come first.

RungWhat it means
0: LeaveHuman judgement, low volume, high accountability. Explicitly out of scope, and saying so is part of the value.
1: EliminateThe step adds no value. The highest-return intervention available, and the one AI enthusiasm routinely skips past.
2: StandardiseRemove variation before adding technology. A precondition for everything above it.
3: IntegrateThe work is rekeying between two systems. Fix the interface, not the human.
4: AutomateDeterministic rules, structured inputs. Conventional automation; no model required.
5: AugmentAI drafts, summarises, extracts or recommends; a named human decides and remains accountable.
6: DelegateAn agent executes end to end within defined boundaries, with human escalation, monitoring and a complete evidence trail.

In a firm of the size we work with, rungs 1 to 3 typically carry the majority of the realisable benefit and require no AI at all. We expect to tell clients that, and we would rather say it in week three than have them discover it after a build.

Then prove it

A score is a hypothesis, not a finding. Every high-ranked candidate gets a designed test: the smallest experiment that could disprove the benefit claim, with the success threshold agreed in writing before it runs.

In practice that is usually a concierge or Wizard-of-Oz test over sixty real historical cases, run in two weeks by one analyst at partial capacity, with no build. It settles the question for a fraction of what building the wrong thing costs. A test that kills a weak candidate has done its job.

Agreeing the threshold beforehand matters more than it sounds. An ambiguous result will otherwise be bent into whichever story the room already preferred.

Where this comes from

The Prove stage is drawn from The Lean Pivot, on evidence-led strategic change, written by PinpointProof's founder. The method in the book is the method inside the product. The goal is validated learning, not building.

FAQ

Common questions.

How do you decide which processes to apply AI to?

Score every step of a validated process map against consistent factors: volume, standardisation, input structure, judgement intensity, error and rework rate, and systems accessibility. Then apply constraint modifiers for regulatory sensitivity and change readiness. The score says how amenable the work is to machine execution; the modifiers say what you are permitted to do about it and in what order.

Why should regulatory sensitivity not reduce the score?

Because conflating "how amenable is this to a machine" with "how risky is it to change" is the most common error in AI opportunity work. A highly regulated step can be an excellent automation candidate. It simply requires a stronger control pattern around it. Folding risk into the score hides good opportunities and produces a ranking nobody can interrogate.

What is a disposition ladder?

A seven-rung scale from leave alone, through eliminate, standardise and integrate, to conventional automation, AI augmentation and finally delegation to an agent within defined boundaries. The cheapest rungs come first, because automating a process that should have been eliminated is the most expensive mistake available.

Should you automate a process before standardising it?

No. Automating variation makes the variation permanent and expensive. If a requirement matrix has three versions in circulation, standardising it is a prerequisite, not an alternative. It frequently delivers most of the available benefit on its own.

How do you prove the benefit before building?

Design the smallest experiment that could disprove the benefit claim, and agree the success threshold in writing before it runs. A concierge or Wizard-of-Oz test over real historical cases will usually settle in two weeks, for a fraction of a build, whether the claimed benefit is real.

Score one of your own processes.

Bring one process to a thirty-minute walkthrough and see the factors, the modifiers and the ladder applied to real work.

Thirty minutes. Bring one process.