The Delegability Test: Five Criteria, Ten Minutes
Whether a task can be delegated to AI is not an opinion to debate in a meeting; it is a test to score in ten minutes. The test has five criteria, fits one sheet per task, and its result lands in one of three tiers: delegate, do together, never delegate. At a consulting table, in a team meeting or alone at a desk — the same five questions, the same ruler.
Below is the test itself: its preparation, the five criteria, the conversion of score into tier, and the rule that outranks all others — every verdict carries a date. The aim is not to end the debate but to lower it to task level, where debates produce decisions.
What Is the Problem?
BU BÖLÜMÜN ÖZETİ
- The fight happens at occupation level
- The decision is made by demo
- Verdicts are issued once and fossilise
Delegation decisions today are made by three bad methods.
The fight happens at occupation level
“Can accounting be delegated” polarises the moment it is asked, because an occupation is a weave of dozens of tasks and a weave takes no single answer. As the activity-occupation distinction shows, the correct unit is the task — the test is the instrument that lowers the debate to that unit.
The decision is made by demo
An impressive demonstration gets watched, and “let’s delegate this too” gets ruled on the spot. A demo shows the tool’s finest moment; the test scores the work’s nature. Without a ruler drawn independently of any tool, every demo persuades — with the ruler in place, a demo is merely one data point.
Verdicts are issued once and fossilise
“That work is never delegated here” gets said once and enters institutional memory as law. Yet the task length machines can finish alone doubles on a regular clock; an undated verdict stays in force even after it expires, and misleads the business in both directions — hoarding the delegable, freeing the undelegatable.
Why Does It Happen?
BU BÖLÜMÜN ÖZETİ
- Without criteria, volume wins
- Fear and fever fill the same vacuum
- The person who does the work is missing from the table
The testless habit has three roots.
Without criteria, volume wins
Absent a shared ruler, the delegation meeting belongs to the loudest voice or the most recently read article. Five written criteria convert voice into score — scores can be argued, voices cannot. The ruler’s worth lies less in its precision than in its commonality: when everyone looks at the same thing, the conversation shortens.
Fear and fever fill the same vacuum
Where no test exists, two emotions compete: the “we’ll lose our jobs” fear writes every task onto the never list, the “we’ll miss the era” fever writes every task onto the delegate list. Neither has looked at the task itself. The test does not outlaw emotion; it queues it — score first, commentary second.
The person who does the work is missing from the table
Delegation gets decided, in most places, in the absence of the person who performs the task; yet only they know its exceptions, its traps and its true hours. A delegation border drawn without the doer is a border drawn without the map — neat on paper, wrong on the ground.
How Is It Done?
BU BÖLÜMÜN ÖZETİ
- Preparation: reduce the task to one line
- Five criteria, scored 0-2
- Convert score to tier, stamp the date
Preparation is one line, the test five criteria, the result three tiers.
Preparation: reduce the task to one line
What gets tested is never a field like “reporting” but a single task: “drafting the weekly sales summary”. The line must show input and output — what is given, what is expected. A task that will not fit one line is not yet testable; it gets split first, because work that cannot be split is usually work that has not been understood.
Five criteria, scored 0-2
One — duration: how many hours of a competent person? Short scores 2, up to half a day 1, longer 0. Two — definedness: can input, output and “what good looks like” be written down? Three — verifiability: is checking the output distinctly faster than doing the work? Four — reversibility: if an error surfaces, can it be fixed harmlessly, or does it touch a customer, the ledger or the law? Five — judgment load: does the task contain signature, commitment or a value choice? If it does, it scores 0 — and that zero admits no bargaining.
Convert score to tier, stamp the date
A total of 8-10 is a delegation candidate: defined narrowly and handed over with an audit routine. 4-7 is do-together: machine drafts, human decides — an approval flow gets built. 0-3 is not delegated today; if the judgment criterion scored zero the verdict is standing, otherwise the task joins the “not yet” list. And at the bottom of every sheet: the test date plus the tool generation it was scored against. An undated verdict is void in this test.
How Long, at What Cost?
BU BÖLÜMÜN ÖZETİ
- Ten minutes per task, half a day per round
- Who fills it in: doer and owner, together
- The criteria travel; the thresholds localise
The test’s economics are deliberately small.
Ten minutes per task, half a day per round
Listing a desk’s weekly tasks and scoring the first twenty takes half a day and leaves three lists behind: delegate, together, never/not-yet. That half day replaces months of “AI strategy” meetings — the strategy is nothing more than the sum of the three lists.
Who fills it in: doer and owner, together
The person who does the work scores; the person who owns the work approves. Where their scores differ, the difference gets discussed — it usually springs from two different definitions of the task, and the discussion sharpens the definition. The test thereby yields a by-product: the work’s first written form.
The criteria travel; the thresholds localise
The five criteria hold in every industry; the score bands flex with the business’s error price. A cautious sector narrows the 8-10 band, a fast one widens it — common ruler, local calibration. For the duration criterion’s grounding, METR’s task-horizon measurement is the reference point.
The Common Mistake
BU BÖLÜMÜN ÖZETİ
- Opening the fifth criterion to negotiation
- Hiding behind the average
- Calibrating the test to the tool
The test runs; it gets punctured in three places.
Opening the fifth criterion to negotiation
Scoring a signature-bearing task “because the judgment is only slight” removes the test’s fuse. The judgment criterion works as a binary: present or absent. The judgment-bearing slice of the task is separated and stays on the human side; the remainder is retested — not bargaining, but slicing.
Hiding behind the average
A task scoring full marks on four criteria and zero on reversibility averages into “together” while in truth it is undelegatable: one irreversible error erases four conveniences. A single low score sets the tier, not the total — the ruler’s fine print.
Calibrating the test to the tool
“Our tool can handle this, so raise the score” surrenders the ruler to the demo. The test measures the work’s nature; the tool’s current skill only says when the “not yet” list gets re-scored.
Frequently Asked Questions
Sık Sorulan Sorular
The “never” list, when grounded in judgment, needs no repeat; the “not yet” list is re-scored every six months. As the horizon curve compounds, the not-yet list will melt — tracking the melt on a calendar protects the business from early fever and late arrival alike.
It can and it should; the error-price criterion especially moves with the client. That is the test’s power in consulting: the same ruler, filled with each client’s own reality, yields a personalised result — a reasoned decision instead of a template recommendation.
First locate the resistance’s address: an objection to the scoring, or a fear of unaudited handover? If the former, the scores get re-struck together; if the latter, the problem is not the test but the audit routine not yet built — and until it is built, the fear is right.
