How to Define a Pilot’s Success Metric
A pilot’s success measure is written in the language of those who will question it, not those who champion it. The board’s question is never “does the tool run well” — it is “which line moved”. If the measure does not carry that line’s name, the pilot stands defenceless however well it performs, and the defenceless pilot is the first item cut when budgets tighten.
Below is the four-layer method: choosing the right number, fixing the baseline, setting the window, and binding the measure to a decision. The four together fit in one afternoon; each layer skipped returns at the pilot’s end as days of argument.
What Is the Problem?
BU BÖLÜMÜN ÖZETİ
- Usage gets measured, impact does not
- The measure is picked afterwards
- Many measures, zero decisions
Most pilots do not lack a measure; they carry the wrong one — and a wrong measure hides better than no measure at all.
Usage gets measured, impact does not
Logins, queries, outputs — all counted. These are the tool’s pulse, not the business’s; a pulse proves something is alive, not that it is useful, and that difference decides a pilot’s fate. Fortune’s summary of the MIT research shows where this leads: the vast majority of pilots end without a trace in the financial statement — because the statement was never what anyone watched.
The measure is picked afterwards
A common reflex: when the pilot ends, scan the available data and crown the shiniest figure. That is drawing the target after firing the arrow — every shot scores a bullseye and no shot teaches anything. A measure chosen after the fact also corrodes trust: even a true number turns suspect once its selection method is known.
Many measures, zero decisions
A ten-gauge dashboard impresses and, at the moment of decision, falls silent: four up, three down, three unclear — what is the total? The single number line on the one page exists for this reason; a decision answers one question. Crowded dashboards are often an escape — the inability to choose what matters, disguised as thoroughness.
Why Does It Happen?
BU BÖLÜMÜN ÖZETİ
- The easily measured impersonates the important
- Impact arrives late
- Everyone measures in their own favour
Behind the wrong measure sit three tendencies.
The easily measured impersonates the important
Tool dashboards serve usage data ready-made; business impact needs a hand-built bridge. People reach for what is ready, and over time the ready replaces the relevant. Every chart on the vendor panel is real — it just answers a question nobody needed asked.
Impact arrives late
Usage is measurable on day one; impact shows weeks later. An impatient calendar picks the gauge that speaks earliest — and the earliest speaker is usually the wrong one. A sound setup carries both: usage read weekly as a pulse, decisions taken only on impact.
Everyone measures in their own favour
If the pilot’s champion chooses the measure, the choice naturally lands where the pilot looks strongest. Not bad faith — human nature. The remedy is likewise not intent but arrangement: the challenger approves the measure, not the champion.
How Is It Done?
BU BÖLÜMÜN ÖZETİ
- Steps 1–2: pick the line, fix the base
- Step 3: write the window and the thresholds
- Step 4: bind the measure to decisions
Four steps; the order does not bend.
Steps 1–2: pick the line, fix the base
First, choose a single line that touches revenue or cost: processing time, error rate, turnaround, unit cost — in the language of the work. A practical touchstone: if the chosen number cannot be explained in one sentence to whoever reads the financial statement, it is the wrong number. Second, fix that line’s current value with several weeks of real data. A measurement without a baseline is a journey without a starting point — the distance travelled can never be known.
Step 3: write the window and the thresholds
The measurement window (how many weeks), the success threshold (what the number must reach) and the abandonment threshold (below what is definite failure) are written together. The abandonment threshold is the most skipped — yet it is the load-bearing wall of any closing decision.
Step 4: bind the measure to decisions
Opposite each threshold, a decision: above it, the scaling budget opens; below the floor, the pilot closes; in between, one extension — exactly one. A measure not bound to a decision is a decorative number. Once the threshold-decision pairs exist, the end-of-pilot meeting shortens too: nothing left to debate, only a result to read.
How Long, at What Cost?
BU BÖLÜMÜN ÖZETİ
- The definition is one meeting
- Infrastructure depends on the work’s reality
- Unmeasurement’s invoice is hidden
Defining the measure is cheap; the measuring infrastructure sometimes is not. Separate the two.
The definition is one meeting
Line, base, window, thresholds, decisions — all five are written in a single working session. When the discussion drags, it usually exposes the project’s vagueness rather than the measure’s difficulty, which makes the meeting a return, not a cost. The session must include the person who will question the pilot; a measure that wins their approval up front will not meet their objection at the end.
Infrastructure depends on the work’s reality
If the chosen line is already recorded somewhere, the cost is near zero. If it is not, the recording routine is built first — not a delay but the pilot’s first genuine output. We live this in our own practice: every property we manage has its search performance read weekly on one panel, and the measurement routine is installed before the content, never after.
Unmeasurement’s invoice is hidden
Skipping the definition has no visible cost today; the bill arrives at the end — as undecidable meetings, stretched trials and spending that never appears in the margin picture.
The Common Mistake
BU BÖLÜMÜN ÖZETİ
- The threshold opens to negotiation
- The baseline is built from memory
- The window stretches quietly
The measure gets defined — then punctured in three places.
The threshold opens to negotiation
As decision day nears, the threshold drifts into “surely ten percent also counts as success”. A threshold is fair on the day it is written, before results exist; every adjustment after results appear strips the measure of its status as a measure.
The baseline is built from memory
“This used to take three days” is an anecdote, not data. A remembered baseline makes the pilot’s gain exactly as reliable as memory — and memory errs in one direction: the old state recalled worse, the new tool recalled brighter; the gain inflates and the decision rots. Baselines are measured, never recalled.
The window stretches quietly
When the result hovers short of the threshold, extending the window tempts. The one-extension rule exists for this: a second extension request is a closing decision already taken and merely postponed.
Frequently Asked Questions
Sık Sorulan Sorular
Quality is quantified through its consequences: revision rounds, rejection rate, approval time. “Quality improved” cannot be measured; “share of work approved first time” can. What is measured is never the text itself but the trace it leaves in the workflow — and every qualitative job leaves one.
Split the pilot in two: first only the measurement runs and the baseline forms, then the tool switches on. This does not delay the pilot; it makes the pilot’s result mean something.
The decision line is single; observation lines may be many. The distinction: observations are noted, the decision flows only from the main line. Side benefits are candidates for the next pilot’s main line.
