Measuring an AI Investment: Four Numbers
“Is the AI working?” 📊 Most businesses answer this by feel: “I think we’re faster”, “I don’t think much changed”. And between those two positions there’s no referee.
Measurement ends that argument. But bad measurement is worse than none — because a wrong number gives a wrong decision a foundation that looks justified. ⚖️ Once a number is on the table nobody objects; questioning what it measures rarely occurs to anyone.
This guide covers doing it properly: which numbers matter, which mislead, when to measure and how to read a report. 📋
Our starting guide said “measure first”; here we set out exactly what to measure and how. Measurement is the cheapest and most frequently skipped part of any project. 📐
Why It’s Hard to Measure 🔍
BU BÖLÜMÜN ÖZETİ
- The gain is indirect
- The baseline is usually unknown
- The effect changes over time
Measuring an AI investment differs from measuring advertising. Three difficulties, and not knowing them leads to the wrong conclusion.
Knowing them determines the right method. 🧭
The gain is indirect
Advertising produces sales directly; AI produces time. ⏱️ Whether that time becomes revenue depends on it being redirected to other work — undirected, the gain stays on paper. So the second question in any measurement is: where did the saved time go?
The baseline is usually unknown
“How long did this task take before?” is a question most businesses can’t answer. 📉 Without a comparison point improvement can’t be demonstrated — and undemonstrated improvement is the first line cut when budgets tighten.
The effect changes over time
The early weeks are a slowdown period. 📈 Measure too early and the result looks negative; the real effect appears after month three, and whoever doesn’t know that gives up early. Explain the curve up front and the month-one dip becomes an expectation rather than a panic.
Four Numbers That Matter 📊
BU BÖLÜMÜN ÖZETİ
- Time: the most direct indicator
- Volume: the capacity equivalent
- Errors: the overlooked indicator
- Usage: the most honest indicator
Dozens of metrics can be produced but four drive decisions. The rest are decoration — and decorative metrics obscure the real number as the report fills up.
Measure all four monthly, the same way each time. ✅
| Number | What it shows | How to measure |
|---|---|---|
| Time | How much faster the work is | Minutes per task |
| Volume | How much work in the same time | Weekly count |
| Errors | Whether quality held | Number of corrections |
| Usage | Whether the system was adopted | People and sessions |
Time: the most direct indicator
How many minutes does the task take now? ⏱️ Review time must be included; leave it out and the gain looks larger than it is. A few days of sampling is enough; you don’t need to count every task.
Volume: the capacity equivalent
If time held steady but output rose, that’s also a gain. 📦 In content and customer communication especially, volume says more than time. But never read alone: if volume rises while quality falls, the net result is negative.
Errors: the overlooked indicator
Speed rose, but if errors rose too, the net gain fell. ⚠️ Untracked, this produces an illusion of improvement. It’s simple to measure: how many outputs were corrected, how many used as-is. That ratio should improve over time.
Usage: the most honest indicator
How many people, how many times? 👥 With low usage the other three numbers are meaningless — you can’t measure the performance of an unused system. It’s also the earliest signal of team resistance.
Misleading Numbers ⚠️
BU BÖLÜMÜN ÖZETİ
- Content produced
- Query volume
- Satisfaction scores
- Resolution rate alone
Four metrics look good but shouldn’t drive decisions. All are accurate; the problem is what they don’t show. They tend to surface when good news is required.
If a report is full of them, be careful. 🚩
Content produced
“80 pieces this month” tells you nothing. 📄 How many were published, how many worked? Count doesn’t substitute for quality. This metric exists to fill a report and usually covers another gap.
Query volume
How many questions went in? 🔢 A high number can be good or bad: a user who isn’t getting the right answer asks repeatedly and inflates the count. What matters is how many were resolved in one go.
Satisfaction scores
User surveys are distorted by politeness. 😊 Most people avoid giving a negative score; behaviour is more reliable than self-report. Rather than “are you satisfied”, look at whether they came back.
Resolution rate alone
A high chatbot resolution rate looks good but it may not be handing over what it should. 🔄 Always read it alongside total contact volume: if resolution rises while enquiries fall, customers are leaving.
When to Measure 📅
BU BÖLÜMÜN ÖZETİ
- Before starting
- Week four
- Month three
- Monthly thereafter
The measurement calendar determines whether results get read correctly. Measuring at the wrong time produces the wrong decision — early measurement especially can get a working system scrapped.
Four points. 🗓️
Before starting
One week recording the current state. 📋 Skip this and everything after is meaningless; retrospective “it used to be about…” estimates are always optimistic — memory favours the new method. A simple table for a week is more than sufficient.
Week four
The first comparison. 🎯 Keep expectations low: this measurement doesn’t tell you whether it worked, only whether to continue. The one thing to check is usage: if the team is using it, continue; if not, find out why.
Month three
Where the real effect shows. 📈 Adoption has settled and early corrections are done; the substantive decision should rest on this measurement. Widening scope also belongs here at the earliest.
Monthly thereafter
The same four numbers, the same way. 🔁 Change the method and comparison breaks — so the measurement approach should be written down at the start. Once a year, run an annual comparison: monthly noise disappears there and the real trend shows.
How to Read a Report 📖
BU BÖLÜMÜN ÖZETİ
- “Compared to what?”
- “What isn’t this number showing?”
- “What happens next month?”
- Let’s build the measurement together
A good report is one page and opens with four numbers. Long reports usually hide a weak result: as page count rises readership falls, and everyone knows it.
Ask three questions. 🔍
“Compared to what?”
A number without comparison is a sentence without context. Every figure should sit against the previous period and the baseline. 📊 “40% improvement” is meaningless without stating against what.
“What isn’t this number showing?”
The best question there is. 💭 Did time fall while errors rose? Did volume rise while quality fell? If the report doesn’t say, asking falls to you. A good report answers it before you ask and records negative movement too.
“What happens next month?”
A report is a decision document, not an archive. 🎯 If the action arising from the numbers isn’t written down, the report was read but not used. Every report should end with at most three actions.
Let’s build the measurement together
Setting up measurement takes a week but shapes the project’s entire life. It’s the cheapest and most valuable part: AI Consultancy. To review your current measurement setup: Digital Audit. 🚀
Frequently Asked Questions 💬
Sık Sorulan Sorular
Measure before starting. 🎯 Setting aside a single week produces the data that ends the argument for the project’s entire life. Not spending that week is buying months of uncertainty.
With four numbers: time, volume, errors and usage. All four measured monthly, the same way each time.
Three reasons: the gain is indirect, the baseline is usually unknown and the effect changes over time. The early weeks are a slowdown period.
Usage. With low usage the other three are meaningless — you can’t measure an unused system’s performance.
Four: output count, query volume, satisfaction scores and resolution rate alone. All accurate, none sufficient for a decision.
Because they’re distorted by politeness. Most people avoid negative scores; behaviour beats self-report.
Four points: before starting, week four, month three and monthly thereafter. Skip the baseline and the rest is meaningless.
After month three. The early weeks are a slowdown; measuring early looks negative and prompts premature abandonment.
Three questions: compared to what, what isn’t this showing and what happens next month?
Usually not. A good report is one page and opens with four numbers; length often hides a weak result.
