The Success Measure That Makes a Business Case Falsifiable
Five success measures in the wording documents use, put through one question: what result would mean this did not work. Plus the table the survivors go into.
A project I worked on stopped in its ninth month. Nobody cancelled it. The sponsor changed roles, the successor arrived with a different list, and within a few weeks the work stopped being mentioned in status meetings. The business case had a section headed Success Measures. Three bullets, all approved by everyone who had to approve them. Not one of them could have come back false.
A success measure is a statement about a future observation, written so that at least one possible result counts as failure. A line that no result can contradict is decoration, and decoration is what most of these sections are made of.
There is one question that separates the two, and I ask it out loud, of every line, in the room where the document is being reviewed: what result would mean this did not work? Not what we hope to see. What we would have to see to admit the thing failed. Lines that get an answer stay. Lines that get a pause and a rewording of the same sentence come out.
In short
- A project without a measure that can fail does not get cancelled. It ends when attention moves somewhere else, and nobody has to say so.
- Three of the five measures below fail the test, and one of the three fails while carrying a percentage, which is how the appearance of rigour survives review.
- One of them is perfectly falsifiable and still does not belong in a business case, because it measures whether we shipped rather than whether it worked.
- A missing baseline calls for writing the measurement first and the target second, with a name and a date against the target, rather than for softer wording.
Why does the test ask for failure rather than success?
Because success never lacks witnesses. Any project that runs long enough produces a screen that exists, a launch date that happened and somebody willing to say the new thing is better than the old one. If your measure is satisfied by that, it was satisfied before you wrote it.
A measure earns its line by ruling something out. It has to name a world in which the money was wasted, in terms specific enough that a person who was not in the room can look at a report and say yes, that happened. The same property is what makes a non-functional requirement worth reviewing: a threshold with no way to miss it never gets built to.
And it explains why these sections sail through approval. A measure that can fail is a measure somebody will be asked about in eleven months, so it gets argued over now. A measure that cannot fail costs nobody anything, so it gets a nod. Read the silence in that review as the absence of anything at stake rather than as agreement.
What this example is built from
The returns case published here covers returns and complaint handling at a mid-size online retailer. The material is the published returns policy, the help centre articles, and ten public complaint threads I selected myself on a date I wrote down. There was no engagement and no client. I have never seen a queue export, an analytics account, a support rota or a cost line from that company, and I am not going to pretend otherwise in the middle of an article about honest measurement.
So every figure below is one of two things: a count of my own reading, labelled as one, or an assumption of mine with the source it would have to come from and the person who would confirm it written beside it. None of them is a fact about that retailer. That constraint is not a disclaimer bolted on at the end, it is the situation you are in every time you write a success measure before anybody has given you access, which is most of the time.
Five measures, one question each
These are the five as they arrive, in the wording documents actually use.
1. "Improve the customer experience around returns."
What result would mean this did not work?
There is not one. Any post-launch survey supports it, any single grateful customer supports it, and a flat result gets explained as too early to tell. There is no observation that this sentence forbids.
Out. Not because the intention is wrong, but because the intention is all it is. It can stay as the goal line above the measures. It cannot occupy a row in the table.
2. "Reduce operational costs by 15%."
What result would mean this did not work?
Costs come down 9%. That answer arrives quickly, and it is why this line survives most reviews. Then the follow-up: 15% of what, on which cost line, over which period, produced by whom? On this case nobody could name the report, because I have never seen one. If nobody can produce the number, nobody can produce the failing number either, and the measure is exactly as testable as the first one while looking like arithmetic.
Out as a measure, in as an open question, with an owner and a due date before build starts. When the answer arrives it comes back as a measure, carrying the name of whoever pulled the figure and the date they pulled it.
3. "Deliver the returns portal by the end of Q3."
What result would mean this did not work?
It ships in November. Instant answer, unambiguous, checkable by anyone with a calendar. This measure is falsifiable and it still does not belong in the success measures section, because it tests whether we did the thing, not whether the thing worked. A business case that counts delivery as success is finished on the day the code ships and can never be wrong after that.
Out of this section, on to the plan, where dates belong and where nobody mistakes them for value.
4. "Fewer damaged-delivery cases end up with nobody handling them."
What result would mean this did not work?
The share of damaged-delivery cases with no named owner after the first contact is the same or higher two months after launch than in the period we measured before it. Specific, produced by an export, and an answer somebody would have to live with.
Survives, and needs rewriting. As written it has no baseline, no target, no date, no owner and no counting rule.
5. "Reduce the number of support contacts per refund case."
What result would mean this did not work?
Contacts per case flat or higher after three months. Also a real answer. This one carries a second problem worth keeping visible: it can be met by making support harder to reach. A measure that can be satisfied by damaging the thing it stands for needs a companion that catches the cheap win, otherwise the honest team and the team gaming it file the same report.
Survives, paired with a guardrail.
Three of the five are out, and the one carrying a percentage was among them. That ratio is normal once the question gets asked of every line rather than only of the ones that already look weak.
The two survivors, written out
SM-1 Share of damaged-delivery cases with no named owner after
first contact.
Counting rule: cases opened under the damaged-delivery
reason in the month, divided into those that have a named
handler recorded against the case at the point of first
response and those that do not.
Baseline: not measured today. The nearest thing I have is
that four of the ten public complaint threads I selected
end with nobody named on the case after the first contact.
That is a count of ten threads I chose myself, it is not a
rate, and it says nothing about how often this happens in
the queue. The real baseline comes from two weeks of queue
export before build starts.
Target: at or below one in ten, held for two consecutive
months. This figure is my assumption and nothing else. It
is a placeholder until the two-week measurement exists,
and the number that ships is the one the returns queue
owner sets against that measurement.
This has failed if the share in either measured month is
at or above the baseline, or if the queue is still not
being measured at all two months after launch.
SM-2 Support contacts per refund case.
Counting rule: contacts logged against a case id, divided
by cases closed in the same month.
Baseline: unknown. No public source carries contact
counts, and I will not invent one. Same two-week export
as SM-1, same owner.
Target: below the measured baseline by month three,
by a margin the support lead sets when the baseline
exists. My placeholder margin is one fifth, marked as
mine, and it has no evidence behind it.
This has failed if contacts per case are at or above the
baseline in month three.
SM-2G Guardrail on SM-2: share of started refund cases
abandoned before a decision.
This exists because SM-2 can be met by hiding the contact
channel. If contacts fall and abandonment rises above its
own baseline, SM-2 has not been met, it has been gamed,
and the result is read as a failure regardless of what the
contact ratio says.
What makes these usable is the last sentence of each one. If a measure does not end with a line beginning "this has failed if", it has not finished being written.
The artifact: six columns
Every surviving measure gets a row. Six columns, because those are the six things people argue about eleven months later.
| Measure | Baseline | Target | Measured on | Measured by | Counts as failure |
|---|---|---|---|---|---|
| SM-1 Damaged-delivery cases with no named owner after first contact | Not measured. My own count of ten public threads I picked is not a rate and is not the baseline. Two weeks of queue export before build starts. | Author assumption: at or below 1 in 10, two consecutive months. Placeholder until the baseline exists; the figure that ships is the queue owner's. | Second and third full month after launch | Owner of the returns queue export. Slot left open: no public source names the role holder. | Share at or above baseline in either month, or the queue is still unmeasured in month two. |
| SM-2 Support contacts per refund case | Unknown, not published anywhere. Same two-week export. | Author assumption: one fifth below baseline by month three, confirmed or replaced by the support lead. | Monthly, months one to three after launch | Support lead. Slot open for the same reason. | Ratio at or above baseline in month three. |
| SM-2G Started cases abandoned before a decision | Unknown. Same export. | No rise above baseline. | Alongside SM-2, monthly | Same person as SM-2 | Abandonment above baseline while contacts fall. That result reads as a failed SM-2. |
And the same table empty, with what belongs in each cell.
| Measure | Baseline | Target | Measured on | Measured by | Counts as failure |
|---|---|---|---|---|---|
| [one countable thing, with the counting rule: what is in the numerator and what is in the denominator] | [the number today, with source and date, or "not measured" plus the export that will produce it and when] | [the number, who set it, on what evidence; if it is yours, write "author assumption, placeholder until baseline"] | [a date or a named month, not "after launch"] | [a person, not a team; leave the slot empty rather than filling it with a plausible job title] | [the observation that ends this as a failure, written so a stranger could confirm it from the same report] |
Two of those columns do most of the work and both are usually missing. "Measured by" turns a number into somebody's job. "Counts as failure" is what keeps the other five honest, because a target with no failure condition beside it drifts on its own.
What goes into the table when there is no baseline?
This is the situation nearly every guide skips, and it is the normal one. The current figure is not collected, or it is collected in a way nobody trusts, or it sits with a person who will not hand it over until the case is approved. The tempting fix is to soften the measure until the missing number stops mattering, which is how you end up back at improving the customer experience.
Do the opposite. Make the measurement itself the first deliverable, and write it with the same six columns: two weeks of export, a stated counting rule, a named person, a date it lands. The target line then reads as a placeholder with an owner attached, in plain words: set on this date, from the first measurement, by this person, and my figure until then is X, which is an assumption and carries no evidence. When the baseline comes back and it is nothing like your placeholder, that is information the case needed before anyone spent money, not an embarrassment to be quietly reworded. The one move that is never available is a target with no baseline and no note, because the first real measurement then turns into a negotiation about whether the number was ever meant literally, and that negotiation is won by whoever is closest to the deadline.
One more thing is worth saying out loud in this position. If the queue is not measured at all today, the first result of the project is that it can be, and that belongs in the business case as an outcome rather than in a technical annex. It is also the honest answer to the sponsor asking why you cannot simply put a number in: I can, it will be mine, and the document will say so on its face.
Where to start
Open the last business case you wrote and read only the success measures. Ask the one question of each line and write the answer beside it. In my experience a third survive intact, a third turn into open questions with an owner, and a third were goals wearing a metric's clothes.
Give the survivors the six columns. That is twenty minutes of typing. The hour after it is the conversation about who owns the baseline export and when it starts, and that conversation is the actual work, the same way the counted problem statement in eleven edits is worth more than the ten edits around it.
The standard I hold my own documents to, success measures included, is the scorecard: a fixed list of checks across the six artifacts, versioned, free, no email address asked for. A measure that can come back false is the only kind that can also come back true, and a project that can be shown to have worked is a project that gets a second release.
Read next
A Business Case a CFO Can Argue With
One page of business case, taken apart line by line: where each number came from, what halving it does to the decision,
Acceptance Criteria That Survive a Sprint Review
What an acceptance criterion actually is, six weak ones shown as they arrived and as they went out, and the three object
Get the Proof Pack
Six blank templates and one worked case. Free, one email. The scorecard is a separate file and needs none.
Send it to meAll field notes · The portfolio guide · More in The Document That Gets Read