Fog Is a Race
Between Two Clocks
Why better strategic search keeps making your option space bigger — and what actually shrinks it
The best strategy year in the firm's history produced fifty coherent futures and eliminated nothing. The process was sound. The criterion it was missing had never needed to exist before.
Scott Farrell · LeverageAI · August 2026
What a board leaves with
- ✓ A ratio you can measure this quarter — branching inputs against evidence outputs, one analyst, one week
- ✓ A six-step sequence that ends in a purchased observation rather than a conclusion, with reflexive counterplay made compulsory
- ✓ Two portfolios with different ledgers, and an eight-field option row whose last field is what makes killing affordable
- ✓ A four-part terminal-value decomposition you can underwrite — and the audit question that prices the term sitting in your portfolio
TL;DR
- •Fog is not weather. It is a ratio. The branching clock is how fast plausible strategic moves multiply market-wide. The evidence clock is how fast your firm can force reality to eliminate one. Fog thickens when the first outruns the second — which means the same environment is bewilderment for one firm and a sequence of cheap reversible experiments for another. Only one of the two clocks is yours.
- •The scarce capability is falsification throughput. Options eliminated per quarter, each carrying the evidence that killed it — not candidates generated, not decisions minuted. Almost no board can state that number, and finding that out is worth more than the number would have been.
- •A boundary case aims the search; only a probe closes it. Markets reprice when examined; a brick wall does not. So every case must run through counterplay to a smallest discriminating probe with kill and scale triggers — or it is mutation theatre: fifty coherent futures, nothing eliminated, and more fog than the firm started with.
The Year Nothing Died
A partnership ran the best strategic thinking in its history and finished the year less certain than it started. Nothing about the process was wrong. Something about the scoreboard was.
The offsite was, by every measure anyone in the room had, excellent.
Four practice areas, eleven partners, two days offsite, and an AI-assisted programme of strategic search running underneath it that would have been unaffordable three years earlier. They generated futures at a rate the partnership had never managed: agentic procurement arriving on the buyer's side of the table; a regulator requiring named human sign-off on work that had never carried a signature; a class of AI-native entrant that did not exist eighteen months before; three separate models of what their largest client might bring in-house and in what order.
The futures were good. That matters, and it needs saying before anything else in this book: they were plausible, internally consistent, argued rather than asserted, and several of them are visibly beginning to happen. Someone said "that changes how I think about this" and meant it. Someone photographed the whiteboard on the way out.
Twelve months later, the partnership could not name a single thing it had ruled out.
Not one future eliminated. Not one option closed. Not one assumption taken to a customer, a price, a competitor or a regulator and found to be false. The pricing sheet was unchanged. The hiring plan was unchanged. The graduate intake was unchanged. What had changed was the number of things eleven partners now had to hold in mind when they made any decision at all, which had roughly tripled.
They had spent a year making their own fog thicker, using the best strategic thinking they had ever done.
Here is the part worth sitting with, and it is the reason this book exists rather than a memo: they were not doing it wrong. They were doing exactly what the strategy profession taught them to do, with better instruments than the profession had when it wrote the lesson. The process was sound. The criterion it was missing had simply never needed to exist before.
What does the year look like as arithmetic?
Not as a judgement — as a count. Four quarters back, two columns, and no interpretation required.
The audit, four quarters (composite firm)
Candidate futures generated, clustered and argued
Scenarios carried into the board narrative
Options formally closed, with the evidence that closed them
Composite figures from a generic professional-services firm assembled for this book. They are illustrative of a pattern, not a client result.
Read the third number twice. It is not small. It is empty. And an empty denominator is not a rounding problem; it is a category problem. The firm had no mechanism whose job was to make a possibility go away.
Which produces an effect nobody budgets for. Cognition got cheap; attention did not. There are perhaps four people in that partnership who can actually change a decision, and their calendars did not triple when the option set did. Every future generated is a future someone now carries, and carrying has a cost — it just does not appear as a line item. It appears as hesitation. It appears as the pricing decision that gets deferred for a third consecutive quarter because there is now a coherent story in which either answer is wrong.
Key Insight
Generation got cheap. Elimination did not. Nobody moved the instrument.
Why did nobody notice?
Because the comparison that feels natural is us-today against us-last-year, and by that comparison the year was outstanding. More analysis. Better analysis. Faster analysis. Broader coverage of the possibility space than the partnership had ever achieved. Every instrument in the room was pointed at production, and production was up.
That comparison is the wrong one, and there is a specific structural reason it is wrong which the next chapter is entirely about. For now it is enough to notice that it is the only comparison the instruments can make. You cannot see a missing kill on a dashboard that counts outputs.
The year, read two ways
What the year produced
- • Fifty-plus argued candidate futures
- • A board narrative nobody was embarrassed by
- • Three genuinely new ways of describing the firm's exposure
- • Vocabulary the partnership now shares
- • Two days of the best conversation in years
What the year eliminated
- • Nothing
The asymmetry is the whole book. The left column is real value. It is simply not the same class of thing as the right column, and the firm's instruments could not tell them apart.
This is not one partnership's bad quarter
It is what happens to any firm that improves its generation rate without touching its elimination rate. And generation is precisely what got cheap.
I have been describing the felt version of this for a while:
You can't predict the future as well any more. The solution space expanded — what people are doing, what your competitors are doing, the products coming out. Everything accelerated, and on top of that you can't quite grasp what's going to happen.
And the sharper version, which lands like an insult until you check it: thinking is table stakes now. It is the floor. And if your business is the floor, you are out of business. That claim has already been narrowed properly elsewhere — it is generic, reproducible cognition that becomes the floor, not thinking as such, and the distinction matters enough that an entire book was written to make it . This book takes that as settled and asks the question underneath it.
If generation is cheap and the possibility space keeps growing, what exactly is the scarce thing?
What you get by the end
Four things, and they are the whole book:
- A ratio you can measure this quarter. Fog is not weather. It is the relationship between two rates, one of which is not yours and one of which almost entirely is.
- A sequence that ends in evidence rather than in a conclusion. Six steps from a boundary case to an option state, with the step most processes skip made compulsory — for a reason that has to do with the difference between a brick wall and a market.
- A portfolio whose rows close. Two portfolios, actually, with different ledgers, different sponsors and different burdens of proof — because a single portfolio always collapses into the one that has metrics.
- A terminal-value decomposition you can underwrite. Four terms instead of one alarming scalar, and a direct link between one of those terms and the portfolio, which is what turns this from a strategy practice into a board obligation.
"But our scenario work does change decisions"
Then you already own half the instrument, and the diagnosis will take a week to confirm it. Count four quarters of candidate futures generated, count four quarters of options formally closed with the evidence that closed them, and put the two numbers beside each other. If the right-hand column has entries in it, you are further along than most firms and the rest of this book is about making that rate deliberate rather than accidental.
If it is empty, the diagnosis is already complete and there is nothing more to measure.
The argument in one paragraph
AI makes variation cheap, not truth. AI Fog deepens when possible moves multiply faster than reality eliminates them. Boundary cases aim cognition at the assumptions that matter; adversarial counterplay exposes reflexive responses; paid and deployed probes let the world decide. Productivity creates enterprise value only when the company captures the cognition dividend in a commercial unit that survives cheap cognition. The board's job is not to predict the clearing of the Fog, but to operate an evidence-producing portfolio of options until the successor earns the right to be built.
Every clause in that paragraph is doing work, and most of them are contested. The rest of this book is the argument for each one.
It starts with the thing the partnership's instruments could not see. If the process was sound, and the thinking was good, and the futures were real — what precisely was missing?
Two Clocks, One Ratio
Fog is not a property of your market. It is the relationship between how fast possibilities arrive and how fast you can kill them — and only one of those two rates is yours.
Before the answer, the world in which the question never came up.
You watched the industry, you went to the events, you saw the releases coming. Make the team ten per cent bigger, aim for eleven or twelve per cent more revenue, and away you go.
That is not nostalgia. It is a description of a functioning method, and it worked for decades — which is worth being precise about, because the profession drew the wrong conclusion from its own success.
Planning was never clever. It was cheap. The rate at which genuinely new moves appeared in a market was slow enough that watching was sufficient. A competitor launched something once a year and you heard about it at a conference. A new entrant took eight years to become consequential, which gave you seven to respond. Software categories turned over on a cycle you could see from the far end. Under those conditions, the set of things that might happen stayed small on its own, and nobody needed a mechanism for making possibilities go away because the environment was not manufacturing many.
Strategy departments looked like they had a method. What they had was a slow environment, and a method calibrated to it.
The two rates
Change the environment's speed and you do not get the same problem faster. You get a different problem, with a different shape, and it has two terms rather than one.
Definition — the two clocks
The branching clock. How quickly new competitors, offers, architectures and strategic possibilities appear in your market. It runs market-wide. It does not need your permission, your budget cycle, your adoption curve or your change-management plan.
The evidence clock. How quickly your firm can put something into the world, observe the response, and eliminate or revise a possibility. It runs firm-local. It is almost entirely yours.
Fog thickens when branching outruns evidence.
That is the whole mechanism, and everything else in this book is machinery hanging off it. But the first consequence is the one worth slowing down for, because it moves the problem from a place where you are a spectator to a place where you are an operator.
Fog is relative
Take two firms in the same market, in the same quarter, facing precisely the same set of external unknowns. Same regulators. Same customers. Same set of entrants raising money down the road.
The first runs an annual planning cycle. Questions raised in March are answered — if that is the word — in the following February's board pack, by which time the market has added a dozen moves that were not in the frame when the question was asked. This firm experiences the condition as bewilderment. Its people are not less intelligent; their instrument simply reports at a frequency lower than the phenomenon it is measuring, which in any other domain we would immediately recognise as an aliasing problem.
The second firm can put a price, an offer or a scoped commitment into the world and read the response inside a quarter. It experiences the same environment as a sequence of cheap, reversible experiments. It does not know more about the future than the first firm. It knows more about the present, continuously, and it has learned to spend money converting uncertainty into observation.
Key Insight
The uncertainty outside is identical for both firms. The difference is how quickly each can make that uncertainty answerable.
This is why "the market is uncertain" is such an unhelpful thing for a board to conclude. It is true, it is shared with every competitor, and it implies nothing about what to do on Monday. The ratio implies quite a lot.
The scarce capability, named
If fog is a ratio, the strategic capability that matters is not the one every firm has been buying.
Definition — falsification throughput
The rate at which your firm eliminates strategic options, each elimination carrying the named evidence that killed it.
Options closed per quarter. Not candidates generated. Not initiatives launched. Not decisions minuted. Eliminations, with receipts.
The economic reason this is worth money, rather than merely virtuous, is a rule that holds everywhere: value migrates to the bottleneck, not to the activity producing the most volume. For most of the history of corporate strategy the bottleneck was generation — finding the options, doing the analysis, assembling the case. Firms organised around that bottleneck, staffed for it, and built instruments that counted it.
Then generation stopped being the bottleneck. The instruments did not move. So an entire profession is now optimising the abundant input and reporting it as progress, which is the exact behaviour that looks most like diligence and produces least.
Where this sits in what came before
The condition itself already has a name and a definition, and this book will not rebuild them. The AI Fog is the simultaneous compression of the credible forecast horizon and expansion of the plausible solution space: less time, more possibility, less clarity. It is not a synonym for uncertainty — uncertainty has a single direction, in that you cannot see as far, while the Fog has two, because you cannot see as far and there is more behind it to see . That is the diagnosis, it is taught properly elsewhere, and a reader who wants it should go there.
Half of it has reached the boardroom already. Planning horizons have measurably compressed.
The half everyone is managing
Of CEO planning effort now sits inside a horizon of less than one year1
The same figure one year earlier
Oliver Wyman's own caution is more interesting than the headline: the CEOs taking the longest view see the most opportunity, and compressed horizons, while understandable, may come at the cost of strategic clarity.
The standard response to horizon compression — shorter cycles, re-forecast often, commit late — is correct. If that were the only thing happening, it would also be sufficient.
The second term is the one that gets treated as weather, and the closest antecedent to everything in this chapter is the argument that it is nothing of the kind. Solution-space expansion has a source, and the source is the same cost collapse firms celebrate internally. Which produces an asymmetry that answers the question Chapter 1 closed on: your cheap cognition compounds inside your walls, at the speed of your change programme — gated by adoption, training, data access, the security review, the two teams that have not started. Everyone else's compounds across the whole market at once. No adoption curve applies to the aggregate. No internal resistance slows it .
Which means a firm can have the best internal AI adoption in its sector and still be losing ground — and the comparison that feels natural, us today versus us last year, is the wrong comparison and the only one most dashboards can make.
That is the prior book's contribution and it should be credited as such: the asymmetry is theirs. What this book adds is narrower and more operational. Make both terms countable. Then govern the ratio.
What the board is actually asking
Once the ratio is in view, the standing question changes, and the change is grammatical before it is strategic.
That changes the strategic objective from: How do we understand all the possible futures? to: Which consequential uncertainty can we force the world to answer next?
The first question has no subject who acts and no point at which it is finished. The second has both. It names an agent, implies a budget, and comes with a deadline attached whether you write one or not.
Three sets of clocks, and why they must not be blurred
This is now the third pair of clocks in this body of work, and a reader who knows the others deserves the map. Say it once, clearly, and then never blur it again.
| Clocks | Altitude | Question they answer |
|---|---|---|
| Production, authority, evidence | Inside one engagement | What is this delivery actually waiting on? |
| Runway and proof | The firm in transition | Can we prove a successor before harvest cash runs out? |
| Branching and evidence | The market search | Which bets are worth carrying at all, and how fast can we kill the wrong ones? |
The first two pairs belong to the professional-services companion, which is disciplined about exactly this problem and says so in one line: same word, different altitude .
The search clocks sit above the firm clocks rather than beside them, and the relationship is worth stating precisely because it determines what each instrument is for. The branching-to-evidence ratio decides what the successor bets should even be. The runway-to-proof inequality decides whether the firm lives long enough to run them. A firm can be excellent at one and fatal at the other in either direction: solvent, busy, and betting on the wrong three things; or exactly right about the future and out of cash.
That companion book stops deliberately at this seam. Its management claim ends at the inequality, and it flags the gap in a parenthesis: running a portfolio of successor bets — how many, how governed, how killed — is a later book's machinery.
This is that machinery. And one sentence carries straight up from there without modification, because it is exactly right one altitude higher:
Remember
Activity does not move this clock. Evidence moves it.
"You can't measure the branching clock"
You can, from public information, in about a week, with an analyst. The next chapter does it. Then the chapter after that argues the harder and more useful half — that the evidence clock is not a fixed property of your firm, that it has become physically faster in a way almost nobody has claimed at strategy altitude, and what it costs to move.
One clock you count. The other you buy.
The Clock You Don't Control
Four inputs, one tally sheet, one analyst, one week. The branching clock is countable — and the counting matters more than any individual number in it.
I have been reciting the collection method for two years without noticing I was doing it:
If this keeps going — more entrants, more software, more releases, more AI use, cheaper and smarter — what does my business look like then? If you haven't got a good answer, two things are true: your fog is thicker, and your business is more likely heading to zero.
Read the first half of that again as a list rather than as a lament. Entrants. Software. Releases. Adoption. Price and capability of the underlying models. That is the input schedule. It was spoken as a feeling, and it is a tally sheet with the columns already named.
So take the four columns seriously, and write down what qualifies for each — because the discipline that makes this useful is mostly the discipline of what you refuse to count.
The four inputs
1. Entrant cadence
Counts: genuinely new firms selling into your buyers; funded entrants who were not previously in your set; incumbents crossing in from an adjacent category with a real offer.
Collect from: funding announcements, your own lost-deal notes, buyer conversations, category directories.
Does not count: rebrands, a further round by someone already in your set, a competitor's press release about a partnership.
2. Offer and pricing-model cadence
Counts: new commercial units offered in your market. An outcome price where there was a day rate. A subscription where there was a project. A free tier under a floor you thought was structural. A guarantee nobody previously underwrote.
Collect from: competitor pricing pages, RFP responses you lose, what buyers ask you to match.
Does not count: discounting. A discount is a price move inside the same unit; a new unit changes what is being bought.
3. Capability cadence
Counts: what became purchasable that was not. A model release matters here only if it changes what a competitor or a customer can now do without you.
Collect from: what your own delivery teams stopped doing by hand; what clients now arrive holding.
Does not count: benchmark improvements with no route into your market's work. Most releases are not events; a few are.
4. Regulatory cadence
Counts: moves opened or closed by rule change — including guidance and a visible shift in enforcement posture, which changes behaviour long before the rule does.
Collect from: regulator publications, professional-body guidance, your clients' compliance teams.
Does not count: a consultation with no legal effect and no observable behavioural change.
Is the branching rate actually moving?
Fair question, and the answer should be sourced rather than asserted. Here is what the public record shows — each figure with its caveat attached, because a number without its caveat is decoration.
Capital formation, one quarter
Global venture investment in Q1 2026, across about 6,000 startups2
Of that going to AI companies — $242 billion
The same share one year earlier
The caveat belongs in the same breath as the number. Four rounds — OpenAI, Anthropic, xAI and Waymo — accounted for $188 billion, or 65% of the entire quarter. So this is evidence of funded branching and extreme concentration. It is not evidence of six thousand new threats to your business, and anyone who presents it that way is selling something.
What it does establish is that the capital that funds new commercial architectures is arriving at a rate the strategic planning calendar was never built to absorb. Money is not the same as a competitor. It is the rate at which competitors become possible, which is precisely what a branching clock measures.
On the capability side, the count is less dramatic and more useful. Epoch AI tracked 87 notable model releases from industry in 2025, against seven from all other sources — and industry now accounts for more than 90 per cent of notable models, up from just under half in 2015 and none at all in 20033. The count is interesting; the trend line is the point. Capability now arrives on a commercial release schedule rather than a research one, which means it arrives at the cadence of a product roadmap you do not see.
And some of it lands directly in your competitors' hands as a new thing they can sell. On one agentic benchmark, the success rate on real-world tasks moved from 20 per cent to 77.3 per cent inside a year4. That is a benchmark result, not a market outcome, and it should be labelled as one. But it is exactly the shape of thing that adds moves to somebody else's board: work that could not be sold as a reliable service last year, and can be argued for this year.
Then the two figures that show the branching arriving as commerce rather than as capability. Best-in-class AI-native companies are reported to be compressing the road to $100 million of recurring revenue from a five-to-ten-year gold standard into one to two years5. And on the buy side, 35 per cent of enterprises have already replaced at least one SaaS tool with a custom AI build, with a further 78 per cent planning to do more6. That second one is the constructor channel with the customer's hands on it.
"We have three to five years to respond to a new entrant" has expired. Nobody repealed it. The assembly time underneath it simply moved.
Occasionally a whole category arrives that was not on last year's map in any form. Agentic commerce is forecast to orchestrate up to a trillion dollars of US retail by 20307. Treat that as a category marker and nothing else. It is a forecast, this book does not run on forecasts, and importing one through the side door would be exactly the behaviour Chapter 1 promised to avoid. What the figure legitimately shows is that somebody serious is budgeting for a category that did not meaningfully exist two years ago — which is one more row on the tally sheet.
The curve is not smooth
Here is the second reason to hold a rate rather than a forecast, and it is the more important one.
One frontier model release was reported to have "demonstrated a striking leap in cyber capabilities relative to prior models, including the ability to autonomously discover and exploit zero-day vulnerabilities in major operating systems and web browsers"8 — a capability the lab itself noted it had not trained for, arriving downstream of general improvements in reasoning.
The coverage that followed moved within weeks from alarm to de-escalation, and the de-escalation is the part boards misread. A story that goes from "this is dangerous" to "this is improving vulnerability discovery enough that banks are patching" in a month is not a relief. It is a boundary case becoming structural inside a single quarter. This has a name in the parent doctrine — Domain-Spike Risk: the risk that AI capability leaps vertically inside a single domain, shattering whole-industry financial models without warning .
The planning consequence is specific rather than atmospheric. A rhythm built on smooth progression cannot see this class of move — not because it is looking in the wrong place, but because interpolation is the wrong instrument for a step function. Which is why the honest thing to hold is a rate, updated, rather than a projection, defended.
Key Insight
A rate is comparable. A forecast is not. You can hold a rate beside your own elimination count and get an action out of it; you cannot do anything with a date except be wrong about it later.
Sorting what you counted
Four columns produce a pile, and a pile is not a diagnosis. Sort each counted item by the route it travelled, using the taxonomy the parent book established: customers, competitors, constructors .
The sort is not administrative. A firm heavily exposed on the customer route — clients arriving with the first pass already done — has a completely different problem from one exposed on the constructor route, where the threat is that the thing it sells becomes economical to build. Both raise the branching count. They call for different boundary cases, and eventually for different probes. If your tally sheet cannot tell them apart, it is measuring weather again.
What number is bad?
There isn't one, and this book is not going to invent one.
Nobody publishes a benchmark for branching rate by industry, because nobody counts it. If you see a figure claiming to be that benchmark, it was made up. The count is useful in precisely two comparisons, and both of them are internal:
- Against your own prior quarters. Is the rate rising, and on which of the four rows?
- Against your own evidence output. Which is the entire point, and which is the column you almost certainly do not have yet.
That second comparison is the instrument. Neither number means anything alone. An entrant count read by itself is a vanity metric — impressive, alarming, and unactionable. Read against eliminations, it becomes the only strategic ratio on the page.
"These are global figures. My market is small."
Correct, and that is the point rather than an objection to it. The published numbers in this chapter do two jobs: they establish that the phenomenon is real, and they show you the shape of the collection method. Neither is the thing you govern by.
What you govern by is four rows, four quarters, for your market — assembled by one analyst in about a week from public information and your own lost-deal notes. That is the whole cost. The barrier to running this has never been money or capability. It is that nobody has been asked to.
The branching tally sheet
| Input | Q-3 | Q-2 | Q-1 | Q0 | Route |
|---|---|---|---|---|---|
| Entrant cadence | competitor | ||||
| Offer / pricing-model cadence | competitor / customer | ||||
| Capability cadence | constructor / customer | ||||
| Regulatory cadence | all three |
Fill it in and you will have something most boards have never seen: a number for the clock you do not control.
And an empty column beside it, for the one you do.
The Clock You Do
The evidence clock is not a fixed property of your firm. It got physically faster, for a reason already documented in this corpus — and almost nobody has claimed that speed at strategy altitude.
Put the two cadences side by side and the problem stops being philosophical.
A market that adds a consequential move somewhere between weekly and monthly. A strategy process that raises a question in March and reports on it in the following February's board pack. That is not a disagreement about rigour. It is a sampling-rate error, and an engineer shown those two numbers would say so before saying anything at all about the people involved.
A question raised in March and answered in February was answered to a different firm.
Not metaphorically. The context that made the question consequential — the client who was wavering, the competitor who had not yet moved, the price that was still defensible — has been replaced. The answer arrives correct and inert.
Three things to count
The right-hand column of the diagnosis has three rows, and each one has a qualifier that does most of the work.
What counts, and what only looks like it counts
Counts as evidence output
- • A probe shipped that could have returned a result you did not want
- • An option closed, with a named source, a dated observation and a stated reason
- • A behavioural observation: what a customer, competitor or regulator actually did
- • A specified silence — where you named the audience and the expected response first
Does not count
- • A pilot that could only succeed
- • "We decided not to pursue it" — a decision wearing evidence's clothes
- • A workshop conclusion, however well argued
- • Enthusiasm, attendance, a good meeting, a positive reply
- • A survey of stated preference where behaviour was observable and nobody looked
The third row is the most diagnostic and the one nobody has: median elapsed time from question raised to evidence received. Most firms cannot compute it, not because the data is hard to assemble but because nobody dates the question. Strategic questions enter the organisation undated and leave unresolved, which makes them impossible to age and therefore impossible to manage.
Start dating them. That single administrative change surfaces more than the count does — you discover that four of your live questions have been carried for over two years, and that two of them were answered by the market some time ago without anyone writing it down.
Why is this suddenly available?
Because the environment changed underneath the practice, and the change was documented in this corpus for an entirely different purpose.
Key Insight
AI has shortened world-response latency below the decay time of the original thought.
That inequality was written about a personal working rhythm: previously the world answered after you had cognitively moved on, and now the answer collides with a live thought, and the collision produces another thought . It was an argument about idea rate.
It is also, unmodified, an argument about strategy. A firm has a decay time too. It is the interval over which the people who could act on an answer still hold the context that made the question worth asking — the pipeline they were looking at, the client conversation that prompted it, the pricing decision it was going to inform. Beyond that interval, evidence arrives as history. Inside it, evidence arrives as a decision.
What changed is that the interval on the world's side collapsed. You can put a price into the market and read the response in a quarter. You can publish a claim to a specified audience and read the silence in a fortnight. You can stand up a scoped build with a named acceptance event in weeks rather than in a budget cycle. None of that was economically available to a mid-sized firm five years ago, and all of it is now — which means the evidence clock stopped being a constant and became a decision.
One more thing worth stealing from that chapter: its honesty about its own numbers. It offers a rough shape — a couple of ideas a month then, ten or twenty a day now on a hot stretch — and then refuses to defend the arithmetic, defending only the regime change. Treat that as a shape claim, not a paper result. This book will behave the same way about its own central claim, and the reason to say so here rather than in the appendix is that the credibility has to be earned before it is spent.
One clock, three parts
The evidence clock is a single number in the diagnosis, but it is a composite, and knowing which part of it is slow tells you what to fix. This is a decomposition of the evidence clock — not a fourth pair of clocks, and it should never be used as a rival framework.
Epistemic — a hypothesis changes
Speed: hours to days. Cost: near zero, now.
What it proves: that a belief was internally inconsistent, or that someone had not thought about a mechanism.
The characteristic error: stopping here and calling it learning. This is the clock that has been sped up a hundredfold by cheap cognition, which is exactly why a firm can feel like it is learning continuously and be unable to name anything it now knows.
Behavioural — money, authority or attention moves
Speed: weeks to a quarter. Cost: real, and usually the cost of putting something into the market.
What it proves: the first honest thing. Somebody outside your building did something differently, in a direction you predicted or did not.
The characteristic error: accepting a stated preference in place of an observed one. What a client says about pricing in a workshop and what their procurement function does in a tender are different measurements of different things.
Outcome — reality reports whether the move worked
Speed: quarters to years. Cost: mostly organisational patience, which is scarcer than money.
What it proves: whether the thing you believed was true was also load-bearing.
The characteristic error: abandonment. The row is closed at the behavioural clock, the outcome is never backfilled, and the firm carries a conclusion it never actually checked.
The third one is where the discipline has to live, and it costs almost nothing except memory. A row stays open past the quarter in which it was fashionable. A named person owns backfilling the outcome. The board pack carries open rows as well as closed ones — because an open row is a statement about what you are still waiting to learn, and that is more useful than a closed row full of confidence.
What does a probe cost, and how many can we afford?
My own line about margin was made about survival, and it transfers exactly:
It isn't how good your work is. It's your margin — margin is how long you can tread water while the rain keeps coming.
Read at this altitude: margin is how many probes you can afford before the answer matters. A firm with thin margin does not get to run a leisurely evidence programme, and pretending otherwise produces the worst of both — a slow clock and a spent balance sheet.
The sizing logic does not need a formula, and it certainly does not need a number this book has no source for. It needs a comparison:
A probe should cost less than the option's carrying cost over the interval it would otherwise stay open. And the carrying cost of an unresolved strategic option is mostly not the analysis — it is the decisions the option defers. A pricing structure not changed. A hire not made or made defensively. A capability not built because two of the live futures do not need it. Those are real costs, they compound quarterly, and they are invisible precisely because they are omissions.
Most firms discover, doing this comparison honestly for the first time, that a probe they refused to fund at the price of a modest project has been carrying a deferred decision worth considerably more for three years.
"Our decisions are too big to test"
Some are. Not many, and rarely all the way down.
The sub-clock decomposition is what rescues this. You usually cannot test the decision. You can nearly always move the behavioural clock on one component of it — one practice area, one segment, one price, one commitment — and a component result frequently kills a whole branch, because the branch depended on the component behaving differently. The design of that move is Chapter 7's subject, and it is more constrained than it sounds.
The related objection — speed at the expense of rigour — has the relationship exactly backwards. A discriminating probe is a rigour instrument. What it refuses is the pretence that an argument settled a question that only an observation could settle. Rigour without a clock is a paper that arrives after the decision, and there is nothing rigorous about that.
And if you cannot move it?
This is the paragraph that stops this book being advice for fast firms only.
A firm that measures its evidence clock and finds it cannot be made to beat the branching clock is not excused from the ratio. It has learned something specific and actionable: it is navigating on inference. The correct response is not to abandon the measurement; it is to size commitments to match. Smaller. More reversible. Explicitly labelled in the portfolio as carried on inference rather than on evidence — which is a completely different posture from carrying them and not knowing.
There is also a real reason not to leave it there, and it belongs to the parent doctrine's argument about a search engine with no connection to the world: a sealed engine in permanent fog does not merely stagnate. It confidently generates ever more internally-consistent futures, each one further from reality . A slow evidence clock is not neutral. It is the condition under which a good analytical apparatus becomes an increasingly sophisticated way of being wrong.
Both columns, one page
You now have the instrument. Branching inputs on the left, four rows, four quarters. Evidence outputs on the right, three rows, same period. Chapter 9 builds the page for a real firm and reads it two ways: as a ratio, and as a comparison of intervals — median time to evidence against the interval at which the market adds a consequential move.
What Part I has not supplied is the thing that decides what to test. A firm with a fast evidence clock and no aim will run beautifully instrumented experiments on questions that do not matter, which is an expensive way to be busy.
The instrument that does the aiming is the best one we have, it has two thousand years of lineage behind it, and it has a limit that this corpus had not stated.
Where the Spear Stops Working
The best instrument we have for aiming a strategic search has a limit, and it is not a limit of rigour. It is a limit of what kind of system is being examined.
Two thousand years ago, Lucretius settled an argument about the size of the universe without leaving his chair.
The question was whether space has an edge. The move was not to reason about edges. It was to push one variable to its limit and see what the geometry did: what happens if a man walks to the alleged boundary and throws a spear at it? Either the spear flies through — in which case there is no edge, only more space — or it strikes something, in which case there is something on the other side, and again no edge.
Argument over.
Hold those two words up to the light for a moment, because everything in this chapter is underneath them.
The instrument, working properly
The move Lucretius made has a name in this corpus and it is the human half of a two-part partnership: push one variable in the stuck argument to an absurd extreme until the geometry of the situation has no choice but to force a structural answer . It is not the question-answering move and it is not the question-reframing move. It is a third thing: a reshaping of the problem so that its structure is forced to declare itself.
And it is stubbornly human. Three things it needs that a model does not have: lived friction — the gut sense that something is wrong here, which comes from having skin in this particular game; taste for which extreme — most problems have seven knobs and pushing the wrong one to infinity produces eloquent noise; and willingness to imagine the absurd — sitting inside a deliberately impossible version of reality long enough to derive a result and carry it back. Frontier models run a boundary case beautifully once you name the variable. They are noticeably worse at choosing it.
None of that is under attack in this chapter. The instrument is excellent. The question is what it is an instrument for.
Why "argument over" was available
Space does not care that a spear was thrown at it.
That sentence is easy to nod at and worth sitting with, because the whole disanalogy lives in it. The system Lucretius examined is passive. It has no strategy. It has no pricing. It does not convene a meeting when it notices it is being examined. It has no interest whatsoever in the outcome of the argument, and — this is the load-bearing part — it does not anticipate the examination and pre-position itself.
The same is true all the way down the lineage. Galileo's falling bodies do not confer about what to do when tied together. The beam of light in the lift does not have a view on relativity. Every canonical thought experiment in physics works on something that cannot respond strategically, and the reason the results transport so cleanly into the real world is precisely that the real world, in those domains, is not paying attention.
Key Insight
"Argument over" is available to Lucretius because the thing he is arguing about has no strategy. Every business use of the same move is borrowing a guarantee that does not come with it.
A thought experiment in physics acts on a passive system. The brick wall does not respond strategically. A beam of light does not change its pricing model because Einstein is examining it. A market does.
Watch it happen
Take the boundary case this book will use for the rest of its length: competent generic first-pass advice in our sector becomes effectively free to the client.
Push the variable. Now watch what a passive system would never do.
- Customers may buy dramatically more analysis, because the price of asking collapsed and there are questions they never bothered to ask. Or dramatically less, because they now produce the first pass themselves. Both are rational. Which one happens depends partly on what you do next.
- Regulators may require a licensed human to approve what the free tool produced — which does not reduce the compression at all. It relocates the scarcity from analysis into authority, and changes who can sell.
- Competitors may bundle the free first pass into a broader service and use it as a distribution wedge, so the thing that became worthless becomes the shopfront.
- Platforms may capture the discovery layer entirely, at which point being cheap and good is irrelevant because the customer's agent never reaches you.
- You respond visibly, and your response changes what customers believe is normal — which changes what your competitors must match.
Now the second-order point, which is the one that actually breaks the analogy. Those five responses are not independent. The participants respond to each other. A competitor's bundling move changes what customers expect as standard, which changes what regulators are asked to permit, which changes which of your responses remains legal. The answer is not a function of the boundary condition. It is a function of the boundary condition and a game played by parties who can all see the same boundary coming.
The correction
My published version of this claim was blunter than it should have been. In my own words:
I took my boundary cases and thought experiments, Einstein-style, and said: that's how you pierce the AI fog.
That is the version I would now change, and the change is small enough to fit in a sentence and consequential enough to have produced this book:
The correction
Thought experiments aim the search. Discriminating experiments pierce the Fog locally.
Worth noting how that correction was arrived at, because the method is the same one the book argues for. A second reviewer working blind on the same material — no access to the first reviewer's conclusions — landed on a different phrasing of the identical point: boundary cases chart structural exposures inside permanent fog. Two routes, same destination.
And then the honest discount, which belongs in the same paragraph: both reviewers read the same transcript, so that is convergence rather than replication. It is weak confirmation. Weak confirmation stated as weak is worth more than strong confirmation asserted, and a book about evidence discipline should behave that way about its own claims.
What survives — and it is most of it
It would be a disaster if a reader closed this chapter believing thought experiments had been discredited. They have been located, and the territory they own is large.
What a boundary case can and cannot establish
Can
- • Which of your assumptions are load-bearing — the ones whose falsity changes what you are worth
- • Which revenue unit fails first, and in what order the rest go
- • Where value migrates, and to whom
- • What observation would tell you the mechanism has started
- • What is reversible now and expensive later
- • Timing certainty traded for structural certainty — you do not need to know whether it arrives in one year or five to learn that a large share of your proposition disappears when it does
Cannot
- • Tell you what the participants will do
- • Settle which of several coherent outcomes actually occurs
- • Establish who captures a cost reduction
- • Survive its own publication — once your view is visible, it becomes an input to the game
- • End the argument
The filter still holds too, and it is the entry gate for everything that follows: an extreme earns strategic attention only when it changes the shape of your legal moves rather than their size. If the answer at the extreme is "we'd be busier" or "we'd be poorer", it is a contingency, not a pivot . That work is done, it is done well, and this book uses it as an input rather than rebuilding it.
Myth vs Reality
Myth
"The thought experiment proved it. We ran the boundary case, the answer was structural, and now we know what the market does."
Reality
The thought experiment aimed it. It identified the assumption worth attacking and the observation worth buying. What the market does is a separate purchase, and it has a price.
The step this makes compulsory
If the system reprices when examined, then any sequence that runs boundary → breakage → conclusion has skipped the step at which the answer is actually determined. Not enriched. Skipped.
That step is counterplay: what the customer, the competitor, the regulator and the platform do next — including in response to each other, and to you. It is not an optional deepening for teams with time. It is the difference between a conditional map and a claim, and its absence is the most common reason strategy work feels rigorous and changes nothing.
Two ways to run the same boundary case
❌ Boundary → breakage → conclusion
- • Push the variable, identify what fails
- • Write it up as a finding
- • Present the coherent future to the board
Produces: a persuasive document, an enlarged option set, and no way to tell which of four outcomes you are in.
✓ Boundary → breakage → counterplay → economics → probe → state
- • Push the variable, identify what fails
- • Play out the four responding parties, including against each other
- • Find where the capture condition flips
- • Buy the observation that separates the survivors
Produces: a smaller option set, a named disposition, and an asset that survives whichever way it resolves.
Two objections, answered here
"Physics has observer effects too." It does, and they are a different phenomenon. Measurement disturbance is not strategic response. A photon does not consult its lawyers, brief its board, or pre-position itself in anticipation of being measured next quarter. The disanalogy is about intent and anticipation, not about disturbance — and it is the anticipation that does the damage, because it means the participants are already responding to a boundary case you have not run yet.
"So thought experiments are useless." Read the two-column table again. They are the aiming instrument, aiming is a large job, and no amount of evidence-gathering rescues a firm that is testing the wrong thing quickly. This book is not an argument against boundary cases. It is an argument that they were being asked to do a second job they were never able to do.
What happens if nothing closes
The failure mode has a shape, and it is worth naming before Part II builds the machinery:
Without the real-world ending, a thought-experiment engine can become a Fog factory: producing increasingly coherent futures while the company ships nothing.
Which is the same warning the parent doctrine gives about any sealed search apparatus — that it does not merely stagnate, it confidently generates ever more internally-consistent futures, each further from reality . Coherence is not a truth signal. It is what a good generator produces by construction.
One last honesty, before the machinery. Reflexivity does not only break the thought experiment — it also degrades the probe. Put a price into the market and you change the market you were measuring; a competitor may respond to your move rather than to the underlying condition, and the counterfactual is not available for inspection. That is a real limit, it belongs in the record as a confounder, and Chapter 7 handles it directly rather than hoping the reader does not notice.
It is still not an argument for the scenario pack. A probe buys locally valid, decaying information. A scenario pack buys none.
So: if the boundary case only aims, something has to close. The sequence that runs from an aimed question to a closed option is the whole of the next chapter — and step three of it is the one this chapter just made compulsory.
Boundary to Option State
Six steps, in an order that is not arbitrary. Most strategy processes run the first two and present the result as a conclusion.
Here is the whole machine, and then the parts.
The sequence
- 1. Boundary. What if competent advice, software or coordination approaches zero cost?
- 2. Breakage. Which revenue unit, cost structure or customer behaviour fails first — and in what order do the rest go?
- 3. Counterplay. What do the customer, competitor, regulator and platform do next — including in response to each other, and to you?
- 4. Capture economics. Who keeps the cost reduction, and at what threshold does that answer change?
- 5. Probe. What is the smallest real offer, price or operating experiment that distinguishes between the surviving explanations?
- 6. Option state. What evidence causes us to scale, defer, redirect or kill — decided before the evidence arrives?
The useful way to read that is not as a list but as a chain of handoffs. Each step produces a specifically-shaped object that the next step needs and cannot manufacture for itself. Skip one and the chain does not shorten; it breaks.
Step 1 — Boundary
Produces: a named load-bearing assumption — a belief about your business whose falsity changes what you are worth.
The entry gate is already built and this book uses it as an input: an extreme earns attention only when it changes the shape of your legal moves, not their size. "Our largest client becomes insolvent" is a serious risk and a bad boundary case, because at the extreme you do the same things harder with less money. "Competent generic advice becomes free" is a good one, because options die and new options appear .
Disqualified: a trend ("AI keeps improving") — it has no extreme. A contingency ("our biggest client fails") — it has an extreme and it is size, not shape. A boundary that carries a date — the date is a forecast wearing a boundary case's clothes.
Step 2 — Breakage
Produces: a failure ordering, not a failure list.
The question is not "what would be difficult." It is: at that extreme, which revenue unit stops being sellable at any price, which cost structure stops being supportable, which customer behaviour stops making sense — and crucially, in what order. The ordering is what makes the next two steps tractable, because it tells you which single thing to play the counterplay against.
Disqualified: "everything gets harder." That is a mood. If every unit fails simultaneously and equally, the analysis has not been done; real businesses have a first domino.
Step 3 — Counterplay
Produces: a set of rival explanations — and this is the object the whole downstream machine runs on.
Key Insight
Counterplay produces the rival explanations. Without rivals, a probe has nothing to discriminate between — which is why a firm that skips step three cannot design a useful experiment even with unlimited budget.
Four parties, and the discipline is to give each one its own paragraph rather than a bullet. What is it optimising? What move does it make that firms consistently fail to anticipate? And what would you observe if it had already started?
The customer
Optimising: the total cost of getting the outcome — not your price. Your fee is one term in their equation, alongside internal effort, verification, integration and the risk of being wrong.
The move firms fail to anticipate: partial substitution. Not "we will stop buying" but "we will keep the parts we can now do and buy only the difficult residue" — which strips the profitable middle out of an engagement and leaves the expensive tail, priced for a mix that no longer exists.
Observable: scope shrinking while the headline price holds. Clients arriving with the first pass already drafted. Requests to be priced against a baseline they produced.
The competitor
Optimising: the survival of their own unit — which may make a move that is bad for the category rational for them.
The move firms fail to anticipate: passing the compression through deliberately, as a weapon rather than as a concession. Every partnership models competitors absorbing the saving as margin, because that is what they intend to do. A competitor with a weaker book and less to protect will hand it to the customer and use it to take accounts.
Observable: a competitor publishing a fixed price where the category previously published a range. Public price transparency is nearly always an attack.
The regulator
Optimising: accountability, not efficiency. This is the actor firms model worst, because they model it as an obstacle to the technology.
The move firms fail to anticipate: relocating scarcity into authority rather than restricting the capability. The tool stays legal; the signature becomes mandatory. Compression is untouched and the question of who is allowed to sell changes completely — often in the incumbent's favour, which is why this leg is worth modelling as an opportunity and not only as a threat.
Observable: guidance that names a person rather than a process. Enforcement posture shifting before any rule changes.
The platform
Optimising: position in the flow. It does not want your margin; it wants to be unavoidable.
The move firms fail to anticipate: capturing discovery, so that quality becomes invisible. You are not out-competed; you are not reached. A firm that competes on being demonstrably better has no answer to an intermediary that decides who gets demonstrated.
Observable: buyers arriving through a route that did not exist last year. Enquiries that arrive pre-structured, in someone else's format.
Step 4 — Capture economics
Produces: the thresholds that make the rival explanations distinguishable.
A cost reduction does not evaporate. It goes somewhere: to you as margin, to the customer as price, to a competitor through imitation, or into a successor business through reinvestment. Step four asks which, and then asks the more useful question underneath it — at what point does the answer flip?
Thresholds are what convert a set of stories into a set of measurable claims. "If more than roughly a third of the market publishes fixed prices, the absorb-as-margin branch collapses" is a claim you can watch. "Competitors might pass it through" is not.
Step 5 — Probe
Produces: an observation that at least one live explanation cannot accommodate.
That is the entire specification, and it is stricter than it sounds. Chapter 7 is about how to meet it. For now the only thing that matters structurally is what the probe consumes: it consumes the rival explanations from step three and the thresholds from step four. A probe designed without those inputs is a demonstration with a budget.
Step 6 — Option state
Produces: a disposition, and a residue.
The disposition — scale, kill, defer to a named trigger, keep collecting, transfer to operations — is decided before the evidence arrives, because a trigger written afterwards is a rationalisation with a date on it. The residue is what the probe leaves behind whichever way it resolves, and it is what makes killing things affordable. Both get their own chapters in Part IV.
What breaks if you reorder it?
The order is not a preference. Two reorderings are common and both fail in specific ways.
Two broken orderings
❌ Probe before counterplay
- • The team is energetic, the clock is fast, so they go and test something
- • The test returns a clean result
- • The result is consistent with three of the four explanations, none of which were written down
You bought an observation that cannot discriminate. Money spent, evidence clock unmoved, and a false sense of having learned.
❌ Economics before counterplay
- • The team models who captures the saving
- • To model it, they must assume how the other parties behave
- • The assumption becomes an input rather than a question
You have assumed the capture condition you were trying to discover — the most elegant way to confirm what you already believed.
✓ In order
- • Counterplay enumerates who can determine the answer
- • Economics finds where their behaviour flips the outcome
- • The probe is aimed at the flip, not at the topic
Each step hands the next one something it could not have produced alone.
What you have after three steps
Not an answer. A conditional map — and it is worth naming its contents precisely, because a team that knows what it is holding stops mistaking it for a conclusion:
Every item on that list is genuinely valuable and none of them is knowledge about what will happen. The last two are the ones that hand forward: the observation that would signal the mechanism, and the action whose reversibility is decaying, are between them most of the input to a probe design.
The same machine, derived from money
Worth noting that this sequence has an independent derivation which arrives at the same structure from a completely different direction — real-options reasoning rather than epistemics. In that version the steps read: push a structural variable to expose the threatened assumption; identify the earliest observable signal; make a bounded investment that creates evidence and reusable assets; predefine the gate for increasing, redirecting or killing; preserve the option not to proceed; repeat as the world changes.
Different vocabulary, same skeleton. One version emphasises what you learn and the other emphasises what you paid and what rights you bought. They agree on the shape, and the agreement is mildly reassuring — a structure two different framings converge on is more likely to be about the problem than about the framing.
Run it on a cadence, or don't bother
The sequence is not an annual set-piece. Boundary cases stop being a one-off stress test and become a recurring navigation instrument — the chart refreshed rather than framed .
Which imposes a design constraint on the whole thing: every step has to be cheap enough to run again next quarter. A sequence that takes six months is a project, and projects do not iterate — they conclude, get presented, and are superseded by the next project. If your version of this takes half a year, the fault is almost always step three being run as a research exercise instead of as a two-hour argument between people who know the market.
Steps one to four are half a day of the right four people. Step five costs money. Step six costs courage.
That is a more honest cost breakdown than any budget line, and it locates the real barrier. The analytical work is cheap and always was. The reason firms stop at step two is not that steps three and four are hard. It is that steps five and six are the ones where somebody has to spend something and then live with what comes back.
Which leaves the step with actual design content in it. Steps one to four can be done well by any capable team in an afternoon. Step five is where firms with an excellent process still learn nothing, and the reason is almost always the same: they designed a test that could only agree with them.
The Smallest Discriminating Probe
The design rule is a refusal, and it is the reason the right probe is usually smaller, cheaper and considerably more uncomfortable than the one that gets funded.
A probe's job is not to show that you were right.
That sentence sounds obvious and is violated almost universally, because the institutional incentives all point the other way. A sponsor proposes an initiative. The initiative gets a business case. The business case has success criteria. The success criteria are written by the person who wants it to succeed. And at no point in that chain does anyone write down the result that would make them stop.
The discrimination test
Before you run it, write down every outcome the probe could produce and, beside each one, which live explanation it kills.
If no outcome kills anything, you have not designed an experiment. You have designed a demonstration with a budget.
Call the resulting artefact the kill map, and make it a required attachment. It takes twenty minutes and it is the highest-yield twenty minutes in the whole sequence, because it is where most proposed probes quietly die — not through rejection, but because the person writing the map discovers there is no row in which anything is eliminated.
A well-formed kill map, for a probe with four rival explanations in play, looks like this:
| If we observe… | This dies | This survives |
|---|---|---|
| Outcome A | Explanations 2 and 4 | 1 and 3, now distinguishable by their thresholds |
| Outcome B | Explanation 1 | 2, 3, 4 — a weaker result, and still a result |
| Outcome C | Nothing | All four — and this row is the design warning |
A probe with one blank row is acceptable; the world does not always cooperate. A probe with three blank rows is a press release.
This is sixty years old
None of it is new, which should be reassuring rather than deflating. In 1964 the physicist John Platt wrote an article in Science asking why some fields progress faster than others, and concluded it had nothing to do with the intelligence of the people in them. It had to do with a method, applied systematically rather than occasionally:
Devising alternative hypotheses; devising a crucial experiment (or several of them), with alternative possible outcomes, each of which will, as nearly as possible, exclude one or more of the hypotheses; carrying out the experiment so as to get a clean result… It is like climbing a tree.10
Note what Platt was arguing against: the single-hypothesis habit — the researcher with one favoured theory, gathering support for it. In business that habit has a name and a template. It is called the business case, and it is structurally incapable of eliminating anything, because it has only one candidate.
The tree image is the useful part for a portfolio. You are not trying to reach the truth in one move. You are at a fork, and the question is only ever which branch to stop climbing.
Why smallest?
Because size and discrimination are different properties, and firms consistently buy the wrong one.
Key Insight
Discrimination buys information. Size buys confidence. They are not the same commodity, and only one of them shrinks your option set.
A small probe that separates two explanations beats a large one that confirms the favourite, and it beats it three times over. On cost, obviously. On the evidence clock, because a small thing reports back inside the interval where the question is still live. And politically — which sounds like a soft consideration and is the decisive one.
A large probe carries a sponsor, a budget line and a reputation. When its result is unwelcome, the result gets reinterpreted rather than accepted: the conditions were unrepresentative, the team was learning, the market was distracted. A small probe can survive its own result, because nobody staked anything on it except the answer. Design probes that can afford to embarrass you, or the organisation will do the embarrassment-avoidance for you, invisibly, at the interpretation stage.
Five kinds of probe
1. Price or scope change on a bounded set
Discriminates: capture conditions. Who keeps the saving; whether buyers will pay for an outcome rather than for effort; whether your unit is defensible.
Cannot discriminate: long-run demand, or what happens once competitors respond. It measures this quarter's willingness in this segment.
2. A published offer with a stated hypothesis and a specified audience
Discriminates: whether a proposition is legible and wanted — whether the thing you think you are selling is a thing anyone recognises as a purchase.
Cannot discriminate: willingness to pay at scale, or delivery feasibility. Interest is not revenue and everybody knows that, right up until the interest arrives.
3. A scoped build with a named acceptance event
Discriminates: deliverability, and whether a buyer will formally accept the new unit — which is the hardest thing to fake and the most expensive to discover late.
Cannot discriminate: competitive response, or whether the second one is cheaper than the first. One acceptance proves the unit exists, not that it transfers.
4. A deliberate refusal
Discriminates: who values what — cheaply, and with unusual clarity. Stop offering a class of work and watch precisely who objects, how hard, and what they do instead.
Cannot discriminate: much, if repeated. This is a one-shot instrument per relationship, and it spends goodwill. Use it where the hypothesis is genuinely load-bearing and the alternative is another year of speculation.
5. Instrumented non-response
Discriminates: whether a market that could have responded chose not to — which is real information and the cheapest kind available.
Cannot discriminate: anything at all, unless the exposure was specified in advance. See the pitfall below, which is the most abused evidence type in strategy.
What a probe must leave behind
The result is the least durable thing a probe produces. The residue is what compounds, and it has nine fields — a schema that descends directly from the experiment discipline developed for marketing work, where the same failure was diagnosed first :
The standard that goes with it is uncompromising and worth adopting verbatim: if a test does not produce these artefacts, it was entertainment with metrics.
Fields three and seven are the ones that separate this from any strategy process you have run before. Recording disagreement means the awkward client who said the offer made no sense is in the record rather than in someone's memory. Backfilling the delayed outcome means the row stays open past the quarter in which it was interesting — which is administratively annoying and is where nearly all the value is, because the epistemic clock is cheap and the outcome clock is the one that tells you whether you were actually right.
The probe changes what it measures
Chapter 5 flagged this and it needs settling rather than acknowledging.
Put a fixed price into your market and you have not run a clean experiment on market preference. You have made a move. Competitors may respond to your move rather than to the underlying condition; customers may read it as a signal about your confidence; and the counterfactual — what would have happened had you not moved — is permanently unavailable for inspection. The instrument perturbs the system, and unlike in physics the system perturbs back on purpose.
Three practical consequences, none of which is "therefore don't."
- Write it in as a confounder at design time. Field five of the schema exists for exactly this. A confounder named in advance is a caveat; a confounder discovered at review is an argument.
- Prefer probes whose signal arrives before the response can. Speed is not only an operational property here — it is an epistemic one. A result that lands before competitors have repriced is measuring something closer to the underlying condition.
- Treat the decay rate as an input. How long does a probe's information stay valid in your market? That interval is itself a two-clock measurement, and a market where probe results decay in six weeks is telling you something important about your branching rate.
A probe buys locally valid, decaying information. That is a real limitation, honestly stated. It remains infinitely more than a scenario pack buys, which is nothing that decays because nothing was purchased.
What if there is no probe?
Sometimes there genuinely is not. Some structural questions cannot be tested inside a horizon that matters — whether a regulator will act in three years, whether a platform will vertically integrate, whether a technology curve continues.
Do not invent one. A manufactured probe on an untestable question produces a result that everyone quietly knows is decorative, and it corrodes the discipline faster than having no probe at all. Instead do three things: name the trigger — the observable event that would reopen the question and change the answer; assign an owner who watches for it; and mark the option explicitly as carried on inference rather than on evidence.
The failure is not carrying an untested option. It is not knowing which of your options are untested.
A portfolio in which four rows are evidenced and three are explicitly labelled as inference is in excellent health. One in which all seven look identical is not a portfolio; it is a list of opinions with a governance wrapper.
How often are you wrong?
More often than feels possible, and the only reliable data comes from the one domain where organisations actually measure elimination at scale.
What controlled experimentation reveals about confident ideas
Of changes at Microsoft had a negative or neutral effect on the metric they were designed to improve11
Of changes have a positive effect on target metrics in well-optimised domains such as Bing and Google
These are product experiments run by expert teams with instrumentation, not strategic probes. The transfer is an analogy and should be read as one.
Take the analogy caveat seriously, and then take the argument that follows it more seriously still. Those figures come from teams that test constantly, in domains they know intimately, with measurement infrastructure most companies will never build. Their intuitions are the best-calibrated commercial intuitions available anywhere — because they get corrected weekly.
Strategic intuitions about market structure are corrected approximately never. There is no reason to believe they are better calibrated than a Bing product manager's, and every structural reason to believe they are considerably worse.
Which reframes what the whole apparatus is for. The point of an experimentation capability was never any individual result; it is the rate at which trustworthy tests can be run at all12. That is falsification throughput, argued from the product floor. This book is carrying the same claim one altitude up, to the place where the unit of sale itself is the thing under test.
Everything in this chapter describes what a firm does instead of a failure that has already been named — not at board level, where it is invisible, but one floor down, in a department that hit it first and gave it a name.
Mutation Theatre at Strategy Altitude
The marketing floor hit this failure first, when generation got cheap for them. It has the same shape at board level, with scenarios instead of creatives and capital instead of click-through.
Picture the Monday meeting, two floors down from the boardroom.
Two subject lines, two creatives, two landing-page layouts. Traffic split cleanly. Variant B beat variant A on the metric the dashboard was built to optimise. Someone screenshots the green uplift. Someone else opens a brief that says make more like B. By Wednesday the calendar is full of surface cousins — same colour family, same structure, same length band — and the organisation calls that a culture of testing.
It is a culture of selection. And selection is not nothing: randomised assignment genuinely controls for seasonality, traffic mix and the other gremlins that make before-and-after stories dishonest. Comparing beats asserting. The failure is not in the testing.
The failure is in what the organisation stores afterwards. It stores a winner label. It almost never stores a mechanism. Which is why the next test is a clone — clones are what you invent when the only thing you learned is that this object scored.
When generative systems made variant production nearly free, that should have made learning explode. Mostly it did not:
You can now mint a hundred near-misses before lunch. In practice it often makes mutation theatre explode: more objects, same ignorance, prettier decks.
Key Insight
The scarce resource moved. It is no longer the ability to produce another candidate. It is the ability to design a candidate that would falsify a mechanism rather than redecorate a winner.
That diagnosis was written about creative testing . Read it again with the nouns changed.
The same meeting, eight floors up
A board asks AI for fifty futures. The model returns fifty persuasive futures. Executives discuss them, combine three, put one in a deck and call the exercise foresight. The solution space has expanded, but nothing has been learned.
Identical pathology. Cheap generation, abundant candidates, a selection ritual, a stored label, no stored mechanism. The differences are all in the direction of worse.
The stakes are larger: the currency is capital allocation rather than click-through, and the objects being cloned are commitments about what the firm is for. The feedback is slower by orders of magnitude: a creative test reports in days, a strategic bet in years. And — this is the reason the failure has gone unnamed at this altitude for so long — the marketing team at least had a dashboard that reported something. It reported the wrong thing, but it reported. The board has no equivalent instrument at all. There is no screen anywhere in the building showing strategic options eliminated this quarter, which means the failure produces no signal of any kind.
Four enemies, translated
The original diagnosis named four enemies so the argument would not have to keep re-litigating the mood. Each has a board-level twin.
The same four enemies, two altitudes
On the marketing floor
- Ship-the-winner analytics — the tournament ends at the trophy.
- Surface cloning — "more like B" as a substitute for isolation.
- Dashboard amnesia — learning that lives in a vendor UI dies with the campaign and the person.
- Heat-as-truth — engagement treated as authority over what stays in the canon.
In the boardroom
- Deck-the-winner strategy — the search ends when the narrative is chosen.
- "More like that scenario" — cloning the visible features of a future that resonated, without knowing which mechanism made it resonate.
- The whiteboard photograph — the year's best thinking in somebody's camera roll between a parking sign and a restaurant menu.
- The good room — the quality of the discussion treated as evidence about the world.
Heat is the seductive one
The fourth enemy deserves its own section, because it is the failure that feels most like success and arrives with a smile.
A client leans forward. A prospect says the framing is exactly right. A partner reports that the meeting was the best they have had in that account for years. An article draws unusual response. Every one of those is real information about something. None of them is information about whether the underlying thesis is true.
The qualification from the original work is exact and transfers without modification: a local gradient tells you the direction of informative travel, not the guaranteed prize . Walk the heat. Do not legislate from it.
To be clear: generate more, not less
A careless reading of this chapter produces exactly the wrong instruction, so state the right one plainly.
Generation should stay lavish. Cheap cognition on the generation side is the premise of this entire book — it is why the branching clock runs fast, and it is also the thing that makes a serious search affordable for a mid-sized firm for the first time. A partnership generating fifty candidate futures a quarter is doing something valuable. A partnership generating fifty and killing four is in outstanding health.
The problem is never the numerator. It is that the denominator was never installed.
Two quarters, same budget
❌ Fifty futures, zero kills
- • Option set at quarter end: larger than at quarter start
- • Stored artefacts: a narrative, a deck, shared vocabulary
- • Stored mechanisms: none
- • Decisions the board can now make that it could not before: none
Net contribution to fog: positive. The firm paid its best people to thicken it.
✓ Eight futures, three kills
- • Option set at quarter end: smaller, and the survivors are sharper
- • Stored artefacts: three killed branches with the evidence that killed them
- • Stored mechanisms: why each died, and what would reopen it
- • Decisions the board can now make: the ones the dead branches were blocking
Net contribution to fog: negative. Which is the only quarter-end result that means anything.
The antidote is a receipt requirement
The cure for mutation theatre is not restraint. It is a standing rule about what a strategic probe must leave behind — the observation; the competing causal explanations; the unresolved confounders; the alternative that was rejected; and the next experiment most capable of hurting the leading explanation.
Chapter 7 gave that residue nine fields. The standard that enforces it is one sentence, borrowed intact: if a test does not produce these artefacts, it was entertainment with metrics.
Myth vs Reality
Myth
"We ran fifty scenarios." Read as thoroughness. Presented as coverage. Received as diligence.
Reality
"We eliminated none." The option set grew by fifty, the attention available to carry it did not grow at all, and the firm is measurably foggier than it was in January.
Where the metric came from
Falsification throughput is not invented here. Its direct ancestor is a failure-mode tell in the parent book, in the list of ways a search discipline hollows out: rows are being added faster than decisions are being changed. Track the ratio; it is the metric that matters .
This book makes two changes to that ratio, and each buys something specific.
- The numerator changes from decisions changed to options eliminated, with evidence. A decision can change for many reasons, including a persuasive partner and a bad quarter. An elimination has to name what killed it, which makes it auditable and makes it impossible to claim by restating the status quo in new language.
- The denominator changes from rows to time. A per-row ratio is a quality measure of the ledger. A per-quarter rate is a capability measure of the firm — and capability is what you invest in, staff for and compare against a branching rate.
That is the whole promotion: from a tell you notice when something has gone wrong, to a standing instrument you run when nothing has.
Why this is urgent rather than merely untidy
Because a firm in mutation theatre is not standing still.
A search apparatus that works well manufactures fog as a by-product of working well — and the parent book generalised that beyond your own engine to every other actor's search running simultaneously .
So the firm running fifty futures a quarter with no elimination mechanism is contributing to the condition it is trying to navigate. At its own expense. Using its most expensive people. And reporting it upward as strategic capability.
The strategy function is the last place in the business still measuring itself on production volume.
"But we do write things down." Probably. Check field six of the residue schema — decision changed or branch killed. Recording what happened is not the same as recording what it eliminated, and a minute that says "the board discussed three scenarios and noted the analysis" is a record of attendance.
That is the doctrine complete. Two clocks, a ratio, a sequence with counterplay made compulsory, a probe designed to discriminate, and a named failure mode to steer around.
The rest is proof. One firm, one revenue unit, five chapters — starting with the diagnosis you were told to run two chapters ago, done properly, including the parts where the collection is awkward and the judgement calls are real.
A Two-Clock Diagnosis, Worked
One firm, four quarters, two columns. Including the collection that was awkward, the entry that was counted wrongly, and the moment the exercise terminated early.
A note on the firm. It is a composite, assembled for this book from the shape of several situations rather than from one. Its numbers are illustrative arithmetic, not client results, and they are labelled as such wherever they appear. What is real is the collection method, the disqualification rules and the failure modes — those are transferable and they are the point of the chapter.
Professional services. Four practice areas. Around forty delivery staff. The revenue unit is the four-to-twelve-week analytical engagement, billed by the hour.
Chosen because it is the case most readers can map onto their own, not because any of this is about professional services. The ratio is domain-general. This is simply the most legible specimen — the arithmetic is visible, the unit of sale is unambiguous, and the compression is already underway.
The left column: four quarters of branching
Four rows, collected from public information and the firm's own records over about three days.
| Branching input | Q-3 | Q-2 | Q-1 | Q0 | Route |
|---|---|---|---|---|---|
| Entrant cadence | 1 | 2 | 2 | 4 | competitor |
| Offer / pricing-model cadence | 0 | 1 | 1 | 3 | competitor / customer |
| Capability cadence | 1 | 1 | 2 | 2 | constructor / customer |
| Regulatory cadence | 0 | 0 | 1 | 1 | all three |
| Total consequential moves added | 2 | 4 | 6 | 10 |
Behind each cell, the collection:
- Entrant cadence came from three sources cross-checked: funding announcements filtered to firms selling to buyers the firm also sells to; its own lost-deal notes, which turned out to name two entrants nobody in the partnership had heard of; and three buyer conversations, which named a fourth.
- Offer and pricing-model cadence was the hardest to collect and the most informative. Two competitor fixed-price announcements in the final quarter, and — more consequential — a free product tier covering work the firm had scoped as a project six months earlier.
- Capability cadence counts only what changed what a competitor or client can now do without this firm. Applying that filter honestly discards most releases, which is exactly what it is for. Six were considered; two qualified.
- Regulatory cadence caught one guidance note in two of the four practice areas which named a person as accountable rather than a process. That is the observable Chapter 6 flagged, and it arrived without anyone treating it as strategic news.
The entry that was counted wrongly
This is worth showing rather than smoothing over, because a reader who has never watched the correction made will not know how to make it.
The first pass counted five offer-cadence events in the final quarter, not three. Two of them were a single competitor's rebrand, which came with new positioning, a new website, a new three-word category name and a very confident launch. It felt like an event. Two people in the partnership had discussed it.
It does not count. The commercial unit did not change. The same firm was selling the same engagements, priced the same way, to the same buyers, with better adjectives. Nothing new became purchasable and no move that was previously illegal became legal.
Read the corrected row and the trend is unambiguous. Two, four, six, ten. The firm's market added as many consequential moves in the last quarter as in the previous three combined — and the partnership, asked to guess before the count was run, said "about the same as always, maybe busier."
The right column: four quarters of evidence
Three rows. This took about half a day, and most of that was arguing about what counted.
| Evidence output | Four-quarter total | How it was audited |
|---|---|---|
| Probes shipped that could have returned an unwanted result | 1 | Eleven initiatives reviewed. Nine could only succeed by construction. One was a genuine test whose result was reinterpreted. One qualifies. |
| Options closed with named evidence | 0 | Two closures were recorded. Both, examined, restated the status quo. |
| Median time from question raised to evidence received | n/a | Not computable. No strategic question in the firm carried a date. |
Each row has a story, and the second one is the one most firms will recognise about themselves.
The reinterpreted probe
Eighteen months earlier the firm had quoted one engagement on a fixed price. It went over. The post-mortem concluded that fixed pricing was unworkable in their kind of work because scope discovery is inherently uncertain, and the option was informally shelved.
That was a real test with a real result, and it should count. Except that when the diagnosis examined it, three things had happened. The engagement chosen was the most technically ambiguous in the pipeline, selected precisely because "if it works there it works anywhere" — which is a demonstration criterion, not a discrimination criterion. No one had written down, in advance, what overrun would have meant versus what it did mean. And the conclusion drawn — fixed pricing does not work — was considerably broader than the observation supported, which was that fixed pricing did not work on the firm's single hardest engagement with no scope instrument.
So it does not count as a closed option. It counts as a probe that was run and then not interpreted, which is the most expensive category of all: the money was spent and the information was discarded.
The two false closures
Both were recorded in board minutes as strategic decisions. Apply the counter-test — what would we otherwise have done? — and both dissolve.
The first: "we will focus on our core sectors rather than diversifying." The firm had not been diversifying. There was no live diversification proposal, no budget attached to one, and no partner advocating it. The decision displaced nothing, which means it recorded the status quo with a new label.
The second: "we will monitor developments in agentic procurement." Monitor is the tell. Nobody owned it, no observable was named, no trigger would reopen it, and it appeared in the next three board packs in identical words.
What four quarters produced (composite firm)
Consequential moves added to the market
Probes shipped that could have returned an unwanted result
Options closed with named evidence
The ratio, and where the exercise stops
Key Insight
An empty denominator is a complete diagnosis. There is nothing left to measure and everything left to decide.
Do not refine the numerator. Do not commission a better branching count. Do not benchmark against the industry, which has no benchmark. The exercise has already produced its finding, which is that this firm has no mechanism whose job is to make a possibility go away — and every hour spent improving the left-hand column now is an hour spent making the diagnosis more precise rather than acting on it.
What follows is not more measurement. It is a decision about whether to install an elimination mechanism at all.
The second reading: intervals, not counts
The same page supports a different and blunter comparison. Not how many against how many, but how often against how long.
Ten consequential moves in a quarter is roughly one every six weeks. Against that, the firm's median time from question to evidence is not merely long — it is undefined, and the four carried questions it could partially reconstruct had been open for between eleven months and two and a half years.
You are not slow. You are navigating on inference.
That distinction is not a consolation. It is an instruction, and it changes what the firm should do next week regardless of whether it ever fixes the clock. A firm navigating on inference should be sizing its commitments to match: smaller, more reversible, and explicitly labelled as inference rather than evidence in whatever record it keeps. The failure is not the slow clock. The failure is a slow clock and commitments sized as though the clock were fast.
What a normal strategy review would have missed
Three things came out of a week's counting that four years of strategy offsites had not produced.
1. The carried questions
Four strategic questions had been in continuous circulation for more than a year; two for more than two years. Nobody had a list, because nobody dated anything. Each had been discussed repeatedly, each had generated analysis, and none had ever been closed or explicitly deferred to a trigger. They simply recurred, like weather.
2. The false closures
Two recorded conclusions that displaced nothing. Both had appeared in board minutes as evidence that strategy was being done. The counter-test that dissolved them takes one sentence and had never been asked.
3. The unwatched route
Three of the four practice areas were exposed on the customer route — clients arriving with the first pass already done — and nobody was watching that row. The firm's competitive intelligence was entirely organised around competitors, because that is what competitive intelligence has always meant. The most consequential branching in the firm's market was happening inside its clients' offices and arriving as smaller scopes at the same headline price.
None of those is subtle. All three were invisible for the same reason: the instruments counted production, and these are all failures of elimination.
What the week cost
One analyst for five days. Four hours of partner time, spread across two sessions, for the judgement calls — which entries counted, which closures were real, whether the reinterpreted probe qualified.
That is the whole invoice, and it is the most persuasive fact in the chapter. The barrier to this diagnosis has never been cost, capability or tooling. Every firm reading this could have run it in any of the last twenty quarters. The barrier is that nobody has been asked for the number, because no board pack has a row for it.
Choosing what to run
Four carried questions came out of the diagnosis. Run them through the entry gate — an extreme earns strategic attention only when it changes the shape of the legal moves rather than their size — and they do not survive equally:
- "Should we open in a second city?" — at the extreme, the firm does the same things in more places. Size, not shape. It is a legitimate investment question and it is not a strategic search question.
- "Will our largest client insource?" — a contingency with a named party. It belongs in risk management, with an owner and a trigger.
- "Do we need an AI capability?" — not a boundary case at all. It is a budget question wearing a strategy costume, and it has already been answered by the operating portfolio.
- "What happens if competent first-pass advice becomes free?" — at the extreme, an entire category of promise stops being sellable at any price, a pricing structure becomes indefensible, the graduate pyramid inverts from engine to liability, and a service that was previously unstaffable becomes viable. Options die and options appear.
One of four. That is a normal yield, and the three rejections are not failures — they are three things the firm now knows it does not need to carry in this particular portfolio.
The survivor gets run properly, in the next chapter, through the first three steps of the sequence. Including the step that nobody skips on purpose and almost everybody skips.
The Boundary Case, Walked
Steps one to three, run in full. By the end there are four live explanations and no way to choose between them from the armchair — which is the analysis working, not failing.
Step 1 — Boundary
Competent generic first-pass advice in our sector is effectively free to the client, delivered with tools they already own.
No date attached, and the absence of the date is deliberate rather than evasive.
Before spending a morning on it, run the gate. At the extreme, does the set of things this firm is allowed to do change shape, or only size?
Two candidates through the same gate
✓ "Competent first-pass advice is free"
- • An entire category of promise becomes unsellable at any price
- • The hourly pricing structure becomes indefensible, not merely uncomfortable
- • The graduate pyramid inverts from engine to liability
- • A standing advisory commitment — unstaffable while first-pass work was expensive — becomes viable
Options die and options appear. Shape. Passes.
✗ "Our largest client becomes insolvent"
- • The firm does the same things it does now
- • With less money and more urgency
- • Nothing that was legal becomes illegal
- • Nothing impossible becomes possible
Painful, survivable, familiar. Size. Fails — and belongs in risk management with an owner, a trigger and a mitigation.
The second one is not a bad question. It is a serious risk that deserves serious planning. It is simply not a strategic search question, and putting it in the portfolio is how a portfolio fills up with twelve rows and produces no strategy.
Step 2 — Breakage
Now the arithmetic, walked rather than tabled, because the interesting number is not the one everybody quotes.
An engagement that required 100 consultant-hours now takes 80. The firm still bills by the hour and sells the same number of engagements. Revenue per engagement falls 20 per cent.
Payroll is largely fixed in the short run — you cannot shed a fifth of a delivery team in a quarter, and you would not want to, since the same people deliver the remaining work. So margin falls by considerably more than 20 per cent, which is the first thing partnerships get wrong when they model this on a whiteboard.
The second thing they get wrong is the capacity number. The freed capacity is not 20 per cent. It is:
The arithmetic nobody runs before the AI programme
Consultant-hours per engagement
Effective capacity created (100 ÷ 80)
More engagements needed merely to restore billed-hour revenue
Of prior revenue if the firm sells 20% more work
Composite arithmetic, fully derivable from the 100 → 80 assumption. No external source required and none implied.
Read the last two figures together. A firm that grows sales volume by a fifth — a genuinely good year in most professional-services markets — still ends up below where it started, because the compression outran the growth. And it will report the year as a successful AI programme, because delivery times improved and clients were pleased.
You bill twenty per cent fewer hours for the same work, you don't reprice, you don't restructure — and you call that an AI programme. You're cutting off your own arm. The customer gets the whole benefit. You've just helped yourself go out of business faster.
The failure ordering
Step two is not really about the revenue number. It is about which thing goes first, because that is what tells you where to aim step three.
- First-pass analytical scope. Roughly 30 per cent of billed hours on a standard engagement in this firm. It stops being sellable — not because it stops being useful, but because the client can now produce a version of it that is good enough to negotiate against.
- The leverage model. The economics of the engagement depended on juniors doing the first pass at a rate that carried the seniors. Remove the first pass and the leverage ratio does not shrink gracefully; it inverts.
- The graduate intake. Which existed to feed the leverage model. It is the first line item to be cut, and cutting it is rational quarter by quarter and destroys the firm's ability to manufacture senior judgement over a decade.
- Senior utilisation. Which the pyramid was quietly subsidising. The seniors are left with the difficult residue, priced for a mix that no longer exists.
That fourth consequence — what happens to a firm left holding only the hard cases — is a substantial argument in its own right and it belongs to the professional-services companion, which works it properly . Here it is one line in a failure ordering, and the ordering is what step three needs.
Step 3 — Counterplay
Everything so far could have been done by a competent analyst alone in a room. This is the step that cannot, and it is where the boundary case stops being a thought experiment and starts pointing at something purchasable.
Customer — the move is partial substitution
Not "we stop buying." The rational move is to keep the parts they can now do and buy only the residue — which strips the profitable middle out and leaves the expensive tail, priced for a mix that has gone.
Evidence it has started, from the firm's own records: two of the last five prospects arrived holding their own baseline analysis. Nobody had logged this as competitive intelligence, because it did not arrive as a competitor.
Evidence from an adjacent market where the data exists: 67% of corporate legal departments and 55% of law firms expect AI to change how hours are billed, and 71% of buyers already prefer a flat fee for an entire matter13. Buyers in a comparable professional market are already holding a view about the unit of sale.
Competitor — the move is weaponised pass-through
The partnership's instinct is to hold rate and absorb the compression as margin for three to five years. That plan requires two things to be true simultaneously: that no material competitor passes the compression through to clients, and that clients cannot observe that it happened.
Both are fragile, and the second is nearly gone — a client who has produced their own first pass knows roughly what the work now costs.
Evidence on the first: in the adjacent market, firms collect roughly the same amount per hour whether they discount aggressively or hold firm on realisation14. Rate discipline is weaker in practice than partnerships believe it is in principle — and a competitor with a thinner book and less to protect has every reason to hand the saving to the client and take accounts with it.
Regulator — the move is relocation, not restriction
The guidance note the diagnosis picked up in two of four practice areas named a person as accountable rather than a process. That does not reduce the compression at all. It moves the scarcity out of analysis and into signature.
Which is potentially favourable to this firm, and modelling it only as a threat is the standard error. If the analysis is free but somebody has to be accountable for it, a firm that can carry accountability credibly has just been handed a defensible position by a party it treats as an obstacle.
Platform — no material capture yet
Honest answer: in this firm's segment, nothing. Buyers still arrive directly or through referral, and no intermediary sits between the firm and its market.
The observable to watch is specific: buyers arriving pre-structured — an enquiry that turns up in somebody else's format, with the requirements already framed by a tool or an intermediary the firm did not choose.
Not every leg fires. Saying so is what makes the other three credible, and a counterplay analysis in which all four actors happen to be moving decisively is usually a counterplay analysis that has been written to a conclusion.
And now the part a four-bullet list hides
Those actors are not independent. Walk one chain concretely:
A mid-sized competitor with a weak book publishes a fixed price for the same category of engagement. Within two quarters, three of this firm's buyers have seen it — and the effect is not that they switch. The effect is that a fixed price becomes the normal thing to ask for. Procurement functions start requesting it as a matter of course. At which point one of this firm's clients asks its own compliance team whether a fixed-price engagement changes who is accountable for the analysis, and the compliance team asks the regulator, and the regulator's answer — whatever it is — now applies to a market that has already changed its default.
Nothing in that chain was set off by the boundary condition directly. Each step was set off by someone else's response to it. That is the reflexivity from Chapter 5 arriving as a specific sequence of events rather than as a philosophical caution, and it is why the answer to "who captures the saving" cannot be derived from the boundary condition alone.
Where three steps leave you
With four explanations, all live, all coherent, and all consistent with everything the firm can currently see.
| Explanation | What it predicts |
|---|---|
| 1. Customer capture | The saving passes to clients; the firm's unit shrinks; scope falls at the same headline rate. |
| 2. Competitive pass-through | A weaker competitor forces the saving into market price; the whole category reprices downward. |
| 3. Authority relocation | Scarcity moves into signature and accountability; credentialed incumbents strengthen. |
| 4. Demand expansion | Cheap cognition makes previously uneconomic services viable; total demand for useful cognition rises. |
All four are consistent with everything we currently observe. That is the finding.
It is worth being clear that this is a success. The firm began with an ambient dread about AI and a question it had carried for two years. It now has four named, mutually exclusive, structurally distinct accounts of its own future, each with a prediction attached. That is a conditional map, and it is exactly what steps one to three are for.
It is also, on its own, worth nothing at the bank. Four explanations that all fit the evidence are four explanations you cannot act on, and a firm that stops here has produced the most sophisticated possible version of not knowing.
The four differ in one respect before they differ in anything observable: they have different economics. They imply different answers to who keeps the money, at different thresholds. That is where they first become distinguishable — and where the next chapter starts.
Four Ways the Money Moves
Step four. The claim that started this whole line of work is true under exactly one of four capture conditions — and the difference between them is not something you can reason your way into.
Here is the slogan, in the form its author has used it:
Making every member of staff twenty per cent more productive doesn't get you through the fog. On its own it gets you to zero.
It is directionally right and analytically lazy, and this chapter exists to qualify it. The qualification is not a retreat — the corrected version is considerably more useful than the slogan, because it tells you which firms it applies to and what would have to be true for it not to.
Four outcomes follow from the same 100-to-80 compression. Walk all four.
1. Unsold capacity
What happens: the entire productivity gain passes to customers as fewer billed hours. The freed capacity does not sell. Payroll stays.
The arithmetic: revenue per engagement down 20 per cent, engagement volume flat, and margin down by more than 20 per cent because the cost base did not move. The firm has funded a transfer of value to its clients out of its own investment budget.
The part everyone gets wrong: this is not the pessimistic case. It is the passive case. It is what happens when nobody decides anything — when the delivery teams adopt the tools, the timesheets fall, the invoices follow the timesheets, and no other part of the firm changes. Partnerships model it as the bad-luck branch. It is the branch you get by default.
2. Fully sold capacity
What happens: the firm sells 25 per cent more engagements, restores its billed-hour revenue, and reports a good year.
What has actually been achieved: the old economics defended, and nothing else. The AI investment has been spent entirely on standing still. There is no new unit, no new reason for a client to choose this firm, and no change in what happens when the next compression arrives.
Why it is unstable: competitors with the same models can run the same play, and in a market with finite demand they must. The equilibrium is everybody chasing volume into the same pipeline, which is the classic route to price falling anyway — arriving a year later and looking like a market condition rather than a consequence.
3. Outcome pricing
What happens: the firm holds the engagement price while delivery effort falls. The productivity gain becomes margin rather than a discount.
Two conditions, both of which can fail:
— Competitors must not force the saving into market price. This is not a matter of collective restraint; it takes one firm with a thinner book to reprice the category.
— Verification costs must not eat it. Somebody has to check the machine-produced work, and the cost of checking rises with the stakes. A firm that removes 20 hours of analysis and adds 8 hours of senior review has captured a smaller gain than its timesheet suggests, and in high-consequence work it may have captured none.
4. Demand expansion
What happens: cheap cognition makes a service possible that could not previously exist. Continuous advice instead of periodic engagements. Every transaction examined instead of a sample. Every anomaly investigated instead of only the ones that escalated. Every account analysed instead of the top twenty.
Why it is different from the other three: it is the only outcome that creates a new reason to pay rather than defending an old one. The others redistribute an existing pool; this one changes its size.
What it requires, and the test that keeps firms honest: the service must be one that could not have been offered at all at previous cost structures — not a cheaper version of what you already sell. "The same review, faster" is outcome three wearing outcome four's clothes.
| Outcome | Who captures the saving | Verdict |
|---|---|---|
| Unsold capacity | The customer, entirely | Self-disintermediation. The default if nobody decides. |
| Fully sold capacity | Nobody — it is absorbed by volume | Old economics defended, no advantage created, unstable at equilibrium. |
| Outcome pricing | The firm, conditionally | Margin expands if competitors hold and verification costs stay contained. |
| Demand expansion | The firm, via a larger pool | The only outcome that creates a new reason to pay. |
Two of those need working, not summarising
The table above is where most treatments stop. Two of its rows contain machinery that decides whether the outcome is available at all, and both are routinely modelled wrongly.
Why fully-sold capacity is not an equilibrium
Outcome two looks like the sensible middle path, and in a single firm's spreadsheet it is. Run it forward across a market and it stops being available.
Follow the sequence. Every firm in the category gains roughly the same delivery compression at roughly the same time, because they are buying the same capability from the same suppliers. Each one now needs to sell about 25 per cent more engagements to hold revenue flat. In a market where total demand is not simultaneously rising by a quarter — and there is no reason it would be — those firms are competing for volume that does not exist.
The mechanism from there is ordinary and fast. Competing for volume means competing on price, because there is nothing else to compete on when everyone has the same tools and roughly the same delivery time. Price falls. The firm that thought it was defending its economics by chasing volume discovers that the volume was only available at a discount, which means outcome two decays into outcome one over two to three cycles — arriving late, looking like a market condition, and producing exactly the wrong post-mortem: demand softened.
Demand did not soften. Supply expanded, everyone chased it, and the compression reached the client by a slower route. What makes this worth writing down at design time is that it changes what outcome two is: not a stable landing place, but a delay mechanism that converts a fast visible problem into a slow invisible one.
Verification cost, which is where outcome three quietly dies
Outcome pricing is the outcome most firms want and the one they model least carefully, because the modelling requires admitting a cost that did not previously have a line.
Work an engagement. Twenty hours of first-pass analysis are removed. The firm holds the price. On the timesheet that is twenty hours of margin.
Except somebody now has to check the machine's work — and checking is not the same activity as producing, done faster. It is a different activity with a different skill profile: a senior person reading for the errors a confident generator makes, which are not the errors a junior analyst makes and are considerably harder to spot precisely because the output is fluent. Call it eight hours of senior time. The captured gain is not twenty hours; it is twenty junior hours minus eight senior hours, and depending on the rate differential that is a materially smaller number and occasionally a negative one.
Now scale the stakes. In low-consequence work, verification is light — a skim, a sense check, a spot audit. In work where being wrong is expensive, verification rises faster than production falls, because the cost of an undetected error is unchanged while the volume of plausible-looking output has gone up. The uncomfortable implication is specific: the higher the stakes, the less of the compression the firm keeps, which is the opposite of what most partnerships assume when they reason that their premium work is protected.
That does not kill outcome three. It bounds it, and the bound is measurable: the ratio of verification hours to hours removed, tracked per engagement type. A firm that has never computed that ratio does not know how much of its productivity gain is real.
What demand expansion actually looks like
Outcome four is the only one that creates something, so it deserves the specificity the others do not need. The test — could this have been offered at all before? — sorts it cleanly.
Passes the test, fails the test
Could not have been offered before
- • Continuous rather than periodic. A standing commitment that examines every change as it happens, rather than an annual review. Unstaffable at human cost because the work arrives in small pieces at unpredictable times.
- • Census rather than sample. Every transaction, every contract, every asset — where the previous offer was a risk-weighted sample and everyone knew the sample was a budget.
- • Every exception investigated rather than only the ones that escalated — which changes what the client can claim about its own controls.
- • Account-specific analysis for all accounts, not the top twenty, which changes who inside the client can use the output.
Outcome three in disguise
- • "The same review, faster."
- • "The same report, with better visuals."
- • "The same engagement, with a smaller team."
- • "The same advice, available more often" — unless the increased frequency changes what the client can do, in which case it is genuinely the first column.
The tell is the client's side of the sentence: if the client cannot name something they can now do that they could not do before, it is a cheaper version of the existing unit.
There is a harder question underneath, and it is the one that stops outcome four being a comfortable landing place: who else can offer it? A service that becomes possible because cognition got cheap becomes possible for everybody at the same moment, including the client. Which is why outcome four survives only where the expanded service still requires something the client cannot self-supply — accepted evidence, independence, accountable sign-off, cross-client pattern recognition that no single client can lawfully assemble.
The corrected claim
Bottom Line
Productivity applied to a depreciating commercial unit, with no way to capture the gain or redirect the capacity, accelerates decline. The mistake is not making staff faster. The mistake is calling that a strategy while leaving the revenue unit unchanged.
That is a narrower claim than the slogan and it does more work, because it names the three things that have to be simultaneously true for the damage to occur: a compressible unit, no capture mechanism, and no redirection of the freed capacity. Break any one of them and outcome one does not happen.
It also stops the argument being about AI adoption, which it never was. The three-condition test for whether your commercial unit is exposed at all — is it compressible, is the compression visible or performable by the customer, is entry into the layer above open — belongs to the parent book and is worth running before any of this . A firm that fails the test has a productivity programme and no terminal-value problem, which is a good position and a boring chapter.
Where the answers flip
Each outcome is stable only inside a range, and the boundaries between them are thresholds. This is the most useful thing step four produces, and it is also where an author has to be careful, because the temptation to supply a number is enormous and there is no source for one.
The observation that changes what happens next
Look again at what separates the four outcomes.
It is not the quality of the firm's analysis. It is not whether the partnership is clever. Two firms with identical delivery capability, identical tools and identical clients can land in different outcomes — and the thing that decides which is what customers and competitors actually do.
The difference between them is not analytical. It is a question about what people will do — and that has a price.
Follow the chain, because it is the hinge of the entire book and it has now arrived with a worked case behind it rather than as an assertion:
- The four outcomes are distinguished by the behaviour of other parties.
- Behaviour is empirical — it is a fact about the world, not a conclusion from premises.
- Empirical facts about markets can be purchased: you put something into the world and read the response.
- Therefore the question "which outcome are we in" has a price, and the price is almost certainly lower than the cost of not knowing for another year.
This is what Chapter 5 meant by the boundary case aiming rather than proving. The thought experiment took a vague dread and reduced it to four named, mutually exclusive economic conditions. That is an enormous compression of the problem. And the last step of it — deciding which one is true — is not available to thought at any price.
What to do before you know
The probe takes a cycle. Meanwhile the firm has to operate, and "we are waiting for evidence" is not an operating posture. So there is an interim question worth answering explicitly rather than by default: what do you do that is right in more than one of the four outcomes?
Run the four against each candidate move and some of them survive everywhere:
- Build the outcome-sizing instrument. Needed in outcome three, needed in outcome four, useful in outcome two as a defence, and in outcome one it is the only thing that lets you exit the hourly unit at all. Four for four — do it now, before the evidence, because it is the rare move with no downside branch.
- Measure the verification ratio. Required to price outcome three honestly, required to know whether outcome two's volume is profitable, and it costs a timesheet field.
- Stop hiring the graduate intake on the old ratio. Right in three of four — and wrong in exactly one, the demand-expansion case, where you may need more people rather than fewer. Which makes it a decision to defer with a trigger rather than to make now, and that is a different and more honest answer than "pause hiring while we think".
- Reprice everything immediately. Right in one outcome, actively destructive in two. This is the move partnerships make when they have read the first half of the argument, and it is the reason the sequence has six steps rather than two.
That exercise takes an hour and it produces the interim posture: act on what is invariant across the surviving branches, defer what is branch-specific, and be able to say which is which. A firm that can do that is already navigating better than one waiting for certainty, and it has not pre-committed to any of the four.
One adjacency worth naming and leaving: pricing a bounded commitment when delivery variance is real is its own discipline, with its own instruments for underwriting the variance rather than hoping it averages out. That is a sibling subject and this book will not attempt it.
Key Takeaways
- • "Productivity gets you to zero" holds under one capture condition of four — and that one is the passive default, which is why the slogan feels true.
- • Fully-sold capacity is not an escape; it spends the entire AI investment on defending economics that competitors can defend just as easily.
- • Outcome pricing has two failure conditions, and verification cost is the one firms forget to model.
- • Only demand expansion creates a new reason to pay, and only if the service could not have been offered at all before.
- • Which outcome you are in is a behavioural fact about other people, not a conclusion you can reach by thinking harder.
So the question for the next chapter is exact, and it is the first genuinely commercial question this book has asked: which of the four are we in, and what is the smallest thing we can put into the world to find out?
The Probe That Settles It
Steps five and six, for real. The right probe turns out to be smaller, cheaper and considerably more uncomfortable than any of the three obvious candidates.
Easier to see by elimination. Here are the three probes the firm would have run, and what is wrong with each.
Three probes that would have taught nothing
❌ "Run an AI pilot in delivery"
Discriminates nothing. All four explanations are entirely consistent with delivery getting faster — that is the shared premise, not the question in dispute. It would return a clean, encouraging result and eliminate zero branches.
❌ "Survey clients about pricing preference"
Measures stated preference where behaviour is observable. What a client says about fixed pricing in a relationship conversation and what their procurement function does in a competitive tender are different measurements of different things, and only one of them pays.
❌ "Try fixed price with our friendliest client"
Selection on friendliness is selection on the outcome. It works, and the result is uninterpretable — it cannot distinguish "the market will accept this" from "this client would accept anything from us." Uninterpretable in the other direction too: if it fails there, the firm learns only that it has a problem everywhere.
Each of those is a probe a competent partnership would fund without argument. That is what makes them dangerous.
The probe
PRB-01 — fixed-scope pricing, two practice areas
The move. Two of the four practice areas move to fixed-scope pricing at the next contract cycle, sized on the outcome rather than on estimated hours. First-pass analytical scope is removed from the billable estimate and supplied instead as a pre-built input the firm brings to the table.
Selection. One practice area where the firm is strong, one where it is contested. Chosen for representativeness, explicitly not for the friendliness of the client base.
Owner. Managing partner. Not a committee, not the strategy lead, and not the practice heads whose numbers it affects.
Duration. One contract cycle.
What makes it small: two areas of four; existing clients; one cycle; no new hires; no platform build; nothing that cannot be reversed at the next renewal.
What makes it real: the price is genuinely fixed — no change-request escape hatch that quietly restores hourly billing. The removal of first-pass scope from the estimate is visible to the client, which is the part that makes it a probe rather than an internal efficiency. And the firm commits not to revert mid-cycle, which is the commitment that makes the result mean something and the one partnerships are most tempted to soften.
The discrimination table
Written before the probe runs. This is the artefact that separates an experiment from an initiative, and it takes about twenty minutes.
| If we observe… | This dies | This survives / what it means |
|---|---|---|
| Clients accept at a margin-preserving fixed price, both areas | Customer capture (1) | Outcome pricing (3) strengthens; competitive pass-through (2) weakens for now — it remains the medium-term threat. |
| Clients accept only with the compression handed back as a discount | Outcome pricing (3) | Customer capture (1) confirmed. The branch the firm least wanted is the one that survives, which is precisely why the probe was worth running. |
| Competitors announce equivalents inside a quarter | Nothing directly | Competitive pass-through (2) reporting in, and the threshold for (3) moves. A result that changes a threshold rather than killing a branch is still a result. |
| Clients ask to be priced against their own AI-produced baseline | Nothing | A different boundary has already arrived. The row needs re-cutting rather than resolving — and finding that out early is worth the whole probe. |
| Clients ask who signs | Nothing | Authority relocation (3 in Ch10's list) is live. The scarce complement may be accountability rather than analysis — which changes what the firm should be building. |
| Demand appears for something the firm has never offered | Nothing | Demand expansion (4) has a signal. Rare, easily missed, and the most valuable of the six — which is why the row exists in advance rather than being noticed in hindsight. |
Note honestly what this table does not do. No single outcome kills three of four explanations. Two of the six rows kill nothing at all. That is normal and it is not a design failure — Platt's tree gets climbed one fork at a time, and a probe that promises to resolve everything in one cycle is promising something no experiment has ever delivered.
What the table does guarantee is that two of the six outcomes eliminate a branch outright, and the other four either move a threshold or reframe the question. There is no row in which the firm learns nothing.
Triggers, written before the result
What survives either way
Here is the field most portfolios do not have, demonstrated rather than described.
Whatever this probe returns, the firm ends the cycle holding two things it did not have when it started. The first is an outcome-sizing instrument: a repeatable way of converting a client's desired end state into a bounded scope with a price, which is a genuinely hard piece of commercial machinery that this firm has never had to build because the hour absorbed all the uncertainty. The second is a set of acceptance tests: an agreed definition, written in advance with the client, of what "done" means — which the firm has also never needed, for the same reason.
What a kill costs, with and without residue
Without a residue field
- • The branch dies
- • The probe's full cost is written off
- • The sponsor has nothing to show
- • The next probe is harder to fund
- • Within two cycles, nothing gets killed
With a residue field
- • The branch dies
- • The firm keeps the outcome-sizing instrument
- • The firm keeps the acceptance tests
- • Both get used in the next probe and in ordinary delivery
- • Killing remains affordable, so killing continues
Key Insight
A killed option is a full return on the probe — and a firm that cannot say that out loud will quietly stop killing things.
The future the firm rejected
One more thing gets written down before the probe runs, and it costs more than any other line in the row.
Rejected: "hold rate and volume, and absorb the compression as margin for three to five years."
The reason it is rejected is specific rather than rhetorical. That future requires two things to be true at once — that no material competitor passes the compression through, and that clients cannot observe that it happened — and both are already false in this firm's own pipeline. Two of the last five prospects arrived holding their own baseline analysis, so the second condition has gone. And the first was never a strategy; it was a hope about other people's restraint.
Writing it down costs something real: it removes an option the partnership was quietly relying on. It was nobody's stated plan and it was everybody's fallback — the thing partners assumed when they did not want to think about the question. Once it is on the page as rejected, with a reason, nobody can drift back into it without arguing .
That is what a rejected future buys, and it is why a search in which nothing died did not happen.
The confounders, named now rather than argued later
- The firm's own move changes what competitors do. If a competitor announces a fixed price two months after this probe launches, the firm cannot cleanly tell whether that was a response to the underlying compression or a response to them. The counterfactual is not available. Recorded as a confounder at design time, this is a caveat; discovered at review, it is an argument that consumes a partners' meeting.
- Two practice areas is a small sample. A consistent result across both is meaningful. A split result is not a weak signal — it is no signal on the main question, and it means the two areas differ in some way that was not modelled, which is itself worth chasing.
- Existing clients are not a random sample of the market. They have chosen this firm before. A fixed price accepted by an incumbent relationship says less about winning new work than it appears to.
- One cycle is a behavioural clock, not an outcome clock. It will show whether clients accept. It will not show whether the economics hold over two years. The outcome row stays open with a named owner and a date to backfill it, exactly as Chapter 7 requires.
Reading a result that does not cooperate
The discrimination table assumes the world returns one of six clean answers. Sometimes it does. More often it returns something the table did not anticipate, and how a firm handles that determines whether it runs a second probe or quietly gives up on the discipline.
Three cases, each with a rule.
The split result
Clients in the strong practice area accept the fixed price and hold margin. Clients in the contested area demand the compression back. Two areas, two answers.
The tempting reading: "it works where we are strong" — which sounds like a finding and is a restatement of the input. The firm already knew it was strong there.
The rule: a split result is not a weak signal on the main question. It is no signal on the main question, and a strong signal that the two areas differ in a way the counterplay analysis did not model. Chase that difference rather than averaging the result — the unmodelled variable is now the most interesting thing on the page, and it is usually the scarce complement: something the firm holds in one area and not the other.
The unanticipated response
Clients accept the fixed price and then behave differently in a way nobody predicted — they start bringing more scope to each engagement, or they begin asking for the pre-built input as a standalone purchase.
The rule: an outcome outside the table is a better result than a row in it, because it means the world contained a mechanism the counterplay missed. Record it in full, resist explaining it immediately, and open a new row. The one thing not to do is force it into the nearest existing row so the table can be marked complete.
The genuinely ambiguous result
Two clients accept, one renegotiates, one delays the decision past the cycle. Nothing is eliminated and nothing is confirmed.
The rule: the honest state is continue collecting, with an end date, and the honest record is that the probe was underpowered — two practice areas and one cycle produced too few observations to separate the branches. That is a design finding worth having, and the correct response is usually a second, differently-shaped probe rather than a longer version of the same one. What it must not become is a row that stays open indefinitely because nobody wants to say the money produced nothing.
Notice what all three rules have in common. None of them permits the result to be interpreted toward the sponsor's preference, which is the single failure that turns a probe back into a demonstration after the fact. The discrimination table's real function is not to predict the outcome. It is to constrain the interpretation afterwards, when the firm is under pressure to find the reading that lets it keep going.
Six ways this can end
And what each obliges the firm to do next. The governance of these states — how many can be live, who decides, how they are reported — is Part IV's subject; here they are simply the honest set of endings for this row.
| End state | What it obliges |
|---|---|
| Scale | Move the remaining two practice areas; commission the successor-unit design as a separately funded piece of work. |
| Kill | Record customer capture as the live explanation; open a new row on where the scarce complement actually sits; keep the two instruments. |
| Stand pat | A legitimate scored outcome, not a euphemism: the evidence supports no change to the current unit this cycle, and that is written down as a finding. |
| Defer to a trigger | Name the observable, assign the watcher, and mark the row as carried on inference until it fires. |
| Continue collecting | Only with a stated reason and an end date — this is the state that silently becomes permanent. |
| Re-cut the row | The boundary case was wrong or has been overtaken. Close it and open the replacement, with the reason recorded. |
Notice that scale is one of six and not the goal. A portfolio in which scale is the only celebrated ending will produce probes designed to be scaled, which is where this whole discipline quietly turns back into a business case. The parent framing is exact: what is being bought here is the right, but not the obligation, to make a better-informed investment later — and firms that only celebrate exercised options will pressure their people into recommending builds .
What this book is not claiming
It is not claiming to know how the probe resolves. That would be the same error the whole book is written against — asserting an answer that only the world can supply.
What has been demonstrated is the design: a probe small enough to fund, real enough to matter, with a table showing what every outcome eliminates, triggers written in advance, confounders named at design time, and two instruments that survive a kill. That is the object. Running it is the firm's job, and the answer belongs to their market rather than to this book.
Meanwhile, in the same firm, in the same quarter, a second workstream was running. It cost about the same. Everybody enjoyed it considerably more. It produced fifty futures, a genuinely good discussion, and a board narrative that survived contact with the partnership.
It eliminated nothing, and the comparison is the most useful page in this book.
The Exercise That Killed Nothing
The other workstream, narrated fairly. It cost about the same as the probe, everybody enjoyed it more, and it was genuinely good work. Then the audit.
Six weeks. Two facilitated sessions, a great deal of AI-assisted generation between them, and three partners giving it more attention than they gave anything else that quarter.
Fifty candidate futures produced. Clustered into nine themes. Argued down to three that the partnership considered genuinely consequential. One combined into a board narrative that ran to fourteen pages and survived contact with eleven partners, which is not nothing.
And the futures were good. This has to be said plainly, because the chapter only works if a reader recognises their own best exercise in it rather than a straw one. They were plausible. They were internally consistent. They were argued from mechanisms rather than asserted from headlines. Several of them have since visibly begun. Nobody was lazy, nobody was posturing, and the discussion in the second session was the best the partnership had had in four years.
Now audit it.
Three questions, one scenario at a time
For each of the three surviving scenarios: what would have had to be true for it to be false; what observation would have distinguished it from its neighbour; and what the firm did differently on the Monday.
Scenario B — "The verified advisor"
The scenario, as written: as generation costs collapse, clients increasingly produce their own first-pass analysis but cannot verify it. Trust in machine-produced conclusions becomes the constraint. Professional firms reposition from producers of analysis to verifiers of it, and the profession's value migrates toward accountable sign-off. Firms with strong reputations and credentialed staff are advantaged; firms competing on throughput are exposed.
What would have had to be true for this to be false? Nobody asked. Reading it now, at least three things: that verification stays expensive; that clients want it from an external party rather than building it internally; and that credentials transfer to machine-produced work. None was stated, so none was checked.
What observation would have distinguished it from Scenario A? Nothing in the document. Scenario A described clients insourcing routine work and buying only complex judgement — which produces the same observable world for this firm, from a different mechanism. Both predict smaller scopes and more senior-weighted engagements. They are distinguishable only by who asks for the sign-off, and nobody wrote that down.
What did the firm do differently on the Monday? Nothing. The scenario was compatible with continuing exactly as before while feeling better informed about why.
The other two produce the same three answers, and it is worth walking them rather than asserting the pattern — because the ways they fail are different, and a reader auditing their own exercise will meet all three.
Scenario A — "The insourced routine"
The scenario, as written: clients progressively bring routine analysis, reporting and documentation in-house as their own tooling improves. External spend concentrates on complex, novel and politically sensitive work. Professional firms shrink in headcount but rise in average seniority; the pyramid flattens into something closer to a partnership of specialists.
What would have had to be true for this to be false? That clients find internal production harder than expected — because of data access, accountability, or simply nobody owning it. Plausible, checkable, and unasked. The scenario asserted the client's capability without asking what would stop it.
What observation would have distinguished it from Scenario B? None available. Both predict smaller scopes and more senior-weighted engagements; they differ only in why, and the why does not show up in any number the firm collects. Two scenarios that predict the same observables are one scenario with two explanations attached, which is precisely the situation a discriminating probe exists to resolve.
What did the firm do differently on the Monday? Nothing. The implied action — move upmarket — was already the stated strategy, had been for three years, and appeared in the previous two board packs. The scenario ratified an existing direction, which feels like confirmation and is structurally indistinguishable from having learned nothing.
Scenario C — "Bifurcation"
The scenario, as written: the market splits. At one end, cheap tooling absorbs the commodity layer entirely and price becomes the only variable. At the other, premium assurance and accountable execution command higher rates than the category has ever paid. The middle — competent generalist delivery at a defensible day rate — disappears. Firms must choose an end.
What would have had to be true for this to be false? That the middle persists — which happens whenever switching costs, relationship inertia or procurement convenience are large enough to sustain it. That is an empirical question about this firm's buyers, and there are eleven of them who could have been asked.
What observation would have distinguished it from A and B? This one is actually distinguishable, and that is what makes it the most frustrating of the three: it predicts something the others do not — rate dispersion widening across the market, with the middle of the distribution thinning. That is measurable from public rate data and from the firm's own win/loss record on price. Nobody measured it, because the exercise had no step at which a scenario was asked to name its own observable.
What did the firm do differently on the Monday? Nothing. The implied action — protect the brand and move to the premium end — was, again, what the firm already believed it was doing. Note the pattern across all three: each scenario recommended the thing the partnership was most comfortable with. That is not a coincidence, and it is not dishonesty. It is what happens when a generator is asked for plausible futures by people whose sense of plausibility was formed by the current strategy.
Three scenarios, three distinct failures. One had no falsifier. One was observationally identical to its neighbour. One had a genuine observable that nobody asked it for. And all three recommended what the room already intended.
Key Insight
Every scenario was compatible with every present decision. Which is not a criticism of the scenarios — it is the exact condition under which the search should have stopped.
That condition has a name in the parent book, and it names the point at which a search is finished: a search ends when the next candidate future would not change any present decision .
Applied honestly to this exercise, the stopping line was crossed somewhere around future number twelve. The team then produced thirty-eight more, clustered them elegantly, and presented the result as thoroughness. Every future after the twelfth was decision-neutral on arrival, and the process had no instrument capable of noticing.
Could they have caught it in the room?
Yes, and cheaply, which is the part worth taking away. The stopping rule is not a post-mortem instrument; it is a question you can ask live, and it takes about forty seconds per candidate future.
"If this were true, what would we do differently on Monday?"
Ask it of every future as it is generated, before it goes into a cluster. Three things happen. Most futures produce "nothing" — and those go straight into a discard pile rather than into a theme, which is where the process currently loses its discipline, because clustering makes weak candidates look substantial by association. A handful produce an actual answer, and those are the ones worth the room's attention. And a small number produce an argument, which is the most valuable outcome of all: two people disagree about what the firm would do, which means there is a live question underneath and it has just identified itself.
Run that filter on the fifty and the exercise would have finished in two days with eight candidates, three of which had a Monday answer and one of which started an argument. The six weeks bought the other forty-two, and the six weeks is not really the cost — the cost is that clustering fifty futures into nine themes produces a document so coherent that nobody thinks to ask whether any of its parts survive contact with a decision.
Note that this is not a rule about how many futures to generate. Generate five hundred if generation is cheap — it is. The rule is about what gets carried, and the filter costs nothing because it runs at the point of generation rather than at the point of review.
Side by side
Same firm, same quarter, roughly the same cost. One page.
| The scenario exercise | PRB-01, the fixed-scope probe | |
|---|---|---|
| Elapsed time | Six weeks to the board narrative | One contract cycle; design in two days |
| Cost | Facilitation, generation, three partners' attention | Comparable, plus margin genuinely at risk |
| Futures considered | 50 | 4 |
| Options eliminated | 0 | 1 rejected future at design; at least 1 explanation on any result |
| Evidence produced | None. No party outside the building did anything. | Behavioural: what clients did when the price changed |
| Residue left behind | A narrative, shared vocabulary, a fourteen-page document | An outcome-sizing instrument and a set of acceptance tests |
| What changed on Monday | Nothing | Two practice areas' pricing; one hiring assumption under review |
| What the firm can now do that it could not | Describe its exposure more articulately | Price an outcome; agree a definition of done; rule out one future in writing |
The exercise wins one row, and it is the row nobody should have been counting.
What the exercise was genuinely for
Now the part that stops this chapter being a hatchet job, and it matters more than the table.
Two of the nine clusters became boundary cases in the firm's portfolio — including the one Chapter 10 walked all the way through. The question what if competent first-pass advice becomes free did not arrive from nowhere. It arrived from cluster four of a scenario exercise, where three of the fifty generated futures shared a mechanism nobody had previously named.
Without the exercise, the firm would have aimed at the wrong question. Probably at the one everybody was worried about — whether a large competitor would enter their region — which is a contingency, not a pivot, and would have failed the gate in Chapter 10.
Worth being precise about how the useful clusters differed from the rest, because it is the only part of the exercise that transfers as a technique.
Cluster four contained three futures that shared a mechanism rather than a theme. Different surfaces — one about procurement, one about junior hiring, one about a competitor's pricing — and underneath all three, the same load-bearing assumption: that clients pay for the first pass. That shared mechanism is what made it a boundary case rather than a topic. The other seven clusters were grouped by subject matter: everything about regulation, everything about pricing, everything about talent. Subject clusters produce chapter headings. Mechanism clusters produce boundary cases.
Which suggests one change to how the exercise is run, and it is the only prescription this chapter offers about generation: cluster by what has to be true, not by what it is about. The nine themes would have become four or five, most of the surface variety would have collapsed, and the two mechanism-clusters would have been visible in week one rather than week five.
Which means the correct relationship between the two workstreams is sequential, not competitive. Generation feeds the portfolio; the portfolio closes rows; closed rows tell the next generation cycle where to look. A firm that reads this chapter and cancels its scenario programme has read it exactly backwards, and will find itself with a fast evidence clock pointed at nothing in particular.
The fair defence, and the narrow charge
Scenario planning has a research literature and it is more careful than most of its practitioners.
So the charge here is narrow and should stay narrow. It is not that scenario work is worthless; its own literature says specific techniques show impact, and this chapter's own specimen produced the boundary case that made the probe possible. It is not that the futures were wrong; several were right.
The charge is that an exercise which eliminates nothing has produced no evidence, whatever it produced in insight. That is a claim about evidence, not about value — and the distinction is the reason the two workstreams belong in different portfolios with different scoreboards.
The rule underneath
There is a version of this failure one level down, at the ledger rather than the portfolio, and it has a name: scenario theatre — a wall of futures, a satisfied room, and an unchanged budget. The prescription there is a single required column, the decision delta, and the rule that an entry does not close until it is non-empty .
This is its portfolio-level twin, and the rule generalises in one sentence:
Takeaway
The artefact is never the problem. The absence of a closing mechanism is.
Part III is finished. The firm now holds one designed probe, one future rejected in writing, two reusable instruments it did not have in January, three carried questions with dates on them for the first time, and a diagnosis showing a market that added twenty-two consequential moves while the firm eliminated nothing.
And nowhere to put any of it. The strategy pack cannot hold it — the pack is a document, and what the firm now has is a set of live positions with money attached and triggers pending. The operating plan cannot hold it either, and the reason the operating plan cannot hold it is the subject of the next chapter.
Two Portfolios, Two Ledgers
A single portfolio always collapses into the operating one, because the operating one has metrics and metrics win. The separation is the only structural defence option work has.
Watch a probe die. It does not take long and nobody has to be against it.
The quarterly investment review has one page. On it, in the same table, sit an operating initiative and an option row. The operating initiative is a document-processing automation with a fourteen-month payback, a run-rate saving, a confident owner and a vendor quote. The option row is PRB-01: two practice areas repriced, margin genuinely at risk, a stated probability of being killed, and a residue — an outcome-sizing instrument, some acceptance tests — that nobody can value because nothing like it has been valued before.
Nobody argues against the option row. Somebody asks what the return is. The honest answer is "we will know less wrong things", which is true, correct, and unsayable in a budget meeting. The row gets deferred to next quarter — not rejected, deferred, which feels like courtesy. Repeat that four times and the option portfolio is empty; repeat it for two years and the firm has concluded, without ever deciding, that it does not do this kind of work.
The comparison was made on the operating portfolio's terms, and on those terms the option row loses every time, correctly. Which is why the answer cannot be advocacy. It has to be structural.
Two portfolios
The operating portfolio
Contents: workflow efficiency, quality, cost reduction, staff augmentation.
Funded from: OPEX.
Governed by: ordinary operating metrics — payback, run-rate, adoption, error rates.
Good looks like: the work gets done faster, better, or with less rework.
Never described as: transformation. This is the discipline that costs nothing and changes the most, because the mislabelling is how a board comes to believe it has a terminal-value strategy when it has a cost programme.
The terminal-value option portfolio
Contents: asset conversion, successor offers, customer-agent threats, new commercial units, governed execution capability.
Funded from: a separate allocation, sized deliberately, not from underspend.
Governed by: falsification throughput — options closed with evidence, per quarter.
Good looks like: a branch died and the firm knows why; or one earned the right to serious capital.
Failure looks like: rows added, nothing closed, everyone impressed.
Different sponsor, different ledger, different burden of proof, different page in the board pack. The last of those sounds trivial and is not — a portfolio reported alongside operating initiatives will be judged by operating metrics within two cycles regardless of what its charter says.
Where the productivity budget should actually go
This book has spent thirteen chapters arguing that productivity is not strategy, and it now owes the other half of that argument — because there is an enormous amount of genuinely valuable work that this framework says to do immediately, with no boundary case, no probe and no ledger entry at all.
The test that sorts it is three conditions, and it belongs to the parent book:
Key Insight
Cheap cognition reduces uncertainty when the population is enumerable, the evaluation function is stable, and the action set is closed. Break any one of the three and you are in the paradox.
The clearest case is audit, and it is worth walking because it makes the distinction physical. Auditing has sampled for its entire professional history — not because anyone believed sampling was a good way to find problems, but because examining everything was impossible at human cost. An elaborate, defensible apparatus grew up around that constraint: materiality thresholds, risk tiers, sampling schedules, confidence intervals. All of it downstream of the fact that a person had to read each item.
That constraint is dissolving, and when it does, uncertainty falls. The population of transactions is finite and enumerable — you can count what must be examined before you start, and the count does not change because you got better at examining. The test is stable: a duplicate payment is a duplicate payment, and it does not become something else because a competitor changed their pricing. And more thinking converges: each additional unit of cognition reduces the unexamined remainder rather than creating new categories of thing to examine .
Sampling was never a methodology anyone loved. It was a budget wearing a methodology's clothes. Dissolve the budget and the uncertainty falls with it.
Now run strategy through the same three conditions and it fails all three. The population of futures is not enumerable. The evaluation function is not stable, because what counts as a good strategy is defined partly by other people's moves — which is what makes it strategy rather than optimisation. And the action set is emphatically not closed: every discovery creates new kinds of option.
The same technology that clears an audit thickens a strategy, and it is not behaving inconsistently. The questions have different structure.
Three artefacts, three jobs
A reader who knows this body of work will by now be holding two other instruments and wondering whether this is a third or a replacement. It is a third, and the division is clean enough to state once and never revisit.
| Artefact | Governs | An entry closes when… |
|---|---|---|
| Question Ledger | Search quality — was the space actually searched, and where are the gaps? | The search is inspectable and its gaps are named. |
| Strategic Search Ledger | Decision consequence — did anything actually change? | The decision-changed column is non-empty, with an owner and a date. |
| Option portfolio | Capital and disposition — what did we pay to learn, and what died? | The option reaches an explicit state and the residue is booked. |
What the option portfolio adds that neither predecessor carries is specific: money, a trigger, and a compounding residue. A ledger row can be excellent and cost nothing. A portfolio row has a probe attached with a price on it, a pre-committed condition that would kill it, and a named asset that survives the kill.
Not a warehouse of clever scenarios, but a capital-allocation memory showing what was tested, rejected and reopened.
Three artefacts, three jobs. A firm that merges them builds one thing that does none of them — usually a spreadsheet with twenty columns that is filled in once and abandoned, because no single instrument can be simultaneously a research record, a governance minute and a capital account.
What only the board can do
My own version of why this matters is blunter: if the conversation at the top of the company is trinkets and workflow optimisation and making individual staff faster, there is no plan. That is right, and left there it would license exactly the wrong conclusion.
Why this is not an argument for top-down strategy
Because the board cannot supply the input the whole machine runs on.
Choosing which variable is live requires lived friction with the business — the specific knowledge of which assumption is load-bearing here, which comes from sitting with the situation rather than from having read about situations like it . That friction lives at the customer edge: in delivery failures, pricing objections, the exceptions that keep recurring, and the work staff quietly decline to do because it never works.
In this book's own worked case, the single most consequential observation — two of the last five prospects arrived holding their own baseline analysis — was known to two account leads and to nobody above them. It was not withheld. There was simply no route by which it could travel, because it did not arrive as a competitor, a lost deal or a complaint, and those are the three categories the firm had.
So build the route, and make it a mechanism rather than a value. One that works: each practice head brings two observations per quarter in a fixed form — "a client did something we did not expect" — with no interpretation attached and no requirement that it be significant. Two per head per quarter is small enough to actually happen and large enough that patterns appear within a year. Those observations are the raw material for candidate boundary cases, and they cannot be generated at board level at any price.
The board owns the wager. The operating edge keeps it honest.
Neither half works alone. A board with no friction supply attacks the assumptions it finds interesting, which are usually the ones it has read about. An operating edge with no board mandate generates observations that go nowhere, learns that they go nowhere, and stops.
That is the container. What goes in it is a row with eight fields — seven of which most firms can already half-fill from what they know, and one of which almost nobody has ever written down.
The Option Row
Eight fields. Seven of them a good firm can half-fill today. The eighth is the one that determines whether the portfolio ever kills anything.
Start at the end of the row, because the last field is the one that makes the other seven operable.
Field eight: the asset that compounds even if the option is never exercised.
Consider what happens without it. A probe runs. The result is unwelcome. The branch dies — which is the outcome this entire book has been arguing is valuable — and the ledger records a cost with nothing on the other side. The sponsor has spent money and has a dead idea to show for it. The next probe is harder to fund, because the last one "didn't work". Within two cycles the portfolio has quietly learned that killing is expensive and stops doing it, and nobody ever made that decision.
Now consider it with field eight populated. The branch still dies. And the firm still holds the outcome-sizing instrument and the acceptance tests it had to build to run the probe — two pieces of commercial machinery it did not have, which get used in the next probe and in ordinary delivery regardless. The kill cost the difference between the probe's price and the residue's value, which is a much smaller number and sometimes a negative one.
Key Insight
Field eight is not bookkeeping. It is what makes a kill affordable — and a portfolio that cannot afford to kill will stop killing within two cycles, without anyone deciding to.
The eight fields
Every row carries
- 1. The revenue unit being tested
- 2. The boundary case
- 3. The customer, competitor, regulator and platform counter-moves
- 4. The scarce complement the firm believes it owns
- 5. The smallest paid or deployed probe
- 6. The evidence received
- 7. The kill, scale and revisit triggers
- 8. The asset that compounds even if the option is not exercised
| Field | What goes in it | The tell of a bad entry |
|---|---|---|
| 1. Revenue unit | A unit, not a theme. A thing a customer buys, described the way an invoice would describe it. | It names a market ("our advisory business") or a capability ("data engineering") rather than a purchase. |
| 2. Boundary case | One sentence, pivot-qualified, no date. | It contains a year. Or the word "increasingly", which is a trend pretending to be an extreme. |
| 3. Counterplay | Four actors, their moves, and their moves against each other. | All four are moving decisively in the same direction — a counterplay written backwards from a conclusion. |
| 4. Scarce complement | A belief about what makes your cognition usable, stated so it could be false. | "The client relationship." See below. |
| 5. Smallest probe | The move, with its discrimination table attached. | No row of the table kills anything. |
| 6. Evidence received | Behavioural observations, including specified silence and recorded disagreement. | The field contains adjectives. "Strong interest", "positive reception", "encouraging". |
| 7. Triggers | Kill, scale, revisit — each observable, dated at creation, with an owner. | The trigger was written after the evidence arrived, and remarkably was not met. |
| 8. Compounding residue | A named asset that exists and is usable after a kill. | The field says "learnings". |
Two of those deserve more than a table row.
Field four is where portfolios lie to themselves
The scarce complement is the thing that makes your cognition worth buying once cognition itself is cheap. It is the most consequential field in the row and the most self-flattering, and a portfolio in which every row claims the same complement is a portfolio that has stopped thinking.
The test is one question: what would the customer have to do for this belief to be false, and has anyone looked?
Three candidate complements, tested
"The client relationship"
Falsified by: the client running a competitive process anyway. Which most of them do, most years, and which the firm already knows.
Verdict: not a complement. It is a preference that survives until price differences get large enough — which is precisely the condition the boundary case describes. It belongs in the row only if someone can name a procurement decision it actually changed.
"Proprietary context — we know their estate"
Falsified by: the client's own systems being able to describe the estate to a model as well as your consultants can. Increasingly cheap to check, and checkable this quarter.
Verdict: a real complement with a visible expiry. Worth holding as long as it is dated — and dangerous held as a permanent assumption, because it decays without announcing itself.
"The right to sign"
Falsified by: the requirement for an accountable signature being removed, or extended to parties who currently cannot provide one.
Verdict: a genuine complement where it exists, because it is conferred by someone other than the customer and cannot be self-supplied at any budget. Note that it is also the one most firms under-claim, because it feels like bureaucracy rather than advantage.
The pattern is worth extracting. A scarce complement is something a customer cannot obtain elsewhere at reasonable cost, which makes cognition usable. An asserted advantage is something the firm believes about itself. The first has a falsifier that someone could go and check this quarter. The second has admirers.
Field six, and what "evidence" excludes
Behavioural means somebody outside your building did something differently with money, authority or attention. That is the whole standard, and it excludes a great deal that currently gets recorded as evidence: a positive meeting, an enthusiastic reply, a well-attended session, a prospect who said the framing was exactly right.
It includes two things people often leave out. Specified silence — a named audience who could have responded and did not, where the expected response was written down first. And recorded disagreement — the client who said the offer made no sense to them, which is the single most information-dense thing that happens in most probes and the thing most reliably lost between the meeting and the write-up.
A completed row
OPT-01, from the composite firm walked through Part III. Every field populated, including the parts that are uncomfortable.
OPT-01 — first-pass advice
1. Revenue unit under test. The standard four-to-twelve-week analytical engagement, billed by the hour. Roughly 30 per cent of billed hours on a standard engagement sit in first-pass analytical scope.
2. Boundary case. Competent generic first-pass advice in our sector is effectively free to the client, delivered with tools they already own. No date attached.
3. Counterplay. Customer — partial substitution: keep the routine middle, buy the difficult residue; two of the last five prospects already arrived with their own baseline. Competitor — a weaker firm passes the compression through as a weapon rather than absorbing it as margin. Regulator — guidance in two of four practice areas now names a person rather than a process, relocating scarcity into signature. Platform — no material capture in this segment yet; watch for buyers arriving pre-structured. Interaction — a competitor's published fixed price makes fixed pricing the default request, which pulls the regulator's answer into a market that has already changed.
4. Scarce complement we believe we own. Accepted evidence and the authority to sign — not analysis. Stated as attackable: falsified if clients begin accepting unsigned machine-produced analysis for consequential decisions, or if the signature requirement extends to parties who currently cannot provide one.
5. Smallest paid probe. PRB-01: two practice areas move to fixed-scope pricing at the next contract cycle, sized on outcome; first-pass scope removed from the billable estimate and supplied as a pre-built input. One strong area, one contested. Owner: managing partner. Discrimination table attached — six observable outcomes, two of which kill a branch outright.
6. Evidence received. To date: two of the last five prospects arrived holding their own baseline analysis (behavioural, dated). Delivery time on the last six engagements fell about 18 per cent and the difference was billed away (internal, behavioural). Pending: the probe result. Outcome row open — twenty-four-month economics, owner named, backfill scheduled.
7. Triggers. Scale if fixed-scope wins hold margin across a full cycle in both areas. Kill if clients in both areas systematically require the compression returned as a discount. Revisit if a competitor announces fixed-fee equivalents, or any client asks to be priced against their own AI-produced baseline. All dated at creation; owner, managing partner.
8. Compounding residue. The outcome-sizing instrument — a repeatable way to convert a client's desired end state into a bounded, priced scope. The acceptance tests — an agreed, written definition of done. Both reusable whichever way the pricing question resolves; both currently non-existent in the firm.
Future rejected. "Hold rate and volume, absorb the compression as margin for three to five years." Rejected: requires that no material competitor passes compression through and that clients cannot observe it. Both already false in the pipeline.
Confounders, named at design. Our own move changes competitor behaviour and the counterfactual is unavailable. Two practice areas is a small sample and a split result is no result. Existing clients are not a random sample of the market. One cycle is a behavioural clock, not an outcome clock.
Option state. Live; probe running; behavioural evidence expected at end of cycle; outcome evidence backfilled at twenty-four months.
Read that row as a document and it looks like admin. Read it as a position and it is something no strategy pack has ever contained: a named bet, with a price, a condition that would end it, an owner, an honest account of what could make its result meaningless, and a description of what the firm keeps if the whole thing fails.
The same row, a different industry
The fields do not change. Their weight changes completely.
Same eight fields. In the professional-services row, counterplay is dominated by customer and regulator and the probe is a price move. In the software row, counterplay is dominated by platform and competitor and the probe is a publishing move. A firm that copies another industry's row will get it wrong; a firm that copies the fields will not.
Where the row comes from
Not from strategy practice. It descends from an experiment-discipline schema built for a learning system — belief at risk, probe selected, observed response, behavioural evidence, confounders, branch killed, delayed outcome, next falsifying probe, disposition — and from the leftover checklist developed for creative experimentation, which insisted that a test leave behind an observation, candidate mechanisms, confounders, a falsifying next test and a promotion decision .
Two things were added in the transfer up to board altitude, and they are exactly the two things a learning system does not need and a capital instrument does: money (field five is a purchase, not an enquiry) and a named surviving asset (field eight, which no epistemic schema has because knowledge does not need a salvage value).
The row is the product. Everything before it is preparation; everything after it is governance.
Key Takeaways
- • Field eight is what makes killing affordable; without it a portfolio stops killing without deciding to.
- • Field four is where portfolios flatter themselves — a scarce complement has a falsifier, an asserted advantage has admirers.
- • Field six excludes adjectives and includes recorded disagreement.
- • The fields are industry-independent; their weight is not.
- • A row is a position with a price, not a document.
Which leaves the question the row cannot answer about itself. A row that never closes is a row that never existed — and closing turns out to be a governance problem with failure modes of its own.
States, Triggers and the WIP Cap
A row that never closes is a row that never existed. Closing is a governance problem, and the cap is the discipline that makes the rest of it work.
The portfolio in its second year. Thirty-one rows.
Every field populated. Boundary cases that pass the pivot test. Counterplay written out for four actors. Probes specified. Triggers written. A non-executive director has described it as the best strategy documentation she has seen from a firm this size, and she is right.
None of them has closed.
This is Chapter 1's scene with better paperwork, and it is the failure this book is most likely to cause. A firm that adopts an instrument and reproduces the disease inside it is worse off than one that never adopted it, because the instrument now provides cover: we have a rigorous process, and rigour is exactly what the thirty-one rows demonstrate.
Six ways a row can end
Every live option is in exactly one of six states, and each state is a claim with evidence behind it and consequences attached.
| State | What justifies it | What it obliges next |
|---|---|---|
| Kill | A discrimination-table outcome that eliminates the branch. | Book the residue. Record the reason in a form someone could disagree with. Open the successor question if one exists. |
| Stand pat | The probe returned and supports no change to the current unit this cycle. | Write it down as a finding, not as an absence, and set the revisit trigger. |
| Defer to a trigger | No probe exists inside a horizon that matters. | Name the observable. Assign a watcher. Label the row as carried on inference, visibly. |
| Continue collecting | Partial evidence, with a specific gap. | A stated reason and an end date. This is the state that silently becomes permanent. |
| Exercise into a bounded proof | The scale trigger fired. | Separate funding, separate governance, explicit hand-off out of the portfolio. |
| Transfer to operations | It works, it is ordinary, it no longer tests anything. | Remove it from the option portfolio entirely. |
Two of those are routinely mishandled in opposite directions.
Transfer to operations gets forgotten, and the effect is a portfolio slowly clogged with successes. An option that has been exercised, proved and normalised is no longer an option — it is a business line, and it belongs in the operating portfolio where it can be judged on operating metrics. Leaving it in place inflates the WIP count with things that are not being tested and crowds out the ones that should be.
Stand pat gets read as weakness, which is the more damaging error and deserves its own defence.
Triggers, and the quiet corruption
Three rules, and the third is the one that matters.
- Observable by someone other than their author. "If the market moves against us" is not a trigger. "If two or more competitors in this category publish fixed prices" is.
- Dated at creation. Not dated at review, which is the same thing as not dated.
- A changed trigger is itself a decision requiring a reason.
Be clear about what the third rule is not. Triggers legitimately need changing — the world moves, a threshold turns out to be wrong, a competitor does something that makes the original observable meaningless. Refusing to revise triggers would be a different kind of foolishness.
The corruption is not the change. It is the change made silently, after the evidence arrives, in the direction that avoids a kill. It never looks like dishonesty from the inside; it looks like a sensible refinement in light of new information, proposed by the person with the most context, which is the sponsor.
The cap
Cap the number of active options. Otherwise the mandate rewards intellectual WIP and ensures that nothing reaches evidential closure.
The precedent from the parent book is specific: three to five live rows is a working ledger; a firm running twenty is doing research rather than governance and will close none of them . The same number holds one altitude up, and for a reason worth stating rather than asserting.
Key Insight
Adding a row is free. Closing one is not. A cap is the only mechanism that makes adding cost something.
That asymmetry is the whole mechanism. Under no cap, a board never has to rank its questions — any interesting one can be added, and the cost of adding is deferred indefinitely to the people who will eventually have to resolve it. Under a cap, adding requires displacing. Which question comes out? And that forced comparison is the only moment at which a board is obliged to say, out loud, which of its uncertainties actually matters most.
Firms find that conversation uncomfortable, which is precisely why the cap has to be structural rather than aspirational. A cap that can be exceeded "just this quarter" is not a cap; it is an intention, and intentions are exactly what the fifth column of a ledger fills up with.
Open rows are honest rows
Which portfolio is healthier?
Three rows, all open
- • Everyone can see what each is waiting on
- • Each has an owner and a date
- • Two are carried on inference, and say so
- • The board knows exactly what it does not know
Twelve rows, all closed
- • Each closure reads "monitor", "explore", "continue to assess"
- • No evidence is attached to any of them
- • Nothing was displaced by any of them
- • The board believes it has resolved twelve questions
Openness is not failure. False closure is — and it is worse than an open row, because an open row knows it is unresolved.
The practice that makes this real is small and specific: open rows go in the board pack. Not as an appendix of unfinished business, but on the same page as the closures, with their age and what they are waiting on. A pack that shows only closures teaches everyone that closure is what gets reported, and the fifth column fills up with intentions within two cycles.
Knowing when to stop opening
The parent book supplies the back-end rule: a search is finished when the next candidate future would not change any present decision . Run until the marginal future is decision-neutral, then close.
At portfolio level there is a cheaper corollary, and it works at the front end rather than the back: an option whose every possible resolution leaves the capital plan unchanged should never have been opened.
Test it before the row is created. Take each state the row could reach — kill, stand pat, scale — and ask what the firm would do differently in each case. If the answer is the same in all three, the row is decorative, and you have saved a quarter of somebody's attention by not opening it.
The quarterly page
One page, once a quarter
Prepared:
- • Both clock counts — branching inputs by route, evidence outputs by type
- • Every state change since last quarter, with the evidence that caused it
- • The WIP count against the cap
- • Open rows, with their age and what each is waiting on
- • Any trigger changed, with its reason and whether evidence had already arrived
Asked:
- • What did we eliminate, and what killed it?
- • Which consequential uncertainty are we forcing the world to answer next?
Closed:
- • At least one row — or an explicit statement of why none, which is itself a finding
All of this treats the portfolio as a management instrument — a way of running a search so that it terminates. It is also something else, and the something else is what turns this from a good practice into an obligation.
One of the terms in your firm's valuation is sitting in this portfolio, unaudited.
Underwriting Zero
"Our terminal value is zero" is usually a statement about the quality of a firm's search. Decomposed into four terms it becomes auditable — and one of those terms is the portfolio.
The sentence that started this entire line of work was mine, and it was blunt:
Everyone knows the terminal value's zero now. You can't divest.
It sounds like a calibrated forecast. It is not one, and I am the one saying so.
A genuine forecast requires four things: a time horizon, an industry base rate, a competitive model, and a probability distribution over outcomes. None of them was present. What was present was a strong intuition arriving immediately after an unsuccessful attempt to describe the firm's future — which is diagnostic of something, just not of the thing it was taken to diagnose.
Key Insight
A company's inability to articulate a satisfying AI future is evidence of weak strategic search. It is not evidence, by itself, that the company's terminal value is probably zero.
Which is a considerably more actionable finding, because search quality is something a firm controls and a terminal value is not.
Zero as a stress condition
Zero is not the conclusion. It is the stress condition that forces management to reveal whether it owns anything transferable.
The difference between a forecast and a stress condition is not rhetorical, and the analogy from banking is exact. A bank running a stress test is not forecasting a market collapse. Nobody in the room believes the scenario is likely; that is not what it is for. The exercise demonstrates survival under it — and what it produces is not a probability but a capability statement.
Apply the same posture. The question is not will our terminal value be zero, which invites an unresolvable argument about likelihood in which the optimists and pessimists both have good material. The question is: if the scarcity behind our current unit of sale disappears, which assets survive the destruction and can be migrated into whatever remains scarce?
That question has an answer, the answer is auditable, and producing it takes a quarter rather than a philosophy.
Four terms, not one scalar
The decomposition
future value = legacy runoff + convertible-asset value + successor option value − stranded liabilities
Four terms behave differently, decay at different rates, and are audited by different questions. Treating them as one number is how a board ends up with a mood instead of a position.
| Term | What makes it real | What makes it decorative |
|---|---|---|
| Legacy runoff The existing book, delivered excellently, at a declining and increasingly variable cost base. |
Known contract durations, known concentration, a cost base that actually varies with volume. | An assumption of renewal at historic rates, which is the thing the boundary case was about. |
| Convertible assets Context, evidence, relationships, methods, tests, exception patterns. |
Someone can name a conversion that has already happened and the engagement that used it. | "Our methodology." Value here is strictly conditional on conversion actually occurring; unconverted, these ride the runoff curve to zero. |
| Successor option value The bets on what the firm becomes. |
Rows with probes attached, triggers written in advance, and a closure history. | Enthusiasm, pilots, a transformation programme, a slide titled "our AI future". |
| Stranded liabilities Leases, restructuring costs, commitments that outlive the revenue. |
Provisioned, with numbers. | Described as "manageable" — which is the word boards use for the term they have not computed. |
Declining value and live cash are not contradictory
The professional-services companion works this decomposition properly, in its own vocabulary, and supplies the evidence that the first term is real even in categories everyone has written off .
Dead categories, live cash
IBM's mainframe business posting its highest annual revenue in two decades — sixty years into the category18
Average US COBOL developer salary — above the median software developer, in a language routinely declared dead19
What the Notes/Domino installed base sold for, two decades after losing its category war20
None of that contradicts the doctrine. It prices the harvest half of it, and it corrects a specific error that follows from treating terminal value as a single number: "zero terminal value" does not mean "nobody can buy the business." A consolidator or a management buyer can rationally acquire a declining firm at a harvest price. What no informed buyer will pay for is the pretence that the old unit recovers.
The term sitting in your portfolio
Here is what this book adds, and it is the reason this chapter exists rather than being a pointer to the companion.
The successor-option-value term is the terminal-value option portfolio. Not analogous to it. Not supported by it. It is the same object, viewed from the valuation side instead of the management side.
Follow what that implies. An option is worth something only if it can be exercised on evidence — that is what distinguishes an option from a wish. Exercising on evidence requires a mechanism that produces evidence. A firm with no probes has no such mechanism. Therefore a firm with no probes is carrying a successor-option term whose value cannot be established by any procedure it currently operates.
A partnership carrying a large successor-option term with an empty portfolio is carrying an unaudited asset.
And the same property holds one term to the left. Convertible assets are worth something strictly conditional on conversion actually happening; unconverted, they ride the runoff curve to zero. Which means two of the four terms in your firm's valuation are functions of whether you run this discipline — not correlated with it, not improved by it. Functions of it.
That is what turns a strategy practice into a board obligation. A board that does not know its probe count does not know the value of two of its four valuation terms, and is therefore reporting a number it cannot defend.
Option allocation, not prediction
The relieving consequence of the decomposition is that nobody has to know the answer in advance.
Four questions a board can ask cold
One per term. None of them requires preparation, and the quality of the answer is the finding.
- Runoff: what is the contracted duration and client concentration of the book we are running off, and how much of the cost base varies with it?
- Convertible: name one asset converted in the last twelve months, and the engagement that used it without the person who built it.
- Successor options: how many rows have a probe attached, and how many closed last year with the evidence that closed them?
- Stranded: what have we provisioned for, and what have we been calling manageable?
Question two is the one that produces silence, in the authors' experience of asking it. "Our methodology" is not an answer; "our people's experience" is the opposite of an answer, since it describes an asset that walks. A firm that cannot name one conversion in a year is carrying its second term at an assumption.
One adjacency, and then this part is finished. Whether the firm survives long enough to exercise any of it is a different inequality — harvest cash against time-to-proof — and it belongs to the professional-services companion, which makes it the whole of its management chapter . This book governs what the bets should be. That one governs whether you last long enough to settle them.
The instrument now exists, and it is auditable — which means it can be gamed, misapplied and hollowed out. The ways it will fail are specific enough to name, and naming them is what the rest of the book is for.
Six Ways This Goes Wrong
The most likely failure is not the firm that never adopts the portfolio. It is the firm that adopts it and hollows it out — and that failure is worse, because the artefact now provides cover.
Picture a portfolio that satisfies every rule in Part IV.
Five rows, respecting the cap. Every field populated. Boundary cases that pass the pivot test. Counterplay written for four actors each. Probes specified with discrimination tables. Triggers written and dated. A quarterly page, prepared on time, presented cleanly.
Then read the state column. All five say continue collecting.
Nothing has been violated. Every rule has been followed. And the instrument has been adopted and inverted, which is a more sophisticated failure than never adopting it and considerably harder to see — because everyone involved can point at their compliance.
Six ways this happens. Three are inherited from the ledger discipline one level down and are listed with their tells rather than re-argued; three are new to the portfolio level and get more space.
1. Portfolio theatre (inherited)
Tell: the state column reads "monitor", "explore", "continue to assess", "keep a watching brief".
Mechanism: these states are true of everything at all times, which makes them permanently defensible. Nobody can be criticised for monitoring.
Fix: the row stays visibly open with its age displayed, open rows go in the board pack, and "continue collecting" requires an end date at the moment it is assigned.
2. The probe that cannot lose (new)
Tell: every outcome in the discrimination table is consistent with proceeding. The table exists, it is filled in, and no row eliminates anything.
Mechanism: the sponsor writes the table, and the sponsor wants the probe to run. This is not dishonesty — it is the completely ordinary human difficulty of enumerating the outcomes that would embarrass you, about a thing you proposed.
Fix: write the kill map before funding, and have someone other than the sponsor read it with one question: which row here kills something? It takes two minutes and catches most instances, because the failure is visible the moment a second person looks.
3. Heat mistaken for evidence (inherited)
Tell: field six contains adjectives — strong interest, positive reception, encouraging conversations.
Mechanism: enthusiasm arrives faster than behaviour and feels more like a result, especially when the people bringing it are senior and pleased.
Fix: behavioural evidence only, with an explicit "what did they actually do" line. Money, authority or attention moved, or it did not happen.
4. Retrofitted triggers (new)
Tell: the kill trigger was written after the evidence arrived and, remarkably, was not met.
Mechanism: motivated refinement, proposed by the person with the most context — who is also the person with the most invested. It always arrives as a reasonable observation about why the original threshold was too crude.
Fix: the trigger change log from the previous chapter. Old trigger, new trigger, date, reason, whether evidence had already been received. Nothing is forbidden; the pattern simply becomes visible.
5. Over-search (inherited)
Tell: rows are being added faster than options are being closed. Track that ratio; it is the metric that matters.
Mechanism: generation is cheap and closure is not, so any unconstrained system drifts toward opening. This is the failure mode a good firm reaches.
Fix: the cap, and the stopping rule — a search is finished when the next candidate future would not change any present decision .
6. Wrong altitude for the instrument (new)
Tell: falsification throughput is being demanded of a bounded question. A team is designing a discriminating probe for something it could simply go and count.
Mechanism: a discipline adopted enthusiastically gets applied everywhere, and this one has the additional hazard of sounding rigorous wherever it is used.
Fix: the next chapter, which is entirely about where this method does not apply. And note that the cost here is real rather than merely aesthetic: it delays work that should have been flooded with cognition six months ago.
Where each one comes from
Inherited from the ledger discipline
- • Portfolio theatre — the states that commit nobody to anything
- • Heat mistaken for evidence — the fluent answer, accepted
- • Over-search — the failure a working apparatus produces
Named and diagnosed in the parent book. Listed here with their tells; the arguments are there.
New at portfolio level
- • The probe that cannot lose — because probes cost money and sponsors write tables
- • Retrofitted triggers — because triggers are pre-commitments and pre-commitments can be edited
- • Wrong altitude — because this instrument is more expensive to misapply than a ledger is
All three appear only once real money and real pre-commitment enter the process.
What is underneath all six
Bottom Line
Nobody owns the terminal-value question.
Efficiency reports to operations. Delivery reports to the practice leads. Pricing reports to the commercial director. Client relationships report to the partners. And the question of whether the unit survives — the one this entire book is about — reports to nobody.
So the portfolio becomes whoever's side project it is, reviewed by people whose day jobs reward continuity, and side projects close rows with intentions. All six failure modes are downstream of that one structural gap, which is why fixing them individually does not work: each fix is a rule, and rules with no owner decay.
The fix is not a new committee. It is three things:
- A named owner with the standing to displace a row. Not to advocate for one — to take one out. That authority is what makes the cap real, and it cannot be delegated to someone who has to ask permission.
- A cadence in the calendar. Quarterly, on the standing agenda, with the page from Chapter 16. Not "when we next review strategy".
- A board that asks for the portfolio before it approves the strategy. One sentence at the top of the discussion: where is the portfolio, and which rows moved since last quarter? This costs nothing and changes what gets prepared, which is the only reliable way to change what gets done.
People will game the closure column
Of course they will — by recording a decision that was going to happen anyway and attributing it to the portfolio. It is the path of least resistance and it is usually not even cynical.
One counter-test catches nearly all of it: a closed row should be traceable to something the firm would otherwise have done. Name what it displaced. If it cannot name anything, it is a record of the status quo with a new label.
The counter-test, applied to the two closures from Chapter 9
"We will focus on our core sectors rather than diversifying."
What did it displace? Nothing. There was no live diversification proposal, no budget
attached to one, and no partner advocating it. The firm was not diversifying before the decision
and is not diversifying after it. Fails. The row records the status quo and
dresses it as a choice.
"We will monitor developments in agentic procurement."
What did it displace? Nothing, and it also fails three other tests: no owner, no named
observable, no trigger that would reopen it. It appeared in three consecutive board packs in
identical words, which is the clearest possible evidence that nothing was ever going to happen
as a result of it. Fails.
A row that cannot name what it displaced is a record of the status quo with a new label.
Both of those closures were made in good faith by capable people, which is the point. Nobody set out to inflate a column. They wrote down what the firm had concluded, and what the firm had concluded was what it already believed.
And the failure one floor down
My own version, which applies to the grassroots case as much as the boardroom one: grassroots, ground-up, no strategy — which is really no planning. Fail to plan, plan to fail.
A hundred people using AI well, reported upward as adoption metrics, is a genuine capability and worth having. Presented to a board as a strategy, it is portfolio theatre one floor down: an impressive artefact, honestly produced, that answers a question nobody's valuation depends on.
Five of the six failures above are failures of discipline, and discipline is a solved problem once someone owns the question. The sixth is different. It is a failure of scope — using the instrument where it does not belong — and that raises a question this book has been deferring since Chapter 2.
Where does none of this apply?
Where This Method Does Not Apply
A method that cannot say where it does not apply is a sales pitch. Here is the test, the markets where this bites weakly, the falsifier for the central claim, and what the evidence base does not contain.
Straight answer, in the register the parent book used for the same move: the test, not the thesis, tells you which world you are in.
The boundedness test
Cheap cognition reduces uncertainty when three conditions hold together:
- The population is enumerable. You can list the things to be examined before you start, and the count does not change because you got better at examining.
- The evaluation function is stable. You know what a good answer looks like, and nobody else's move redefines it.
- The action set is closed. Finding more does not create new kinds of action; it changes which of a fixed set you take.
Break any one and you are in the paradox — where more cognition produces more options. Hold all three and there is nothing to discipline: go and flood it.
Run four questions from an ordinary firm through it, because the test is only useful if you have watched it discriminate.
| Question | Enumerable? | Stable evaluation? | Closed actions? | Verdict |
|---|---|---|---|---|
| "Which of our contracts contain a clause that becomes unenforceable after the rule change?" | Yes | Yes | Yes | Bounded — flood it. No boundary case, no probe, no row. |
| "Which accounts are at churn risk?" | Yes | Mostly | Yes | Bounded — flood it, with care. "At risk" drifts with context; anything scoring "mostly" deserves a second look before a large budget follows. |
| "Should we build an agent-addressable version of our service?" | No | No | No | Unbounded — discipline it. A portfolio row with a probe. |
| "What if our largest platform partner vertically integrates?" | No | No | No | Unbounded, and possibly untestable inside a useful horizon. Defer with a named trigger; label as carried on inference. |
Misapplying this is expensive, not merely untidy
Pointing falsification throughput at a bounded question is a category error with three costs, and the third one is the one that does lasting damage.
It delays work that should already be done — a census that could have been run last quarter sits behind a probe design. It imposes experimental machinery on a counting problem, which is expensive and produces nothing the count would not have produced. And it teaches the organisation that the discipline is bureaucracy, which is the durable harm: the next time someone proposes a portfolio row on a genuinely unbounded question, the room remembers the four months spent designing a discriminating probe for something they could have simply looked up.
Markets where this bites weakly
Do not argue with a reader in that position. Hand them the test. A firm that runs it and finds itself in the bounded world has had full value from this book — the finding is that its productivity budget should be spent aggressively and its strategic search should be small — and it should go and spend accordingly.
One nuance, because the clean version overstates it. Almost every firm in those categories has one or two genuinely unbounded questions, usually about the layer above or below its protected position: what happens if the licence regime changes; whether the physical asset's output becomes commoditised by someone else's cheap cognition. The right posture is not "no portfolio". It is a very small one, honestly scoped, which is a much easier thing to sustain than a large one.
Key Insight
The test, not the thesis, tells you which world you are in.
What would show this book is wrong
The central claim of this book is that falsification throughput is the scarce strategic capability — that a firm which raises its rate of evidenced elimination will navigate better than one which raises its rate of generation, holding capital constant.
That claim is an argument, not a finding. It has not been measured, and this is the point at which a book either says so or quietly hopes nobody asks.
If firms that count kills navigate no better than firms that do not, holding capital constant, the thesis fails.
The evidence that would show it is specific and someone could run it: a longitudinal comparison of matched firms — same sector, similar size, similar exposure — on rate of evidenced elimination against subsequent enterprise value, successor-unit revenue, or survival of the commercial unit. Nobody has run it. Until somebody does, this book is a mechanism argued from first principles with supporting evidence from an adjacent domain, and it should be read as one.
What the evidence base does not contain
The closest available proxy is the product-experimentation data from Chapter 7 — around two-thirds of changes at Microsoft having a negative or neutral effect on the metric they were designed to improve, and only 10 to 20 per cent having a positive effect in well-optimised domains. That is a product analogy. It was labelled as one there and it is labelled as one here, and the transfer argument is inferential: strategic intuitions are corrected far less often, so there is no reason to expect them to be better calibrated.
And the other portfolio's evidence, so this book does not flatter its favourite
The operating portfolio is also a bet
Of organisations getting zero return from generative AI initiatives21
Of integrated pilots extracting real value — a divide the authors attribute to approach rather than to model quality or regulation
Cite this carefully: it is a preliminary-findings study, about pilots, measuring P&L impact. It is not a claim that AI does not work.
The reason it belongs in this chapter rather than in a chapter arguing for the option portfolio is that it undercuts a comfortable assumption on both sides. The operating portfolio is not the safe one and the option portfolio the speculative one. Both are bets. Only one of them is currently reported to boards as though it were not.
"Isn't this scenario planning with a kill column?"
Scenario planning enumerates plausible futures and weights them. The boundary case deliberately selects an implausible extreme, because implausibility is what forces the structure to declare itself — and the output is not a set of futures at all, it is a list of which assumptions are load-bearing.
But the sharper difference is downstream, and it is where the two practices genuinely part. This framework routes the result through reflexive counterplay, to a paid probe that must return an observation the firm did not already hold, to a state with a trigger. And it reports a number — options eliminated, with the evidence — that scenario planning has never reported, because it does not produce one.
Chapter 13 made the fair case for scenario work and it stands: its own literature says specific techniques show impact, and the specimen in this book produced the boundary case that made its probe possible. The charge here remains narrow. An exercise that eliminates nothing has produced no evidence, whatever it produced in insight.
"You can't experiment on strategic questions"
Most of the time you can, once you stop demanding that a single probe settle the entire future. The probe must discriminate between two live explanations — not prove one, and not resolve the decade.
A price change on two engagements. A published offer with a stated hypothesis and a named audience. A scoped build with an acceptance event. A deliberate refusal, to see who objects. Each of those is available to a mid-sized firm inside one quarter, and each kills something.
Where you genuinely cannot — and the platform-integration question above is a real example — the answer is not to manufacture a probe. Name the trigger, assign the owner, and label the row as carried on inference. The failure is not carrying an untested option. It is not knowing which of your options are untested.
"This is lean startup for boards"
Build-measure-learn optimises a product hypothesis inside a known business model. It assumes the unit of sale and iterates the offer.
This operates one level up, where the unit of sale itself is the thing under test — and where the counterplay step has no analogue in a product loop, because a landing page does not have a regulator, a competitor who reprices in response to it, or a platform that can decide it is never seen.
The closer relative is real options in the plain commercial sense: pay a known amount to buy information and rights that change the decision set, with proceed, defer, reduce, transfer and do-nothing all remaining legitimate outcomes. Not a Black–Scholes exercise, no invented option premiums, no spreadsheet cosplay . The economics here are about optionality and information, not modelling.
That is the honest boundary of the method: a test that tells you whether to use it, a set of markets where it barely applies, an unproven central claim with a named falsifier, and four gaps in the evidence.
Which leaves one test sharper than any of them. This book was written by someone whose business is strategic search — who sells, in other words, the thing it has just argued produces fog when it does not terminate in evidence.
The standard has to survive being pointed at its author, or it is not a standard.
The Adviser Inside the Blast Radius
This is an argument for buying elimination, published by someone who sells strategic search. The standard has to survive being pointed here.
Start with the structure, without softening it.
A search apparatus that works well manufactures fog as a by-product of working well. An adviser whose product is that apparatus therefore enlarges the client's option set — and is then the obvious person to help the client manage the enlarged set. Every individual step in that chain is competent, well-intentioned and genuinely useful. The aggregate is a dependency, and the adviser is paid at every stage of it.
Permanent Fog can become a self-licking ice cream cone. A discovery engine manufactures more Fog by finding more possibilities; an adviser can then monetise navigation through the Fog it helped enlarge. The antidote is to price and measure disposition, not idea volume.
Notice what that mechanism does not require. No bad faith. No manufactured urgency. No strategic vagueness. An adviser who is simply very good at finding possibilities will produce this outcome by being good at it . Which is why the defence cannot be integrity. It has to be an instrument.
Scoring a quarter the way this book scores one
Four questions, applied to the firm that published this argument rather than to a client.
The self-audit
What did we kill? Not what did we produce — what belief, held at the start of the period, is no longer held, and what killed it. This book itself contains one: the claim that boundary cases pierce the fog, narrowed to aim the search, on the strength of the reflexivity argument in Chapter 5. That is a real elimination and it cost a chapter of a previous position.
What did we publish that could have embarrassed us — and did it? The honest answer for most publishing programmes is "nothing", because most publishing is designed to be agreeable. A claim that cannot be wrong in public is not a probe; it is marketing with footnotes.
Which claims moved from argued to evidenced — and which moved backwards? The second half of that question is the one nobody asks. A claim can lose support, and a corpus that only ever accumulates evidence in one direction is not learning.
How many client rows closed, and how many did we open? This is the ratio, applied to the adviser's own book of work rather than to the client's portfolio. An adviser whose clients' option sets grow every quarter is not neutral in that outcome.
Where this book's own claims actually sit
Publishing each claim at its honest rung, rather than at the rung the title implies.
| Claim | Rung | What would move it up |
|---|---|---|
| The branching rate is rising and measurable; planning horizons have compressed; most confident product changes fail their metric | Externally evidenced | Nothing — these are published, sourced and cited. They are the strongest material in the book. |
| The diagnosis method, the six-step sequence, the probe design, the eight-field row | Specimen-backed | Named firms running them, with results — not a composite. A composite demonstrates that an instrument can be operated, not that operating it works. |
| Falsification throughput is the scarce strategic capability; the two-clock ratio is the right instrument | Argued | The longitudinal comparison named in the previous chapter. Nobody has run it, including us. |
| The magnitude of any of it — how much better a high-throughput firm does | Unevidenced | Nothing available. Which is why no number appears anywhere in this book attached to that question. |
The title claim sits on the third rung. It is argued from a mechanism, supported by an analogy from a domain that measures elimination properly, and demonstrated on a specimen. It is not established. A reader who wants an established version of this should wait for somebody to run the comparison, and should be suspicious of anyone who claims to have.
Is this book a probe?
Only if it meets the rule from Chapter 7, and the rule is strict: a non-response is informative only when the audience who could have responded and the expected response class were both named in advance. So name them.
Stating that in advance costs nothing and changes what the response means. Without it, an enthusiastic reception would be indistinguishable from a well-marketed one, and silence would be indistinguishable from indifference, non-exposure and a busy quarter.
The conflict, and the structural answer
An adviser whose product is strategic search has an interest in the client carrying more options rather than fewer. More options means more navigation, more cycles, more mandate. That is a real conflict and it does not go away because both parties are decent.
The answer is not disclosure, which changes nothing about the incentive. It is a measurement inversion.
Key Insight
Measure the adviser on the client's closure rate, not on the adviser's output volume.
Concretely, a mandate scored on: options closed per period, with the evidence that closed them; residue delivered, meaning named instruments the client now operates without the adviser present; and option WIP, which should be flat or falling rather than growing. With idea volume, artefacts produced and workshops run explicitly excluded from the scorecard — not de-emphasised, excluded, because anything on a scorecard becomes a target.
Note what that does to the adviser's incentives. It makes a quarter in which the client kills three options and needs less help a good quarter, and a quarter in which the client generates twelve exciting new possibilities a bad one. That is uncomfortable to sell and it is the only version of this that is not self-serving.
What would make this wrong, in the first person
Two things, and the second is specific to the reflexive position.
The first is the falsifier from the previous chapter, owned rather than described: if firms that count kills navigate no better than firms that do not, holding capital constant, then this book's mechanism is a description of nothing and its author has spent a year being articulate about an artefact.
The second is narrower and closer to home. If this discipline makes clients more dependent on the adviser rather than less, it has failed on its own terms — regardless of whether the mechanism is true. A discipline that only works when the person who wrote it is in the room is not a discipline; it is a service with a framework attached. The observable is the residue field: after four quarters, is the client running the diagnosis, designing the probes and closing the rows themselves? If the answer is no, the transfer failed and the mandate is an expensive conversation subscription.
If my quarterly output is more options and no closures, I am selling the disease.
What actually transfers
The conclusions do not. They decay at the branching rate — a boundary case that was live last year may have been settled by the market since, and an adviser's map of your sector ages exactly as fast as everyone else's.
The discipline does not decay. The two-clock count, the pivot filter, the counterplay step, the discrimination table, the eight fields, the cap, the triggers: none of that goes stale, because none of it is a claim about the world. It is a way of finding out about the world, and a client who owns it can run it after the adviser leaves — which is the only defensible thing to have sold.
What a continuing relationship is for, once the discipline has transferred, is a real question with a real answer, and it is a sibling subject rather than this book's.
The standard survives being pointed at its author — not comfortably, and with two claims sitting a rung lower than the title implies. Which leaves the only thing still missing: what a reader actually does on Monday.
What Changes on Monday
Three moves. None of them needs a budget approval, and the first two can be done this quarter by people who already work for you.
The standing question in most boardrooms is what is our AI strategy?, and it will reliably be answered with the operating portfolio — because that is the portfolio with metrics, owners and a payback period. The question is not badly intentioned. It is badly specified, and the specification determines the answer.
Three changes fix it. Two are administrative and one is a sentence.
Move 1 — Measure both clocks, once
One page, four quarters back, one analyst for a week and about four hours of partner time for the judgement calls.
The two columns
Left — branching inputs
- • Entrant cadence. Not rebrands, not further rounds by firms already in your set.
- • Offer and pricing-model cadence. New commercial units, not discounts.
- • Capability cadence. Only what changed what a competitor or customer can do without you.
- • Regulatory cadence. Rule changes, guidance and enforcement posture — not consultations with no effect.
Sort each entry by route: customer, competitor, constructor. Sources: funding announcements, competitor pricing pages, lost-deal notes, buyer conversations, regulator publications.
Right — evidence outputs
- • Probes shipped that could have returned an unwanted result. A pilot that could only succeed does not count.
- • Options closed with named evidence. A decision is not evidence.
- • Median time from question raised to evidence received. Usually uncomputable at first, because nobody dates the question.
Audit each candidate against the question: could this have embarrassed somebody? If not, it belongs in the operating portfolio.
If the right-hand column is empty, stop. The diagnosis is complete. Do not commission a better branching count, do not benchmark against an industry that has no benchmark, and do not refine the numerator. An empty denominator has told you everything the exercise can tell you, and every further hour spent measuring is an hour spent not deciding.
Then read the same page a second way, as intervals rather than counts. How often does your market add a consequential move, and how long does it take you to learn anything? A market adding one every six weeks against a median time-to-evidence measured in quarters is not telling you that you are slow. It is telling you that you are navigating on inference — and a firm that knows that should be sizing its commitments smaller and more reversible, starting immediately, regardless of whether it ever fixes the clock.
Move 2 — Convert one carried question into a probe
The diagnosis will have produced the list of carried questions. Nobody has that list before they run it — the questions recur like weather, undated and unowned, and seeing them written down with ages attached is usually the most uncomfortable page of the exercise.
Take the one that has been carried longest. Run the six steps.
One question, six steps, one probe
- Boundary. Push one variable to its limit. Apply the gate: at the extreme, does the set of legal moves change shape or size? If the answer is "we'd be busier" or "we'd be poorer", pick a different variable.
- Breakage. Which unit fails first, and in what order do the rest go?
- Counterplay. Insist on this one. Customer, competitor, regulator, platform — and their responses to each other. This is where most conviction quietly dies, and skipping it is the reason the last three attempts produced documents instead of decisions.
- Capture economics. Who keeps the saving, and at what threshold does that flip?
- Probe. Write the discrimination table before funding — every outcome, and what each one kills. Have someone other than the sponsor read it and ask which row eliminates anything.
- Option state. Write the kill trigger before running. Name the residue: what do we still hold if this dies? Assign it to a person, not a committee.
Steps one to four are an afternoon with the right four people. Step five costs money. Step six costs courage, and it is the one that determines whether any of the rest of it mattered.
Move 3 — Change one standing board question
This one is free, takes a sentence, and does more than the other two combined.
Retire what is our AI strategy? and replace it with two questions:
What did we eliminate last quarter, and what killed it?
Which consequential uncertainty are we forcing the world to answer next?
Key Insight
The first cannot be answered with a deck. The second cannot be answered with a forecast.
That is the entire mechanism. Neither question can be satisfied by the artefacts the current process produces, so the process has to produce different artefacts. Within two cycles the pack contains a count instead of a narrative, and the people preparing it have started dating their questions — because they know they will be asked how old each one is.
Changing what gets asked is the only reliable way to change what gets prepared, and changing what gets prepared is the only reliable way to change what gets done. Everything else is exhortation.
The cadence after that
Quarterly, on one page, on the standing agenda:
- • Both clock counts, by route and by type.
- • Every state change since last quarter, with the evidence that caused it.
- • Option WIP against the cap — three to five live rows, and adding one means displacing one.
- • Open rows, with their age and what each is waiting on.
- • Any trigger changed, with its reason and whether evidence had already arrived.
- • At least one row closed — or an explicit statement of why none, which is itself a finding.
Three roles, lightly
Someone owns the terminal-value question, with the standing to displace a row rather than merely advocate for one. Without that, the cap is decorative and the portfolio drifts back to twenty rows within a year.
Someone owns the branching count — most naturally whoever already owns competitive intelligence, with the brief widened from competitors to all four routes. That widening is usually the single highest-yield change in the whole programme, because the customer route is where most firms are most exposed and least instrumented.
Probes need a budget line that is not the pilot budget. The pilot budget is governed by success, which is precisely the wrong governance for something whose job is to be able to fail informatively. A small separate allocation, sized deliberately rather than taken from underspend, is enough.
The honest close
The fog is not lifting. Your own search apparatus is part of the reason — it manufactures candidate futures as a by-product of working well — and so does everybody else's, continuously, without needing your permission or noticing your budget cycle .
My own version is less consoling and more accurate than most strategy writing allows itself to be:
You never really find the dry land. All you can do is be faster than your slower opponents — and stay alive long enough to find out what the next step is.
Which would be a bleak place to end, except that the ratio does something the fog metaphor never did. It locates agency precisely.
You cannot slow the branching clock. Nothing you do affects the rate at which the world invents new moves — not your adoption programme, not your market position, not the quality of your thinking. That clock belongs to everyone and to no one.
The other one is entirely yours. It is a budget line, a design discipline and a standing question, and it is available this quarter to any firm willing to spend something to find out that it was wrong.
You cannot slow the rate at which the world invents new moves. You can decide, this quarter, how fast you are willing to pay to find out which of them are false.
The argument, in one paragraph
AI makes variation cheap, not truth. AI Fog deepens when possible moves multiply faster than reality eliminates them. Boundary cases aim cognition at the assumptions that matter; adversarial counterplay exposes reflexive responses; paid and deployed probes let the world decide. Productivity creates enterprise value only when the company captures the cognition dividend in a commercial unit that survives cheap cognition. The board's job is not to predict the clearing of the Fog, but to operate an evidence-producing portfolio of options until the successor earns the right to be built.
One ask
Run the two-clock count for a single revenue unit before your next board meeting — branching inputs on the left, options killed on the right — and take the ratio into the room.
If the right-hand column is empty, you have just found the most valuable thing on the agenda.
References & Sources
The evidence base behind every claim — primary research, industry analysis, and technical specifications
Research Methodology
This ebook draws on primary research from standards bodies, independent research firms, enterprise technology vendors, and consulting firms. Statistics cited throughout have been cross-referenced against primary sources.
Frameworks and interpretive analysis developed by Scott Farrell / LeverageAI are listed separately below — these represent the practitioner lens through which external research is interpreted, and are not cited inline to avoid self-promotional appearance.
LeverageAI / Scott Farrell — Practitioner Frameworks
The interpretive frameworks, architectural patterns, and practitioner analysis in this ebook were developed through enterprise AI transformation consulting. The articles below are the underlying thinking behind those frameworks. They are listed here for transparency and further exploration — not cited inline, as this is the author's own analytical voice.
Scott Farrell — Cheap Thinking Makes Strategy Harder
Generic cognition defined by reproducibility rather than difficulty becomes the market floor; the two-term AI Fog and its endogenous second term
https://leverageai.com.au/wp-content/media/articles/227-cheap-thinking-makes-strategy-harder.html
Scott Farrell — The Terminal Value Doctrine — Stop Optimising the Horse
The AI Fog defined as simultaneous horizon compression and solution-space expansion; uncertainty has one direction, the Fog has two; the Question Ledger as the evidence pack for search quality
https://leverageai.com.au/wp-content/media/articles/61-terminal-value-doctrine.html
Scott Farrell — The Terminal Value Doctrine: Professional Services
Three delivery clocks price an engagement; the runway and proof clocks govern the firm; "same word, different altitude"; the survival inequality time-to-successor-proof < runway-under-erosion; activity does not move the proof clock, evidence moves it
https://leverageai.com.au/wp-content/media/articles/231-terminal-value-doctrine-professional-services.html
Scott Farrell — Cycle Compression: The Breathing Flywheel
The enabling inequality — AI has shortened world-response latency below the decay time of the original thought; the world's answer now collides with a live thought rather than arriving after the network has gone cold
https://leverageai.com.au/wp-content/media/articles/121-cycle-compression.html
Scott Farrell — The Cognition Dimension Ladder
Permanent Fog: the discovery engine manufactures fog as a side-effect of being good at its job; a sealed engine confidently generates ever more internally-consistent futures, each one further from reality, and world-loop closure is the only tether
https://leverageai.com.au/wp-content/media/articles/62-cognition-dimension-ladder.html
Scott Farrell — The Reshape — A Field Guide to Thought Experiments in the Age of AI
The pivot: push one variable in a stuck argument to an absurd extreme until the geometry of the situation forces a structural answer; AI cannot supply the pivot because it requires lived friction, taste for which extreme, and willingness to imagine the absurd; Lucretius' spear and Mach's variation principle
https://leverageai.com.au/wp-content/media/articles/60-the-reshape.html
Scott Farrell — Semantic Experiment Graph
The leftover checklist a successful experiment must produce — observation with exposure, candidate mechanisms, concepts implicated, confounders, falsifying next test, promotion decision; "If a test does not produce these artefacts, it was entertainment with metrics"; a local gradient tells you the direction of informative travel, not the guaranteed prize
https://leverageai.com.au/wp-content/media/articles/157-semantic-experiment-graph.html
Scott Farrell — Buy Certainty First
A real option on transformation: the right but not the obligation to make a better-informed investment later, in the plain commercial sense rather than as a Black–Scholes exercise; correct non-exercise can be the highest-value resolution, and firms that only celebrate exercised options will pressure consultants to recommend builds
https://leverageai.com.au/wp-content/media/articles/204-buy-certainty-first.html
Major Consulting Firms
Oliver Wyman Forum — The CEO Agenda 2026 [1]
CEOs now devote half of all planning effort to horizons of less than one year, up from 43% the previous year; 96% report increased board involvement in at least one area
https://www.oliverwymanforum.com/ceo-agenda/how-ceos-navigate-geopolitics-trade-technology-people.html
McKinsey & Company — From AI Table Stakes to AI Advantage: Building Competitive Moats [7]
Agentic AI could orchestrate up to $1 trillion in US retail by 2030; AI agents mediating discovery
https://www.mckinsey.com/capabilities/quantumblack/our-insights/from-ai-table-stakes-to-ai-advantage-building-competitive-moats
Industry Analysis & Vendor Research
Crunchbase News — Q1 2026 Shatters Venture Funding Records As AI Boom Pushes Global Startup Investment To $300B [2]
Investors poured $300 billion into 6,000 startups globally in the quarter; AI took $242 billion — 80% of total global venture funding — up from 55% in Q1 2025; OpenAI, Anthropic, xAI and Waymo collectively raised $188 billion, 65% of global venture investment in the quarter
https://news.crunchbase.com/venture/record-breaking-funding-ai-global-q1-2026
Sapphire Ventures — 2026 Outlook: 10 AI Predictions Shaping Enterprise Infrastructure [5]
"Achieving $100M in ARR over 5–10 years used to be the gold standard in SaaS. The best-in-class AI-native companies are now compressing that timeline into 1–2 years, demonstrating truly historic growth rates."
https://sapphireventures.com/blog/2026-outlook-10-ai-predictions-shaping-enterprise-infrastructure-the-next-wave-of-innovation
Retool — The Build vs. Buy Shift: AI, Shadow IT, and the SaaS Replacement Era (2026 Build vs. Buy Report) [6]
Survey of 817 builders and customers: 35% have already replaced at least one SaaS tool with a custom build, and 78% expect to build more of their own tools in 2026
https://retool.com/blog/ai-build-vs-buy-report-2026
Clio — What's Driving Legal AI Pricing in 2026? [13]
67% of corporate legal departments and 55% of law firms expect AI to change how hours are billed; 71% of buyers already prefer flat fees for an entire case
https://www.clio.com/resources/ai-for-lawyers/legal-ai-tool-pricing
Data Center Dynamics (Q4 2025 earnings) — IBM's mainframe business sees highest annual revenue in 20 years [18]
IBM Z posted its highest annual revenue in twenty years, with IBM arguing the mainframe is the lowest unit-cost platform available for certain workloads
https://www.datacenterdynamics.com/en/news/ibms-mainframe-business-sees-highest-annual-revenue-in-20-years
Hypercubic, citing ZipRecruiter (March 2026) — COBOL Job Postings Over Time: Salary Trends [19]
Average US COBOL developer earns $115,475 a year, comfortably above the median for all software developers; "Pundits frame COBOL as a relic. Hiring managers frame it as a staffing crisis."
https://www.hypercubic.ai/insights/cobol-job-postings-over-time-salary-trends-and-which-industries-are-still-hiring
TechCrunch (7 December 2018) — IBM selling Lotus Notes/Domino business to HCL for $1.8B [20]
The final Lotus components sold for $1.8 billion two decades after the category defeat
https://techcrunch.com/2018/12/07/ibm-selling-lotus-notes-domino-business-to-hcl-for-1-8b
Primary Research & Standards Bodies
IEEE Spectrum, reporting the Stanford HAI AI Index 2026 and Epoch AI data — Stanford's AI Index for 2026 Shows the State of the Industry [3]
Epoch AI tracked 87 notable model releases from industry in 2025, compared to just seven from all other sources; models released by industry now make up over 90 percent of notable models, up from just under 50 percent in 2015 and zero in 2003; US organisations released 50 notable models in 2025
https://spectrum.ieee.org/state-of-ai-index-2026
Stanford HAI — Inside the AI Index: 12 Takeaways from the 2026 Report [4]
The success rate of agents handling real-world tasks improved from 20% in 2025 to 77.3% today, according to Terminal-Bench
https://hai.stanford.edu/news/inside-the-ai-index-12-takeaways-from-the-2026-report
Anthropic — Claude Mythos Preview System Card [8]
The model "demonstrated a striking leap in cyber capabilities relative to prior models, including the ability to autonomously discover and exploit zero-day vulnerabilities in major operating systems and web browsers"; the capability was not trained for and was downstream of general reasoning improvements
https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf
George Soros, Journal of Economic Methodology 20(4), December 2013, 309–329 — Fallibility, Reflexivity, and the Human Uncertainty Principle [9]
"The two functions connect the participants' thinking (subjective reality) and the actual state of affairs (objective reality) in opposite directions. In the cognitive function, the participant is cast in the role of a passive observer: the direction of causation is from the world to the mind. In the manipulative function, the participants play an active role: the direction of causation is from the mind to the world. Both functions are subject to fallibility."
https://www.georgesoros.com/2014/01/13/fallibility-reflexivity-and-the-human-uncertainty-principle-2
John R. Platt, Science 146, no. 3642 (16 October 1964): 347–353 — Strong Inference [10]
"1) Devising alternative hypotheses; 2) Devising a crucial experiment (or several of them), with alternative possible outcomes, each of which will, as nearly as possible, exclude one or more of the hypotheses; 3) Carrying out the experiment so as to get a clean result; 1') Recycling the procedure, making subhypotheses or sequential hypotheses to refine the possibilities that remain; and so on. It is like climbing a tree."
https://dunnlab.ucsf.edu/sites/g/files/tkssra15436/files/wysiwyg/Platt%201964.pdf
Iavor Bojinov and Somit Gupta, Harvard Data Science Review, Issue 4.3 (Summer 2022), citing Kohavi et al. (2020) — Online Experimentation: Benefits, Operational and Methodological Challenges, and Scaling Guide [11]
"For example, in a study at Microsoft, around two-thirds of all changes were found to have a negative or neutral effect on the metric they were designed to improve; in well-optimized domains, the number is even lower." And: "in well-optimized domains like Bing and Google, only about 10% to 20% of changes have a positive effect on the target metrics."
https://hdsr.mitpress.mit.edu/pub/aj31wj81
Ron Kohavi and Stefan Thomke, Harvard Business Review 95, no. 5 (September–October 2017): 74–82 — The Surprising Power of Online Experiments: Getting the Most Out of A/B and Other Controlled Tests [12]
Experimentation treated as an organisational capability rather than a technique; the value arises from the rate at which trustworthy controlled tests can be run
https://web-docs.stern.nyu.edu/executive/The%20Surprising%20Power%20of%20Online%20Experiments.pdf
Thomson Reuters Institute, as reported by SignalFire — Law Firm Rates Report 2026 [14]
Firms collect roughly the same amount per hour whether they discount aggressively or hold firm on realization
https://www.thomsonreuters.com/en-us/posts/legal/law-firm-rates-report-2026/
Shardul Phadnis, Chris Caplice and Yossi Sheffi, MIT Sloan Management Review 57, no. 4 (Summer 2016): 21–24; and Phadnis, Caplice, Sheffi and Singh, "Effect of Scenario Planning on Field Experts' Judgment of Long-Range Investment Decisions," Strategic Management Journal 36, no. 9 (2015): 1401–1411 — How Scenario Planning Influences Strategic Decisions [15]
A study of how scenario planning affects executives' strategic choices, run through Future Freight Flows workshops with private-sector and public-sector transportation planners
https://sloanreview.mit.edu/article/how-scenario-planning-influences-strategic-decisions
K.E. Cordova-Pozo and E.A.J.A. Rouwette, Futures (2023) — Types of scenario planning and their effectiveness: A review of reviews [16]
"while there is indeed a scarcity of research into effectiveness of scenario planning, on the level of specific techniques there is evidence of impact"
https://www.sciencedirect.com/science/article/pii/S0016328723000575
Paul J.H. Schoemaker and Shardul S. Phadnis, MIT Sloan Management Review, Winter 2026 — How to Make Scenario Planning Stick [17]
"Developing future scenarios can deepen leaders' strategic insights. Establishing scenario planning as an ongoing capability and reaping its full benefits require linking it to other processes."
https://sloanreview.mit.edu/article/scenario-planning-how-to-use
MIT NANDA (Aditya Challapally, Chris Pease, Ramesh Raskar, Pradyumna Chari), July 2025 — The GenAI Divide: State of AI in Business 2025 [21]
"a surprising result in that 95% of organizations are getting zero return... Just 5% of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact. This divide does not seem to be driven by model quality or regulation, but seems to be determined by approach." Based on a systematic review of over 300 publicly disclosed AI initiatives plus structured interviews; research period January–June 2025
https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
About This Reference List
Compiled August 2026. All URLs verified at time of compilation. Regulatory documents and standards specifications are subject to revision — check primary sources for the most current versions.
Some links to academic papers and vendor research may require free registration. Government and standards body publications are freely accessible.
Run the count on your own firm
Four quarters. Branching inputs on the left — entrants, offers, capabilities, rules. Evidence outputs on the right — probes shipped that could have gone badly, options closed with the evidence that closed them, median time from question to answer.
If the right-hand column is empty, the diagnosis is complete and there is nothing left to measure.
Scott Farrell · LeverageAI · leverageai.com.au