Strategy under AI Fog
Fog Is a Race Between Two Clocks
Why better strategic search keeps making your option space bigger — and what actually shrinks it.
TL;DR
- Fog is not weather. It is a ratio. The branching clock is how fast plausible strategic moves multiply in your market. The evidence clock is how fast your firm can force the world to eliminate one. Fog thickens when the first outruns the second — which means the same environment is bewilderment for one firm and a sequence of cheap reversible experiments for another.
- The scarce capability is falsification throughput, not thought throughput. Options killed per quarter, each with the evidence that killed it. Almost no board can state that number, which is itself the diagnosis.
- A boundary case aims the search; only a probe closes it. Markets reprice when you examine them; a brick wall does not. So every boundary case must run through counterplay to a smallest discriminating probe with kill and scale triggers — or it is mutation theatre: fifty coherent futures, nothing eliminated, more fog than you started with.
A firm I will keep deliberately generic — a professional-services business, four practice areas, a partnership that reads its market carefully — had the best strategy year in its history.
It ran an AI-assisted scenario programme. It generated futures at a rate the partnership had never managed before: agentic procurement, regulatory intervention, a category of AI-native entrant that did not exist eighteen months earlier, three separate models of what its largest client might insource. The offsite was genuinely excellent. People said things like that changes how I think about this. The deck was beautiful and, more unusually, it was right — the futures in it were plausible, internally consistent and well argued.
Twelve months later, the partnership could not name a single thing it had ruled out.
Not one future eliminated. Not one option closed. Not one assumption tested against a customer, a price, a competitor or a regulator and found to be false. The pricing sheet was unchanged. The hiring plan was unchanged. What had changed was the number of things the partnership now had to hold in mind, which had roughly tripled.
They had spent a year making their own fog thicker, using the best strategic thinking they had ever done. And here is the part worth sitting with: they were not doing it wrong. They were doing exactly what the strategy profession taught them, with better tools than the profession had when it wrote the lesson. The process was sound. The criterion it was missing had simply never needed to exist before.
The clock nobody put on the board pack
Half of the diagnosis is already in the boardroom. Planning horizons have measurably compressed: CEOs now devote half of all planning effort to horizons of less than a year, up from 43 per cent the year before6. Everyone can feel the three-year forecast nobody in the room believes. The standard response — shorter cycles, more frequent re-forecasting, commit later — is correct, and if horizon compression were the only thing happening, it would be sufficient.
The other half gets treated as ambient. Turbulence. Unprecedented pace of change. A climate, implying no action beyond resilience.
It is not a climate. It has a source, a direction, and a rate you can count.
Consider what a single quarter now looks like. Investors put roughly US$300 billion into about 6,000 startups globally in Q1 2026; AI companies took $242 billion of it — 80 per cent of all global venture funding, up from 55 per cent in the same quarter a year earlier1. Capital concentration means most of that went to a handful of frontier labs, so this is not 6,000 new threats to your business. But it is the funding of new commercial architectures at a rate the strategy calendar was never designed to absorb. On the capability side, Epoch AI counted 87 notable model releases from industry in a single year2, and one agentic benchmark moved from a 20 per cent success rate on real-world tasks to 77.3 per cent inside a year3. On the assembly side, the best AI-native companies are reported to be compressing the road to $100 million ARR from five-to-ten years into one-to-two4. On the buy side, 35 per cent of enterprises have already replaced at least one SaaS tool with a custom AI build and 78 per cent plan to do more5. And entire categories arrive that were not on last year's map at all: agentic commerce, forecast to orchestrate up to a trillion dollars of US retail by 203022 — a number to treat as a category marker rather than a plan, since it is a forecast and this argument does not run on those.
The rate is also not smooth, which is the second reason a rate is the right instrument and a forecast is not. One frontier model release was reported to have "demonstrated a striking leap in cyber capabilities relative to prior models, including the ability to autonomously discover and exploit zero-day vulnerabilities in major operating systems and web browsers"21 — a capability the lab noted it had not trained for. Whatever you think of the coverage that followed, the planning lesson is structural: capability can jump vertically inside a single domain, and a planning rhythm built on smooth progression cannot see that class of move coming.
Each of those is a number you can put on a page and update quarterly. Together they are an instrument, and the instrument measures something specific.
Definition — the two clocks
The branching clock: how quickly new competitors, offers, architectures and strategic possibilities appear in your market. Market-wide. It does not need your permission, your budget cycle or your change-management plan.
The evidence clock: how quickly your firm can put something into the world, observe the response, and eliminate or revise a possibility. Firm-local. Almost entirely yours.
Fog thickens when branching outruns evidence.
That reframing does one important thing immediately: it makes fog relative. A slow incumbent experiences the condition as bewilderment because its annual planning cycle cannot keep up. An operator with a fast evidence clock experiences the same environment as a sequence of cheap, reversible experiments. The external world is equally uncertain for both. The difference is how quickly each can make that uncertainty answerable.
Which is why the comparison that feels natural — us today versus us last year — is the wrong one, and it is the only comparison most dashboards can make. Your cognition compounds inside your walls at the speed of your change programme, gated by adoption, training, data access and the security review. Everyone else's compounds across the whole market at once, with no adoption curve applying to the aggregate23. A firm can have the best internal AI adoption in its sector and still be losing ground.
Once you hold the ratio, the strategic objective changes shape. It stops being how do we understand all the possible futures? and becomes:
Which consequential uncertainty can we force the world to answer next?
And that question has a metric attached to it, which the first one never did.
Definition — falsification throughput
The rate at which your firm eliminates strategic options, each elimination carrying the named evidence that killed it. Options closed per quarter, not candidates generated per quarter.
Value migrates to the bottleneck, not to the activity producing the most volume. Generation stopped being the bottleneck some time around 2024. Almost nobody moved their instruments.
Three clock families, and why the distinction matters
The corpus now has three sets of clocks and a reader deserves the map, because collapsing them destroys exactly the distinction that makes each useful.
| Clocks | Altitude | Question they answer |
|---|---|---|
| Production, authority, evidence | Inside one engagement | What is this delivery actually waiting on?24 |
| Runway and proof | The firm in transition | Can we prove a successor before harvest cash runs out?24 |
| Branching and evidence | The market search | Which bets are worth carrying at all, and how fast can we kill the wrong ones? |
These sit above the runway/proof pair rather than beside it. The branching-to-evidence ratio determines what the successor bets should be; the runway/proof inequality determines whether the firm survives long enough to run them. The professional-services companion to this argument stops precisely at that seam and says so — running a portfolio of successor bets, how many, how governed, how killed, is explicitly named there as a later book's machinery24. This is that machinery.
One discipline transfers upward verbatim and should be nailed to the wall: activity does not move this clock; evidence moves it.
Where the Einstein analogy runs out
Boundary-case reasoning is the best instrument we have for aiming a strategic search, and I have spent several books arguing for it. Push one variable to an absurd extreme until the geometry of the situation has no choice but to force a structural answer30. Lucretius asked what happens when you throw a spear at the alleged edge of space: it either flies through, in which case there is no edge, or it bounces off something, in which case there is something on the other side and still no edge. Argument over.
Note those two words. Argument over.
They are available to Lucretius because space does not care that a spear was thrown at it. The system under examination is passive. It has no strategy, no pricing, no counsel, and no interest in the outcome of the thought experiment.
A market is not like that.
A thought experiment in physics acts on a passive system. The brick wall does not respond strategically. A beam of light does not change its pricing model because Einstein is examining it. A market does.
This is the limit I did not state clearly enough when I first made the argument, and it changes what a boundary case is for. Push the variable — suppose competent generic first-pass advice becomes effectively free — and every participant responds, including to each other's responses. Customers may buy far more analysis, or far less. Regulators may require a licensed human to sign it. Competitors may bundle it away. Platforms may capture the distribution. Your own visible response changes what customers expect. Soros's name for the underlying structure is the two-way feedback loop between thinking and reality: participants are simultaneously observers of a situation and participants in it, and both functions are fallible11. Physics gets the observer role only. Strategy gets both.
So the honest formulation is narrower than the one I published, and it is the whole hinge of this piece:
Thought experiments aim the search. Discriminating experiments pierce the Fog locally.
A second reviewer, working blind on the same material, arrived at the same correction in different words: boundary cases chart structural exposures inside permanent fog. Two independent routes to the same place is about as much confirmation as this kind of claim gets.
And it forces a step that most strategy processes skip entirely.
The sequence — boundary to option state
- Boundary. What if competent advice, software or coordination approaches zero cost? (Filtered first: at the extreme, does your set of legal moves change shape, or only size? If the answer is "we'd be busier" or "we'd be poorer", it is not a pivot23.)
- Breakage. Which revenue unit, cost structure or customer behaviour fails first?
- Counterplay. What do the customer, the competitor, the regulator and the platform do next — including in response to each other, and to you?
- Capture economics. Who keeps the cost reduction, and at what threshold does that answer change?
- Probe. What is the smallest real offer, price or operating experiment that distinguishes between the surviving explanations?
- Option state. What evidence causes us to scale, defer, redirect or kill — decided before the evidence arrives?
Steps 1 and 2 are where most strategy work stops. Step 3 is what reflexivity forces into existence. Steps 5 and 6 are the evidence clock. A process that ends at step 2 has produced a conditional map and called it a conclusion.
Mutation theatre reaches the boardroom
There is a precise miniature of this failure one floor down, in a department that had it first and already named it.
When generative tools made creative production nearly free, marketing teams could mint a hundred near-misses before lunch. Learning should have exploded. Mostly it did not. Teams generated more variants, crowned a winner on the dashboard, and cloned its visible features without knowing which underlying mechanism produced the result. More objects, same ignorance, prettier decks. The name for it is mutation theatre27, and the diagnosis that goes with it is that the scarce resource moved: it is no longer the ability to produce another candidate, it is the ability to design a candidate that would falsify a mechanism rather than redecorate a winner.
Now read that at strategy altitude.
A board asks AI for fifty futures. The model returns fifty persuasive futures. Executives discuss them, combine three, put one in a deck and call the exercise foresight. The solution space has expanded, but nothing has been learned.
Same pathology, different altitude, larger stakes: the artefacts are scenarios rather than creatives, and the currency is capital rather than click-through.
The antidote is not to stop generating. Generation is cheap and should be lavish. The antidote is a residue requirement — a fixed list of things a strategic probe must leave behind, or it did not happen:
- the observation, dated, with what was exposed to whom;
- the competing causal explanations, ranked;
- the confounders you refuse to pretend you controlled;
- the alternative that was rejected, and why;
- the next experiment most capable of hurting the leading explanation.
Two items on that list are worth pausing on, because a scenario pack structurally cannot produce either. Silence counts as a response — but only if you specified in advance who was exposed and what a response would have looked like; otherwise a non-response is merely ambiguous. And delayed outcomes get backfilled: the row stays open until reality reports, which means someone has to own it past the quarter in which it was fashionable.
One more discipline transfers directly, and it is the one that fails most often in practice. Client enthusiasm, a striking meeting, or praise for a well-sourced answer is heat. Heat may justify another probe. It cannot promote a thesis into doctrine. A local gradient tells you the direction of informative travel, not the guaranteed prize27.
An uncomfortable aside, since this argument implicates its author. A search engine that works well manufactures fog as a by-product of being good at its job26. An adviser can therefore monetise navigation through fog he helped enlarge — a self-licking ice cream cone with a strategy practice attached. The only honest defence is to price and measure disposition rather than idea volume. If my quarterly output is more options and no closures, I am selling the disease.
Walking one boundary case all the way through
Abstractions about evidence are easy to agree with and hard to act on, so here is the whole sequence run end to end on the case most service firms are actually living inside. The arithmetic below is deliberately transparent; the numbers are illustrative and the shape is what transfers.
1. Boundary
Competent generic first-pass advice in our sector is effectively free to the client, delivered with tools they already own.
Check it against the filter before spending a morning on it. At the extreme, does the set of legal moves change shape or size? An entire category of promise becomes unsellable at any price. A pricing structure becomes indefensible. The graduate pyramid turns from an engine into a liability. And something previously uneconomic becomes viable — a standing advisory commitment priced as continuity rather than as a project, which nobody could afford to staff while first-pass work was expensive. Options die and options appear. It passes23.
Note what the probe does not claim: not that advice will cost zero dollars, not that this happens on any date, not that anyone's profession ends. It exchanges timing certainty for structural certainty, which is the only trade available under these conditions.
2. Breakage
An engagement that required 100 consultant-hours now takes 80. If the firm still bills hours and sells the same number of engagements, revenue per engagement falls 20 per cent while payroll is fixed in the short run — so margin falls harder than the efficiency number suggests.
And the freed capacity is not 20 per cent. It is 100/80 = 1.25. The firm must sell 25 per cent more engagements merely to restore the old number of billed hours. Sell 20 per cent more and you land at 96 per cent of the previous billed-hour revenue — having successfully delivered an AI programme.
3. Counterplay
Here is the step the scenario pack skips, and it is where the boundary case stops being a thought experiment and starts pointing at a test.
- Customer. They are not passive recipients of your efficiency. In an adjacent professional-services market where the data exists, 67 per cent of corporate legal departments and 55 per cent of firms expect AI to change how hours are billed, and 71 per cent of buyers already prefer a flat fee for an entire matter15. The customer arrives already holding a view about your unit of sale.
- Competitor. The instinct is to hold rate and absorb compression as margin for a few years. That requires two things to be true simultaneously: that no material competitor passes the compression through, and that clients cannot observe it. Rate data in the same market suggests firms collect roughly the same amount per hour whether they discount aggressively or hold firm on realisation16 — which is to say the defence is weaker than partnerships believe.
- Regulator. May require an accountable human to sign, which does not reduce the compression but relocates the scarcity into authority.
- Platform. May capture distribution entirely, so that being cheap and good is irrelevant if the customer's agent never reaches you.
Every one of those moves changes the answer to the next step. That is precisely why the boundary case cannot settle the question by itself.
4. Capture economics — the four outcomes
"Productivity gets you to zero" is a claim I have made in the blunt form, and it is true only under one specific capture condition. Walk all four.
| Outcome | What happens | Verdict |
|---|---|---|
| Unsold capacity | The whole gain passes to customers; the freed capacity does not sell; payroll stays. | Direct self-disintermediation. This is the case that gets you to zero, and it is the default rather than the exception. |
| Fully sold capacity | The firm sells 25% more engagements and preserves revenue. | The old economics defended, no advantage created. Competitors with the same models can run the same play. |
| Outcome pricing | Engagement price holds while delivery effort falls. | Margin expands — conditional on competitors not forcing the saving into market price, and on verification costs not eating it. |
| Demand expansion | Cheap cognition makes a service possible that could not previously exist: continuous advice, exhaustive transaction analysis, every-anomaly investigation. | The quantity of useful cognition demanded grows. This is the only outcome that creates a new reason to pay. |
So the accurate claim is narrower and more useful than the slogan:
Productivity applied to a depreciating commercial unit, with no way to capture the gain or redirect the capacity, accelerates decline. The mistake is not making staff faster. The mistake is calling that a strategy while leaving the revenue unit unchanged.
Notice what has just happened structurally. Four outcomes, one boundary case, and no way to tell from the armchair which one you are in. The difference between them is not a matter of analysis. It is a matter of what customers and competitors actually do — which is an empirical question, which means it has a price, which means you can buy the answer.
5. The smallest discriminating probe
This is the step that separates a strategy from a document, and the design rule is stricter than it looks.
A probe's job is not to demonstrate the favoured explanation. It is to discriminate between the survivors — to produce an observation that at least one live explanation cannot accommodate. The formal name for this discipline is sixty years old. John Platt called it strong inference: devise alternative hypotheses, devise a crucial experiment with alternative possible outcomes each of which excludes one or more of them, carry it out cleanly, then recycle with the possibilities that remain. It is like climbing a tree.9
The organisational version of the same idea has been argued at length by Kohavi and Thomke, whose central claim is that experimentation is a capability rather than a technique — the value comes from the rate at which an organisation can run trustworthy tests, not from any single result8. Falsification throughput is that claim carried up to the altitude where the unit of sale is what is being tested.
Applied to the four outcomes above, the discriminating probe is not "run an AI pilot" and it is not "survey clients about pricing." It is something closer to: take two practice areas to fixed-scope pricing at the next contract cycle, sized on outcome rather than estimated hours, with first-pass analytical scope removed from the billable estimate and re-cast as a pre-built input we bring to the table.
Now check that against the discrimination test. If clients accept the fixed price at a level that preserves margin, the outcome-pricing branch is live and the unsold-capacity branch weakens. If clients demand the compression back — "your costs fell, so should our price" — the capture goes to the customer and the branch that survives is the one you least wanted. If competitors announce equivalents within a quarter, that is the competitor-counterplay branch reporting in. Each result kills something. That is the property that makes it a probe rather than an initiative.
Two honest constraints, because the design is harder than the principle.
Probes decay, and reflexivity is why. Your probe changes the market it measures. A competitor may respond to your fixed-price move rather than to the underlying condition, and you cannot observe the counterfactual. This is a real limit and it should be written into the row as a confounder. It is also not a reason to prefer a scenario pack, which buys no information at all. Probes buy locally valid, decaying information — and the decay rate is itself a measurable input to your two-clock diagnosis.
Sometimes there is no probe. Some structural questions cannot be tested within a horizon that matters. The correct move then is not to invent one; it is to name the trigger — the observable event that would reopen the question — assign an owner, and say plainly that the option is being carried on inference rather than evidence. A portfolio that knows which of its rows are unevidenced is in far better health than one that does not distinguish.
What makes this economically available at all is a change most firms have not claimed at strategy altitude: world-response latency has fallen below the decay time of the original thought28. Evidence now arrives while the question that prompted it is still cognitively alive. That is an engineering fact about the environment, not an exhortation — and it is the reason "raise your evidence clock" is a plan rather than a wish.
6. Option state
Decided before the evidence arrives, because a trigger written afterwards is a rationalisation. Kill. Scale. Defer until a named trigger. Continue collecting evidence. Exercise into a bounded build. Transfer into ordinary operations.
And the discipline that makes the whole apparatus honest: a killed option is a full return on the probe. Firms that only celebrate exercised options will pressure their people into recommending builds29. The same pathology, one level up, produces a portfolio where nothing ever dies.
Two portfolios, two ledgers
Everything above collapses without a place to keep it, and the place cannot be the operating plan — because the operating plan's metrics will quietly strangle it.
Run two explicitly separate portfolios.
The operating portfolio
Workflow efficiency, quality, cost reduction, staff augmentation. Useful work. Ordinary operating metrics. Funded from OPEX. Never described as transformation.
This is also where the honest good news lives. Where a question has an enumerable population, a stable evaluation function and a closed action set — every transaction examined instead of a sample, every anomaly investigated instead of the escalated ones — cheap cognition simply reduces uncertainty, and no boundary case or ledger is required23. Flood those. There is enormous, unglamorous value there.
The terminal-value option portfolio
Asset conversion, successor offers, customer-agent threats, new commercial units, governed execution capability. Different sponsor, different ledger, different burden of proof. Judged on falsification throughput, not on activity.
Every row carries eight fields:
- the revenue unit being tested;
- the boundary case;
- the customer, competitor, regulator and platform counter-moves;
- the scarce complement the firm believes it owns;
- the smallest paid or deployed probe;
- the evidence received — including silence, if exposure was specified;
- the kill, scale and revisit triggers;
- the asset that compounds even if the option is never exercised.
The eighth field is the one people skip and the one that changes behaviour. It is what makes a killed option affordable. If a probe leaves behind a pricing instrument, an acceptance test, a reusable evidence pattern or a qualification rule, then the branch died and the firm still got paid in capability. Without field eight, every kill is a pure loss and the portfolio will quietly stop killing things.
This is not a warehouse of clever scenarios. It is capital-allocation memory: a record of what was tested, what was rejected, and what would reopen it.
How this relates to the artefacts you may already run
Three instruments now exist across this body of work and they are not substitutes. The distinction is worth stating once, cleanly, because a firm that conflates them will build one artefact that does none of the three jobs.
| Artefact | Governs | Closes when |
|---|---|---|
| Question Ledger25 | Search quality — was the space actually searched? | The search is inspectable and its gaps are named. |
| Strategic Search Ledger23 | Decision consequence — did anything change? | The "decision changed" column is non-empty, with an owner and a date. |
| Option portfolio | Capital and disposition — what did we pay to learn, and what died? | The option reaches an explicit state and the residue is booked. |
Cap the number of live rows. Three to five is a working portfolio; a firm carrying twenty is doing research rather than governance and will close none of them23. Option WIP is a quantity to limit, not a scoreboard to grow.
What a completed row looks like
OPT-01 — first-pass advice
Revenue unit under test: the standard four-to-twelve week analytical engagement, billed by the hour. Roughly 30 per cent of billed hours on a standard engagement sit in first-pass scope.
Boundary case: competent generic first-pass advice is effectively free to the client. No date attached.
Counterplay: customers arriving with their own baseline analysis and a preference for fixed fees; competitors announcing equivalents; regulator requiring named human sign-off in two of four practice areas; platform capture not yet material in this segment.
Scarce complement we believe we own: accepted evidence and the authority to sign — not analysis.
Smallest discriminating probe: two practice areas move to fixed-scope pricing at the next contract cycle, sized on outcome; first-pass scope removed from the billable estimate and supplied as a pre-built input. Owner: managing partner. Runs one contract cycle.
Evidence received: two of the last five prospects arrived holding their own baseline analysis. Internally, delivery time on the last six engagements fell about 18 per cent and the difference was billed away.
Triggers: Scale if fixed-scope wins hold margin across a full cycle. Kill if clients systematically demand the compression back as a discount. Revisit if a competitor announces fixed-fee equivalents, or a client asks us to price against their own AI-produced baseline.
Compounding residue: the outcome-sizing instrument and the acceptance tests, which are reusable whichever way the pricing question resolves.
Option state: live, probe running. Future rejected: "hold rate and volume, absorb the compression as margin for three to five years" — rejected, because it requires that no material competitor passes compression through and that clients cannot observe it, and both are already false in the pipeline.
The rejected future is what makes the row real. Anyone can list futures. Rejecting one costs something, because it removes an option the firm was quietly relying on — and once it is written down as rejected, with a reason, nobody can drift back into it without arguing.
Underwriting zero
Which brings us to the sentence that started all of this, and which I now think needs qualifying rather than repeating.
"Terminal value is probably zero" sounds like a calibrated forecast. Without a horizon, an industry base rate, a competitive model and a distribution, it is not one. What it usually reflects is something different and more diagnostic:
A company's inability to articulate a satisfying AI future is evidence of weak strategic search. It is not evidence, by itself, that the company's terminal value is probably zero.
Zero is a stress condition management must be able to underwrite, not a probability to publish. And the way you underwrite it is to stop treating terminal value as one distant scalar and break it into four terms:
The decomposition
Future value = legacy runoff + convertible-asset value + successor option value − stranded liabilities
The professional-services translation of this — runoff cash, tail scarcity and migration income, convertible assets, successor options, restructuring liabilities — is worked in detail elsewhere, along with the evidence that declining terminal value and strong harvest returns can comfortably coexist: a mainframe business posting its highest annual revenue in twenty years18, COBOL developers earning above the median software salary in a language routinely declared dead19, and an installed base selling for US$1.8 billion two decades after losing its category war20. So "zero terminal value" does not mean "nobody can buy the business." What no informed buyer will pay for is the pretence that the old unit recovers24.
What this piece adds is the link between two of the terms, and it is the sentence that turns a strategy discipline into a board obligation.
The successor-option term in that decomposition is the terminal-value option portfolio. It is worth whatever your evidence clock says it is worth — because an option is worth something only if it can be exercised on evidence, and a firm running no probes has no mechanism for producing that evidence. A partnership carrying a large successor-option term with an empty portfolio is carrying an unaudited asset. The convertible-asset term has the same property: convertible assets are worth something strictly conditional on conversion actually happening. Unconverted, they ride the runoff curve to zero24.
Which is why the board does not have to know which successor wins before acting. It can harvest the current revenue stream, convert context and evidence and relationships into reusable assets, run several bounded successor experiments, and release serious capital only when one earns it. That is strategy as option allocation rather than prediction — and the option allocation is only as good as the evidence clock underneath it.
Six ways this goes wrong
The most likely failure is not the firm that never adopts the discipline. It is the firm that adopts the artefact and hollows it out, because a hollow artefact provides cover.
- Portfolio theatre. Tell: the option-state column reads "monitor", "explore", "continue to assess". Fix: the row stays visibly open, and open rows go in the board pack too. Openness is not failure; false closure is.
- The probe that cannot lose. Tell: every possible result is consistent with proceeding. Fix: before running it, write down which explanation each outcome kills. If no outcome kills anything, you have designed a demonstration.
- Heat mistaken for evidence. Tell: the evidence field contains enthusiasm, attendance or praise. Fix: behavioural evidence only — what did they actually do with money, authority or attention?
- Retrofitted triggers. Tell: the kill trigger was written after the result came in and, remarkably, was not met. Fix: triggers are dated at creation, and a changed trigger is itself a decision requiring a reason.
- Over-search. Tell: rows are being added faster than options are being closed. Fix: track that ratio; it is the metric that matters23. A search apparatus that works will produce plausible futures forever until someone tells it to stop — a search is finished when the next candidate future would not change any present decision.
- Wrong altitude for the instrument. Tell: falsification throughput is being demanded of a bounded question — one with an enumerable population, a stable evaluation function and a closed action set. Fix: flood those with cognition instead. The test, not the thesis, tells you which world you are in23.
Underneath all six sits one condition: nobody owns the terminal-value question. Efficiency reports to operations, delivery to the practice leads, pricing to the commercial director, relationships to the partners — and whether the unit survives reports to nobody. So the portfolio becomes somebody's side project, reviewed by people whose day jobs reward continuity, and side projects close rows with intentions.
Three objections worth answering
"This is scenario planning with a kill column."
Scenario planning enumerates plausible futures and weights them; the boundary case deliberately selects an implausible extreme, because implausibility is what forces the structure to show itself. But the sharper difference is downstream: this framework routes the result through reflexive counterplay to a paid probe that must return an answer the firm did not already hold, and it reports a number — options eliminated, with the evidence — that scenario planning has never reported because it does not produce one.
That is not a dismissal of the practice. Scenario planning has been shown to change how field experts judge long-range investment decisions12, and its own literature is more careful than its practitioners: a review of reviews concludes that while research into scenario planning's effectiveness is scarce, specific techniques do show evidence of impact13, and the most recent MIT Sloan treatment argues that reaping the full benefits requires establishing it as an ongoing capability linked to other processes14 — which is, almost word for word, this argument. The charge here is narrow: an exercise that eliminates nothing has produced no evidence, whatever it produced in insight.
"You can't run experiments on strategic questions."
Most of the time you can, once you stop demanding that a single probe settle the entire future. The probe must discriminate between two live explanations, not prove one. A price change on two engagements. A published offer with a stated hypothesis and a specified audience. A scoped build with a named acceptance event. A deliberate refusal, to see who objects.
Where genuinely no probe exists inside a horizon that matters, say so in the row and carry the option on inference with a named trigger. The failure is not carrying an untested option. The failure is not knowing which of your options are untested.
"This is lean startup for boards."
Build-measure-learn optimises a product hypothesis inside a known business model. This operates a level up, where the unit of sale itself is the thing under test — and the counterplay step has no analogue in a product loop, because a landing page does not have a regulator.
The closer relative is real options in the plain commercial sense: paying a known amount to buy information and rights that change the decision set, with proceed, defer, reduce, transfer and do-nothing all remaining legitimate outcomes17. Not a Black–Scholes exercise, no invented premiums, no spreadsheet cosplay29.
What the evidence outside our walls suggests
One caution before the Monday list, because this piece's central claim is not proven and should not be dressed as if it were.
Nobody publishes a benchmark for strategic probes or option kills per quarter, by industry, because nobody counts them. There is no measured comparison of firms that raised falsification throughput against firms that raised generation rate. That comparison is the claim, and it is currently an argument rather than a finding. Its falsifier is straightforward: if firms that count kills navigate no better than firms that do not, holding capital constant, the thesis fails.
What we do have is strongly suggestive from the one domain where organisations actually measure elimination at scale. In a study at Microsoft, around two-thirds of all changes had a negative or neutral effect on the metric they were designed to improve; in well-optimised domains like Bing and Google, only about 10 to 20 per cent of changes have a positive effect on the target metrics7. Read that as a statement about confidence: where organisations do buy the answer, the majority of their considered, expert-designed ideas turn out to be wrong. There is no reason to think boardroom intuitions about market structure are better calibrated than product intuitions about a search results page — and every reason to think they are worse, because they are tested less often. These are product experiments, not strategic probes, and the transfer is an analogy rather than a proof. It is still the most honest available estimate of how much of what you currently believe is wrong.
Meanwhile the operating-portfolio side has its own sobering number. MIT's NANDA research reports that 95 per cent of organisations in its sample are getting zero return from generative AI initiatives, with just 5 per cent of integrated pilots extracting real value — a divide the authors attribute to approach rather than to model quality or regulation10. Cite it carefully: it concerns pilots and measurable P&L impact, from a preliminary-findings study. But it does undercut the comfortable assumption that the operating portfolio is the safe one and the option portfolio is the speculative one. Both are bets. Only one of them is currently being reported as though it were not.
What to do on Monday
Three moves. None of them requires a budget approval, and the first two can be done this quarter by people who already work for you.
1. Measure both clocks, once
One page, four quarters of history.
Branching inputs (count them from public information): new entrants in your segment; new offers or pricing models launched by incumbents; platform or model capabilities that became purchasable and were not before; regulatory changes that opened or closed a move. You are not forecasting. You are counting how many new moves appeared on the board.
Evidence outputs (count them from your own records): probes shipped that could have returned a result you did not want; options formally closed, with the evidence that closed them; the median elapsed time from question raised to evidence received.
Then divide. Most firms doing this for the first time will find the denominator is zero, at which point the ratio is undefined and the diagnosis is complete. That is not a failure of the exercise; it is the exercise working.
2. Convert one existing row into a probe
Take the strategic question your team has been carrying longest — the one that reappears in every planning cycle and never resolves. Run it through the six steps. Insist on step three; the counterplay is where most conviction quietly dies. Then write the smallest thing you could do in the world that would produce an observation at least one of your explanations cannot survive.
Write the kill trigger before you run it. Assign the row to a person, not a committee.
3. Change one standing board question
The current question is what is our AI strategy?, and it will be answered with the operating portfolio because that is the portfolio with metrics.
Replace it with two:
What did we eliminate last quarter, and what killed it?
Which consequential uncertainty are we forcing the world to answer next?
The first cannot be answered with a deck. The second cannot be answered with a forecast. Between them they will reorganise what gets prepared for the meeting, which is the only reliable way to reorganise what gets done.
The fog is not lifting. Your own search apparatus is part of the reason — it manufactures new candidate futures as a by-product of working well, and so does everybody else's, and none of theirs needs your permission26. You will not find dry land, and the strategy industry should stop selling maps of it.
But the ratio locates agency precisely, and that is the useful part. One of the two clocks belongs entirely to you. You cannot slow the rate at which the world invents new moves. You can decide, this quarter, how fast you are willing to pay to find out which of them are false.
AI makes variation cheap, not truth. Boundary cases aim cognition at the assumptions that matter; adversarial counterplay exposes the reflexive responses; paid and deployed probes let the world decide. The board's job is not to predict the clearing of the fog, but to operate an evidence-producing portfolio of options until the successor earns the right to be built.
One ask: run the two-clock count for a single revenue unit before your next board meeting — branching inputs on the left, options killed on the right — and take the ratio into the room. If the right-hand column is empty, you have just found the most valuable thing on the agenda.
References
- Crunchbase News. "Q1 2026 Shatters Venture Funding Records As AI Boom Pushes Global Startup Investment To $300B." — "Overall, AI shattered records last quarter, with $242 billion — 80% of total global venture funding in Q1 — going to companies in the sector. The previous record was set in Q1 2025, when AI accounted for 55% of global venture funding." news.crunchbase.com/venture/record-breaking-funding-ai-global-q1-2026
- IEEE Spectrum. "Stanford's AI Index for 2026 Shows the State of the Industry." — "Epoch AI tracked 87 notable model releases from industry in 2025, compared to just seven from all other sources." spectrum.ieee.org/state-of-ai-index-2026
- Stanford HAI. "Inside the AI Index: 12 Takeaways from the 2026 Report." — "the success rate of agents handling real-world tasks improved from 20% in 2025 to 77.3% today, according to Terminal-Bench." hai.stanford.edu/news/inside-the-ai-index-12-takeaways-from-the-2026-report
- Sapphire Ventures. "2026 Outlook: 10 AI Predictions Shaping Enterprise Infrastructure." — "Achieving $100M in ARR over 5–10 years used to be the gold standard in SaaS. The best-in-class AI-native companies are now compressing that timeline into 1–2 years." sapphireventures.com/blog/2026-outlook-10-ai-predictions-shaping-enterprise-infrastructure-the-next-wave-of-innovation
- Retool. "The Build vs. Buy Shift: AI, Shadow IT, and the SaaS Replacement Era" (2026 Build vs. Buy Report). — Survey of 817 builders and customers: 35% have already replaced at least one SaaS tool with a custom build, and 78% expect to build more of their own tools in 2026. retool.com/blog/ai-build-vs-buy-report-2026
- Oliver Wyman Forum. "The CEO Agenda 2026." — CEOs now devote half of all planning effort to horizons of less than one year, up from 43% the previous year; "Compressed time horizons, while understandable, might come at the cost of strategic clarity." oliverwymanforum.com/ceo-agenda/how-ceos-navigate-geopolitics-trade-technology-people.html
- Iavor Bojinov and Somit Gupta. "Online Experimentation: Benefits, Operational and Methodological Challenges, and Scaling Guide." Harvard Data Science Review, Issue 4.3 (Summer 2022), citing Kohavi et al. (2020). — "in a study at Microsoft, around two-thirds of all changes were found to have a negative or neutral effect on the metric they were designed to improve"; "in well-optimized domains like Bing and Google, only about 10% to 20% of changes have a positive effect on the target metrics." hdsr.mitpress.mit.edu/pub/aj31wj81
- Ron Kohavi and Stefan Thomke. "The Surprising Power of Online Experiments: Getting the Most Out of A/B and Other Controlled Tests." Harvard Business Review 95, no. 5 (September–October 2017): 74–82. web-docs.stern.nyu.edu/executive/The%20Surprising%20Power%20of%20Online%20Experiments.pdf
- John R. Platt. "Strong Inference." Science 146, no. 3642 (16 October 1964): 347–353. — "Devising a crucial experiment (or several of them), with alternative possible outcomes, each of which will, as nearly as possible, exclude one or more of the hypotheses... It is like climbing a tree." dunnlab.ucsf.edu/sites/g/files/tkssra15436/files/wysiwyg/Platt%201964.pdf
- MIT NANDA (Challapally, Pease, Raskar, Chari). "The GenAI Divide: State of AI in Business 2025." July 2025. — "95% of organizations are getting zero return... Just 5% of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact." mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
- George Soros. "Fallibility, Reflexivity, and the Human Uncertainty Principle." Journal of Economic Methodology 20, no. 4 (December 2013): 309–329. — "In the cognitive function, the participant is cast in the role of a passive observer: the direction of causation is from the world to the mind. In the manipulative function, the participants play an active role: the direction of causation is from the mind to the world. Both functions are subject to fallibility." georgesoros.com/2014/01/13/fallibility-reflexivity-and-the-human-uncertainty-principle-2
- Shardul Phadnis, Chris Caplice and Yossi Sheffi. "How Scenario Planning Influences Strategic Decisions." MIT Sloan Management Review 57, no. 4 (Summer 2016): 21–24; and Phadnis, Caplice, Sheffi and Singh, "Effect of Scenario Planning on Field Experts' Judgment of Long-Range Investment Decisions," Strategic Management Journal 36, no. 9 (2015): 1401–1411. sloanreview.mit.edu/article/how-scenario-planning-influences-strategic-decisions
- K.E. Cordova-Pozo and E.A.J.A. Rouwette. "Types of scenario planning and their effectiveness: A review of reviews." Futures (2023). — "while there is indeed a scarcity of research into effectiveness of scenario planning, on the level of specific techniques there is evidence of impact." sciencedirect.com/science/article/pii/S0016328723000575
- Paul J.H. Schoemaker and Shardul S. Phadnis. "How to Make Scenario Planning Stick." MIT Sloan Management Review, Winter 2026. — "Establishing scenario planning as an ongoing capability and reaping its full benefits require linking it to other processes." sloanreview.mit.edu/article/scenario-planning-how-to-use
- Clio. "What's Driving Legal AI Pricing in 2026?" — 67% of corporate legal departments and 55% of law firms expect AI to change how hours are billed; 71% of buyers already prefer flat fees for an entire case. clio.com/resources/ai-for-lawyers/legal-ai-tool-pricing
- Thomson Reuters Institute. "Law Firm Rates Report 2026." — Firms collect roughly the same amount per hour whether they discount aggressively or hold firm on realization. thomsonreuters.com/en-us/posts/legal/law-firm-rates-report-2026/
- Timothy A. Luehrman. "Investment Opportunities as Real Options: Getting Started on the Numbers." Harvard Business Review, July–August 1998. — A corporate investment opportunity is the right, but not the obligation, to acquire something. hbr.org/1998/07/investment-opportunities-as-real-options-getting-started-on-the-numbers
- Data Center Dynamics. "IBM's mainframe business sees highest annual revenue in 20 years." (Q4 2025 earnings.) datacenterdynamics.com/en/news/ibms-mainframe-business-sees-highest-annual-revenue-in-20-years
- Hypercubic, citing ZipRecruiter (March 2026). "COBOL Job Postings Over Time: Salary Trends." — Average US COBOL developer $115,475/yr, above the median software developer; "Pundits frame COBOL as a relic. Hiring managers frame it as a staffing crisis." hypercubic.ai/insights/cobol-job-postings-over-time-salary-trends-and-which-industries-are-still-hiring
- TechCrunch (7 December 2018). "IBM selling Lotus Notes/Domino business to HCL for $1.8B." techcrunch.com/2018/12/07/ibm-selling-lotus-notes-domino-business-to-hcl-for-1-8b
- Anthropic. "Claude Mythos Preview System Card." — "demonstrated a striking leap in cyber capabilities relative to prior models, including the ability to autonomously discover and exploit zero-day vulnerabilities in major operating systems and web browsers." www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf
- McKinsey & Company. "From AI Table Stakes to AI Advantage: Building Competitive Moats." — Agentic AI could orchestrate up to $1 trillion in US retail by 2030. mckinsey.com/capabilities/quantumblack/our-insights/from-ai-table-stakes-to-ai-advantage-building-competitive-moats
- Scott Farrell, LeverageAI. "Cheap Thinking Makes Strategy Harder." — The two-term Fog and its endogenous second term; the pivot test (shape, not size); the Strategic Search Ledger and its decision-delta closure rule; the boundedness test; over-search and the decision-neutral stopping rule. leverageai.com.au/wp-content/media/articles/227-cheap-thinking-makes-strategy-harder.html
- Scott Farrell, LeverageAI. "The Terminal Value Doctrine: Professional Services." — The runway and proof clocks and the survival inequality; the four-part valuation decomposition for a declining category; harvest without denial. leverageai.com.au/wp-content/media/articles/231-terminal-value-doctrine-professional-services.html
- Scott Farrell, LeverageAI. "The Terminal Value Doctrine — Stop Optimising the Horse." — The AI Fog defined as simultaneous horizon compression and solution-space expansion; the Question Ledger as the evidence pack for search quality. leverageai.com.au/wp-content/media/articles/61-terminal-value-doctrine.html
- Scott Farrell, LeverageAI. "The Cognition Dimension Ladder." — Permanent Fog: the discovery engine manufactures fog as a side-effect of being good at its job; boundary cases as a recurring navigation instrument. leverageai.com.au/wp-content/media/articles/62-cognition-dimension-ladder.html
- Scott Farrell, LeverageAI. "Semantic Experiment Graph." — Mutation theatre; the scarce skill is designing a candidate that would falsify a mechanism; the residue checklist a real experiment must leave behind; heat is not truth authority. leverageai.com.au/wp-content/media/articles/157-semantic-experiment-graph.html
- Scott Farrell, LeverageAI. "Cycle Compression: The Breathing Flywheel." — The enabling inequality: AI has shortened world-response latency below the decay time of the original thought. leverageai.com.au/wp-content/media/articles/121-cycle-compression.html
- Scott Farrell, LeverageAI. "Buy Certainty First." — A real option on transformation in the plain commercial sense; correct non-exercise as a legitimate, high-value resolution. leverageai.com.au/wp-content/media/articles/204-buy-certainty-first.html
- Scott Farrell, LeverageAI. "The Reshape — A Field Guide to Thought Experiments in the Age of AI." — The pivot: push one variable to an absurd extreme until the geometry of the situation forces a structural answer; why AI cannot choose the variable. leverageai.com.au/wp-content/media/articles/60-the-reshape.html
