The Client Owns Yesterday
The Evolution Mandate

The Client Owns Yesterday

Recurring strategy revenue without dependence

Permanent fog creates a permanent strategic function inside your client. It creates no entitlement for you to be its permanent supplier.

So the recurring fee has to attach to something a board can hold, a buyer can audit — and a supplier can lose.

By the end of this book you can

  • ✓ Say what your recurring fee is attached to, and what changed this cycle
  • ✓ Cap an option portfolio so a cycle has to close something before it opens something
  • ✓ Separate who proposes, who disposes, who constructs and who verifies — in a firm of one
  • ✓ Design a termination boundary that preserves operation and ends evolution
  • ✓ Split a monthly number into components that each name what they refuse
  • ✓ Write the two-quarter conditions under which you would tell a client to stop buying

Scott Farrell · LeverageAI

01
Part I: What the Fee Is For

A Billing Cadence Is Not a Product

I named a number and a justification, and mistook them for a commercial object. Here is what was actually missing.

TL;DR

  • A monthly figure is an invoice frequency. If you cannot say what the client holds at any moment, you have not sold them anything — you have sold them a standing appointment.
  • Two of my own published books fenced this question and named it a later volume. This is that volume, and the price label does not get upgraded on the way in.
  • The three shapes a recurring advisory offer usually collapses into — the availability retainer, the hostage, the shadow principal — are all failures, and each one gets a chapter.

The Sentence

I said it out loud in a strategy conversation and regretted it in about two seconds.

"I charge them an ongoing fee — say, a hundred thousand a month. Not because I've made myself invaluable or indispensable. It's because they're still in their fog, and they need my input on the next part."

Two things were wrong with that sentence, and the first one is easy.

The number has never been paid. It has no market evidence behind it at all. It appears in this book as a designed hypothesis and nothing else, and it carries that label every single time it appears. The reason is not caution. It is that publishing a price I have never charged as though it were a price would be exactly the status inflation I spend my working life telling other firms to stop committing.

The second thing wrong is the one this book is about.

A hundred thousand a month is a billing cadence, not a product.

I had named an invoice frequency and a justification. I had not named a commercial object. There was nothing in that sentence a client could hold, a board could inspect, or a competitor could be compared against. And that gap is not a wording problem you fix in the proposal. It is the reason recurring advisory revenue keeps collapsing back into the shape everybody says they don't want.

Notice, though, what else happened in that sentence, because the rhythm recurs through this entire book and it is the method rather than a flourish. I asserted the fee, and in the same breath I amputated the reason that would have made it comfortable. Not because I've made myself indispensable. Assert, then remove the comfortable justification, then find out whether anything is left standing. That move is uncomfortable to perform on your own revenue, and it is the only way I know to design a recurring offer that survives its second renewal.

Designed hypothesis, not a price

Every appearance of the A$100k/month figure in this book carries this label. It is a design object — a number chosen to make the shape of an offer thinkable — and it has never been validated by a transaction.

It was already published under that label before this book existed. In the offer ladder in The Terminal Value Doctrine for Professional Services, the rung-three row reads: Ongoing kernel access — order of A$100k / month — continuing — "Designed hypothesis: the rung this book fences to a later volume."

The Baton

I have been circling this question for two books without landing it, and saying so is not throat-clearing — it is the reason this one exists.

In The Terminal Value Doctrine for Professional Services I published a three-rung offer ladder. Rung one is a bounded decision: fixed price, fixed clock, decision-complete, with a stand-pat verdict as a successful outcome rather than a failed funnel stage. Rung two is a bounded construction: one commercial offer, one working delivery vessel, one production proof, one launch and transfer package — all four or it isn't done. And rung three got exactly one sentence, on purpose: beyond the build lies a continuing mandate to keep the successor current as models and markets move, and its governance, rails and economics are a later book's subject.

Then, in Fog Is a Race Between Two Clocks, I spent a book arguing that what transfers to a client is the discipline and not the conclusions — because conclusions decay at the branching rate and an adviser's map of your sector ages exactly as fast as everyone else's. That book's reflexive chapter, the one that points its own standard at its author, closes on this:

"What a continuing relationship is for, once the discipline has transferred, is a real question with a real answer, and it is a sibling subject rather than this book's."

This is that volume. And the price label does not get upgraded on the way in.

There is one rule I inherit along with the question, and it constrains everything that follows. The ladder binds all three rungs with an anti-funnel rule: each rung is a complete purchase, priced on its own value, with its own acceptance. The moment rung one is priced as a loss-leader for rung two, it stops being a decision product and becomes a sales document with a fee. Rung three does not get an exemption. Whatever the recurring object turns out to be, it cannot be justified as the thing that makes the first two engagements worth selling — which rules out the most commercially attractive way to describe it.

What Happens When the Engagement Ends Well

Your version of my two-second regret arrives at a specific moment. The bounded engagement has landed. The client is pleased, in the particular way clients are pleased when something they were told would be hard turned out to be finishable. Somebody says: so — what does ongoing look like?

And you have two instincts that do not fit together. Recurring revenue is good, and you would like some. You also have a settled conviction, probably arrived at by watching somebody else do it badly, that you do not want to be the person a client cannot leave.

In the absence of a designed object, that moment collapses into one of three shapes. All three are failures. Each gets a chapter later in this book, so here they are named crisply and dropped.

Three shapes, three failures

The availability retainer
  • Unit of sale: hours of access to a scarce person.
  • What makes it renew: having been responsive.
  • From the buyer's chair: "we're paying a lot and I can't explain to the board what changed."
The hostage
  • Unit of sale: continued access to machinery, history or data only you can reach.
  • What makes it renew: the cost of leaving.
  • From the buyer's chair: excellent renewal metrics, right up until it ends badly.
The shadow principal
  • Unit of sale: a strategy engine you both diagnose with and sell through.
  • What makes it renew: the options you found, that you then build.
  • From the buyer's chair: nobody is lying, and the gradient still points at you.

The first is the one everybody names. The second is the one everybody swears they would never build. The third is the one almost nobody has a word for, which is why it is the most common.

Three Comfortable Answers, All Foreclosed

The question underneath all of this is why do they keep paying? There are three answers close to hand, and I have written myself out of every one of them.

"Because it's hard to leave."

This one contradicts everything I have published about copy-path economics. A sophisticated partner does not experience switching cost as loyalty; they experience it as a number, and they cost the copy path quietly while telling you the relationship is going well. An offer whose renewal depends on that number is an offer with a countdown on it.

"Because I know how it all works."

That is founder dependence wearing a nicer suit. And it is precisely the failure my own anti-dependency work exists to prevent. The test cuts both ways — and the half most readers skip is the half about the creator, not the buyer.

"The partner becomes capable of delivering the current product. You remain the fastest, safest and most credible way to improve the product, govern its evolution and create the next five."

If you remain the execution runtime, you never get paid for the factory, and every next offer is still your calendar. Chapter 5 does the work on this; here it is enough to say that indispensability is not a moat I am allowed to reach for, because I have already published the argument against it and I have to live under my own doctrine.

I said it more bluntly at the time: I've already written that you can't make yourself indispensable. I can't put myself in the direct path. I have to give them the ability to execute their own tools. That is a design constraint that costs me money, and I am not going to abandon it the moment a recurring invoice appears on the horizon.

"Because they're in fog."

This is the answer that feels most defensible and is the most quietly hollow. Fog is a condition, not a deliverable. It has no boundary, no acceptance test and no completion state. And it is exactly the argument a self-interested adviser would reach for, which means a competent buyer should discount it.

Worse, my own doctrine already establishes that a firm which becomes genuinely capable can internalise a great deal of what it used to buy — and that a published body of doctrine is precisely the artefact a client-side AI consumes best. Subscription advisory sold to a client who can run the same engines over the same kernel is not automatically a destination. It may be another temporary migration form, dissolving on a schedule the client sets.

Chapter 2 takes that argument seriously, because it is the strongest thing anyone can say against this book and it deserves to be answered rather than absorbed.

The Answer That Survives

Strip out the three foreclosed answers and something narrow is left. It is not a slogan. It is a division:

Key Insight

Yesterday's solution has transferred. Tomorrow's solution hasn't been discovered yet.

Everything the client has already learned, already operates, already decided — that is theirs, and no part of my fee should be gating it. What is being bought is the next thing: the frontier that has not been mapped, the option that has not been formed, the doctrine release that has not shipped. Not access. Not availability. Not the accumulated corpus. The next thing.

That object has a name in this book. I call it an evolution mandate: a maintained portfolio of strategic options plus a licensed living kernel, governed so that renewal is earned rather than structural. One sentence is all it gets here, because a definition is not a specification, and the specification is the remaining fourteen chapters.

What This Book Is Not About

One paragraph of fence, because three adjacent arguments will otherwise walk in and take over.

This is the recurring layer for any expertise vendor — not a consulting-industry account. What AI pressure is doing to the economics of professional services is a sibling book's subject and it is already published. So is the machinery for transferring a founder's judgment into an organisation, which is a different problem with different instruments. So is the question of how a finished engagement becomes higher-order strategic learning. Each of those gets named once more in this book, at the point where a reader might reasonably expect it, and then dropped.

What you should be able to do by the end is narrower and more useful than any of them: say what your recurring fee is attached to, who is allowed to decide what, what survives cancellation, and what evidence would prove the whole arrangement is not worth buying.

That last one is the difficult part, and it is why the book ends where it does rather than with a close.

02
Part I: What the Fee Is For

Permanent Function, No Permanent Supplier

Why the strategic function never closes — and why that argument, made honestly, says nothing about who should staff it.

Chapter 1 named an object and promised a specification. Before any specification, one question has to be settled, because a product with no defensible reason to recur will drift back into availability no matter how carefully it is drawn.

Permanent Fog creates a permanent function. It does not entitle you to be the permanent supplier of that function.

That sentence looks obvious for about four seconds. Then you notice that every recurring advisory offer in the market is an implicit claim about its second half, and that almost none of them make the claim explicitly — because making it explicitly invites the question this chapter has to answer.

So the chapter has two jobs. First, prove the function is permanent using an argument that does not depend on my commercial interest. Second, show what — if anything — makes a supplier legitimate once that is established.

Why "You're in Fog" Is Not a Product Argument

The fog position is mine and I stand by it. Its strong form is worse than the version most people carry: the discovery engine does not dispel the fog. It manufactures fog as a side-effect of being good at its job. Every cycle surfaces new candidate moves; new candidates expand the solution space; an expanded solution space deepens the fog. Which is why the right posture is not to run the engine until conditions clear, but to build an engine that is a productive operator inside permanent fog, refreshing its chart on a schedule the way a sailor does rather than waiting for dry land.

All of that is true, and none of it is a product argument.

"You are in fog" is a condition. It has no boundary, no acceptance test, and no completion state. A customer cannot accept it in the ordinary commercial sense of accepting something, because there is nothing to accept. Sell the condition and you are selling weather.

It is also, precisely, the argument a self-interested adviser would reach for — which means a competent buyer should discount it on arrival. That is not cynicism about buyers. It is what makes the next section necessary: the case for recurrence has to survive being made by somebody who profits from it.

The Boundedness Test

Here is the argument that holds whether or not I am in the room.

Cheap cognition reduces uncertainty under three conditions: the population is enumerable, the evaluation function is stable, and the action set is closed. Break any one of them and more cognition does the opposite.

Walk the clearest case, because the distinction only becomes physical when you do.

Auditing has sampled for its entire professional history. Not because anyone believed sampling was a good way to find problems, but because examining everything was impossible at human cost. An elaborate and genuinely defensible apparatus grew up around that constraint — materiality thresholds, risk tiers, sampling schedules, confidence intervals — and every part of it is downstream of the fact that a person had to read each item.

Now check the three conditions. The population of transactions is finite and countable before you start, and the count does not change because you got better at examining. The test is stable: a duplicate payment is a duplicate payment, and it does not become something else because a competitor changed their pricing. And more thinking converges — each additional unit of cognition reduces the unexamined remainder rather than creating new categories of thing to examine.

"Sampling was never a methodology anyone loved. It was a budget wearing a methodology's clothes. Dissolve the budget and the uncertainty falls with it."

Now run strategy through the same three conditions.

The population of futures is not enumerable — you cannot write the list before you start, and the list grows as you work. The evaluation function is not stable, because what counts as a good strategy is defined partly by other people's moves, which is exactly what makes it strategy rather than optimisation. And the action set is emphatically not closed: every discovery creates new kinds of option, not merely more instances of a known kind.

Three conditions. Audit passes all three. Strategy fails all three.

Key Insight

The same technology that clears an audit thickens a strategy, and it is not behaving inconsistently. The questions have different structure.

Be precise about what that proves, because the temptation is to take one step too many. It proves the strategic function does not converge, and therefore does not close. A firm will need that function for as long as its action set stays open, which is indefinitely.

It says nothing whatsoever about who should hold it. That is the hinge of this book, and everything after it is downstream of the turn.

The Market Is Moving Both Ways at Once

Two facts, and they sit together to produce this book's operating condition.

The demand side first. Half of the most expensive attention in the enterprise now sits inside a horizon shorter than most strategy engagements take to deliver, and the share is rising year on year.

Three numbers describing the weather

50%

of CEO time now spent planning for horizons under one year — up from 43% the year before1

$9.39b

Accenture Managed Services revenue in Q3 FY26 — overtaking Consulting at $9.33b2

>40%

of BCG's $14.4b revenue now tech- and AI-focused3

That first figure is not a scare statistic. It is a structural mismatch: a one-off engagement ends before its own consequences arrive, into an attention profile that has already moved on.

The supply side has responded, and it has responded by selling more recurring service, not less. Accenture is the only large advisory firm that publishes the split quarterly, and two quarters before the revenue crossover the bookings gap was wider still: $11.06 billion of Managed Services new bookings against $9.88 billion of Consulting, with $2.2 billion of advanced-AI bookings inside the same quarter.4

BCG, meanwhile, describes its own move in language that should stop any IP-heavy firm mid-stride — the firm has been "embedding proprietary knowledge, data, and proven delivery approaches into reusable, human-led agentic processes that accelerate impact."3 That is kernel compilation, announced in a revenue release. And McKinsey's global managing partner is publicly framed as driving an organisational transformation "to focus less on traditional consulting services and more on delivering outcomes" — a sentence taken from the freely readable editorial framing of the interview, since the interview itself sits behind a paywall.5

Now the fact that makes those two paragraphs interesting rather than reassuring. That conversion is happening into a shrinking pool. On the UK numbers, 2025 was the sector's worst performance since the lockdown period — flat or negative depending on how you define the market.6

Put those together and you get the condition this book's object has to survive. A category converting to recurring revenue while its total demand contracts is a category about to produce a great many dependency products wearing partnership language. Not through bad faith. Through the ordinary pressure of needing next year's number to look like this year's.

The Reflexive Test

Now the uncomfortable version of the question, and it is uncomfortable because I built the machinery that raises it.

If cheap cognition lets clients internalise research, analysis, software and first-pass advice — which is the argument I have been making to them, at length, in print — why would it not also let them internalise strategic navigation?

This is not a hypothetical objection. I have written the design. Point a patient, long-horizon model at three compiled corpora — what a firm believes, what it can prove, and who it knows — and the strategic question changes category. It stops being generation, which is inventing plausible strategy, and becomes search over an actual option space. The division of labour in that design is explicit and it is the part that matters here: the machine proposes with receipts; you dispose. And: "your decision authority is the one component in this whole system that was never for rent." That argument is The Strategy Engine, and it is not published as an article — so it is named here rather than linked.

I wrote that. It applies to me. A client who compiles those three corpora has, by my own argument, acquired most of a strategy function — and the part they have acquired is exactly the part I would otherwise be selling.

So the answer cannot be because they still need me. It has to be a bounded, observable delta produced beyond the client's own apparatus. And it has to be measurable, or it is a feeling with an invoice attached. Everything from Chapter 12 onward exists to measure that delta, including the possibility that it comes back at zero.

So What Is Recurrence Actually For?

Here is the mechanism, and it is the part most recurring offers never articulate because they have never needed to.

Strategic learning runs on at least three clocks.

The epistemic clock

The conversation changes a hypothesis. Fast. This is the clock a good workshop runs on, and the only one most engagements ever close.

The behavioural clock

Money, authority or attention actually moves. Slower, and gated by things an adviser does not control — budget cycles, politics, somebody's willingness to be accountable for a number.

The outcome clock

Reality reveals whether the move worked. Slowest, and the only one that can tell you whether the judgment behind the recommendation was any good.

A bounded engagement almost always ends before the third clock closes. It has to — that is what bounded means, and I have argued at length that bounding is the right way to sell the first two rungs of the ladder. But it leaves something structurally unfinished: the engagement produces a decision and then loses the ability to find out whether the decision was right.

Which gives recurrence a job. A recurring mandate is structurally legitimate because it can carry open questions across the outcome clock, backfill what actually happened, and update the evaluation function that produced the original recommendation. Recurrence is justified by that closure work. Not by availability, and not by the weather.

What a period produces vs what a period must close

✗ What a period usually produces

  • • Analysis, and more of it than last quarter
  • • New options, argued and plausible
  • • A board narrative nobody was embarrassed by
  • • A good conversation everybody remembers

All real value. None of it closes a clock.

✓ What a period must close

  • • An assumption moved, with the evidence that moved it
  • • An option disposed — killed, deferred to a named trigger, exercised or transferred
  • • A question promoted into something the client runs themselves
  • • An earlier decision's outcome backfilled against what was predicted

Each one is checkable by somebody who does not work for you.

And the falsifier the clocks hand you, which Chapter 14 turns into a renewal instrument: if a period closed no clocks, that period was not worth buying. Not "was disappointing". Was not worth buying.

Two Warnings to Carry Forward

The first is the one I have to say about my own product before anybody else says it. A discovery engine that generates more possibilities creates more fog, and the adviser who caused the expansion is then the obvious person to help manage it. Every individual step in that chain is competent and useful. The aggregate is a dependency, and the adviser is paid at every stage of it. That is the problem Chapter 4 answers with an instrument rather than a promise, so I will leave it sitting there uncomfortably until then.

The second is a boundary, and it belongs early so that you distrust this book correctly. A published kernel is exactly the artefact a client-side AI consumes best — legibility is what makes doctrine publishable in the first place. Which means subscription advisory sold to a client who can run the same engines over the same kernel is not automatically a terminal position. It may be another temporary migration form, dissolving on a schedule set by the client's compilation rate.

I ran that attack on my own parent doctrine in an earlier book, and what survived it was narrower than the original claim: the kernel's continuing refresh from cross-client variation — the client can ingest the snapshot, they cannot ingest the flywheel that keeps it current — plus structural independence, accountable consequence, and proprietary machinery kept deliberately ahead, which is a race and must be priced as one rather than booked as a moat.

Everything this book builds has to survive that paragraph. Part VI is where it gets tested rather than argued.

There is a third thing I will name in one sentence and then leave alone, because it belongs to another book: most firms see AI as something that will help their workflow, and I see it as an industry pressure. That distinction changes what kind of question you are answering, and it is the spine of the professional-services account rather than of this one.

Where That Leaves Us

Two halves, cleanly separated. The function is permanent because the questions have a structure cheap cognition cannot converge on — an argument that survives being made by an interested party, which is why it is the one worth keeping. The supplier is not permanent, because nothing about that structure names who should hold it, and my own published work makes the case for the client holding more of it every year.

Which produces the next question, and it is a design question rather than a philosophical one: if recurrence has a job, what is the object the client holds while the job is being done?

03
Part II: The Unit of Sale

The Maintained State

What the client holds at any moment — and the five things a period has to change before it counts as having happened.

At Any Moment, What Do I Actually Have?

That is the buyer's version of the question the last chapter ended on, and it is the only test that matters at the start of a period. Not what will you do for us. Not how will we work together. Just: at this instant, before anybody does anything new, what is it I own?

A retainer answers you have access. That is a real answer, and it is why retainers sell. It is also why they cannot be defended at the second renewal, because access has no state — there is nothing that accumulated, nothing that changed, nothing to point at.

The mandate answers with an object.

At any point, the board has a current, evidence-backed portfolio of assumptions, threats, options, experiments, decisions and reopening triggers — and the organisational machinery to act on it.

Read that as a specification rather than a sentence, because every noun in it is load-bearing and a reader who skims will file it as a slogan.

  • Current — it has a freshness property. Which means it can be stale. Which means staleness is observable, and somebody can be asked about it.
  • Evidence-backed — every entry can name what supports it. Not "we discussed this"; what supports it.
  • Portfolio — a bounded set under management. Not an accumulation, which is the failure Chapter 4 is entirely about.
  • Reopening triggers — the decisions carry the conditions that would reverse them. A decision without a reopening condition is a decision nobody can revisit without losing face, which is how firms end up defending positions they no longer hold.
  • And the organisational machinery to act on it — the state without the capacity to move on it is a document. An expensive, well-argued document.

That sentence is also the map of the remaining thirteen chapters, and it is the only map you need: everything else in this book either produces that object, governs it, prices it, or tests it. Part III decides who is allowed to change it. Part IV decides what happens to it when the money stops. Part V is the interface onto it. Part VI asks whether any of it was worth paying for.

  The availability retainer The evolution mandate
Unit of sale Hours of access to a scarce person A maintained state the client owns
What "a good month" means You were responsive, present and useful The state changed, in a way the board can point at

The second row is harder to sell and easier to defend. That trade is the whole commercial bet of this book, and it is worth being honest that it is a trade rather than a free upgrade — the first row closes faster, every time.

A Cycle Is Defined by What Changed

Which produces the operating rule. A cycle is not defined by what happened in it. It is defined by what changed, and there are exactly five objects that can change.

1. Terminal-value assumptions

Which premise moved, what evidence moved it, and which premise should now be reopened. These are the load-bearing beliefs about where the business's value is going — the ones that, if wrong, make everything downstream wrong too.

The failure tell: an assumptions register that only ever gains rows. If nothing was ever demoted, nothing was ever really tested.

2. The option portfolio

What was created, strengthened, weakened, killed, deferred, or approved for construction. Chapter 4 does the full treatment; here it is one of five, and deliberately not the centre.

The failure tell: growth in both directions at once — more options, and more of them live.

3. Standing questions

Which one-off inquiry became a client-owned recurring capability, and which stale question was retired. This is the transfer object, and the next section walks one all the way through.

The failure tell: a register of questions with my name in the owner column.

4. Decision receipts

What the board decided, against which evidence, which rejected alternatives, and which falsifiers. The rejected alternatives are the part that gets dropped first and matters most — a decision record without them cannot be audited later, and cannot be handed to another adviser without a re-briefing that costs more than the decision did.

The failure tell: minutes rather than receipts.

5. The capability release

Which tests, rules, playbooks, tools or kernel components now let the client handle more without me.

The failure tell: nothing shipped, and nobody noticed.

The fifth is what makes the whole thing honest, and it is the one a supplier has every incentive to let slide.

Key Insight

A period with no capability release is a period that increased dependence — whatever else it produced.

The obvious objection arrives immediately and deserves answering rather than absorbing: some periods genuinely have no release. The work was all evidence-gathering, and the evidence wasn't finished.

Three responses, in order of how often they turn out to be the real answer. Either the release is the evidence instrument — the harness, the test, the collection protocol the client now runs themselves — in which case ship it and say so. Or the period should have been shorter, and the cadence is wrong. Or the mandate is in the state Chapter 14 calls a stop condition, and the two-quarter clock has started. The rule does not bend. The period's shape does.

The Standing Question, Walked

The third object is where transfer actually happens, so it is worth doing properly rather than naming.

A standing question is not a recurring consultant discussion, and it is not a dashboard KPI. It is a promoted artefact with a schema, and the schema is what makes the promotion real rather than rhetorical.

It carries seven things, and each one buys you something specific:

A standing question carries… …so that
The question, and why it matters Its purpose survives the person who asked it
The entities and source types it covers Its scope is explicit and auditable
Its normal baseline and the signals that perturb it A change is legible as a change
Evidence requirements and escalation criteria A finding has a defined bar before it reaches a human
Cadence and a responsible human owner It runs on a schedule and somebody owns it
Conditions for revision or retirement It can be improved or killed, not left to rot

Now populate it, because a schema with nothing in it proves nothing.

Worked: promoting one question

The one-off, as it arrives. Somebody in a cycle meeting asks: why do two of our service lines selling the same offer produce different margin? The first time, that is an investigation. A walk through the evidence, an answer, a disposition, and everyone moves on. Useful, and gone.

The promoted, generalised form. Across all service lines running this offer, identify material divergence in delivery effort, exception handling, discount behaviour and senior intervention.

Scope: the named offer, across every service line running it, reading delivery timesheets, discount approvals, exception logs and the escalation record. Named sources, so somebody can argue about whether the right ones are in there.

Baseline and perturbation: the current spread between lines, plus the signals that would count as a real perturbation rather than noise — a new line taking on the offer, a pricing change, a senior departure.

Evidence bar: divergence above a stated threshold, sustained across two cycles, before it reaches a human. Below that, it is logged and not escalated.

Cadence and owner: monthly, owned by a named person in the client's own commercial team. Not by me. The owner field is where transfer either happens or quietly doesn't.

Retirement condition: kill it when three consecutive cycles produce no divergence above threshold. Written down now, while nobody is attached to it.

What just happened. That question has stopped being something they buy and become something they run. It is small, specific and repeatable — which is exactly why it works where "capability building" does not. Nobody has to be trained. Something has to be handed over, with an owner's name on it.

The rhythm this produces is worth stating in one line, because it is what makes the relationship compound rather than repeat: new uncertainty, bounded inquiry, decision, client capability, next frontier. Each turn should leave the client holding something they did not previously operate.

And the loop underneath it is not mine — it is borrowed intact from assurance work, where a human asks a good question, the machinery conducts a deep walk, humans review the result, the useful question is promoted, it runs repeatedly across the estate, new findings and variations emerge, and the panel improves the question. The line that captures why this matters commercially: traditional consulting answers a question and leaves. This turns a good question into a permanent institutional capability.

Who Decides, and Who Supplies

One correction before Part II goes any further, because a misreading here would run through everything: the maintained state is not my working notes with a client's name on them.

The division of labour comes from the same assurance discipline, and it is clean. The panel — the client's people, with authority — provides which questions matter, significance, accountability, proportionality, the challenge to evidence and to alternative interpretations, the promotion of valuable questions and the retirement of noisy ones, and the authority to act. The machinery provides breadth, persistence, parallelism and exhibits.

I sit with the panel on the challenge. I sit nowhere near the authority. Chapter 7 turns that into six checkable controls; here it is enough to know the boundary exists and which side of it I am on.

Which raises the client-side half of the same boundary, and it is a genuinely non-delegable list. Five things only a board can do:

  • Decide which revenue assumptions deserve to be attacked — because choosing what to doubt is a capital decision wearing an intellectual costume.
  • Insist that productivity and transformation are not reported as the same class of work.
  • Allocate capital between current operations and successor options, explicitly, as a number.
  • Define what evidence would justify commitment — in advance, which is the only time it can be done honestly.
  • Own the consequence of building, deferring or refusing.

None of those can be bought from me, and an adviser who quietly takes any of them on has not been generous — they have taken the client's authority and issued an invoice for it.

Where the State Lives

The state is client-owned. Not shared, not "ours", not held by me on their behalf. Chapter 8 builds the full architecture, but the practical form belongs here, because a reader who finishes Part II thinking the portfolio lives in my system has misread the entire book.

Concretely: the portfolio, the decision receipts and the standing questions live somewhere the client can read without asking me, and can export without asking me. That is the whole test. If the answer to can we see the register? is "I'll send it over", the state is not maintained. It is reported. And a reported state is a retainer deliverable with better vocabulary — it arrives when I choose, in the form I choose, containing what I chose to put in it.

Not a Month of Me

Which brings the chapter to the sentence it has been building toward.

The invoice can arrive monthly. The product is a maintained option-and-question portfolio — not a month of me.

Monthly billing was never the problem. A retainer-shaped product is. The cadence is administrative; the object is commercial; and confusing the two is how a good firm ends up selling its own calendar at a premium and calling the arrangement a partnership.

There is a flaw in what this chapter has just built, though, and it is large enough to have its own chapter. A portfolio that only grows is not a maintained state. It is inventory with a subscription attached — and the person best placed to keep adding to it is me.

04
Part II: The Unit of Sale

Dispositions, and Why the Portfolio Has a Cap

The instrument that stops a maintained state becoming an intellectual work-in-progress farm — and stops me being paid for filling it.

The person best placed to keep adding to that portfolio is me. So let me put the accusation on the table in its strongest form, before anyone else does.

Permanent Fog can become a self-licking ice cream cone. A discovery engine manufactures more Fog by finding more possibilities; an adviser can then monetise navigation through the Fog it helped enlarge. The antidote is to price and measure disposition, not idea volume.

I have written the longer version of that, and it removes every escape route I might have reached for. A search apparatus that works well manufactures fog as a by-product of working well. An adviser whose product is that apparatus enlarges the client's option set — and is then the obvious person to help manage the enlarged set.

Every individual step in that chain is competent, well-intentioned and genuinely useful. The aggregate is a dependency, and the adviser is paid at every stage of it.

And then the part that decides what kind of chapter this has to be:

"Notice what that mechanism does not require. No bad faith. No manufactured urgency. No strategic vagueness. An adviser who is simply very good at finding possibilities will produce this outcome by being good at it. Which is why the defence cannot be integrity. It has to be an instrument."

Everything below is that instrument. Four parts: what a cycle is allowed to do to an option; why "do nothing" has to be scored rather than assumed; where the portfolio sits in the board pack; and what I get measured on. The last one is the one that costs me.

Six Dispositions, and No Seventh

A cycle does one of six things to any live option. The fourth line on each card below is the one nobody writes down, and it is the one that makes the list useful.

Kill

Means: closed, with the evidence that closed it recorded.

Requires: a named falsifier that actually fired.

Costs the client: the sunk exploration, and the discomfort of somebody having been wrong in front of colleagues.

Dishonest form: killing quietly, so nobody can audit whether the kill was justified — or whether it was just the option that lost its sponsor.

Stand pat

Means: the current position is deliberately retained and scored. The next section is entirely about that second word.

Requires: an explicit evaluation of the status quo, entered as a candidate.

Costs the client: nothing, which is precisely why it gets skipped.

Dishonest form: stand-pat as a euphemism for nobody looked at it.

Defer until a named trigger

Means: not "later". A named, observable event, with somebody instrumented to notice it.

Requires: the trigger written down and an owner watching for it.

Costs the client: the attention of keeping one sensor alive.

Dishonest form: a trigger nobody is instrumented to notice. That is a kill with better manners.

Continue evidence collection

Means: with a stated question, a stated cost and an end date. All three.

Requires: naming what observation would move the option, not just "more work".

Costs the client: real money, in the least visible way any line item can.

Dishonest form: see the box below. This is the one that rots.

Exercise into a bounded successor proof

Means: the option earns a build — which opens a separate commercial conversation, not this one.

Requires: a separate authorisation. Chapter 7's firewall depends on that separation existing here, in the vocabulary, from the start.

Costs the client: capital, and the opportunity cost of the options this one displaces.

Dishonest form: exercise as the default happy ending, with the other five presented as ways of not deciding.

Transfer into ordinary client operation

Means: it stops being an option and becomes something they run.

Requires: an owner on their side and the machinery to operate it.

Costs the client: the absorption load of taking something on.

Costs me: money. Every time. Chapter 5 is that problem.

Pitfall — the disposition that rots

"Continue evidence collection" is the state a conflicted adviser reaches for when killing would be cheaper and more useful. It looks diligent. It consumes budget. And it postpones the moment anyone has to be wrong, indefinitely, at a monthly rate.

The fix is one field: put age on the register. Age is the number that embarrasses the adviser, which is exactly why it belongs there and exactly why it will be argued out if you let it.

There is a version of this failure in physical inventory that reads across perfectly: when a decision surface proposes thirty reserves and zero accepts, you are not looking at a high-criticality fleet — you are looking at diligence cosplay, and the right question is which exposures should be knowingly accepted before capital moves. A portfolio carrying six "continue evidence" rows and no kills is the same failure wearing strategy language.

Giving "Do Nothing" a Score

The second disposition needs a mechanism, and the cleanest one I know comes from a chess engine I wrote and then spent months not understanding.

An engine cannot search a position to checkmate, so it searches a few moves deep and evaluates. The trouble is that a fixed depth can chop the board off in the middle of an exchange — stop the instant after you have taken a pawn but before the recapture, and you will record that you are a pawn up while you are actually about to lose a bishop. The standard fix is a quiescence search: at the depth limit, keep searching the forcing moves until the board goes quiet, then score it.

But there is one more line, and I left it out.

"Quiescence isn't 'search all the captures.' It's 'search all the captures against the option of doing nothing.' The current static evaluation is a candidate move that always sits in the list."

Before you look at a single capture, you enter the current board's own evaluation as a candidate — a floor every action has to beat. Leave that line out and your floor starts at minus infinity. Every capture clears it. The side to move is compelled to grab something, and plays the resulting queen donation with total confidence, because by its own arithmetic it just chose the best available move.

One assignment. best = evaluate(board) instead of best = -INF. That is the entire difference between an engine that must capture and an engine that is allowed to keep its queen.

The same loop, read as portfolio governance

✗ The compelled loop

  • • The floor starts at minus infinity
  • • Every plausible option clears it
  • • "Do nothing" is not on the ballot
  • • Something gets exercised every cycle

Nobody had to want this. The loop produced it.

✓ The stand-pat loop

  • • The floor is the value of the current position
  • • An option is taken only if it beats doing nothing
  • • "Do nothing" is a scored candidate that can win
  • • The portfolio declines bad options cleanly

The status quo has to be beaten, not merely improved upon rhetorically.

The exception transfers too, and it is the sharper half of the rule. Stand-pat is disallowed when the engine is in check, because a check is a genuine threat that has to be answered — you cannot assume a quiet option even exists. So a mandate needs its equivalent: a named, pre-committed class of events that genuinely cannot be declined. A regulatory deadline. A contract expiring. A competitor crossing a stated threshold.

Takeaway

Doing nothing is the standing default. Only a genuine, must-answer event removes it. Everything else, you are allowed to leave alone.

Two Portfolios, Two Ledgers

The structural half of the cap is not about options at all. It is about which page of the board pack they appear on — and the mechanism is social rather than analytical, so it is worth watching one die.

The quarterly investment review has one page. On it, in the same table, sit an operating initiative and an option row. The operating initiative is a document-processing automation with a fourteen-month payback, a run-rate saving, a confident owner and a vendor quote. The option row is two practice areas repriced, margin genuinely at risk, a stated probability of being killed, and a residue — an outcome-sizing instrument, some acceptance tests — that nobody can value because nothing like it has been valued before.

Nobody argues against the option row. Somebody asks what the return is.

"The honest answer is 'we will know less wrong things', which is true, correct, and unsayable in a budget meeting. The row gets deferred to next quarter — not rejected, deferred, which feels like courtesy. Repeat that four times and the option portfolio is empty; repeat it for two years and the firm has concluded, without ever deciding, that it does not do this kind of work."

The comparison was made on the operating portfolio's terms, and on those terms the option row loses every time, correctly. Which is why the answer cannot be advocacy. It has to be structural.

Different sponsor, different ledger, different page in the board pack

The operating portfolio
  • Contents: workflow efficiency, quality, cost reduction, staff augmentation
  • Funded from: operating expenditure
  • Governed by: payback, run-rate, adoption, error rates
  • Good looks like: the work gets done faster or with less rework
  • Never described as: transformation
The terminal-value option portfolio
  • Contents: asset conversion, successor offers, customer-agent threats, new commercial units
  • Funded from: a separate allocation, sized deliberately, not from underspend
  • Governed by: falsification throughput — options closed with evidence, per quarter
  • Good looks like: a branch died and the firm knows why; or one earned serious capital
  • Failure looks like: rows added, nothing closed, everyone impressed

One clarification, so nobody merges instruments that do different jobs. A question ledger governs search quality — was the space actually searched, and where are the gaps. A strategic search ledger governs decision consequence — did anything actually change. The option portfolio governs capital and disposition — what did we pay to learn, and what died. What the portfolio adds that neither predecessor carries is specific: money, a trigger, and a compounding residue. A ledger row can be excellent and cost nothing; a portfolio row has a probe attached with a price on it, a pre-committed condition that would kill it, and a named asset that survives the kill. A firm that merges the three builds one thing that does none of them.

The Cap

Now the number that is not a target.

Declare a maximum number of live options. Not an aspiration, not a guideline — a cap, so that opening one requires closing one. That constraint is the only mechanism I know of that makes elimination as cheap as generation, and elimination is precisely what did not get cheap. I have watched a partnership generate more than fifty argued candidate futures in a year and close exactly none of them, while every instrument in the room reported an outstanding year.

The objection arrives instantly and it is a fair one: a cap will make us miss something.

The cap does not limit what may be noticed. It limits what may be carried as live and funded. Anything noticed and not carried becomes a deferred option with a named trigger — which is a disposition, not a loss, and the trigger is what makes it recoverable.

Then the harder half of the answer, which is the one that changes behaviour: a firm that cannot bear to convert a notice into a deferral does not have a portfolio problem. It has an attention problem, and the portfolio is where the symptom happens to show up.

The Academic Version Is Older and Worse

None of this is about AI, and it is worth knowing that the mechanism was documented before anyone was selling AI strategy.

Studying sequential venture investments, Isin Guler found a striking asymmetry: signals of a company's progress — patent counts and the like — reliably predicted investor behaviour for successful companies, and did not for unsuccessful ones. "Signals of failure are more ambiguous and complex; and firm-level differences are more pronounced in management of unsuccessful options."7

The named failure modes are worse than ambiguity. Investors "may prefer to modify project goals or standards instead of abandoning projects, in an effort to create a more favorable outcome" — and there is a documented tendency toward "rational overcommitment… especially when their personal interests are at stake."7

And the thing that separated firms who handled failure well from those who did not: they "interpreted and acted on negative information more swiftly than others."7

"Rational overcommitment, especially when personal interests are at stake" is the academic name for the shadow-principal problem applied to option closure, written up two decades before this argument was commercially urgent. The mechanism is not new and it is not about AI. Cheap cognition only raised the volume.

Two honest caveats. That is 2006 data, published in 2007, and it is venture investment rather than advisory — so it is a foundational finding about how humans manage options under personal exposure, not current evidence about consulting. It earns its place because the mechanism transfers, not because the setting matches.

What I Get Measured On

Which brings the chapter to the part that costs me, and it belongs here rather than in the governance chapter. Chapter 7 fixes authority — who is allowed to decide what. This fixes what I get paid for noticing, which is a narrower and more uncomfortable thing.

"The answer is not disclosure, which changes nothing about the incentive. It is a measurement inversion."

Key Insight

Measure the adviser on the client's closure rate, not on the adviser's output volume.

Three lines on the scorecard, and one exclusion that does the real work:

  • Options closed per period, with the evidence that closed them.
  • Residue delivered — named instruments the client now operates without me present.
  • Option WIP — flat or falling, not growing.

And then: idea volume, artefacts produced and workshops run are excluded. "Not de-emphasised, excluded, because anything on a scorecard becomes a target." That distinction is not fussiness. A metric kept on the page "for context" is a metric somebody will optimise, and the three things easiest to produce in an advisory relationship are ideas, artefacts and workshops.

The Register

All of which reduces to one artefact the client owns and can read without asking me.

Field Why it is there
Option id So it can be referred to across cycles without being re-described
Opened date So age is computable rather than remembered
Current disposition One of six. There is no "in progress"
Evidence that moved it last So a disposition can be challenged on its warrant, not on its confidence
Reopening trigger So a kill or a defer is recoverable without anyone losing face
Disposition owner (client-side) Because the board owns the wager, and the register should show it
Age Because nothing else exposes an option nobody is willing to close

One line to add alongside it, which Chapter 6 turns into a conflict and Chapter 13 into an instrument: track the rate of do-not-build and defer dispositions, and what happened to the relationship afterwards. That is a standard lane KPI in any honest commercial design and it is the only evidence that "no-build is a live outcome" is a fact rather than a clause.

Which leaves the sixth disposition sitting there unresolved. Transfer into ordinary client operation is on the list, it is the right outcome, and every single time it fires it removes a reason to renew.

05
Part II: The Unit of Sale

Transfer the Known, Earn the Frontier

Six directions the relationship has to travel — and why the version of this argument most firms stop at is the version that fails.

So here is the problem stated commercially rather than ethically, because the ethical version is easy and the commercial version is the one that decides what gets built. Every period of this mandate is supposed to make part of me unnecessary. And I still want to be renewed.

The naive reading of that is worth saying out loud, because most people are carrying it whether or not they would admit it: a good adviser who genuinely transfers makes themselves progressively poorer. Which is precisely why so many firms quietly under-transfer. They keep the tricky bits, the tooling, the history, the one spreadsheet nobody else understands — and then call the residue a relationship.

The rule that dissolves it is four words long.

Transfer the known; earn renewal at the frontier.

Transfer and renewal are not opposites. They are the same instrument read from two ends. What transfers is yesterday. What is bought is tomorrow. The relationship becomes indefensible only when those two are collapsed into a single payment — which is what happens by default, because a single monthly number cannot express the distinction. Chapter 8 makes it architectural and Chapter 9 makes it a line item; this chapter establishes that it is real.

The Dependency Gradient

Six commitments, and a mandate is healthy when all six run the right way. Read the third column rather than the second — the middle column is the promise, and the right-hand column is what makes it checkable before anybody is in a renewal meeting.

Over time Direction What it looks like when it inverts
Dependence on me for ordinary execution Falls Their people route every unfamiliar question, proposed framework and article to me. I have rebuilt the labour pyramid with myself as the partner bottleneck.
Known exception escalations Fall The same shapes escalate every quarter. Nothing is fossilising into a rule that somebody else can apply.
Client-owned standing questions Rise Questions keep being asked of me rather than run by them. The owner column has my name in it.
Their ability to defend and operate prior decisions Rises They cannot present last quarter's decision to their own board without me in the room.
My attention on genuinely new options Rises My time goes to retrieval, explanation and rework — and I feel busy and valuable while it happens.
Time to construct the next option Falls Every new offer costs what the first one cost. Nothing compounded; we just did it again.

That table is the product specification, disguised as a direction of travel. The reason to write the inversions down is that each of them is observable in the ordinary run of work, months before it shows up in a decision — and none of them feels like a problem at the time. Being sent every hard question feels like being trusted.

Take the first row and follow it, because it is the one that catches good firms. If a consultancy's people send every unfamiliar question, proposed framework and thought-leadership draft to me, I have not built leverage. I have become the escalation desk of a larger firm. That is a worse position than the one I started from, because it is founder dependence at higher volume with a recurring invoice attached — and from the inside it looks exactly like success. The calendar is full. The client is engaged. The renewal is safe. And the thing I actually sell has stopped being produced.

But Complete Independence Is Not the Target Either

This is where most "we build capability, not dependence" positioning stops thinking, and the chapter has to go one step further than the comfortable position.

A static transfer hands the client yesterday's snapshot while the market and the evaluation function keep moving. If they can reproduce all future evaluation, all doctrine updates and all successor construction from what was transferred, then I did not retain an evolving product. I sold a static method and called it a partnership — which is a different failure from indispensability, and just as terminal, because the snapshot ages at exactly the branching rate of the market it describes.

The end state is neither. It is voluntary asymmetry: they no longer need me to run what has already been learned, and they continue choosing me because my living system remains the fastest and safest way to discover what should exist next.

Said the way I would say it across a table:

You should be able to leave me and keep operating. You should choose not to because staying gives you the fastest path to the next thing.

And then the qualifier that does all the work, and that most versions of this sentiment quietly omit: you will not need me to operate yesterday's solution; you may continue to choose me because I remain the fastest, safest route to tomorrow's. That choice must remain real.

Everything in Parts III and IV exists to keep that third sentence true. A choice is real when the alternative is available, costed, and has been taken by somebody before. A choice that exists only in the contract is a preference the supplier has written down on the client's behalf.

Myth vs Reality

Myth

Transferring capability costs you the renewal. Every thing you hand over is a reason for them to leave, so the commercially rational move is to transfer slowly and keep the interesting parts.

Reality

Transferring capability is the only thing that makes a renewal mean anything. A renewal a client cannot decline is not evidence of value. And the slow-transfer strategy converts your best people into a permanent helpdesk for work you already know how to do.

The Axis I Am Importing, Not Rebuilding

This argument has a parent, and the discipline here is to point rather than re-teach.

The anti-dependency test cuts both ways. The commercial objective is not that the partner can never deliver without me — that makes me the next bottleneck and contradicts the operating model I claim to install.

"The partner becomes capable of delivering the current product. You remain the fastest, safest and most credible way to improve the product, govern its evolution and create the next five."

The half most readers skip is the half about the creator rather than the buyer. "If you remain the execution runtime, you never get paid for the factory, and every next offer is still your calendar. The anti-dependency test is not only buyer protection. It is creator protection against becoming a high-status bottleneck."

And the commercial consequence, which is Part IV's pricing conversation arriving four chapters early: "You stop arguing about whether they are 'allowed' to understand the product. You start arguing about which rungs of work are still worth buying from you after they can run the ordinary cases." That is a better argument to be having. It is also a harder one, because it requires knowing what those rungs are.

One further pointer, and then I will stop borrowing. The honest test of transfer is not what happens mid-install — staff preferring firm-grounded work, opportunity cards appearing with claim boundaries, escalations starting to produce reusable rules. All of that is progress and none of it is proof. The proof is engagement two: whether ordinary staff carry materially more on the second engagement because shared machinery improved, rather than because the same specialist worked late again.

Measuring the Gradient Without Inventing Numbers

Six rows is a claim. A scorecard is an instrument. For each row: what is measured, where the number comes from, the direction required, and what happens if it inverts.

Row Where the number comes from
Ordinary-execution dependence The escalation log — count, and share that were routine
Known exception escalations The same log, filtered to shapes seen before
Client-owned standing questions The standing-question register, counted by owner — theirs versus mine
Ability to defend prior decisions Whether their own people presented the decision to their own board, unaccompanied
My attention on novelty My own time records — honestly kept, including the night work
Time to construct the next option Elapsed time from option approval to a launched successor, offer over offer

The honesty rule that governs all six is a method rather than a caveat: direction and comparison across cycles, never fabricated percentages. Shape, not invented magnitudes. And if a firm cannot yet measure a row at all, that is itself the finding — the offer is not instrumented enough to claim what it is claiming.

The full quantitative form — paid bounded units divided by genuinely scarce expert dispositions — belongs with the trials in Chapter 13, because it is a proof instrument rather than a health check. What sits here is the qualitative version, which a reader can run this quarter with a spreadsheet and an honest hour.

One operating rule to go with it: if a row inverts once, it is noise. Evidence collection is lumpy and outcome clocks are slow. If it inverts twice, something changed in the relationship, and it gets named at the renewal rather than absorbed.

The Version I Have to Live Under

There is a falsifier for all of this that I published before I wrote this book, and it is pointed at me rather than at a client.

"If this discipline makes clients more dependent on the adviser rather than less, it has failed on its own terms — regardless of whether the mechanism is true. A discipline that only works when the person who wrote it is in the room is not a discipline; it is a service with a framework attached."

The observable that goes with it is specific enough to be checked by somebody who does not trust me: after four quarters, is the client running the diagnosis, designing the probes and closing the rows themselves? If the answer is no, the transfer failed and the mandate is an expensive conversation subscription.

That is a falsifier for this book, not only for a client relationship. It is also checkable by anyone who buys the mandate, which is the entire reason to publish it rather than keep it as a private standard.

The Public Evidence, Including the Part That Argues Against Me

There is a dataset that cuts the wrong way for this chapter, and it belongs in the middle of the argument rather than in a footnote.

The ANA and the 4As found that average client–agency relationship tenure had roughly doubled since 2016, to about seven years. Inside that number is the finding that matters here: clients without mandatory review periods averaged 8.1 years, while those with frequent reviews ran as low as 3.8 years.8

Read plainly: removing the periodic falsification test lengthens the relationship. Which is the opposite of what an evolution mandate does on purpose.

I am not going to hide that, and I am not going to reinterpret it into support. What that data measures is relationship survival, not relationship value — and those are different quantities that happen to be easy to confuse, because one of them is easy to count. The mandate deliberately reintroduces the test that shortens tenure, which means I am betting that a relationship designed to be killable is worth more per year than one designed to persist.

That is a bet. It might be wrong. And the ANA numbers are exactly the shape of evidence that would show it was wrong, which is more than most positioning offers you.

The buyer side, meanwhile, is already moving toward the gradient rather than away from it. The UK National Audit Office is unambiguous that consultants "should only be used where they represent best value for money and not to replace capability required inside the civil service"9 — and its recommendations turn transfer from a courtesy into a contract term: "ensuring that civil servants learn from consultants while they are working together by building knowledge transfer agreements into contracts."9

When one of the largest buyers of advisory services in a G7 economy writes transfer into the paperwork, the gradient stops being an ethical flourish and becomes a description of where procurement is heading.

Where That Leaves the Argument

One adjacency to name and leave alone, because it is a real question and it belongs to a different book: whether I personally become more capable while my firm becomes no more transferable is a distinct failure with its own instrument — the founder-multiplier trap, and an ablation test designed to detect it. Different question, different apparatus, named here so you know I have not confused the two.

What this chapter settles is the direction of travel. Six rows, one end state, and a falsifier I have to live under.

What it does not settle is who is steering. The gradient describes where the relationship should go. It says nothing about who is allowed to decide which option gets exercised, which build gets recommended, or whose evidence gets believed — and the person currently best placed to answer all three of those is the person being paid by the outcome.

06
Part III: Who Is Allowed to Decide

The Shadow Principal Is Me

Six legitimate services, one gradient, and why the defence cannot be that I am a decent person.

Here is the list of what I would be doing simultaneously inside this relationship. Set down flat, without commentary.

The six hats

  1. Diagnose the continuing danger.
  2. Maintain the strategic search.
  3. Recommend which option to construct.
  4. Sell the construction.
  5. License the machinery it runs on.
  6. Help produce the narrative explaining why it all matters.

Now the observation that makes this a chapter rather than a confession: every one of those is a legitimate service, and a serious client would reasonably want all six from the same firm. There is no step on that list I would remove from a proposal. Splitting them across four suppliers would produce worse work at higher cost with more handoff loss.

The problem is not that any step is wrong. It is that the six together form a gradient, and the gradient has a direction, and the direction is not neutral with respect to what I recommend.

Which is why I am going to refuse the obvious defence up front. Integrity is not a control. Not because I lack it — because it is not a mechanism. It cannot be audited. It does not survive a bad quarter, a cash-flow squeeze, or the ordinary human capacity to find the argument that happens to suit you genuinely persuasive. A defence that depends on the adviser being a good person is not a defence. It is a hope with a testimonial attached.

Better Words Than Product Culture Has

Agency law got here first, and its vocabulary is sharper than anything the consulting industry uses about itself. In agency law, an agent acts on behalf of a principal and owes that principal a duty of loyalty — to put the principal's interests first in matters connected with the relationship.

Two terms do the work.

Shadow principal

The party who actually owns the objective function while the user experiences the agent as theirs.

Double agent

A system that presents as aligned with you while its incentive gradient points elsewhere. Not a cartoon villain — a faithful agent, of somebody else.

I coined those for consumer AI and recommender systems — for the feed ranker that understands you, predicts you, configures your environment, and was never working for the person holding the phone. And the sentence that made the argument portable transfers to advisory without a word changed:

Excellence and loyalty are orthogonal until the principal is specified.

A recommendation can be brilliant, well-evidenced, correctly reasoned and optimised for continuation. Those are not in tension. "Capability without a loyalty target is not neutral. It is power looking for a gradient" — and a strategy engine with a deep canon behind it is a great deal of capability.

I should be honest about the lineage while I am borrowing it. The label was an editorial sharpening of a plain-language instinct, and the original chapter says so explicitly rather than dressing it up as ancient doctrine. Same discipline here: the vocabulary is a tool for seeing the structure, not a claim to have discovered it.

The Metrics Are the Principal in Numerical Form

The consumer version of this argument turns on a distinction that translates cleanly. Engagement AI asks what will this person do next inside my system? Intent AI asks what is this person trying to have happen in their life? The first produces a north star of interface continuation; the second, of interface termination.

"The metrics are not a reporting detail — they are the principal in numerical form."

Translate the two columns into advisory and the problem becomes visible without anyone having to be accused of anything.

If I instrument this… …I have declared this principal
Renewals. Licence usage. Constructions sold. Months invoiced. "Engagement health." Me
Decisions changed. Options correctly killed. Questions transferred to client ownership. Ordinary work no longer requiring me. Time from signal to authorised disposition. The client

Now notice the asymmetry, because it is not an accident and it is what makes the left column win by default. Everything in the top row is trivially easy to measure and appears on every advisory dashboard in existence. Everything in the bottom row requires an agreement with the client about what counts — which means a conversation, a definition, and somebody willing to be held to it. So the bottom row is usually absent, and its absence is never a decision anybody makes.

The next trap is soft language, and the original chapter names it precisely. "Client value." "Partnership." "Strategic depth." Those words fit comfortably on a slide next to any dashboard at all. "The discipline is harsher: write the question the system is actually answering, then write the numbers that prove it answered well."

And the evasion that follows, translated: adding a "quality of engagement" score or a "strategic depth" rating to a continuation regime does not invert the regime. A more meaningful hour spent inside a dependency is still an hour scored as success, because the client remained available to be scored.

If only the left column is instrumented, the left column is the real principal — whatever the engagement letter says, and whatever anybody sincerely intends.

The Question About My Own Offer

Important

What happens when the client's interest is "stop constructing and operate what we have", while my commercial interest is another build, another licence year and another month?

That is not a rhetorical question and it is not about advisers in general. It is about my firm, and the answer is that nothing in the arrangement as described so far prevents the wrong outcome.

If success is measured through renewal, licence usage and constructions sold, your strategy machinery can manufacture its own necessity while sincerely believing it is helping.

Spend a moment on that final clause, because cutting it turns the whole argument into a morality tale and morality tales are useless here. While sincerely believing it is helping. The sincerity is load-bearing.

A conflicted adviser who knows they are conflicted is a manageable problem — you can watch them, and so can they. A conflicted adviser who is genuinely trying to help, whose incentive gradient quietly shapes which options look interesting and which objections feel like nitpicking, is the actual failure mode. And it produces better-argued recommendations, not worse ones, because the person making them believes every word. That is why the detection cannot be behavioural. There is nothing to observe.

The Broker

The cleanest analogy is a desk everybody already understands.

A broker paid only when the customer trades has an incentive to overtrade. An adviser paid only when options become builds has an incentive to over-construct.

Push the analogy one step further and the shape becomes legible to a board. Navigation resembles an external investment office maintaining a portfolio of possible strategic investments. The build capability resembles the desk that exercises a selected option. Nobody would let one desk do both without controls — the arrangement has a name in finance, and the name is not a compliment.

Advisory does exactly that, every day, without anyone noticing it needs a name. The firm that finds the opportunity is the firm that quotes for building it, and the client's protection is that the firm seems decent.

The fix is visible inside the analogy, which is why the analogy earns its space: pay for dispositions, including non-exercise. Correct non-exercise is a valuable result. Firms that celebrate only exercised options will apply steady, entirely sincere pressure to recommend builds — and the pressure will feel like enthusiasm. Chapter 4's scorecard is where that becomes an instrument; here it is enough to see that the incentive has a shape and the shape is fixable by pricing rather than by character.

Is Any of This Enforceable?

A reasonable reader is now wondering whether this is good intentions with a diagram. It is not, and one profession has already had the argument at national scale.

Sarbanes–Oxley §201 does not ask auditors to disclose conflicts of interest. It makes a list of services unlawful to provide to an audit client contemporaneously with the audit:

"it shall be unlawful for a registered public accounting firm … to provide to that issuer, contemporaneously with the audit, any non-audit service, including— (1) bookkeeping or other services related to the accounting records or financial statements of the audit client; (2) financial information systems design and implementation; (3) appraisal or valuation services…"10

Read item (2) again. The judgement the law encodes is precisely the mandate's problem: you may not both certify the state of the world and build the thing you certified. A legislature looked at that structure, considered disclosure as the remedy, and rejected it.

And note the governance shape, which is the part almost everyone skips. Everything not prohibited still has to be "approved in advance by the audit committee of the issuer."10 The party who disposes is not the party who proposes. That is control two of the map in the next chapter, written into federal law in 2002.

Key Insight

The party who disposes is not the party who proposes. Everything else in Part III is an implementation detail of that sentence.

What an Ungoverned Conflict Surface Costs

Australia has a more recent worked example, and it is worth reading structurally rather than as gossip.

The Department of Finance's 2025 examination of PricewaterhouseCoopers Australia's ethical soundness quotes the Switkowski review's finding directly: "There has not been, and does not yet appear to be, an overarching framework providing clear instructions to partners and staff as to how to identify or manage the various types of actual, potential, or perceived conflicts."11

And separately, the finding that names the actual defect: the firm "appears to lack a process for, or practice of, consolidating all conflicts of interest information. Without a readily obtainable enterprise-wide view of conflicts, the ability to manage conflicts is compromised."11

The Senate inquiry's diagnosis was explicitly structural rather than individual: "structural weaknesses in governance, transparency and accountability have contributed to ethical failures across the consulting sector."11

Let me be exact about what I am and am not claiming. I am not litigating that matter, and it is not an analogy for anybody's honesty. I am pointing at one finding: the defect was the absence of a consolidated view of the conflict surface. Nobody could see the whole of it at once. That is a design defect, and design defects are inherited by anyone who builds the same shape — which is what makes a government-authored review more useful here than a hundred opinion pieces about consulting ethics.

The scale objection answers itself, and it answers in the uncomfortable direction. Nobody reading this is a Big Four partnership. That is not the point. The mechanism scales all the way down, and a one-person firm has more of it, not less — because there is nobody else in the room to notice. No independent audit board. No second partner with a different book of business. No committee whose actual job is to ask the awkward question about the option you are most excited about.

The Failure That Arrives Last

Pitfall

The version of this conflict that shows up latest is a relationship the client stays in because leaving would disable access to their own history. Renewal rate looks excellent. Client satisfaction looks fine. What is being measured is a hostage, and it will read as a strong account right up until the moment it ends badly and publicly.

Chapter 8's termination boundary is the architecture that prevents it. Chapter 14's stop conditions are the instrument that detects it. It is named here so that both can point back rather than re-introduce it.

What I Am Not Arguing

Integrated advise-and-build is not inherently corrupt, and separating it entirely destroys real value. The person who found the option usually knows most about building it. Forcing a clean handoff wastes exactly the knowledge that made the option worth exercising, and replaces one problem with a translation loss that the client pays for twice.

So the claim is not separate the companies. Incorporation is an expensive answer to a design question, and it is unavailable to almost everyone this book is written for.

The claim is separate the authorities — which turns out to be affordable for a firm of one, and is the whole of the next chapter.

Two easy answers refused, then: integrity, and structural separation by incorporation. What is left has to be a set of controls that somebody who does not trust me can check.

07
Part III: Who Is Allowed to Decide

The Conversion Firewall

Six controls, four artefacts, four signatures — and why a firm of one needs more of this than a firm of a thousand.

Controls a sceptic can check. That is the requirement, and it rules out most of what firms currently do about conflict of interest, which is to mention it early and warmly and then never again.

So take the six hats from the last chapter and ask, of each: who else could hold this?

Four of These Walls Already Exist

Not for a mandate — for the rung below it. Importing them rather than re-deriving them saves a third of this chapter and is also just honest about where they came from.

At the first rung of the ladder, the bounded decision, the conflict has exactly the same shape: the firm selling the diagnosis profits if the diagnosis says construct. And the answer is the same too: "Disclosure does not dissolve that conflict. Structure does — and the structure is buildable."

Stand-pat, harvest-only and build-nothing are valid paid outcomes

The client pays for the decision, whichever decision it is — so the diagnosis has no revenue reason to say construct.

Construction is separately priced and separately commissioned

"Two purchases, two signatures, a genuine gap between them in which the client can walk."

The decision pack is portable

Engineered for it, not merely permitted: "authoritative inputs, tests, boundaries, not a teaser."

The review names its own rejection conditions

In writing, before the work starts — what evidence would show the doctrine does not apply to this firm.

Underneath all four sits the rule that binds every rung, and it constrains this book as much as the last one: each rung is a complete purchase, priced on its own value, with its own acceptance. "The moment rung 1 is priced as a loss-leader for rung 2, it stops being a decision product and becomes a sales document with a fee."

Then the turn that makes this chapter necessary. A review is bounded. A mandate is not. A bounded review ends, which limits how far the gradient can bend it — there are only so many weeks in which to be quietly persuaded of your own preference. A mandate runs continuously, which means the conflict compounds, the shared context deepens, and the client's ability to see the arrangement from outside decays exactly as the relationship gets better. So it needs the four walls, plus two more.

Six Controls

Each with what it means in practice, what I am specifically not allowed to do, and how a buyer would check it without taking my word for anything.

Control What it rules out How a buyer checks
1. I challenge and propose Weighting, selection, sequencing. That is the whole of my authority over the option set and it stops there. Does any artefact I produce contain a ranking presented as a conclusion?
2. The client board owns weighting and disposition A rubber stamp. They must receive the rejected alternatives and the residual uncertainty, not just the recommendation. Can the board reconstruct the decision I would have made — and say why they made a different one?
3. Navigation outputs stay portable A register that only makes sense with me narrating it. Hand it to another firm and count how many questions come back.
4. No-build, defer and another-implementer stay live Live means with precedent. If every cycle so far has recommended a build, the outcome is not live regardless of the contract. What is the actual historic rate — and what happened to the relationship afterwards?
5. Construction is separately authorised A build that begins because the option was approved. Separate instrument, separate decision date, separate signature, real gap. Is there a date on which the client could have walked and nothing would have been half-built?
6. Verification does not rest solely on the option's authors Me certifying that my own portfolio is still justified. Who has ever disagreed with the register, and what happened next?

Why the Sixth Needs Its Own Argument

It is the one small firms drop first, and it is the one with the most published machinery behind it.

Put both jobs in one process and you invent a conflict: the same machinery that compressed the ambiguity is asked to certify that no load-bearing ambiguity was lost. A merge that removes a live disagreement makes the graph easier to walk — and destroys the disagreement the audit needed to see.

The clean split comes from knowledge systems and transfers to a maintained option portfolio without modification.

Two jobs, two questions

The janitor — maintains shape
  • Asks: is this world still lean and coherent?
  • • Merges duplicates, splits overloaded nodes, marks supersession, manages expiry
  • • Critically: can preserve contradictions as edges rather than averaging them into false consensus
  • Failure if alone: pretty structure, unjustified claims
The auditor — maintains warrant
  • Asks: is this world still justified?
  • • Reconstructs each consequential claim from its support path, challenges derivation, compares dates and authority, tests coverage, emits findings
  • It does not tidy
  • Failure if alone: findings with nowhere durable to land — or worse, findings silently rewritten as truth

The operation is not are you sure? It is: reconstruct the current claim from its evidence path and report any mismatch.

Key Insight

Having a citation field filled is not the same as having a warrant you can still defend.

Apply that to an option register that has been curated for four quarters. The party who has been maintaining the rows — adding evidence, adjusting confidence, keeping the shape tidy — cannot also be the party certifying that the evidence still supports them. That is not a comment on anybody's honesty. It is a comment on what four quarters of curation does to a person's ability to see a premise that has quietly rotted underneath a row they wrote themselves.

Then the version that catches careful firms rather than sloppy ones. Reviews that share one upstream evidence feed are not multi-line assurance — only a shared root in different coats. Applied here: if my shadow cycle, my kernel and my recommendation all read the same evidence set, then three independent confirmations are one confirmation with three signatures on it. The trial in Chapter 13 has to be designed around that or it proves nothing at considerable expense.

The Client-Side Half of Control Two

A governance map only works if the other party actually holds the authority it has been assigned. Which means saying what that requires of them, and it is more than attendance.

Five things only a board can do:

  • Decide which revenue assumptions deserve to be attacked — "because choosing what to doubt is a capital decision wearing an intellectual costume."
  • Insist that productivity and transformation are not reported as the same class of work.
  • Allocate capital between current operations and successor options, explicitly, as a number.
  • Define what evidence would justify commitment — in advance, "which is the only time it can be done honestly."
  • Own the consequence of building, deferring or refusing.

And the counterweight, without which this becomes top-down theatre: a board cannot supply the input the machine runs on. Choosing which variable is live requires lived friction with the business, and that friction lives at the customer edge — in delivery failures, pricing objections, the exceptions that keep recurring, and the work staff quietly decline to do because it never works. The board owns the wager. The operating edge keeps it honest.

Which has a practical consequence for how a mandate schedules itself: if my only interface is the board, I will attack the assumptions the board finds interesting, and those are usually the ones it has read about. Time in the delivery layer is not relationship-building. It is where the input comes from.

Walking One Option, All the Way Through

"Separate speech acts" is an abstraction until you draw it. Four steps, four artefacts, four signatures — and a column for the thing each artefact is not allowed to contain, which is where the firewall actually lives.

Step Artefact Signature Must not contain
Nomination Option card: hypothesis, evidence, what would kill it, estimated probe cost Mine Any price for building it
Disposition Board decision receipt: chosen state, rejected alternatives, residual uncertainty, reopening trigger Client My recommendation restated as the decision
Scoping Bounded construction brief: outcome, boundary, exclusions, acceptance criteria Client commissions An assumption that I am the builder
Acceptance Independent test against criteria agreed before the build began Not solely mine Criteria written after construction started

Walk it. An option card goes across with a hypothesis, the evidence behind it, the observation that would kill it, and what a probe would cost. It carries no build price, no implementation sketch, no timeline. The board reads it alongside the others, decides, and issues a receipt that records not only what they chose but what they rejected and what would make them revisit. That receipt is theirs and it makes sense without me.

If the disposition is exercise, a separate document opens on a separate date: a bounded construction brief with an outcome, a boundary, exclusions and acceptance criteria — written before anybody knows who is building it. And when it is built, the acceptance test runs against criteria that existed beforehand, assessed by somebody who is not solely the builder.

Pitfall — the most common breach, and the least noticed

An option card that arrives with a build price attached. It looks like helpfulness — the client will ask eventually, so why not save a week. But the moment a nomination carries an implementation estimate, the disposition conversation has already been anchored, and the board is no longer choosing between six dispositions. It is choosing between my option at my price and not my option. The firewall was breached before anybody sat down.

The audit of your own offer takes ten minutes: look for any row where the same name appears three times.

The Questions I Would Least Like to Be Asked

Here is the other side of the table, handed over deliberately. Six of these already exist as a buyer's checklist for an AI-era diagnosis, and they work unchanged for a recurring mandate:

  1. Can stand-pat win — and has it ever?
  2. Am I paying for the decision either way, or only if it is the "right" one?
  3. Is the decision pack portable to another builder?
  4. Are the falsifiers named in writing before the work starts?
  5. Is construction separately commissioned, with a real gap between purchases?
  6. Who accepts completion — and is the oracle independent of the builder?

A mandate needs a seventh that a bounded review does not: who verifies the option register, and have they ever disagreed with it?

I am publishing the questions I would least like to be asked, because a buyer who cannot ask them has no way to tell my offer apart from the one it is trying to displace. Both will sound thoughtful. Both will use the word partnership.

What the Firewall Costs

Independent verification is not free. Portable outputs take longer to produce than working notes, because a working note only has to make sense to the person who wrote it. A separate authorisation step slows a build by weeks, and some of those weeks are pure calendar.

Those are real costs, and they are precisely why firewalls get quietly dropped in the third quarter — not by decision, but by a series of individually reasonable accommodations. So the cost belongs in the fee, explicitly, as a line. That is Chapter 9's job, and naming it here is what stops it being absorbed as overhead and then abandoned.

The objection that arrives at this point is the right one: my client is one person. There is no board.

Then the disposition authority is that person, and the two controls that matter most are portability and independent verification — because a single decision-maker with no portable record is maximally exposed to the adviser's framing. There is nobody to compare notes with, no minute that records the alternative, and no colleague who remembers the version of the argument from before you both agreed.

Takeaway

Small clients need more firewall, not less.

Speech Acts, Not Companies

You do not necessarily need separate companies or teams. You do need separate speech acts, artefacts and authorities.

That sentence is what makes the chapter affordable. It converts a governance requirement into a documentation requirement, and a documentation requirement is something a firm of one can actually meet on a Tuesday. Four artefacts. Four signatures. One column of things each artefact is not allowed to contain.

What the controls decide is who steers while the relationship runs. They say nothing at all about what is left when it stops — and that question is the one a buyer will eventually ask in a room where the answer cannot be improvised.

08
Part IV: What Survives Cancellation

Three Planes

One binary that settles every ownership argument — and the shipped design of my own that fails it.

The question a buyer eventually asks, and it is the one that reveals what you have actually sold them: what happens if we stop paying?

Most advisers answer it badly, and the badness is diagnostic rather than dishonest. They answer with reassurance — we'd work through it, we're not like that — because nobody has ever asked them to design the answer, and because the honest version requires deciding something they have been able to avoid deciding.

The Termination Test

If stopping payment destroys current operation, you are monetising dependence.

If stopping payment preserves current operation but ends future evolution, you are licensing a living capability.

That is the sharpest binary in this book and the chapter organises around it, because of what it does to an otherwise intractable conversation. You do not have to agree about who owns what. You have to agree about what happens on the Monday after cancellation — and that is a single observable, which two parties can settle in an afternoon rather than in a negotiation.

The honest consequence for the seller is why most firms never run the test. The second position is commercially weaker in the short run and far stronger in the long run, and it cannot be adopted halfway. A ninety-day grace period is still position one, with a countdown. So is a data-export-on-request clause. So is anything that ends in "of course we'd help you transition."

Three Planes

The reason the binary resolves cleanly is that "the system" was never one thing. It is three, and they have three different owners.

The state plane — client-owned

Holds: their data, operating state, local knowledge base, receipts, client-specific derivatives, credentials and exports.

The test: can they get all of it out, in a form usable without my software? Not "is there an export button" — can a competent third party read what comes out of it.

The runtime plane — client-operable

Holds: the last accepted solution, with runbooks, documented dependencies and an executable exit path. The commercial form varies with criticality — a perpetual last-state internal-use licence, source escrow, continuation rights, or a fair buyout.

The test is counterfactual: if the supplier disappeared tomorrow, could the organisation still understand and continue the critical software? That is a demonstrated property, not a promised one.

The evolution plane — mine

Holds: the generic compiler, evaluation harnesses, doctrine, cross-client failure shapes, new releases, next-offer construction capability.

The test: this is the plane the fee gates. Nothing else is.

The architectural principle underneath is one I have argued elsewhere and will not rebuild here: rent the model, own the map. What survives a vendor change is the compiled claims and edges, the policy, the receipts, the proposal history including everything rejected and why, and the outcome feedback that improved the map. What does not survive is the weights and whatever the vendor called "memory". Or, compressed: "the model is the engine. The wiki is the memory. The DAG is the law. The receipt is the evidence."

The buyer's fear here is not paranoia, and I want to describe its shape without reaching for a number I cannot stand behind. Vendor-dependency surveys are in wide circulation and I went looking for one with a disclosed methodology I could cite. I did not find one I trust — the most-quoted figures in this space trace back through aggregators to pages that do not contain them. Chapter 15 lists what got dropped and why.

So here is the shape instead, which is checkable against your own experience rather than against a statistic. Organisations consistently believe they could switch a critical AI vendor far more easily than the ones who have actually attempted it report. The gap between the anticipated migration and the executed one is the whole risk, and it is invisible until somebody tries. And the framing from my own earlier argument is better than any percentage would be: once AI stops being an experiment and becomes the backbone of your business, you are no longer adopting software — you are entering a committed relationship with a slightly vague escape clause.

Bottom Line

The hedge against that is not a procurement clause. It is an architecture.

Which is worth stating as a design consequence rather than a warning. A buyer cannot verify a promise about behaviour under termination, because the only test is the event itself. They can verify an architecture — by asking what runs after the licence expires, and then asking to see it run.

"It Runs in Their Cloud Account" Settles Nothing

Five different questions hide inside who owns the app, and treating them as one is how ownership conversations go bad — because a single question produces a single answer, and a single answer means somebody loses.

The question The defensible answer Fought about?
Who controls the account and infrastructure? The client. Their account, their billing, their network boundary. Rarely, once stated
Who owns the data and the encryption keys? The client. Unconditionally, including keys. Yes — and correctly
Who owns the software copyright and generic source? Me, for the generic layer. Theirs, for client-specific configuration and derivatives. Conceded once the others are settled
Who has the right to operate, modify and update? The client operates and modifies their instance; I control the release channel. Sometimes
What survives termination? State and runtime, indefinitely. Evolution, no. Yes — and it is the only one that matters

When I was first asked whether I or the client should own the appliance we would deploy, my honest answer was that I was six of one and half a dozen of the other. That indecision was correct, and I should have said so more precisely: the question was malformed. It has five answers, they do not have to match, and collapsing them is what produces an all-or-nothing negotiation neither party wanted.

My Own Design, Against My Own Test

This is a doctrine part, and I am about to bring my own machinery into it. That is deliberate. A sovereignty chapter that will not criticise its author's own shipped design is decoration.

What exists

The AWS Marketplace Knowledge Appliance is a two-day productisation spike from July 2026. In that interval the repository accumulated a fail-closed privacy tokeniser, a corpus engine, a licence agent, a vendor kernel, a six-image build set, a smoke verifier, a governance renderer and nested infrastructure-as-code. The infrastructure notes record zero template lint errors after the templates were aligned to the actual container ports, health paths and environment contracts; a sandbox deployment exposed those mismatches and forced corrections to firewall handling and image-pull allowlists.

Three things about it are real in a way worth being specific about, because specificity is the only evidence available here.

  • Privacy is a transaction boundary, not a policy. Raw content crosses a mandatory tokeniser before it reaches storage or a model. Analyzer failure, configuration failure, transport failure or vault failure refuses the operation rather than degrading gracefully into the unsafe path. And the live check found a production-only defect the offline tests could not reveal — using a database connection as a context manager closed the persistent vault connection after the first write. The fix is in the code. That is the fail-closed claim surviving a real dependency failure rather than passing a unit test.
  • Evidence and navigation are separated. Immutable tokenised source evidence, with no raw-content staging table and no update path for an accepted document; hash identity makes re-drops idempotent. The generative step is bounded — a model may propose pages and relationships, but the batch is rejected unless source pages cover every supplied document, and deterministic code owns page identity, endpoint validation, inverse labels and reconciliation.
  • The commercial surface is deliberately narrow. The vendor kernel exposes two authenticated routes: health, and slice. No page listing. No export. No pagination. Deterministic code caps page count, excerpt size and total response bytes, checks revocation before accepting a token, applies per-key rate limits, and appends an audit line for allowed and denied requests alike.

Evidence ceiling

The seller registration and the Marketplace listing were never completed. This is a built productisation path, not a shipped product, and there is no customer.

There is a dated deployment receipt for the kernel — a live multi-page slice returned from a working deployment on a named date. A dated deployment receipt is not a market, and a sandbox is not a tenant. The full ledger is in Chapter 15.

The licence design, and the critique

Entitlement is composed from two independent sources: a vendor-signed token, and a marketplace entitlement check. A confirmed lapse from either source immediately disables new ingest, new compilation and kernel access — while ask-over-existing-knowledge survives for a persisted ninety-day read-only sunset. After day ninety, all features are off. Uncertain checks retain the last stable feature set rather than widening permissions.

Say what is good about it first, because a critique of a design you have not credited is worthless. That structure avoids both poor extremes: it does not let an expired installation keep compounding value, and it does not make the customer's already-compiled knowledge unusable on day zero. For a spike, it was a thoughtful anti-hostage move.

Now apply the test at the top of this chapter, and watch it fail.

Two ways to end a licence

✗ Lapse on a timer (what I built)

  • • New value creation stops immediately
  • • Existing knowledge stays readable — for ninety days
  • • Then everything stops

Preserves value on a countdown. Still position one.

✓ Lapse that ends evolution (what the test asks for)

  • • Entitlement gates new value creation — ingest, compilation, kernel access, releases
  • • The last accepted state remains operable indefinitely
  • • What ends is the future, not the past

Commercially weaker on paper. Removes the off switch from every renewal conversation.

The change is commercial rather than technical — the feature matrix already exists, and moving one row is a decision rather than an engineering programme. And it matters more than it looks. As long as a countdown exists, every renewal conversation happens in its shadow, whether or not either party mentions it. The client is not choosing. They are avoiding. And a renewal driven by avoidance tells me nothing about whether next quarter was worth buying.

I'm not charging you because I can turn your business off. I'm charging you because what I add next month is worth buying.

The Regulator Has Already Drafted This Line

The EU Data Act gives customers of in-scope data-processing services a statutory right to switch, and technical cooperation for porting. Switching may be requested at any time during the contract; the provider must begin after a maximum two-month notice; the switch completes within a thirty-day transitional period during which the provider maintains business continuity and security, extendable to seven months where the technical complexity is genuine.

Two clauses matter for a mandate, and together they are the fair-termination boundary written by a legislature rather than by a supplier. Switching charges are "only permitted under narrow conditions and will be prohibited entirely from Jan. 12, 2027" — while the Act still permits "proportionate early termination penalties or fees" where a provider has granted benefits or made substantial up-front investment in a long-term contract.12

Read as a design principle: exit must be possible and cheap; genuinely reserved capacity may still be paid for. That is precisely the position Chapter 9's fee architecture takes, and it is useful to know that a regulator arrived at the same line independently while thinking about cloud infrastructure rather than advisory.

One provenance note, because it matters for how much weight to put on the wording: that is a law firm's description of the Regulation, not the Article text. The official register blocks automated retrieval, so the dates and the shape are reliable and the exact statutory phrasing should be checked against source before anyone drafts from it.

The Planes Coming Apart in Public

There is one specimen of this architecture executing under duress, which is worth more than any amount of design argument.

When PwC Australia divested its state and federal government consulting business to create Scyne Advisory, the divestment conditions allowed PwC Australia to "continue to cover the on-going licensing or sale of PwC proprietary products… so that entities' use of these in their day-to-day operations were not disrupted", while PwC Australia agreed not to compete for contracts in the general government sector for five years.11

Read the shape rather than the scandal. The licensed product kept flowing. The advisory relationship was severed. State and runtime survived; evolution did not.

That is the three-plane architecture, executed in public, by people who had no choice about it and no interest in illustrating my point. Which makes it better evidence than a diagram: the planes are separable in practice, under pressure, with lawyers watching.

Three Verbs, Not One

My original instinct was that the IP migrates as it is consumed. That is directionally right and contractually dangerous, because IP does not become the client's by touching their business.

What is true is narrower and I said it plainly at the time: you cannot lead a firm through a body of doctrine and then pretend they have not learned it. That transfer is part of what they bought. Trying to contract them into un-learning it would be absurd, and the attempt would poison everything else in the agreement.

So replace one dangerous verb with three precise ones.

License the background system

Doctrine, generic frameworks and decision rules, the capability kernel, compiler and orchestration machinery, generic tests and evaluation harnesses, failure classes, cross-client promotion methods, and later releases. The client receives defined rights in a named field of use.

Vest the applied client capability

Their data and evidence, their firm knowledge, their decisions and rejected alternatives, their engagement record, client-specific configurations and mappings, local tests and operating records, the explanation of what was decided and why, and the internal capability created by applying the doctrine. The client does not acquire title to the doctrine. It acquires the ability to act from the judgment already applied to its world.

Promote reusable learning under a protocol

Capture in the engagement record first; classify; de-identify; abstract to a failure shape or design rule; verify transferability; human gate; and leave the client a usable record. As the source puts it: "Automatic climb of client material is a confidentiality incident with a calendar invite from Legal, not compounding."

The compact rule: their evidence and applied capability stay theirs; my background system stays mine; generic improvement crosses only through explicit rights and governed promotion.

The ownership map underneath is imported, not rebuilt — three kernels (capability, firm, client), joined at runtime rather than merged into ownership. "The value is the runtime join, not the transfer of every asset into one party's ownership." And the line that makes the boundary feel less like a negotiation: "The partner owns its territory. You own the map of the new country and the machinery for entering it."

The Termination-Boundary Schedule

All of which reduces to five commitments a buyer can put in a contract.

Commitment Detail
Always retained Data, decisions, outputs, exports — in a form usable without my software.
Remains internally usable Accepted prior deliverables, indefinitely, with no countdown.
Read path to their own history Preserved. This is the clause that prevents the lock-in Chapter 6 named.
What cancellation ends New private-kernel compilation, future releases, new doctrine, support, protected compiler functions, design authority, cross-client learning.
If it becomes operationally critical An expensive but fair source escrow or buyout — priced in advance, not negotiated during separation.

The asymmetry that makes it credible is worth stating plainly: the client leaves with everything that has already been built, and nothing that has not been. If that sentence makes a supplier uncomfortable, the discomfort is the diagnosis rather than an objection.

Which Plane Does Your Fee Gate?

That is the question to take away, and it is answerable this week about an offer you already sell. Not are we fair. Which plane does the money gate?

If the answer is any plane other than evolution, you are monetising dependence — whatever the marketing says, whatever the relationship feels like, and however genuinely you would help during a transition.

The planes say what is being sold. They do not say what each part costs, and a single undifferentiated number cannot express three different objects to a buyer who has just been shown that there are three.

09
Part IV: What Survives Cancellation

Unbundle the Hundred Thousand

Seven components, four licence classes, six terms that make a reserve honest — and no number for any of them.

A hundred thousand a month is $1.2 million a year. Say that annually rather than monthly, once, because monthly framing is how large numbers get approved without being examined.

A single undifferentiated figure at that scale produces two entirely predictable failures, and they arrive in this order. First, unlimited customer expectations — because a number with no internal structure has no edges, and a promise with no edges expands to fill whatever the client needs this month. Second, invisible founder-capacity consumption — because nothing in the arrangement records which promise is eating the scarce thing.

So split it. And make each line say what it funds and what it refuses, because the refusal is what turns a fee architecture into something other than a menu.

The Label, at the Point of Maximum Temptation

This is the chapter where it would be easiest to let the figure quietly become a price. So restate it in full: a designed hypothesis, already published as one, with a status column attached, before this book existed.

The reason for publishing designed numbers at all is worth quoting, because it is the same reason they appear here:

"the reader needs the shape of successor pricing more than the figures: value-anchored rather than day-rate-derived, clock-carrying, firewall-wrapped, with stand-pat priced in. And because a book that tells professional-services firms to take the medicine should show its author's own dosage — labelled, dated, and open to the market's verdict."

Now the arguments the number cannot rest on. Each of these is a real thing somebody will say, usually me, on a bad day:

  • The existence of fog.
  • The quantity of my intellectual property.
  • Access to my attention.
  • The number of agents I run.
  • The fact that building it internally would take longer.

Not one of those is a value argument. They are supply-side descriptions wearing demand-side clothes — statements about what I have, dressed up as statements about what it is worth. What could justify the figure is narrower: a sufficiently valuable and active portfolio of terminal-value decisions. Which makes eligibility part of the design rather than a sales judgement, and that gets its own section below.

Seven Components

Component What it funds What it refuses
Capability licence Maintained doctrine, evaluators, compiler machinery, release updates Being treated as a one-time document handover
Navigation cycle Evidence refresh, option portfolio, board dispositions Unbounded exploration with no closure
Reserved design authority Priority, a response clock, conflict clearance on genuinely novel questions An unlimited helpdesk
Per-engagement participation Use of the licensed machinery in their own client work Making me the delivery bottleneck
Successor construction Separately priced build and proof "We're paying this month, can you also build…"
Exclusivity / field-of-use Competitive protection, bought explicitly Free competitor lockout by implication
Buyout / continuity rights A priced path to broader ownership A negotiation conducted during separation

Three of those are routinely collapsed into one another, and each collapse has a signature failure. Licence into navigation produces a client who believes the doctrine updates are "part of the meetings", so the licence quietly becomes free and the navigation becomes the whole product. Navigation into reserved authority produces a cycle that is really an availability promise with an agenda attached — the disposition work gets squeezed by whatever came up. And reserved authority into everything is the classic: the scarce hours become the elastic term that absorbs every other shortfall, invisibly, until the relationship is unprofitable in a way nobody can locate.

The lineage is worth stating and so is its constraint. This tracks the fee architecture I published for a partner relationship — practice installation, annual capability-kernel licence, per-engagement participation, new-offer build fees, design-authority retainer, certification and renewal, and an explicit buyout or broader field-of-use option. And that chapter carries an honesty constraint I am carrying forward unchanged: it is "a design derived from one specimen relationship… not a published rate card and not a claim that any named firm has signed it. If you need numbers, negotiate from these components against real transfer metrics — not against invented industry averages."

On the last line especially, readers resist. A buyout looks like planning for failure. It is the opposite: "A clear, expensive but fair buyout path is better than pretending permanent exclusivity will hold forever… Without that door, a sophisticated partner treats 'forever exclusive' as fiction and plans the copy path quietly. With the door, the conversation stays commercial." Buyout is a feature of adult commercial design, not a confession that the relationship failed.

Why the Split Is Legible

This is not novel accounting. Any recurring service with scarce capacity inside it splits the same way: subscriptions pay for continuous reconciliation and review; activation fees pay for time-bounded interventions; reservation fees pay for exclusive capital; consumption is still a transaction; and exceptional escalation pays for scarce expert disposition when the typed path runs out. Collapse them and you obscure which promise creates value and which consumes capacity.

Two disciplines from that world transfer directly. First: configure the price from a census of the actual exposure rather than inventing a number. "You are not estimating vibes. You are configuring a product." Second, and this one governs the whole book: any ratio or multiplier that comes out of a source conversation is a shape to collect, not a fact to assert.

Pitfall — three collapsed conversations that kill the price

1. Pricing from the part. "The kernel is just a wiki, so access should be cheap." Answer: the buyer is not purchasing the corpus. Part V is entirely about why, and the evidence there is not mine.

2. Pricing from hours saved. "AI does most of this now, so the fee should fall." Answer: the client is not buying fewer hours; they are buying my willingness and ability to absorb controlled complexity. Low friction is not low price. Predictability is a premium attribute when the supplier possesses machinery capable of holding the risk.

3. Pricing from fear. "Buy this or you'll be disintermediated." Answer: that drifts toward an indemnity nobody can keep, and it is the sales posture this entire book exists to refuse.

The Honesty Correction

Now the paragraph that stops this chapter being a purity exercise.

They may not be buying hours. But they are buying reserved capacity. At $1.2 million a year a buyer will reasonably expect priority, response commitments and limits on conflicts — and pretending that no human capacity has been reserved obscures the very scarcity being priced. Availability is not the unit of sale. It is still a component of the promise, and it should have a line with a number beside it.

Which has a consequence most advisory firms never follow through. Once you admit the reserve is real, it has to be specified like any other dedicated reserve — or it will quietly become an availability retainer with better vocabulary, which is the exact thing Chapter 1 refused.

The six terms that make a reserve honest come from physical capacity, and they transfer without modification because the economics are identical — capital immobilised for one party's exclusive benefit, with an obsolescence risk somebody has to carry.

Term What must be explicit for reserved design authority
Ownership Who holds what while the capacity sits reserved. "Ambiguity here is litigation cosplay."
Annual reservation fee Pays for exclusivity, the immobilisation of scarce attention, and the work not done for anyone else. Separate from what consumption costs.
Consumption How a draw is authorised, priced and documented — and what a draw triggers.
Replenishment What happens after a heavy month, and at whose cost. Standing replenish, or re-approve each cycle.
Expiry and reallocation When exclusivity ends and what the options are: extend, purchase, return to pool, credit.
Cancellation Both directions. Notice periods, and the treatment of unearned fees.

And the pitfall stated in the source is the one advisory firms walk into every year: launch a dedicated reserve without a reservation fee and expiry rules and "you will train customers to demand exclusivity as a free service entitlement, immobilise capital, and then discover obsolescence is a relationship argument rather than a contract clause." In advisory, the immobilised capital is my calendar, and the obsolescence is the work I did not do for anyone else while holding it open.

Four Licence Classes

There is a contradiction sitting inside this offer from day one, and the standard formulation cannot survive it.

"Limited internal use, not resale" collapses the moment I help a firm produce public thought leadership using my frameworks — because that is not internal use. It collapses again when their people embed those frameworks in proposals to their customers, which is client-facing commercial use and not internal use either. Both of those are things I would want to enable. Neither is covered by the words in the agreement.

Four rights, then, and be generous inside each:

1. Internal decision and operating use

The books distributed internally under licence. The frameworks applied to their own business. Client-specific strategy and internal playbooks used perpetually.

2. Use while delivering approved services to their own customers

A different object entirely. This is client-facing commercial use and it belongs in a field-of-use grant with its own terms.

3. Public publication, quotation and co-branding

With agreed attribution and review. Generous, and named.

4. White-labelling, sublicensing, training or resale

Separately priced, or refused. Either is fine. Silence is not.

Walk one artefact through all four, because listing classes proves nothing. Take a single named framework of mine.

One framework, four days

Monday. Their strategy team applies it to their own portfolio and writes an internal memo. Class one. Perpetual, unrestricted, and it does not expire when the mandate does — that is the vesting rule from Chapter 8.

Tuesday. The same framework appears in a proposal to their client, shaping a paid engagement. Class two. Still generous — but it is a field-of-use grant now, because my doctrine has become an input to their revenue, and the commercial terms should say so rather than being discovered later.

Wednesday. A partner presents it from a conference stage under their logo. Class three. Attribution and review, agreed in advance, so nobody is negotiating about a slide deck after it has been shown.

Thursday. They train two other firms in it as part of an alliance programme. Class four. Separately priced, or refused — and the answer should have existed before anybody booked the room.

Those rights can all be generous. They cannot remain implicit.

And the counter-intuitive half: at $1.2 million a year, unclear rights are more dangerous, not less. The generous deal is the one where the client knows exactly what it can keep, publish, modify and commercialise — because ambiguity at that price gets resolved by whoever is more willing to be difficult about it.

Which is the practical form of something I believe: if they are paying top dollar, I do not need to penny-pinch on ownership. Generosity here is a consequence of pricing correctly, not a concession extracted from me. That is a better argument for clear rights than any amount of legal caution, and it produces a better agreement.

What this chapter will not do is draft the licence. The contract-design checklist already exists — background IP, field of use, ownership of firm and client material, generic improvements, permitted reuse, continuity, escrow where needed, and a fair buyout path. Name the classes, point at the checklist, stop. There would be strong contracts around it, and that is not the interesting part.

Thought Leadership as a Sensor, Not a Department

The licence-class question surfaces something that will otherwise eat the mandate, so it gets a page.

The failure shape: if I review every article, framework and market claim their people produce, the recurring product becomes editorial outsourcing at a founder's price. Everyone will be pleased with it for about two quarters.

The pipeline instead, in five lines. Their people create client- and sector-specific propositions using the licensed system. Their own internal editorial and evidence gate handles ordinary publication. Publications are emitted as market probes with a stated hypothesis. Response, non-response and objections return as evidence. Only findings that could change the capability kernel or the option portfolio escalate to me.

The sensor discipline is a sibling's and I will not re-derive it, but one rule is load-bearing here: silence becomes information only when the audience who could have responded and the expected response class were both named in advance. "Null is not 'failure as a creator'. Null is a measurement, provided you had an expectation and some honest sense of exposure class." And the mirror failure — publishing only what you predict will land — turns the sensor into "a mirror with a content calendar", which is why a minority slot stays reserved for the internally important but currently cold.

Keep it downstream of real operating evidence. The best external material is the public exhaust of decisions, experiments and field learning that survived evidence and confidentiality review — not opinion produced to a schedule.

Competitive Rights, Eligibility, and Separate Ledgers

Competitive rights are a priced choice, not a courtesy

A firm paying this much may reasonably object if its field learning immediately improves a product I sell to its direct rivals. That objection is legitimate and it deserves a menu rather than reassurance: no competitor exclusivity; a time-limited lead or embargo; field-of-use or territory exclusivity; a broader licence; a buyout. Price each separately. Generic cross-client compounding is genuinely valuable to me — which is exactly why it is not free.

Not every company can support this

Some clients need the bounded decision and the bounded construction, and should then operate internally. Others have multiple service lines, active customer substitution, substantial migration capital, and enough absorption capacity to justify continuous navigation.

A product that cannot say which is which will be sold to the wrong buyers — and when it fails there, the failure will be blamed on the product rather than on the qualification. Eligibility is a design component. Writing it down is also the cheapest way to stop yourself accepting a mandate you should have declined.

Separate ledgers, even inside one relationship

Capability licence and releases. Navigation and evidence refresh. Bounded build activations. The scarce design-authority reserve. Major constructions. External publication and field-of-use rights.

One invoice is fine. One ledger is not — it is how you discover, two years in, that the licence has been silently subsidising an unlimited helpdesk, and that the most profitable-looking line in the relationship is the one consuming the scarce input.

The Client-Side Value Equation

annual value, as the client reconstructs it
losses avoided through earlier correct stops
+ value of options exercised earlier
+ options preserved
+ successor revenue or margin created
+ client capability accumulated
− fee and absorption cost

One rule makes it honest, and it is a rule rather than a preference: no term in that equation is filled with a number I supplied. Every one comes from client evidence, reconstructed by their own people, in their own systems.

Which produces a consequence I would rather state here than have discovered later: if they cannot reconstruct it from their own evidence, that is a stop condition. Chapter 14 lists it among the seven. It appears here so the equation is read as an instrument rather than as a sales tool, because in most hands it would be a sales tool.

And a note on the term buyers forget. Absorption cost is not the fee — it is what the mandate consumes of their executive attention. A client without enough absorption capacity will buy a portfolio it cannot act on, and will experience that as the product failing. It is a qualification failure, and it belongs on my side of the ledger.

What I Refuse to Publish

No number for any component. Not a range, not an example, not an "order of magnitude for illustration".

Not because it is secret. Because I have not sold it. A designed price presented as a market fact is exactly the status inflation this book's last chapter refuses — and my own earlier volume already said so about this very figure: publishing designed numbers as validated "would be exactly the status inflation" its standards chapter prohibits.

One thing to do with your next recurring proposal

Split the single number into named components. Beside each one, write what it funds and what it refuses.

If a component has no refusal, it is not a component. It is a promise with a fee attached, and it will expand until somebody is unhappy.

The components say what each part of the fee is for. They do not say what the client actually touches when they use the largest one — and the answer to that turns out to remove me from most of the work I currently get paid for.

10
Part V: The Access Rail

Compiled Me, Living Me

Deliberately removing myself from the work I currently get paid for — and the falsifier that comes attached.

The capability licence is the largest component in that fee sheet. So here is what the client actually touches when they use it, starting with the most commercially self-damaging thing I said in the conversation that produced this book.

To some extent I am my wiki. My intellectual property is the part of me that can be recalled properly. And talking to me is the slow path — because my recall is not as good as the machine's, and there is a limit to how well I can explain any of it in one sitting.

That is not modesty. It is a product decision, and it has a price attached in both directions.

Here is your version of it, because most expertise vendors are living inside this without having named it: if every ordinary question about your own doctrine has to route through your calendar, you have made yourself the synchronous retrieval layer of your own intellectual property. Every one of those conversations feels like value being delivered. Collectively they are the reason you have no time to produce anything new.

Three Things, Not One

"Scott is his wiki" goes too far, and the imprecision is expensive because it licenses exactly the wrong product. Separate it properly.

The wiki is externalised memory and doctrine

Frameworks, evidence, relationships, rejection history, accumulated cases. Durable and inspectable. It does not decide anything.

The interface is retrieval and routing

It can search more broadly, descend into sources, and produce a cited answer faster than my unaided recall. It does not know what matters.

I am the evolving evaluation function

Deciding what matters now, detecting when the existing map is inadequate, resolving consequential ambiguity, and taking responsibility for a disposition.

Talk to me when the frontier needs to move. Talk to the compiled version of me when you need the accumulated territory.

Key Insight

Compiled me handles known terrain. Living me changes the map at the frontier.

Be honest about which side wins on which axis, because the flattering version of this is useless as a design input.

Where the machine beats me outright
  • • Exhaustive recall
  • • Joining frameworks written months apart
  • • Retrieving exact distinctions and provenance
  • • Following relationships through the canon
  • • Producing several grounded interpretations quickly
Where it does not automatically beat me
  • • Noticing the perturbation that deserves attention
  • • Deciding which variable is load-bearing
  • • Recognising that an existing framework is wrong or incomplete
  • • Reading political and interpersonal context
  • • Accepting responsibility for a strategic commitment
  • • Creating the worldview patch that does not yet exist

Which produces a rule rather than a sentiment: if I personally answer routine doctrine questions every month, the system has failed. If the rail answers yesterday's questions while my conversations discover tomorrow's framework, the relationship is compounding.

This Is Self-Disintermediation, Not a Feature

The access rail is not another benefit. It is deliberate self-disintermediation.

Its job is to remove the retrieval and explanation work the client should no longer need me for. Set it against the shape it replaces, which is the shape almost every advisory relationship still runs on: client has question → book me → I think → I explain → invoice.

Replace it with: client has ordinary question → the rail. Then only genuinely difficult things rise — conflicting doctrines, novel client situations, new boundary cases, consequential interpretation, strategic perturbations, this doesn't fit anything we have, and does this mean the firm should change direction?

And the falsifier goes on immediately, in the same breath as the claim: the rail should cannibalise my routine explanatory work. If it does not, it is interface theatre.

There is a name for what this fixes, and it is not flattering to how most organisations use their seniors. "Hey Terry, can you write the change management chapter for this RFP?" sounds like teamwork. Structurally it is a lookup addressed to a human. The organisation needs the standard principles, the language that worked last time, and the difference between current practice and last year's template — and instead of querying a system built for retrieval, it queries a person built for judgment, and pays in calendar latency, incomplete recall and a document nobody wants to re-read.

Terry's scarce resource was never the ability to remember the boilerplate chapter. It was the ability to say: this client is regulated and hostile to change theatre, the standard chapter will lose, here is what must be different, and here is what we should refuse to promise. That argument is made in full in Wiki for the Humans, which is not published as an article, so it is named here rather than linked.

Applied to a mandate: the rail exists so that expensive human interaction is about judgment and novelty, and so that "expert" stops being a synonym for "human search".

What the Client Is Actually Buying

Not the corpus. And the counter-position deserves to be stated at full strength before the answer.

A corpus dump is not useful access. The buyer could already read my published books, put them in a competent model, and ask questions. Much of the framework IP is public by design — I publish it. Secrecy cannot carry this proposition, and any pitch that leans on privileged access to a body of writing is selling something that is one download away from being free.

So the promise is not all my ebooks in a chat window. It is: ask a live strategic question and receive the few relevant doctrines, their evidence and relationships, their applicability conditions, and the unresolved difference that requires judgment.

The theory underneath that is the reason it is even possible. When ideas become abundant, scarcity moves to matching, timing and application — a wiki is an IP matching engine, "not a museum. Not a graveyard of PDFs. A graph of claims and edges sitting ready" so that live context can be soft-joined to a latent idea. The architecture is almost embarrassingly simple once you stop optimising the cupboard: raw ideas become a compiled canon of pages, edges and names; live context arrives as person × company × problem × time; the agent routes to the two or three intersecting frameworks; and a human does the matching. That argument is Don't Vault Your IP, Route It, also unpublished as an article, and also named here rather than linked.

One line on why a chat over documents is a categorically different object. An archive of finished objects does not improve by sitting next to itself, and does not recombine when you ask a different question. "What sits in your hand now, if the stack is built properly, is not an archive of finished objects. It is a graph of claims and edges that recombines on every question." The value is finishing the thought while the thought is still warm.

The Hard Limit of Ask

Now the boundary, and it is the one that stops the rail being oversold as the answer to Chapter 2's condition.

Ask works only when someone knows what to ask.

The interface ladder makes the cost structure explicit, and it ships in this order:

Rung 1 — Ask

The typed question with a receipted answer. Ships the day the map exists. Cost to the user: they must know what to ask.

Rung 2 — Browse

The role-scoped map. Catches the question you did not have words for — recognition, where asking required knowing. Cost: a click.

Rung 3 — Ambient

The system sees the context and offers the answer. Cost: nothing at all.

"Each rung ships independently; none depends on the next; all read the same substrate. Each rung is just a cheaper question for the user to ask."

The consequence for a mandate is a real limit on the product, and it should be said rather than finessed: the rail handles accumulated territory. It does not, by itself, pierce anyone's fog. Chapter 2 argued that the strategic function is permanent because its questions do not converge. An interface that requires you to already have the question cannot be the answer to that condition, and any pitch implying otherwise is selling the wrong rung.

Which is why the service has three modes rather than one, and why the boundary becomes a design rather than a caveat:

  1. Self-service orientation — ordinary research, explanation, source retrieval.
  2. Decision-grade inquiry — mandatory counter-cases, falsifiers, source descent, explicit uncertainty.
  3. Frontier evaluation — me, or eventually another proven evaluator, on genuinely novel or consequential cases.

The mandate is justified mostly by the third mode, and by continually improving the first two. That sentence is also a warning about where the money would otherwise drift.

Every Serious Query Is a Disposition

An answer is not the output. A disposition is.

Four query dispositions

Served

Known doctrine adequately resolves the question. Obliges nothing — and that is the point. This is the disposition that should grow as a share of the total.

Qualified

The doctrine applies only after adding conditions. Obliges the conditions to be written back, or the next person hits the same wall and nobody learns anything.

Contradicted

Client evidence challenges an existing position. Obliges a review, not a patch.

Unresolved

A new framework, experiment or offer may be required.

The final category is not system failure. It is frontier detection. And it is the highest-value output the rail produces, because it is the only one that generates work the third mode has to do — which is to say, the only one that generates the thing the fee is actually for.

The Demand Observatory

Repeated questions are not evidence of slow users. They are cache misses in the organisation's knowledge architecture — and question frequency and recency reveal what knowledge matters and should determine compilation order. The owner who logged every question her practice asked her for ten years was never the problem in that story; she was the prototype, and the log was the demand-side map.

Read across a quarter, a query stream reveals:

  • which frameworks clients repeatedly activate;
  • where users cannot formulate the right question at all;
  • which answers still require me despite allegedly settled doctrine;
  • where a framework produces repeated exceptions;
  • which unanswered questions nominate a new offer;
  • which doctrine is no longer earning its place.

Which closes a more valuable loop than usage analytics ever could: client question → contextual answer → decision or contradiction → frontier review → doctrine or machinery change → improved next answer. The value is not that the corpus gets larger. It is that more routine questions terminate correctly, while genuinely novel questions reach the frontier sooner.

One guard, stated as a rule because it will otherwise be violated by accident: query popularity must never become truth authority. A corpus optimised for what gets asked becomes increasingly articulate inside its existing worldview, which is a failure Chapter 11 treats in full.

An Upstream Release Channel

It is not a library card and should not become a hostage mechanism. It is closer to an interface into a maintained upstream system.

The client owns and operates its local world: its evidence, its decisions, its applied configurations, its promoted project learning, and its ability to execute yesterday's accepted answer.

I maintain the upstream: new frameworks, qualifications and deprecations, rejection history, new evaluation cases, improved routing, newly observed failure classes, and the machinery for constructing the next offer.

The rail is where the current upstream meets the client's current state. Stop paying and your build keeps running. What stops is the release stream.

Notice that this is the same boundary Chapter 8 drew between the runtime plane and the evolution plane — restated as a distribution model rather than as a termination clause. Which is why it explains recurring value without requiring indispensability, and why the analogy is worth more than the metaphor it replaces: nobody thinks a software vendor is holding them hostage because the next release costs money.

The economics have a condition attached, though, and it runs both ways. The partner path has to remain economically preferable to bypass. If the client stays because departure disables its history, that is lock-in. If it stays because the maintained path is faster, better governed and more productive than rebuilding the evaluation function, that is an appreciating partnership. And the failure on my side of that trade is just as real: "if every ordinary engagement still requires your hands, the partner will correctly conclude they bought labour, not capability — and the copy path starts looking like liberation."

What I Have Actually Built

Time to be specific about my own machinery, in the middle of the chapter that would most benefit from leaving it vague.

What exists is a deployed ask-and-cite web interface over my compiled IP wiki. It walks a map, opens the relevant pages and source chapters, and cites only material it actually read. Deterministic index routing, full-page reads, two answer registers, citations validated against the run's read set, and a visible execution trace. Authenticated deployment.

Evidence ceiling

In my own note at the time: "askui work done, but it's a demo hard coded to the ip wiki." An audit found the coupling thin and localised — corpus paths already environment variables, prompts already files on disk — and the decision recorded alongside it was explicit: do not generalise until the demo has done its sales job or a client deployment is scheduled.

So it proves conversational access to my own doctrine, with real provenance. It does not prove comprehensive coverage, correct interpretation, decision quality, or that anybody will pay for it. The three-row ceiling table is in Chapter 11; the full ledger is in Chapter 15.

One further honesty note, because it is the kind of detail that usually gets hidden. Repeated versions of the same question can traverse different pages and chapters. That is not inherently a defect — different entrances into the same territory is what free agents do. The property that matters is whether materially different routes recover equivalent load-bearing evidence and reach the same qualifications, which is measurable and belongs to Chapter 13's trial hygiene.

The Division of Labour, Stated as a Purchase

The rail answers what do we already know? My work increasingly answers what don't we know yet? Those are two different products with two different cost structures, and pricing them as one is how a founder ends up subsidising retrieval out of the frontier budget.

There is an objection to all of this that I have been holding off, and it is serious enough to undo the chapter. If a client genuinely becomes AI-native — which is what I have been telling them to do — why would their people leave their working context, visit another website, and hand-formulate questions for my interface?

11
Part V: The Access Rail

Make the Kernel Callable

Why the interface I built may be the wrong primary surface — and what has to be exposed instead.

They mostly would not. That is the answer to the question the last chapter closed on, and it is worth accepting rather than arguing with.

A separate site means my intelligence sits in another silo, waiting for a human to remember that it exists. Which is exactly the failure the application era spent twenty years producing and that the agent era is supposed to end — the useful thing, one context switch away, being used by the people who happen to be conscientious about it.

The awkward implication for the previous chapter should be said rather than dodged: the rail I have actually built is the rung most exposed to this objection. That is not a reason to abandon it — the three modes are real, and orientation, inspection and exception handling are genuine jobs. It is a reason to stop treating it as the endpoint.

The stronger product is that the client's own AI invokes my kernel through a bounded, machine-operable connection. They keep their local context — clients, projects, decisions, internal history — and call me as an external specialist at the point of need. The human interface becomes the research desk and the exception desk. The agent connection becomes the routine distribution surface.

Pixels Are Not a Delegation Surface

That requires a different contract from the one a chat interface satisfies, and the difference is precise enough to be scored.

Two different contracts with the world

A human UI
  • • Pixels and layout
  • • Navigation and taxonomy
  • • Gestures, forms, confirmations
  • • Persuasion and brand
  • • Journey-state for eyes
A delegation surface
  • State — what is true now
  • Actions — what can be done
  • Delegated authority — who, how far, how long
  • Consequences — fees, failures, idempotency
  • Subscribe-to-changes — when the world moves

Shipping AI onto the left column does not create the right column. People are adding language models to pixels and navigation and calling it agent readiness. It is not a smaller version of the right column. It is a different contract.

And "we have an API" is usually not the answer, which is where most firms stop. A backend built for your own client proves the domain is computable. It does not prove that a customer's agent can, under a grant the customer controls: read only the slice of state it needs; invoke only the actions that were authorised; understand the monetary and policy consequences before acting; and learn when something changed without a human refreshing. "Agents fail in the gaps: authority semantics, structured consequences, and change subscription."

The fake solution deserves naming too, because it is what teams reach for under deadline: let the agent drive the human UI. Screenshot, click, scrape, hope the layout does not change. That is RPA cosplay. "Impersonating fingers is not addressability. Addressability is a stable, authorised, machine-legible surface you intended to expose."

One of Those Five Is This Book's Economics

Look again at the fifth element.

Key Insight

Subscribe-to-changes is the release stream. The recurring fee, expressed as a surface rather than as an invoice line.

A client agent that can ask what changed in the doctrine since we last decided this? and be answered without a human in the loop is the recurring product. Chapter 8 called that object the evolution plane. Chapter 10 called it the upstream release channel. This is the same thing with a wire protocol, and the convergence is not a coincidence — it is what happens when a commercial boundary and a technical boundary are drawn from the same principle.

The commercial consequence is larger than it first appears. It converts what am I paying for each month? from a narrative question into a machine-answerable one. The client's own agent can observe the release stream directly, which means the value of the licence becomes something they instrument rather than something I assert at renewal. That is uncomfortable and it is the correct direction.

There is a related requirement that sits one layer up, at the point of sale rather than the point of use: publish enough structure — eligibility, required inputs, price band, valid outcomes, exclusions, evidence returned, acceptance — that an authorised customer agent can determine fit without a discovery call. "If the future buyer arrives through an agent, the amorphous proposition is not just weak. It is invisible."

Insourcing Becomes a Channel

Do not make the client choose between its AI and your expertise. Make your expertise callable by its AI.

Spelled out: the client's AI performs generic cognition internally and purchases specialist external cognition at the point of need. Which moves the externalisation boundary from who performs the analysis? to whose maintained decision machinery does the client's agent invoke?

That is the resolution to this book's central tension, and it is worth being explicit about why. The mandate is supposed to make the client better at strategy than they were when they hired me. Every other framing treats that outcome as an erosion of my position — something to be managed, slowed, or quietly under-delivered. This one treats it as the condition under which my position becomes distributable, because a more capable client agent calls more often, not less. It knows when to call.

The service-side consequence is already well described: when the agent is the customer, UI polish stops being the moat. The human interface becomes "the showroom and the exception desk", and the delegation surface becomes "the loading dock". Addressability is not only defence against bypass — it is distribution into the customer's agent stack. And the board question transfers unchanged: why is the client's AI unable to use my business?

Are the Rails Ready?

Partly, and the gap is exactly where a mandate lives.

The rails themselves are in production. The A2A protocol "moved from initial release to a production-ready open standard" within a year, passing 150 supporting organisations including the major cloud and enterprise vendors, and sells itself explicitly on agents transacting "without being locked into a single vendor's ecosystem".13 MCP's own roadmap describes it running "in production at companies large and small".14

And they are running into precisely the problems a mandate cares about: enterprises deploying MCP hit "a predictable set of problems: audit trails, SSO-integrated auth, gateway behavior, and configuration portability."

The gap that matters most, though, is unusually on-point for this book. A systematic analysis of five agent interoperability protocols against a six-dimension governance taxonomy — membership, deliberation, voting, dissent preservation, human escalation, audit and replay — found that "voting and dissent preservation are universally absent across all five protocols", and concluded that agent community governance "constitutes a missing architectural layer above current interoperability standards, not a missing feature within them."15

Read that against everything in Part II. Dissent preservation is not an exotic requirement here — it is literally the mandate's cycle output. Contradictions. Rejected alternatives. Falsifiers. Reopening triggers. The protocol layer does not carry any of it, which means the mandate's governance has to live in the application above the protocol — and anyone selling "agent-ready advisory" on the strength of protocol support is selling the wrong layer.

The enterprise side of the same gap: 74% of respondents expect their companies to be using AI agents at least moderately by 2027, while only 21% report mature governance structures for them.16

The Three-Field Join

What is actually being sold, once the interface question is settled, is a join across three fields.

1. The capability kernel

Doctrine, successor-offer methods, evaluation methods, failure classes, and subsequent releases. Alone, it risks generic doctrine.

2. The client's firm kernel

Their offers, commercial history, project patterns, margins, methods, relationships and promoted institutional learning. Alone, it risks incumbent blindness and a narrow sample.

3. The live decision world

The present customer, project, evidence, constraint, option and authority boundary. Without deployment, the other two produce elegant advice; and deployment without evaluation produces motion without learning.

Remove any one term and the result collapses, which is what makes it multiplicative rather than additive — and which is why the product is the join rather than any corpus. The ownership architecture underneath was established in Chapter 8: three kernels, joined at runtime, without collapsing ownership territories.

The Evidence Ceiling

Which is the point at which this chapter has to publish its ceiling, in the middle of its most sellable section, where it costs something.

Layer Status
IP-only conversational access Working specimen
Buyer-account and bounded-kernel architecture Built productisation path
Federated client-context product with recurring market economics Still to be proved

Substantiate the middle row rather than asserting it, because "built" is the word most often stretched. The intended architecture is specific: a compiler operating on the buyer's side, reading the buyer's own compiled knowledge, and obtaining bounded, question-specific slices from my kernel. The kernel endpoint has exactly two authenticated routes — health and slice. No page listing. No export. No pagination. Deterministic code caps page count, excerpt size and total response bytes, checks revocation before accepting a token, applies per-key rate limits, and writes an audit line for allowed and denied requests alike. That is a real, built commercial boundary, and it took the Marketplace listing with it when the listing was left unfinished.

And the third row, without softening. No client context has ever been joined. The federated product is unproven. Not "early". Not "in pilot". Unproven.

The reason to write the three rows down is that otherwise the strongest one lends its maturity to the others — a deployed interface starts to sound like a shipped appliance, which starts to sound like a federated product, and nobody ever decided to make that claim. Chapter 15 runs the same instrument over every claim in this book.

Expose the Comparison; Do Not Smooth It

Assume the third row exists one day. There is a design rule that has to be in place before it does, and it is not obvious.

The join must not return one smooth voice that assimilates the client into my worldview. It must expose the comparison.

The response format — five fields, and what an empty one means

Relevant doctrine

Which two or three frameworks bear on this. Empty means: the question is outside the kernel's competence, and saying so is more valuable than a plausible answer.

Client evidence

What their own record says. Empty means: this is generic advice, and it should be labelled as such rather than delivered with their logo on it.

Agreement or contradiction

Do the two align? Empty means: nobody actually compared them, which is the failure the whole format exists to prevent.

Decision consequence

What changes if this is true. Empty means: it is interesting rather than useful.

Falsifier or missing observation

What would show this is wrong. Empty means: the claim is unfalsifiable as stated and should be rewritten before anyone acts on it.

A strategically honest rail must be capable of producing all of these sentences: the client evidence supports the doctrine; the doctrine applies only under narrower conditions; the evidence contradicts the doctrine; no current framework explains this; and this may change the canon. A system that can only produce the first is not a decision surface. It is a compliment.

Compiled Blindness

Because there is a failure mode that the smooth-voice version produces reliably, and it is invisible from inside.

Pitfall — the five-step slide

  1. New evidence is retrieved through existing concepts.
  2. The model produces a fluent explanation in the canon's language.
  3. The explanation feels true because it is recognisable.
  4. An n=1 anomaly is absorbed as confirmation.
  5. Subsequent retrieval makes the interpretation look increasingly established.

That is not compounding judgment. It is compounding confidence. And the better the canon, the better the laundering — which makes a dense, well-maintained corpus more dangerous here than a thin one, not less.

The remedies are named and they are cheap relative to the failure. Three guards: an alien-signal lane for consequential observations with no detectable fit to the current map; a suppression audit that periodically samples what the system dismissed; and a replayable trace. "Missing guards make compounding unsafe."

And the reason to bother, from the same argument, which is also the best short statement of what the licence is actually selling: "After a year, someone can steal the code. They cannot steal the year. The wiki does not just get bigger. It becomes more discriminating."

The institutional version of the same rot has its own name: calcified lore. A conditional pattern hardens into an unconditional rule, and the tell is "a claim with no conditions attached that everyone obeys". At that point you have replaced an opaque model prior with an opaque institutional one — "and the second is worse because it has authority." The fix is to keep the conditions on the claim: scope, conditions, evidence, consequence, escalation rule.

Which produces a falsifier the client can apply to me, and should: if successive cycles always confirm the canon, rarely demote anything, and generate no genuinely alien categories, the system is assimilating evidence rather than learning from it.

Key Insight

Contradiction is not a defect in an IP system. It is one of its highest-value outputs — and it is the output the client is actually paying for.

Comprehensiveness Is Not the Goal

One last correction, because it is the instinct every IP-heavy vendor has and it points the wrong way. "Access to all my IP" should become "access to a bounded, current, evaluated decision surface." Literal comprehensiveness is neither necessary nor necessarily desirable: a larger corpus with a worse evaluation function is a worse product, and it is also more expensive to maintain.

There is a self-test that comes free with that position, borrowed from a neighbouring discipline and sharp enough to run this week: "If every answer requires the full thread in context, you have not compiled — you have wrapped search."

That is the internal version of the trial two chapters from now. It is also the cheapest possible diagnostic on your own system: take a real question, and see whether the answer can be assembled from claims, relationships and pointers, or only from stuffing the corpus into a context window and hoping.

All of which has been assuming something that has not yet been tested anywhere in this book. Every argument in Part V rests on the maintained kernel being an appreciating asset — that it gets more valuable, relative to the alternatives, as the models improve.

I do not currently know whether that is true.

12
Part VI: Proof

Net AI Beta

Does the next model release make my kernel more valuable, or my client's substitute better faster? One instrument, and an answer I do not have.

The question has a shape, at least. When the next frontier model lands, two things improve at once: the maintained kernel, and the thing that would replace it. Only the difference is a business.

This is the shortest chapter in the book, and it is short because it contains one instrument and one admission. That is not the same as thin.

The Case For, Made Properly First

A casual user gets a slightly better chatbot with each model release. An operator with a flywheel of prior frameworks, a semantic corpus, a compiler, an evaluation function and a documented taste kernel gets something categorically different: a free uplift to the entire production function.

I have watched that happen to my own corpus, and the strategic implication is the whole argument for building early:

"the value of a model improvement depends on how much machinery you have built to absorb it — and the time to build that machinery is now, because it cannot be acquired retroactively by spend after the drop lands. The dividend compounds for first movers, and the gap widens with each release."

That is a real mechanism and I am not walking it back. Any weaker version of it would make the rest of this chapter read as false modesty, which is worse than overclaiming because it is harder to argue with.

The Case Against, in the Same Breath

The same model improvement makes the client's substitute better.

A frontier model, plus my published books, plus firm-specific context I will never have, is a genuine competitor to a maintained kernel — and it improves on exactly the same schedule, at none of the maintenance cost, with no licence fee.

This is not symmetric with the dividend, and the asymmetry is the trap. The dividend is almost always measured gross: against the kernel's own prior performance. Which is the one comparison guaranteed to look good, because a prepared apparatus does absorb an upgrade better than an unprepared one. That measurement is true and it is not the business question.

Key Insight

A static corpus can have negative net AI beta even if its query interface improves. The public model gets better at reconstructing the same answer while the corpus ages.

Anybody selling access to a library should find that sentence uncomfortable. Which, stripped of positioning, is what most "IP subscription" offers are — including, until it is tested, mine.

The Instrument

net AI beta
dividend to the maintained kernel
− client substitution gain
− commoditisation of its outputs

Three terms, each of which can be populated by observation rather than assertion.

  • Dividend to the maintained kernel. How much better the kernel path gets when the model improves — more territory reached, better joins across frameworks written months apart, fewer misses on things it should have found.
  • Client substitution gain. How much better the client's own path gets from the same release: frontier model, plus published doctrine, plus their own context. This term is the one an incumbent never measures, because measuring it requires deliberately running the competitor.
  • Commoditisation of outputs. How much of what the kernel produces becomes reconstructable without it. This is the term everyone forgets, and it moves fastest — an answer that required a compiled canon last year can become a competent zero-shot response this year, and nothing about the canon changed.

Which is why the Model Dividend does not settle the question, and it is worth saying directly because I am the one who published the dividend. It establishes that prepared machinery captures more of an upgrade. It says nothing about whether the gap to the substitute widens or narrows. Gross and net are different quantities, and only one of them is a business.

Six Assets That Make Beta Positive

If the answer is going to be positive, it will be because of things a general model and a set of published books cannot automatically possess. There are six, and each one has a maintenance cost that turns this list from a claim into a build plan.

1. Resolved client cases with later outcome receipts

Why a general model cannot hold it: not the recommendation — what actually happened afterwards. None of it is public and most of it is never written down anywhere.

Cost to maintain: an outcome-backfill discipline that runs long after the invoices stop, on work nobody is currently paying for.

2. Rejection history and known failure paths

Why a general model cannot hold it: what was tried and did not work is the single most valuable and least-published class of knowledge in any expertise business. Nobody writes up the thing they abandoned.

Cost to maintain: the discomfort of recording your own wrong turns in a form somebody else can read and cite back at you.

3. Evaluation tests and calibrated escalation rules

Why a general model cannot hold it: this is the machinery that decides what an answer means, and it is the thing an outsider cannot infer from the answers themselves.

Cost to maintain: keeping a test suite for judgment, which nobody budgets for and everybody defers.

4. Continually updated relationships between concepts

Why a general model cannot hold it: not more pages — better edges, including deprecations. The edges are the part that ages badly and invisibly.

Cost to maintain: a janitor function that subtracts as well as adds, which is the least rewarding work in the building.

5. New sensors and offer patterns from live work

Why a general model cannot hold it: it requires being in the room where the pattern first appeared.

Cost to maintain: the promotion protocol from Chapter 8, run properly rather than assumed — classify, de-identify, abstract, verify, human gate.

6. Evidence that the kernel changes real decisions

Why a general model cannot hold it: it is a claim about consequence, and consequence has to be observed rather than reasoned about.

Cost to maintain: the instruments in Chapter 13 — and the willingness to publish what they return.

The through-line is worth stating as a belief rather than leaving implicit. Every one of those six is a flow, not a stock. Which is exactly why a published corpus has negative net beta and a maintained one might not — and it is why the recurring fee is attached to the maintenance rather than to the access. If the fee were for access, the argument in this chapter would already be lost.

A Worked Assessment Across One Model Upgrade

Specified so it can be executed next release rather than admired.

The protocol

Before the release

  • Freeze a question family. A set of consequential questions in one domain, with a fixed decision standard. Freeze the wording — and separately freeze a paraphrased variant of each, so route variance is testable later.
  • Record both paths. For each question: what the kernel path returned, and what the published-books path returned. Keep the walks, not just the prose. The path is the diagnostic; the answer is only the summary.
  • Write down what you expect to happen to each of the three terms. Before, not after.

After the release

Re-run both paths on the frozen family, unchanged. No prompt improvements, no "while we're here" tidying of the corpus. Any change on your side invalidates the comparison and you will not notice.

Then assess three movements — and only three

  1. Did the kernel path get better?
  2. Did the books-only path get better faster?
  3. Did any answer the kernel used to own become reconstructable without it?

Score direction and comparison, not percentages. The honest before-and-after discipline is to hold the goal fixed, vary the orientation condition, and assess time-to-value, relevant breadth, quality and residual human judgment — rather than inventing a productivity figure. The neighbouring argument is blunt about why: corporate value here is "useful joins, not headcount multipliers", and "anyone selling that multiplier is inventing".

One sharpening from the same source, because it defines what the kernel path is supposed to do and therefore what a positive result would look like. The test of a compiled field is whether it can "surface a structurally adjacent project the asker did not know existed" and "separate evidence from inference when the join is surprising". If the kernel path only summarises the document you already named, you did not buy orientation capital — you bought faster archive access, and net beta will be negative whatever the interface does.

Worksheet field What goes in it
Question family Domain, question count, decision standard, and the paraphrase variants
Freeze date Before the release was announced, ideally
Model versions Both paths, before and after, pinned
The three movements Direction and comparison per movement, with the walks attached
Disposition Asset / neutral / eroding
Pre-committed consequence What happens commercially under each disposition — written before the re-run

What I Have and Have Not Run

I have not completed this assessment across a model upgrade for my own kernel.

The protocol is specified. The result is not claimed. And that admission is the chapter rather than a caveat attached to it — because a book arguing that recurring fees must attach to an appreciating asset, published by somebody who has not measured whether his own asset is appreciating, has exactly one honest option: publish the instrument and the gap in the same place.

Myth vs Reality

Myth

Better models make my accumulated IP more valuable. The machinery absorbs the upgrade; the corpus deepens; the moat widens with every release.

Reality

Better models make everybody's substitute better too, including the substitute for you. Only the net matters — and nobody who has not run the frozen-family comparison knows the sign of their own.

What a Negative Result Would Actually Mean

Not "the kernel is worthless". That is the conclusion a reader will jump to, and it is wrong in a way that matters commercially.

A negative net beta means the licence is not the product. The mandate would then rest on the navigation cycle and the frontier work — the disposition machinery in Part II, and the judgment that Chapter 10 said the machine does not automatically outperform. That is a survivable finding. It is a smaller business than the one this book describes, and it is still a business.

What makes it survivable is timing. Knowing it after one model release costs a fee-sheet line. Discovering it at the third renewal costs the relationship, because by then the client has worked it out first and the conversation is no longer about design.

So let me pre-commit the consequence here, and Chapter 13 will inherit it: if net beta is negative across two model upgrades, the capability-licence line in Chapter 9's fee sheet is downgraded to a delivery convenience, and Part V of this book is demoted with it.

The Question to Take Away

Not is my corpus valuable? — which is unanswerable and flattering, and which every expertise vendor answers correctly and uselessly.

Does the next model release widen or narrow the gap between my corpus and my client's substitute?

That one has an answer, the answer changes what you should charge for, and almost nobody has measured it.

Net beta asks whether the asset appreciates. It does not ask whether anybody should pay for it — and that needs a trial with the consequences written down before it runs.

13
Part VI: Proof

The Trials You Commit To Before You Run Them

Three arms, four consequences written in advance, and a scoring rule that only counts differences that changed something.

A demonstration is designed to succeed. That is not a criticism of demonstrations — it is what they are for. But it means everything in Part V has been a demonstration, and a demonstration cannot settle whether anybody should pay.

So this chapter specifies the tests. And it writes down what happens commercially under each outcome, before any of them are run, because a consequence written afterwards is a rationalisation wearing a lab coat.

The Case Against My Own Product

Stated first, at full strength, without rebuttal. If I cannot make the sceptic's argument better than the sceptic can, the trial that follows is theatre.

Six reasons the access rail may be worth nothing

  1. The principal books are already available.
  2. The client can put them in a competent model.
  3. Chat over documents is increasingly cheap to reproduce.
  4. The client already possesses better firm-specific context than I do.
  5. The rail may improve convenience without changing a single material decision.
  6. In that world, the licence is packaging — not a strategic capability.

And then the version that actually stings, because it turns my own doctrine on my own product.

At a shallow level the professional-services question is how do we use AI? Then it becomes how do we sell AI? — trinkets, copilots, licences and add-on products bolted onto what the firm already does. I named that category, and I have been rude about it in print.

If the access rail is simply a polished chat feature attached to a large monthly relationship, it is exactly that. An AI trinket, sold by the person who coined the term, inside a mandate that criticises the practice. The industry version of that argument belongs to a sibling book; what belongs here is the self-application, and the only honest response to it is not a better argument. It is a held-out comparison with consequences committed in advance.

What the Public Evidence Already Says

It cuts both ways, which is why it belongs before the protocol rather than after it.

First: a flat corpus plus a frontier model performs badly on realistic enterprise material. A benchmark for enterprise "deep search" — source-aware, multi-hop reasoning across documents, meeting transcripts, chat messages, code repositories and URLs, over a retrieval pool of 39,190 artefacts, with both answerable and unanswerable queries — found that "even the best-performing agentic RAG methods achieve an average performance score of 32.96", with retrieval named as the main bottleneck: "existing methods struggle to conduct deep searches and retrieve all necessary evidence. Consequently, they often reason over partial context, leading to significant performance degradation."17

Second: structure changes the result, measurably and in production. A customer-service system that stopped treating past tickets as plain text and instead built a knowledge graph capturing intra-issue structure and inter-issue relations "reduced median per-issue resolution time by 28.6%" after roughly six months of live operation.18

Two numbers, pointing in opposite directions

32.96

average score of the best agentic RAG methods on a realistic heterogeneous enterprise corpus. Chat over documents is not a solved problem.

28.6%

median per-issue resolution-time reduction when the same class of corpus was turned into typed structure. Different domain, operational metric.

Put together, they give the sharpest available statement of what the trial is actually testing: the delta lives in the structure, not in the access. Which means arm B below is not testing whether my corpus is good. It is testing whether it is compiled.

Be careful about what the second result does not establish. It is a different domain, a different task, and an operational metric rather than a decision-quality one. It shows the mechanism is real. It does not show that it transfers to strategy, and anybody quoting it as though it does — including me, if I am careless — is doing the thing this book keeps refusing.

The A/B/C Trial

Arm A

A frontier model plus my published books. The substitute the client already has, for free.

Arm B

The access rail over my compiled kernel. What I have actually built.

Arm C

The federated rail over my kernel, the client's firm memory and live project evidence. What I have argued for and never built.

Hold the question and the decision standard constant across all three. Then compare:

  • time to an accepted decision-grade answer;
  • relevant prior projects and doctrines activated;
  • unsupported-claim and correction rates;
  • human re-briefing burden — how much the person had to re-explain to get a usable answer;
  • counter-cases and falsifiers discovered;
  • whether an offer, test, capital allocation or decision actually changed;
  • whether I was required for retrieval, or only for genuine novelty.

Question selection decides whether any of that means anything. Use consequential questions from live work rather than curated showcases. Include questions the kernel should handle badly — a trial with no expected failures is a demo. And include at least one question where the client's own context is decisive, or arm C is untested by construction and you have run a two-arm trial with a third label on it.

Consequences, Committed in Advance

Signed before the run

If B does not materially outperform A

IP-only access is packaging. The capability licence loses its line in Chapter 9's fee sheet.

If C materially outperforms B

The federated join — not corpus access — is the product, and the commercial emphasis moves accordingly.

If C changes important decisions while reducing my routine involvement

The recurring capability is real.

If all three produce similar outcomes

The human frontier work and the build capability may still be valuable, but the access rail carries no significant independent value and must not be priced as though it does.

Pre-commitment is the whole mechanism, and it is worth being explicit about what it costs. It removes the option of reinterpreting a disappointing result as "early signal" — which is the standard move, and the reason most internal trials never change anything. A thesis that names its own demotion conditions is not philosophy. It is a design with a stop rule.

Trial Hygiene

A badly run trial is worse than none, because it produces a number that ends arguments.

Vary the wording and the starting context. Different navigation paths are not a defect — path variance "is what free agents do when the territory has more than one entrance."

The load-bearing property is evidence invariance: whether materially different routes recover the same load-bearing sources, chapters and claims. "A page opened and unused is not evidence. A chapter that carries the warrant is."

And answer invariance is necessary and insufficient, which is the trap a careful team walks into. "Answers can agree while floating free of evidence. Score conclusions and qualifications — two answers that share a headline but drop opposite caveats are not invariant."

What varies What stays stable Interpretation
Path Evidence and answer Healthy route resilience
Path and evidence Answer The dangerous cell. Useful redundancy — or model-prior luck
Answer Evidence Synthesis or judgement instability — not a navigation problem
Everything Nothing The question family is not yet one task

Pitfall — the cell that produces a false positive

Row two is where a careful trial goes wrong, and it looks like success. Paths differ, the opened sources differ in ways that matter, and the answers still sound the same — same headline conclusions, similar confidence, tidy structure. Two rival stories explain that. Either the territory has multiple valid warrants, which is genuine redundancy. Or the answer is stable because the model already believed it and the graph is decoration.

How to tell them apart: remove a source that should be pivotal and see whether the answer flinches. If it survives without an alternate grounded route and without an honest confession, you have model-prior substitution. The fix is grounding discipline and synthesis gates — not a prettier path dashboard.

And repeat the frozen cases after model upgrades, which is where this chapter meets the last one. The two instruments should share a question family, because the question you are asking in Chapter 12 — is the gap widening or narrowing? — is answered by re-running the same three arms.

The External Decision-Delta Trial

A/B/C tests the kernel. This tests the mandate, and it is the harder of the two because the subject is a relationship rather than a system.

The protocol

  1. After transfer, the client team runs a navigation cycle without me.
  2. I run an independent shadow cycle from the same agreed evidence.
  3. Compare the resulting options, falsifications, dispositions, evidence quality and elapsed time.
  4. Count only consequential differences.
  5. Renew the full mandate only where the shadow cycle creates a material delta over the client-owned apparatus.

Step four is where this either becomes an instrument or an argument, so specify it before anybody is invested in the outcome.

Counts as consequential Does not count
A different disposition on a live option Better prose
A different reopening trigger More references
An option killed that the other side kept, or kept that the other side killed Greater confidence
A falsifier the other side missed A more articulate framing of the same disposition
A material evidence gap identified Finding the same thing faster

Who adjudicates is not a detail. I am an interested party in my own renewal, so this is Chapter 7's sixth control applied to me: the adjudication cannot rest solely on the people who produced the shadow cycle. Three practical options, in ascending order of cost and descending order of convenience: a client-side reviewer holding the scoring rules; a rotating panel drawn from their own leadership; or a third party. The first is the only one most relationships will actually run, so design for it — which means the scoring rules have to be simple enough that somebody with a day job can apply them without me in the room.

And the correlated-evidence trap from Chapter 7 applies here with full force. If my shadow cycle reads only the evidence my own kernel curated, the trial compares me to myself and returns a flattering result. The agreed evidence set has to be theirs, and the walk has to be recorded on both sides.

Pitfall — the trial that runs once

This is not a test of the client's team. It is a test of my marginal contribution, and if it is not framed that way in writing, it will be experienced as an audit of their people — which is a thing you get to do exactly once. Put the framing in the protocol document. Sequence it so their cycle completes and is committed before they see mine. Agree who sees what, and when. A trial that damages the relationship it was meant to justify has proved something, but not what you wanted.

Alongside it, five measures worth tracking that cost almost nothing once the register exists: share of ordinary work completed without me; time from market signal to authorised disposition; active option WIP and closure rate; time and cost to launch the next offer relative to the first; and reusable learning promoted or rejected, with receipts.

The Ratio Underneath Both

successor-offer leverage
paid bounded units ÷ scarce expert dispositions

Both halves need discipline or the metric lies.

Numerator discipline. Count only paid, bounded commercial units — navigation cycles sold, licences active, bounded builds commissioned. Not free pilots. Not demos. Not "influenced" pipeline. And not every AI-assisted task inside a legacy day-rate project, as though the unit of sale had changed when it did not. "Inflating the numerator is how consulting with software around it looks like a product on a dashboard."

Denominator discipline. Count only genuinely scarce expert work — material dispositions requiring authorised judgment under consequence. Not every human touch. Not junior assembly the machine should have absorbed. Not project-management theatre. "If the denominator secretly includes all labour, you are back to a labour productivity metric. If it secretly excludes the founder's night work, you are lying about scarcity. The denominator is a design claim about where authority still lives."

What improvement looks like, without invented numbers: cycle n+1 delivers more paid bounded units per scarce disposition than cycle n, or holds units while disposition intensity falls, without quality collapse. Exception classes become typed, and fewer of them require the same principal.

Key Insight

If the ratio does not improve across engagements, you are still selling consulting with software around it.

There is a finding hidden inside the metric that applies to me right now: "If the firm cannot yet measure the ratio, that is itself a finding: the offer is not instrumented enough to claim scalability." That is currently true of my own mandate, and it is in Chapter 15's ledger.

The anti-patterns are worth listing because each is a live temptation in a recurring relationship where nobody outside is checking: counting free work in the numerator; moving scarce work off the books as "sales support" or "R&D"; redefining dispositions mid-stream to protect a narrative; comparing unlike units without banding; and declaring victory after one heroic cycle.

What to Put on the Wall Before the First Sale

  1. What counts as one paid bounded unit.
  2. What counts as a scarce expert disposition — classes, not vibes.
  3. How both will be logged per cycle.
  4. What "improve" means for the next three cycles.
  5. What demotion looks like if the ratio stagnates.

Five sentences. Written before anything is sold, because after the first sale each of them becomes a negotiation with somebody's revenue attached.

One Adjacency, Then the Close

A related instrument asks a different question, and I want to be clear I have not confused the two. It runs temporally frozen cases across four arms — me with the kernel; another practitioner with the kernel; a strong practitioner given the copied questions but no kernel; and the client's own team with its own AI and context — to find out whether the founder rather than the kernel is the product. That ablation belongs to a sibling book.

This trial asks whether the kernel earns its fee.

Query volume is not the proof. Decision altitude is.
14
Part VI: Proof

The Renewal You Should Lose

Six questions, seven stop conditions, seven legitimate endings — and the version of this product I would be proudest of.

"We've been talking about running the next cycle internally."

Your stomach drops. Mine does. And that drop is the diagnostic — not the sentence, the reaction to it. If a client saying they might be capable of doing this themselves registers as a threat rather than as a result, then somewhere in the preceding two years the product quietly became something else.

Decision altitude was the standard. This is where it gets applied to a decision that costs money.

Six Questions, Every Quarter

In writing, answered from evidence rather than from recollection.

The question What answers it
1. Which load-bearing assumption changed? The assumptions register diff, with the evidence that moved it
2. Which strategic option was advanced, killed, deferred or newly created? The disposition register from Chapter 4 — with ages
3. Which question became a client-owned recurring capability? The standing-question register, with owners on their side
4. Which ordinary activity no longer requires me? The escalation log, and the dependency gradient from Chapter 5
5. Which new release or tool materially changed what the client can do? The release notes — and whether anybody used them
6. Why is the next quarter worth buying rather than internalising? No artefact. It has to be argued fresh, every quarter, against a live alternative.

Question six is the centre of this chapter, and it is the one no retainer asks because no retainer survives it.

The honest form is not are we useful? — a question to which the answer is almost always yes and which therefore decides nothing. It is: is the next quarter of this more valuable than the same money spent building it internally? Asked at a moment when the client has just spent a year becoming more capable of doing exactly that, largely because of work I did.

Which means the question gets harder every quarter the mandate succeeds. That is not a flaw in the instrument. It is the instrument working. A recurring product whose renewal argument gets easier over time is measuring accumulated switching cost and calling it trust.

Seven Stop Conditions

Narrow or stop the product if, after two quarters:

The stop conditions
  • • No material decision changed.
  • • No option was exercised or correctly killed.
  • • The outputs are mostly calls, decks or marketing prose.
  • • The client remains equally dependent for ordinary work.
  • • No client capability was promoted into durable operation.

The lock-in detector

• The licence is valuable chiefly because exit would disable access to their own history.

  • • The annual value case cannot be reconstructed from client evidence.

The sixth is separated for a reason. It is the failure Chapter 6 named arriving where it was always going to arrive, and it is the only stop condition that produces excellent commercial metrics right up until the relationship ends badly and in public. Renewal rate strong. Satisfaction fine. What is being measured is a hostage with an invoice.

The seventh is the quietest and the most useful. If the client cannot reconstruct the annual value equation from Chapter 9 using their own evidence, then whatever I have been sending them has been reporting, not evidence. Two quarters is generous for that one.

And on the number: two quarters, not one. One bad quarter is noise — evidence collection is lumpy and outcome clocks are slow, which is the entire argument for recurrence in the first place. Two is a pattern. Naming the number in advance is what stops the conversation becoming a negotiation about whether this particular quarter counts.

Seven Legitimate Endings

Each a designed outcome with a commercial shape, not a failure with a face saved.

1. Renew the complete mandate

Everything continues. The fee is unchanged. This is one of seven, not the default.

2. Narrow to the kernel licence only

Navigation stops; the release stream continues. The fee drops to one component. This is the ending the worked case below lands on.

3. Commission a separate build

The mandate pauses or continues alongside. The build is its own purchase with its own acceptance, per Chapter 7.

4. Take an option to another implementer

The decision pack is portable, and this is what portability was for. The mandate fee is unaffected; the build revenue goes elsewhere.

5. Move navigation inside

The client runs the cycle. I may remain on frontier escalation, priced as a reserve, or not at all.

6. Buy out selected rights

Priced in advance under Chapter 9 rather than negotiated under pressure. A one-off replaces a recurring line.

7. Terminate and continue operating the last accepted state

Chapter 8's boundary, exercised. Everything they have keeps working. The fee goes to zero.

Five of those seven reduce or end my revenue. All seven are outcomes I would sign.

A product with one good ending is not a product. It is a subscription with a renewal form, and everyone in the room knows it — which is why the renewal conversation in most advisory relationships is conducted in a slightly embarrassed register.

The Counter-Case That Is Also the Strongest Chapter

State it without hedging, because hedging it is what makes the rest sound like sales.

After the first two engagements the client may hold: the doctrine; a functioning internal kernel; trained people; its own question ledger; the tools and the appliance; superior company context; and access to the same frontier models I use. Its internal team may then navigate the next frontier more effectively than I can.

If that happens, non-renewal is not betrayal. It is evidence that the recurring offer is unnecessary for that client.

The third product has to earn itself repeatedly. If the client stays only because its history becomes inaccessible, the software stops working, or nobody else understands what I built, then renewal proves lock-in rather than value — and I have built the thing this book was written against, with better vocabulary.

Which is the KPI inversion arriving at its destination. In consumer AI, the tell is what a product does with exit: engagement metrics cannot celebrate a user leaving, so a product scored on time-in-app has already answered the loyalty question in code. As the original puts it: "The KPI set must be allowed to celebrate exit. If exit is scored as churn, the principal was never the user."

Key Insight

If a lost renewal is scored as churn, the principal was never the client.

Most readers will recognise themselves in the consequence. A CRM with a "churned" stage and no "correctly internalised" stage has already answered the principal question, in code, without anybody having decided anything.

A Renewal, Correctly Lost

A designed scenario, not a client engagement. The mandate described in this book has not been sold to anyone. What follows is the shape of an outcome the instruments are built to produce, worked through so it can be argued with — not a case study, and not a composite of real engagements.

A mid-market data and analytics consultancy, two years in

The bounded decision and the bounded construction both delivered. The mandate has run eight quarterly cycles.

What the gradient showed in the four cycles before the renewal

  • Known escalations falling, quarter on quarter. The shapes that used to arrive in my inbox — pricing exceptions on the new offer, scoping disputes on a familiar engagement class — are being closed against rules their own people wrote.
  • Standing questions transferred. Four of the six live questions now have owners in their commercial team, with cadences and retirement conditions, and they run whether or not I am on the call.
  • They present their own decisions. Their people took last quarter's disposition to their own board and defended it, including the rejected alternatives, without me in the room. That was the row I watched hardest, and it moved last.
  • My attention concentrated. Two genuinely novel questions this quarter, rather than fourteen familiar ones. Which felt, from the inside, like being needed less.
  • Time-to-launch fell. The second successor offer reached market materially faster than the first, and the difference was machinery rather than heroics.

What the decision-delta trial returned

Their internal cycle produced two of the three consequential differences my shadow cycle found. The third was a reopening trigger they had not set — a named competitor threshold on a deferred option — which they adopted. Score it honestly: a material delta, and a narrowing one. One trigger is not four cycles of navigation fee.

The renewal conversation

Six questions, answered from their evidence. Questions one through five answered well — the mandate had done what it said. Question six answered by them rather than by me, which is how it is supposed to work: they could argue the navigation cycle was now cheaper to run internally than to buy, and they were right. Disposition: terminal state two — narrow to the kernel licence only.

What happened to the money

I stop invoicing navigation. I keep invoicing the licence and the occasional bounded build. That is a substantial revenue reduction, and there is no way to present it as anything else. It was the correct outcome, and it would have arrived eventually whether or not it was designed for — the only question was whether it arrived as a decision or as a surprise.

What the relationship became

Not smaller. Different. An upstream release channel plus frontier escalation, with the client operating the function they bought the capability to operate. Read against Chapter 4's scorecard — options closed, residue delivered, work-in-progress flat or falling — every line is green. Written up as churn, it is a lost account. Written up honestly, the client internalised the function, which is what the gradient was for, and kept buying the one thing they could not build.

"A supplier confident in its evolving capability should not fear those exits. It should expect the authorised path to keep winning because it is better, not because every alternative was disabled."

The Falsifier I Have to Live Under

There is a version of failure that no renewal conversation would catch, because it looks like success from both sides of the table.

"If this discipline makes clients more dependent on the adviser rather than less, it has failed on its own terms — regardless of whether the mechanism is true."

The observable is specific enough to be checked by somebody who does not trust me: after four quarters, is the client running the diagnosis, designing the probes and closing the rows themselves? If the answer is no, the transfer failed and the mandate is an expensive conversation subscription.

That is the harshest instrument in this book, and it is harsh because of what it declines to ask. It does not ask whether the doctrine is correct. It asks whether the doctrine survives being sold by its author — which is a separate question, and a much less flattering one.

The Renewal Standard

Not "they used the rail a lot". Six observables:

  • known-terrain dependence on me fell;
  • consequential questions were answered earlier;
  • contradictions altered doctrine or client action;
  • options were opened, killed or constructed;
  • the client's own operating capability increased;
  • the maintained upstream made the next decision materially better.

Read that list once against a bad quarter, because the failure mode is not obvious from inside. High activity. Warm relationship. Everybody enjoying the sessions. No contradiction anywhere. Nothing killed. And a client no more capable than they were three months ago. That quarter feels like a good one — better than a quarter spent killing two options and handing over a register — which is precisely why the standard is written down rather than remembered.

Two Channels, Opposite Directions

Channel Function Direction over time
The rail + the living kernel Retrieve and apply known doctrine Client self-sufficiency rises
Me + the build capability Interpret discontinuities, construct new options My attention moves toward novelty

That is the finished product on two rows, and the important word is in the third column. The two rows are supposed to move in opposite directions. A relationship where both rows grow is not a healthy mandate. It is a firm becoming more dependent while being told it is becoming more capable.

Kill Conditions for the Programme Itself

A product that cannot fail is a story. Prefer a programme that can die honestly to one that lives forever as theatre — and watch for the specific rot, which is that "renaming failure as 'culture change' is how archive projects become permanent."

Repair when the gradient is moving and the instruments exist: rules are fossilising, escalation rates are falling, the trend is up even if the level is poor.

Demote when the premises are false: every question still novel after honest compilation, no way to score whether an answer was right, or a product face that can only be kept by heroics.

The negative patterns worth naming for a mandate specifically, because each one is comfortable: the cycle that only produces prose; the escalation that dies in a private message and leaves no rule behind; the dashboard that never measures absence; and the renewal defended with relationship language rather than with the six questions.

Falsifiability is a strength here rather than a hedge. A programme that can name its own death conditions is safer to fund than a permanent knowledge theatre with no stop rule — and a buyer sophisticated enough to be worth having will recognise the difference immediately.

Before the first invoice, while it is still cheap

Write the two-quarter conditions under which you would tell this client to stop buying. Six questions on one page, seven stop conditions on the next.

If you cannot write them, you do not have a mandate. You have a retainer with better vocabulary — and the difference will be discovered by your client, at a time of their choosing.

The Whole Thing, Compressed

The mandate maintains a client-owned portfolio of terminal-value questions and options inside permanent fog. I supply the living capability kernel, disciplined challenge and next-option machinery; the client retains authority, evidence, operating capability and the right to leave. Each cycle transfers yesterday's answer into the client while earning renewal by finding, testing and constructing tomorrow's.

The right to leave is written into the sentence rather than appended to it. That is the whole design, and it is the only part of it I would refuse to negotiate.

Which brings the argument back to where Chapter 2 left it. The adviser must also survive the externalisation boundary. Everything I have been telling professional-services firms about cheap cognition applies to the firm doing the telling, and there is no altitude high enough to escape it:

You do not escape cheap cognition by moving vaguely "up the value chain." You escape only by transferring what has become reproducible, retaining what continues to evolve, and proving — cycle by cycle — that the externally supplied frontier is still worth buying.

So: I deliberately give the client direct access to the portion of my advisory they can and should self-supply. I stop selling myself as a human retrieval interface. And I preserve the scarce part — sensing the perturbation, locating its significance inside a deep canon, challenging the canon when it no longer fits, and constructing the next commercial response before the pressure makes the old one irrelevant.

The client owns yesterday. You have to keep earning tomorrow.
15
Part VII: The Specimen Under Its Own Test

What I Can and Cannot Yet Prove

Every claim in this book at its highest defensible rung — including the ones that are argued and nothing more.

The last chapter specified the conditions under which a client should stop paying me. The same discipline now runs over my own claims, because a doctrine that exempts its own conclusions from its own test is a brochure.

That is not a new position I have adopted for the closing pages. I applied it to my parent doctrine's own successor hypothesis and ran it to failure. I applied it to the fog argument by pointing its standard at its author and publishing each claim at its honest rung rather than at the rung the title implied. This is the same instrument, third time, on a book with less evidence behind it than either of them.

What it means concretely: every claim here gets its highest defensible rung, not the rung of the strongest thing sitting next to it. That failure has a shape worth naming, because it is almost always unintentional. A proposal system producing proposals is not evidence of conversion. A deployed interface is not evidence of repeatable willingness to pay. And a founder's powerful knowledge system is not evidence that capability transfers across a bench.

The Ladder

The commercial status ladder

argued → offer designed → implemented → internally used → externally sold → client accepted → repeated → transferred → economically scaled

The two most commonly collapsed pairs: implemented → internally used (building it is not using it) and externally sold → client accepted (an invoice is not an outcome). "We've built it" and "somebody has paid for it" are four rungs apart.

Each claim below carries five things: current rung, supporting receipt, evidence ceiling, next promotion test, and kill or downgrade condition. The last two are what make this an instrument rather than a disclosure — a ceiling with no promotion test is just modesty, and modesty is not auditable.

The Ledger

Claim Rung Ceiling Kill / downgrade condition
The evolution mandate Argued and offer-designed No paying customer Three qualified conversations that cannot get past the fee architecture to the object
The access rail Implemented, internally used Deliberately hard-coded to my own wiki Arm B ≈ arm A
The AWS Marketplace Knowledge Appliance Implemented Unfinished Marketplace listing and seller registration. No customer If entitlement cannot be changed to preserve last-state operation indefinitely
The federated client-context product Argued Unproven. No client context has ever been joined Arm C ≈ arm B
Net AI beta for my kernel Argued; protocol specified Assessment not run across any model upgrade Negative across two upgrades → Part V demoted
The A/B/C and decision-delta trials Designed; consequences pre-committed Neither has been run The trials are the kill conditions

The evolution mandate itself

Argued, and designed as an offer. The receipt is this book, plus the offer ladder published in a prior volume with rung three fenced to a later one. The ceiling is that nobody has paid for it. The A$100k/month figure is a designed hypothesis that has never been validated by a transaction — and it was already published under that label before this book existed, which is the only reason I am entitled to keep using it rather than quietly upgrading it in the retelling.

The next promotion test is one paid cycle, with the six renewal questions answered from the client's own evidence rather than from my report. The kill condition is narrower and more likely: if the first three qualified conversations cannot get past the fee architecture to the object — if buyers keep hearing "retainer" no matter how the components are drawn — then the components are wrong, and the problem is the design rather than the market.

The access rail

Implemented and internally used. A working specimen.

The receipt is real and specific: a deployed ask-and-cite web application over my compiled IP wiki — deterministic index routing, full-page reads, citations validated against the run's read set, a visible execution trace, and an authenticated deployment.

The ceiling is equally specific. It is deliberately hard-coded to my own wiki, by an explicit decision recorded at the time: "askui work done, but it's a demo hard coded to the ip wiki." An audit found the coupling thin and localised, with a sequencing note not to generalise until the demo had done its sales job or a client deployment was scheduled. So it proves conversational access to my doctrine with real provenance. It does not prove coverage, interpretation quality, decision impact, or that anybody will pay for it.

The AWS Marketplace Knowledge Appliance

Implemented. A built productisation path.

The receipt is a two-day build in July 2026, and the components are worth listing because the specificity is the only evidence available: fail-closed privacy tokenisation in front of both storage and model calls; immutable tokenised source evidence with no raw-content staging table and no update path for an accepted document; bounded generation where a proposal batch is rejected unless source pages cover every supplied document; a vendor kernel with two authenticated routes and no listing, export or pagination surface, with deterministic caps, revocation checks, per-key rate limits and an audit line for allowed and denied requests alike; least-privilege buyer-account infrastructure with empty-egress-by-default security groups; zero infrastructure lint errors after sandbox-driven correction; and a dated deployment receipt for the kernel.

The ceiling is that the seller registration and the Marketplace listing were never completed. There is no customer. A dated deployment receipt is not a market, and a sandbox is not a tenant.

Its next promotion test is a completed listing and one buyer-account installation. Its kill condition is architectural rather than commercial: if the entitlement model cannot be changed to preserve last-state operation indefinitely — the critique in Chapter 8 — then the architecture cannot carry the commercial claim this book makes for it, and the claim goes rather than the caveat.

The federated client-context product

Argued. That is the whole rung.

There is an intended architecture, and it is coherent: the compiler operating on the buyer's side, reading their own compiled knowledge, taking bounded question-specific slices from my kernel. The ceiling is that no client context has ever been joined. Not "early". Not "in pilot". Unproven. Its promotion test is arm C of the trial, and its kill condition is arm C looking like arm B.

Net AI beta, and the two trials

Protocols specified. Nothing run.

The net-beta assessment has not been executed across any model upgrade for my own kernel. The A/B/C trial has not been run. The decision-delta trial has never been run with a client, for the simple reason that there has never been a client under this mandate.

What exists is the protocol and, in the A/B/C case, four commercial consequences committed in advance. That is not nothing — a pre-committed consequence is a real constraint on future behaviour — but it is not a result, and the distinction is the entire point of this chapter.

Why a Ledger Rather Than a Case Study

Because I do not have a case study. And manufacturing the shape of one — a composite, a de-identified "client", an illustrative engagement written in the past tense — would be exactly the behaviour this book spends fifteen chapters arguing against. The worked renewal in Chapter 14 is labelled a designed scenario inside its own box for that reason, not in a footnote where the label could be missed.

Key Insight

A specimen proves that a method can exist. It does not prove that it is your answer, or the industry's destiny.

The discipline here is inherited rather than invented, and both sources are mine, which is the point. The fee architecture this book builds on is explicitly derived from one specimen relationship and is "not a published rate card and not a claim that any named firm has signed it." The diagnostic ladder behind the trinket self-accusation in Chapter 13 is stated by its own author as "a design, not a surveyed industry standard", with its specimen labelled n=1 and its economic claims marked as hypotheses to check against your own proposal cost, win rate, gross margin and reuse — "not as a verdict about any named firm."

Inheriting those constraints is not modesty. It is the only way the doctrine stays usable by somebody else — because a reader can only calibrate how much weight to put on a design if they know what is holding it up.

And the prohibition I published against myself, on this exact figure, before writing this book: publishing designed numbers as validated "would be exactly the status inflation" my own standards chapter prohibits. It applies here unchanged, which is why Chapter 9 contains no number for any component.

The External Evidence, and Its Gaps

The research behind this book produced its own confidence report, and publishing that does more for the argument's credibility than another citation would.

Strong
  • The recurring / outcome shift in advisory. Primary and auditable — a listed firm's own quarterly revenue split, and a private firm's published revenue mix.
  • Conflict of interest. Statutory and government-authored rather than rhetorical.
  • Retrieval versus structure. Peer-reviewed, from named research groups, with methods you can read.
  • Planning horizons. A disclosed survey with a fielding window and a sample size.
Weak, or absent
  • AI-driven insourcing of advisory. No credible, dated, methodologically-disclosed survey found. What exists is procurement-driven and public-sector-weighted — a real and different claim.
  • Knowledge and methodology licensing. A near-total gap. Every search returned vendor marketing about productising consulting, with no data and no named researchers.
  • Retainer churn. Every circulating percentage traced to tooling blogs with no methodology. Dropped entirely.
  • AI vendor lock-in survey figures. The widely-repeated cluster traced to a page that, on inspection, does not contain the figures attributed to it and has no methodology section. Dropped, and treated as a fabrication-risk chain.

Chapter 9 is therefore argued from first principles and from two real licensing specimens — a divestment in which product licensing survived the severing of an advisory relationship, and a large firm publicly describing kernel compilation in a revenue release — rather than from a market study that does not exist.

The rule that produced all of it is worth stating as method rather than as apology: where a figure could not be traced to a primary or methodologically-disclosed source, it was left out rather than softened. Several passages in this book would have been punchier with a number. They are more useful without one, because a reader can check a mechanism and cannot check a statistic they are unable to follow home.

What Would Change My Mind

Four outcomes, each already appearing as a pre-commitment somewhere in the book. Collecting them here is what makes them auditable rather than scattered.

1. If arm B ≈ arm A

The capability licence is packaging, and Chapter 9's fee sheet loses a line.

2. If a client's internal team beats the shadow cycle

The navigation cycle is the wrong product for that client class. Chapter 9's eligibility rule tightens — rather than the doctrine softening, which is the usual move.

3. If net AI beta is negative across two model upgrades

The mandate rests entirely on frontier work and construction, and Part V is downgraded to a delivery convenience.

4. If the discipline makes clients more dependent rather than less

It has failed on its own terms, regardless of whether the mechanism is true. This one is checkable by anybody who buys the mandate — which is the entire reason to publish it.

What to Do With a Book Whose Author Has No Receipts

Take the specification, and run the instruments against your own relationships — where you do have receipts.

You have renewal history. You have escalation logs, whether or not you call them that. You have a record of what your clients stopped needing you for, sitting in the gap between what you did in year one and what you do now. And you have an answer to question six from Chapter 14 that you have never had to write down, which is the single cheapest diagnostic in this book.

Every design decision here is something to falsify locally. That is what a specimen is for: it shortens the distance between this might work and here is what I would have to observe for it to be working — and that distance, rather than the case study, is what a reader is actually short of.

There is a reciprocal obligation, and I would rather state it than imply it. If you run any of these instruments and they return something that breaks a claim in this book, that finding is worth more to me than a renewal. The stop conditions above are how I would be expected to act on it, and they are published precisely so that acting on it is not optional.

Three Boundaries, Named Once Each

A reader who has just been handed fourteen instruments needs to know which problems they do not solve, before trying to use them on the wrong thing.

  • The professional-services industry account — externalisation share, demand-side disintermediation, harvest / migrate / construct, the runway clock. A sibling book owns it. This one borrowed two paragraphs of context in Chapter 2 and nothing else.
  • The post-engagement perturbation review — turning finished engagements into higher-order strategic learning. A sibling owns the instrument. This book named only the promotion boundary the mandate must respect.
  • The founder-multiplier trap and its ablation test — whether better AI makes me more capable while leaving my firm no more transferable. A sibling owns it. This book asks whether the kernel earns its fee, which is a different question with a different apparatus.

The Only Form of the Argument I Am Entitled To

An evidence ledger published beside the doctrine is the only honest version of this book. It is also the most useful version, because a reader can see exactly which parts are load-bearing and which are still hypotheses wearing a diagram — and can therefore decide how much of their own commercial future to hang on each one.

The mandate is designed to be losable. So is the book.

REF
Sources & Evidence

References & Sources

The evidence base behind every claim — primary research, industry analysis, and technical specifications

Research Methodology

This ebook draws on primary research from standards bodies, independent research firms, enterprise technology vendors, and consulting firms. Statistics cited throughout have been cross-referenced against primary sources.

Frameworks and interpretive analysis developed by Scott Farrell / LeverageAI are listed separately below — these represent the practitioner lens through which external research is interpreted, and are not cited inline to avoid self-promotional appearance.

LeverageAI / Scott Farrell — Practitioner Frameworks

The interpretive frameworks, architectural patterns, and practitioner analysis in this ebook were developed through enterprise AI transformation consulting. The articles below are the underlying thinking behind those frameworks. They are listed here for transparency and further exploration — not cited inline, as this is the author's own analytical voice.

Scott Farrell — The Terminal Value Doctrine for Professional Services

The offer ladder, the labelled price table, and the anti-funnel rule, ch15 #de2950

https://leverageai.com.au/wp-content/media/articles/231-terminal-value-doctrine-professional-services.html

Scott Farrell — Fog Is a Race Between Two Clocks

The adviser inside the blast radius; the hand-off of the continuing-relationship question to a sibling volume, ch20 #5bd755

https://leverageai.com.au/wp-content/media/articles/232-fog-is-a-race-between-two-clocks.html

Scott Farrell — Make Copying Irrational

The anti-dependency test cuts both ways; transfer is creator protection against becoming a high-status bottleneck, ch5 #d74c16

https://leverageai.com.au/wp-content/media/articles/211-make-copying-irrational.html

Scott Farrell — The Cognition Dimension Ladder

The Permanent Fog: the discovery accelerator manufactures Fog as a side-effect of being good at its job; refresh the Question Ledger on a schedule, ch10 #ace73b

https://leverageai.com.au/wp-content/media/articles/62-cognition-dimension-ladder.html

Scott Farrell — Elastic Assurance

Standing Questions: purpose, scope, baseline and perturbing signals, evidence bar, cadence and human owner, retirement condition; question compounding; the panel decides and the AI supplies, ch6 #fe6c66

https://leverageai.com.au/wp-content/media/articles/136-elastic-assurance.html

Scott Farrell — Preparedness Is the Product

Three pools and the terms that make them honest; thirty Reserves and zero Accepts is diligence cosplay; ask which exposures should be accepted knowingly before capital moves, ch6 #b62430

https://leverageai.com.au/wp-content/media/articles/214-preparedness-is-the-product.html

Scott Farrell — Stand Pat

The stand-pat score in quiescence search: the current position's own evaluation entered as a candidate the loop may end on, and the check exception, ch2 #485387

https://leverageai.com.au/wp-content/media/articles/101-stand-pat.html

Scott Farrell — Buy Certainty First

Two lanes, full KPI sets: rate of do-not-build and defer dispositions, with relationship outcomes after, tracked as a lane KPI, ch12 #fd57f4

https://leverageai.com.au/wp-content/media/articles/204-buy-certainty-first.html

Scott Farrell — Forward-Deployed Practice OS

Installing the practice OS: what success looks like mid-install is not proof of transfer; engagement two is the only proof that matters, ch5 #9d6797

https://leverageai.com.au/wp-content/media/articles/167-forward-deployed-practice-os.html

Scott Farrell — AI-Native Successor Offer

Scarce-expert elasticity: shape rather than fabricated percentages; if the firm cannot measure the ratio, that is itself a finding, ch9 #75106a

https://leverageai.com.au/wp-content/media/articles/213-ai-native-successor-offer.html

Scott Farrell — The Fiduciary Agent

Double agents and shadow principals: agency-law vocabulary applied to AI; excellence and loyalty are orthogonal until the principal is specified; capability without a loyalty target is power looking for a gradient, ch4 #35cb62

https://leverageai.com.au/wp-content/media/articles/116-fiduciary-agent.html

Scott Farrell — The Engagement Auditor Is Not the Janitor

Why combining the roles fails: the machinery that compressed ambiguity cannot certify that none was lost; correlated assurance is one shared root in different coats, ch3 #e069d4

https://leverageai.com.au/wp-content/media/articles/174-the-engagement-auditor-is-not-the-janitor.html

Scott Farrell — The Model Is Not the Memory

The agent is replaceable and the memory is the asset; rent the model, own the map; the enterprise vendor-dependency survey; calcified lore and the failure gallery, ch14 #9b8545

https://leverageai.com.au/wp-content/media/articles/68-the-model-is-not-the-memory.html

Scott Farrell — Publishing Is an Active Sensor

Null response and the exploration budget: silence is informative only with a stated expectation and exposure class; a sensor that only samples where it expects signal becomes a mirror with a content calendar, ch8 #bd018c

https://leverageai.com.au/wp-content/media/articles/158-publishing-is-an-active-sensor.html

Scott Farrell — Conversation Is the REPL

Not a thousand songs in your pocket: an archive of finished objects versus a graph of claims and edges that recombines on every question; finishing the thought while it is still warm, ch2 #187e72

https://leverageai.com.au/wp-content/media/articles/132-the-conversation-is-the-repl.html

Scott Farrell — Capture Was Never the Bottleneck

The interface ladder — Ask, Browse, Ambient — in shipping order, with the cost to the user falling down the ladder; repeated questions as cache misses and the demand-side map, ch10 #9375a5

https://leverageai.com.au/wp-content/media/articles/84-capture-was-never-the-bottleneck.html

Scott Farrell — Route-Invariant Grounding

Four quantities of route-invariant grounding: path variance is not a failure mode by itself; evidence invariance is the load-bearing property; answer invariance is necessary and insufficient, ch3 #a44a89

https://leverageai.com.au/wp-content/media/articles/182-route-invariant-grounding.html

Scott Farrell — Agent Addressability

Pixels are not a delegation surface: human UI versus the five-element delegation surface; "we have an API" is not the answer; impersonating fingers is not addressability, ch3 #f75bd2

https://leverageai.com.au/wp-content/media/articles/111-agent-addressability.html

Scott Farrell — The Moat Is the Memory

What to build if you only take one thing: instrument the deposit layer, close the feedback loops, and check which of the three guards — alien-signal lane, suppression audit, replayable trace — are missing; after a year someone can steal the code but not the year, ch8 #d02b4e

https://leverageai.com.au/wp-content/media/articles/149-the-moat-is-the-memory.html

Scott Farrell — Succession Product

The ingestion boundary: claims, relationships and pointers as the model-facing layer; the pass condition for representation — if every answer requires the full thread in context, you have not compiled, you have wrapped search, ch4 #5de8c2

https://leverageai.com.au/wp-content/media/articles/216-succession-product.html

Scott Farrell — Orientation Capital

The room you could not afford before: corporate value is useful joins rather than headcount multipliers; no invented detection rates; what leaders should demand in demos, ch7 #5b6bba

https://leverageai.com.au/wp-content/media/articles/161-orientation-capital.html

Scott Farrell — Five Postures of an AI-Native Consultancy

A firm diagnostic and an organisational design — a design, not a surveyed industry standard; the specimen labelled n=1; economic claims to be checked against your own numbers rather than treated as a verdict, ch2 #781847

https://leverageai.com.au/wp-content/media/articles/210-five-postures-ai-native-consultancy.html

Major Consulting Firms

Oliver Wyman Forum — CEO Agenda 2026: How CEOs Navigate Geopolitics, Trade, Technology and People [1]

50% of CEO time is dedicated to planning for less than one year, up from 43% in 2025; survey of 415 CEOs fielded 12 January to 13 March 2026

https://www.oliverwymanforum.com/ceo-agenda/how-ceos-navigate-geopolitics-trade-technology-people.html

Adi Ignatius, Harvard Business Review, January–February 2026 — "We Want to Make Ourselves Better" — The HBR Interview with Bob Sternfels [5]

"And to set McKinsey up for the AI era, he and his partners are driving an organizational transformation to focus less on traditional consulting services and more on delivering outcomes." (Quoted from the free editorial framing; interview body paywalled.)

https://hbr.org/2026/01/we-want-to-make-ourselves-better

Industry Analysis & Vendor Research

Accenture — Accenture Reports Third-Quarter Fiscal 2026 Results [2]

Revenues by Type of Work: Consulting $9.33 billion; Managed Services $9.39 billion

https://newsroom.accenture.com/content/3qfy26-earnings/accenture-reports-third-quarter-fiscal-2026-results.pdf

Boston Consulting Group (PR Newswire) — BCG Reports $14.4 Billion in Revenue, Marking 22nd Consecutive Year of Growth [3]

AI- and tech-focused services now represent over 40% of BCG's total revenue, driven by 25% year-over-year growth in AI services

https://www.prnewswire.com/news-releases/bcg-reports-14-4-billion-in-revenue-marking-22nd-consecutive-year-of-growth-302751073.html

Accenture — Accenture Reports First-Quarter Fiscal 2026 Results [4]

Consulting new bookings were $9.88 billion; Managed Services new bookings were $11.06 billion; Advanced AI new bookings of $2.2 billion

https://newsroom.accenture.com/content/1qfy26-earnings/accenture-reports-first-quarter-fiscal-2026-results.pdf

Consultancy.uk — Consultants point to 'AI-fatigue', and organisational overhauls in their predictions for 2026 [6]

"2025 was a difficult year for the consulting sector. Depending on the definition of the market, consulting in the UK either saw flat growth, or negative growth – and its worst performance since the lockdown period in either case."

https://www.consultancy.uk/news/42610/consultants-point-to-ai-fatigue-and-organisational-overhauls-in-their-predictions-for-2026

UK National Audit Office, 21 November 2025 — Government lacks a clear picture on how much it spends on consultants [9]

"Consultants should only be used where they represent best value for money and not to replace capability required inside the civil service."

https://www.nao.org.uk/press-releases/government-lacks-a-clear-picture-on-how-much-it-spends-on-consultants/

Department of Finance (Australia), August 2025, quoting the Switkowski Review — Examination of the ethical soundness of PricewaterhouseCoopers Australia [11]

"There has not been, and does not yet appear to be, an overarching framework providing clear instructions to partners and staff as to how to identify or manage the various types of actual, potential, or perceived conflicts. There is also insufficient guidance for how to differentiate between various types of conflicts of interest."

https://www.finance.gov.au/sites/default/files/2025-08/examination-of-pwc-australias-ethical-soundness.pdf

The Linux Foundation, 9 April 2026 — A2A Protocol Surpasses 150 Organizations, Lands in Major Cloud Platforms, and Sees Enterprise Production Use in First Year [13]

"In less than a year, A2A has moved from initial release to a production-ready open standard for seamless agent-to-agent communication"; "A2A provides a common semantic model and version negotiation that standardize how agents discover, communicate, and transact with each other, without being locked into a single vendor's ecosystem."

https://www.linuxfoundation.org/press/a2a-protocol-surpasses-150-organizations-lands-in-major-cloud-platforms-and-sees-enterprise-production-use-in-first-year

David Soria Parra (Lead Maintainer), Model Context Protocol Blog, 9 March 2026 — The 2026 MCP Roadmap [14]

"Over the past year MCP has moved well past its origins as a way to wire up local tools. It now runs in production at companies large and small, powers agent workflows, and is shaped by a growing community through Working Groups"; "Enterprises are deploying MCP and running into a predictable set of problems: audit trails, SSO-integrated auth, gateway behavior, and configuration portability."

https://blog.modelcontextprotocol.io/posts/2026-mcp-roadmap/

Andy Bayiates, Deloitte Insights, 24 April 2026, drawing on Deloitte's State of AI in the Enterprise (January 2026), n = 3,235 IT and business leaders across 24 countries — Business and IT leaders report AI agents are scaling faster than their guardrails [16]

"By 2027, 74% of respondents expect their companies to be using AI agents at least 'moderately'"; only 21% of enterprises report having mature governance structures for agentic AI.

https://www.deloitte.com/us/en/insights/topics/emerging-technologies/ai-agents-scaling-faster.html

Primary Research & Standards Bodies

Isin Guler, 24 September 2006; forthcoming in Advances in Strategic Management, 2007 — An Empirical Examination of Management of Real Options in the U.S. Venture Capital Industry [7]

"Signals of a company's progress, such as the number of its patents, are significant predictors of VC investment practices in the case of successful companies, but not in the case of unsuccessful companies… signals of failure are more ambiguous and complex; and firm-level differences are more pronounced in management of unsuccessful options."

https://isinguler.web.unc.edu/wp-content/uploads/sites/15730/2018/04/Guler-AISM.pdf

Association of National Advertisers & American Association of Advertising Agencies, 30 April 2025 — New ANA and 4As Report Reveals Client-Agency Relationship Tenure Has Doubled Since 2016 [8]

"Clients without mandatory review periods (60% of respondents) have significantly longer relationships (8.1 years) than those with frequent reviews (as low as 3.8 years)."

https://www.ana.net/content/show/id/pr-2025-04-tenure

Legal Information Institute, Cornell Law School — 15 U.S. Code §78j-1 — Audit requirements, subsection (g) Prohibited activities (Sarbanes-Oxley §201) [10]

"(g) Prohibited activities — Except as provided in subsection (h), it shall be unlawful for a registered public accounting firm … that performs for any issuer any audit … to provide to that issuer, contemporaneously with the audit, any non-audit service, including— (1) bookkeeping…; (2) financial information systems design and implementation; (3) appraisal or valuation services…"

https://www.law.cornell.edu/uscode/text/15/78j-1

Greenberg Traurig LLP, September 2025 (analysis of Regulation (EU) 2023/2854) — Cloud Switching Under the EU Data Act: Implications for IaaS, PaaS, and SaaS Providers [12]

"switching charges (fees for executing the switching request) are only permitted under narrow conditions and will be prohibited entirely from Jan. 12, 2027"; "the Data Act permits to provide for proportionate early termination penalties or fees"

https://www.gtlaw.com/en/insights/2025/9/cloud-switching-under-the-eu-data-act

Richard Kang, Yudho Diponegoro, arXiv:2606.31498, 30 June 2026 — Governance Gaps in Agent Interoperability Protocols: What MCP, A2A, and ACP Cannot Express [15]

"The resulting gap matrix reveals that voting and dissent preservation are universally absent across all five protocols, deliberation is absent or at most partial, and no protocol encodes the full set of primitives required for governed agent communities… The analysis establishes that agent community governance constitutes a missing architectural layer above current interoperability standards, not a missing feature within them."

https://arxiv.org/abs/2606.31498

Prafulla Kumar Choubey, Xiangyu Peng, Shilpa Bhagavath, Kung-Hsiang Huang, Caiming Xiong, Chien-Sheng Wu (Salesforce AI Research), arXiv:2506.23139, 29 June 2025 — Benchmarking Deep Search over Heterogeneous Enterprise Data [17]

"We release our benchmark with both answerable and unanswerable queries, and retrieval pool of 39,190 enterprise artifacts… Our experiments reveal that even the best-performing agentic RAG methods achieve an average performance score of 32.96 on our benchmark. With further analysis, we highlight retrieval as the main bottleneck."

https://arxiv.org/abs/2506.23139

Zhentao Xu, Mark Jerome Cruz, Matthew Guevara, Tie Wang, Manasi Deshpande, Xiaofeng Wang, Zheng Li (LinkedIn), arXiv:2404.17723, 26 April 2024 — Retrieval-Augmented Generation with Knowledge Graphs for Customer Service Question Answering [18]

The method achieved a 77.6% improvement in Mean Reciprocal Rank over baseline approaches; after roughly six months in LinkedIn's customer service operations it "reduced median per-issue resolution time by 28.6%."

https://arxiv.org/abs/2404.17723

About This Reference List

Compiled August 2026. All URLs verified at time of compilation. Regulatory documents and standards specifications are subject to revision — check primary sources for the most current versions.

Some links to academic papers and vendor research may require free registration. Government and standards body publications are freely accessible.