Fixed Price Is Underwriting
An extension of AI-Native Service Architecture

Fixed Price
Is Underwriting

Earn the square by owning the variance

AI did not make fixed price easy. It changed what the price is attached to — and that turned your firm into an underwriter that has never written a policy.

Scott Farrell · LeverageAI · August 2026

What a practice lead leaves with

  • ✓ A four-class variance schedule you can run on a completed engagement this afternoon — including the calls two reasonable people would argue about
  • ✓ Four cheap forms of an acceptance oracle the producer does not control, and the ratio that prices the third one
  • ✓ An accumulation register that makes correlated exposure visible in ninety minutes — and the arithmetic showing why no reserve covers it
  • ✓ The eight records that separate a product from subsidised heroics, six of which are a column in a file you already keep

TL;DR

  • You did not get better at estimating. The object changed. The old fixed price was attached to forecast labour, which is why it needed an estimate. The new one is attached to measured, machine-absorbable complexity — and the moment a price attaches to a measured exposure, six underwriting jobs stop being prudent and become structurally necessary. Most firms are doing three of them.
  • "Scope" is one word doing four jobs. Provider-owned interior variation, metered residual scarcity, boundary mutation, and external or correlated tails — four owners, four consequences, four speeds. Silent absorption of the first class is not generosity; it is unpriced risk capital, contributed by your firm and recorded nowhere.
  • Your own flywheel correlates your book. Shared kernel, shared parsers, one pinned model: the machinery that makes engagement twenty cheaper than engagement two is the machinery that makes all twenty fail together. Pooling requires independence, and independence used to be free. Version pinning, evals, rollback and shared-component regression are not hygiene — they are commercial underwriting controls with a line in the price.
  • The compounding asset is the loss history of complexity. Not the workflow, not the contract template — competitors rent the same models and copy the offer by Tuesday. What they cannot copy is the record of which census variables predicted effort, which bands stayed profitable, and which client behaviours consumed reserve. Engagement twenty is structurally safer than engagement two only if that record lives in the firm.
01
Part I: The Business You Are Actually In

The Engagement That Worked for the Wrong Reason

It came in on the number, on the date, with no change requests and a client who wanted to talk about the next one. Everybody in the room drew the same conclusion, and the conclusion was wrong.

You know the meeting. The engagement closed cleanly. Fixed fee, fixed clock, nothing escalated, margin held within a percentage point of the model, and the client asked — unprompted — what else we could do on that basis. Somebody said the thing everybody was thinking, which was that we had finally got good at this.

What they meant, specifically, was that we had got good at estimating. The machine had eaten enough of the delivery mess that the number was safe. Do more of that, quote with more confidence, and the old argument about scope goes away.

I have been in that meeting on both sides of the table, and I want to say plainly what I now think happened, because getting it wrong is expensive and the expense arrives late — somewhere around engagement eight, when the pattern has had enough time to become the business model.

What the room concluded

  • We estimated better than we used to.
  • AI absorbed the variance that historically broke fixed prices.
  • Therefore fixed price is now safe for work like this.
  • Therefore: quote more of it, with more confidence.

What actually happened

  • The estimate did not get better.
  • The object the price was attached to changed.
  • Which quietly enrolled the firm in a different business.
  • That business has six jobs. We were doing two.

The object changed

The old fixed price was attached to forecast labour. That is why it needed an estimate, and why the estimate carried the whole risk. Somebody sat down with incomplete information and social pressure and produced a number of days. Everything after that — the plan, the margin, the arguments in month four — was downstream of whether that person had been lucky.

The new fixed price is attached to something else: measured, machine-absorbable complexity. A census of the input surface before anybody quotes. A band the estate falls into. A defined quantity of consequential human judgement inside the price. A reserve for named classes of surprise. None of that is a better guess. It is not a guess at all.

That distinction sounds like semantics until you notice what you can do with each of the two things.

An estimate has exactly two states. It is right, or it is wrong, and you find out which one at the end. You can pad it, you can defend it, and you can be embarrassed by it. That is the complete list of available operations.

A measured exposure has as many states as your machinery can distinguish. You can classify it. You can put it in a band with other exposures that behave like it. You can hold a named reserve against the parts of it that surprise you. You can exclude the parts you cannot observe. You can price the parts you can. And — this is the operation nobody has on an estimate — you can refuse it.

An estimate is something you defend. An exposure is something you can refuse.

Which makes six jobs structurally necessary

Here is the claim this book spends its length paying off. Once the price attaches to measured exposure rather than to forecast effort, a specific set of activities stops being prudent and starts being required. Inspect the exposure before accepting it. Sort exposures into classes that behave alike. Price the fee as the cost of what you expect plus the cost of what you are holding. Hold an explicit allowance for named surprises. Write down what you will not cover. Refuse what you cannot bound. And keep a record of what actually happened, so the next price is better than this one.

That is not a maturity model and it is not a set of good habits. It is the operating requirement of the object you have just sold. A firm that sells a measured exposure and runs none of that machinery has not been careless — it has been doing a different job than the one it agreed to.

The mapping is Chapter 3, and it fits on one page. You can audit your own firm against it in an afternoon, and I would suggest doing that before you read the rest, because the finding is more interesting when it is yours.

The state this book is written against

There is a condition I want to name early, because everything else in this book is an attempt to get a firm out of it.

The dangerous middle state

A firm carrying the commercial form of underwriting — the fixed number, the productised proposal, the confident boundary, the language of bands and inclusions — and none of the machinery underneath it.

Later in this book I give it a name: underwriting-washing. For now the important thing is that it is not a halfway house on the road to the real thing. It is a distinct and worse position than the one it replaced.

Understand why it is worse than honest time-and-materials, because the instinct is to file it as "not quite there yet".

Time-and-materials is a crude instrument and it is not a dishonest one. It puts the variance where the client can see it. Both parties know exactly what they are in: the buyer is funding discovery, the supplier is selling access to people, and when reality turns out to be messier than anyone thought, the mess shows up on an invoice where somebody can argue about it.

The form-without-machinery hides retained variance behind a number that implies it was measured. From outside, it looks like productisation. From inside, it is a naked risk position that nobody has sized. And the firm running it is not primarily deceiving the client — it is deceiving itself, because a fixed price with no records produces no signal at all. Ask that firm how much risk capital it contributed to client projects last quarter and there is no answer available, not because it is confidential but because nobody wrote any of it down.

Two clocks on the failure

When that position fails, it fails on one of two clocks, and they behave completely differently.

The slow clock. Unpriced absorption, one engagement at a time. The team eats a problem here, a delay there, a rework cycle nobody raised. Every individual decision is defensible; most of them are the right call in the room. It never appears on a profit-and-loss line because it is composed entirely of things that were never recorded as events. It is detected, if at all, as unexplained margin drift, and it gets attributed to a difficult year.

The fast clock. A single shared cause — one model upgrade, one parser defect, one connector deprecation — damaging every live engagement in the same week. No per-engagement reserve covers it, because every reserve in the book was sized for events that arrive independently. This one arrives once and draws all of them.

The second clock is newer, it is a direct consequence of the machinery that makes AI-native delivery work at all, and it gets almost no attention. Part III is about nothing else.

What this book inherits, and will not re-teach

Two pieces of my own published work do heavy lifting here, and I am going to reference them rather than re-derive them. That is a deliberate choice about how these books relate to each other, and it is worth stating once so you know what to expect for the rest of this one.

The first is the Square: a stable commercial perimeter around an adaptive, machine-scale production interior. Eight fields on the outside, deliberately rigid. The inside deliberately fluid — search, decomposition, generation, retries, tool choice, regeneration, internal replanning — and the critical property that the interior may be horrendously irregular without any of that irregularity reaching the buyer.

The second is typed uncertainty: the discipline that turns every unresolved item into a bounded terminal state with a precise assertion and a next consumer, so that unknowns become deliverables rather than unbounded labour. And the clause inside it that carries more weight than the rest combined — not observed does not mean does not exist.

Both are published, both are cited, and both appear in this book as organs rather than as chapters. Re-teaching them would spend exactly the pages this book needs for its own contribution, which is narrower and sits underneath both of them.

What this book adds

The risk economics underneath the perimeter. Not where to draw the boundary — that is settled — but what has to be true about your measurement, your classes, your reserve, your exclusions and your records for the boundary to be safe to draw at all.

Two books that stopped exactly here

I am not the only one who noticed this gap; I left it twice in the last week myself.

The Terminal Value Doctrine: Professional Services reached the successor offer, described it as a commercial unit that "knowingly prices the responsibility and variance the supplier retains, with the insurer's vocabulary … doing the work the blended rate used to fake" — and then said, in the same paragraph, that the underwriting machinery beneath that sentence belonged to a later book.

Fog Is a Race Between Two Clocks hit the same seam from the pricing side and named it as an adjacency it would not attempt: "pricing a bounded commitment when delivery variance is real is its own discipline, with its own instruments for underwriting the variance rather than hoping it averages out."

This is that book.

What you should be able to do afterwards

Four things, and they are the test I would like this book judged on. Not whether the argument is interesting — whether you can act on it without me in the room.

  • 1Classify every variance in a live engagement into one of four classes — including the two or three that are genuinely arguable, using rules you wrote down before the argument started.
  • 2Design an acceptance oracle you do not control, using at least one of four forms that cost an afternoon rather than a second delivery team.
  • 3Build the register that makes your correlated exposure visible — every shared component, everything that can change it without asking you, and how long it would take you to notice.
  • 4Run the eight records that tell you whether your square is a product or whether it is being subsidised by heroics — which look identical from outside and diverge exactly when you try to scale.

One correction before we go on, because it is my own and it matters for the vocabulary. When I first talked about this shape out loud, I described AI as the hedge that let you fixed-price amorphous work. The instinct was right and the word was wrong, and I said so in print afterwards: a financial hedge offsets risk, and AI offers no such guarantee. What it does is radically reduce the marginal cost of responding to many forms of variance. That makes it elastic capacity — a variance absorber — which is a different thing, with different consequences, and the difference is most of what follows.

Which raises the obvious objection, and it is a good one. If fixed price is really an underwriting problem, why is any of this new? Firms have been quoting fixed prices for a very long time.

They have. Including mine, in 2004.

02
Part I: The Business You Are Actually In

AI Did Not Invent Fixed Price

A proposal from 2004, a regulation older than that, and a sentence from the largest strategy firm in the world — all saying the same thing about what actually changed.

In 2004 my earlier consultancy put a proposal in front of a client that offered them a choice. They could buy the work on time and materials, or they could buy it fixed price. Both routes were costed, both were on the same page, and the client picked.

What interests me now is not the choice. It is what the fixed-price route required before anybody would write a number on it.

It required a paid requirements analysis phase, followed by a paid statement-of-work phase, with both credited against the later build. It required formal change control with a documented form for every request. And it required continuous comparison of contract value against actual cost for the life of the project — not a review at the end, a standing comparison.

Read that structure back through the vocabulary of this book, one line at a time, and it stops looking like historical trivia.

What the 2004 proposal already had

  • The census. A paid requirements phase is the purchase of the right to know what you are pricing.
  • The commercial bridge. Crediting it against the build is how you sell a census to a buyer who thinks they are buying a project.
  • Loss-ratio monitoring. Contract value against actual cost, continuously, at n=1.
  • A typed reopening. Formal change control with a documented form and an approval path.

What it could not reach

  • The census was human. Every field you wanted measured cost a consultant-week.
  • So it was narrow. You measured what you could afford to measure, not what drove cost.
  • So the territory you could honestly fix-price was small — and everything outside it went on the clock.
  • And the classes never accumulated. One project's actuals did not price the next one, because nobody could afford to keep them comparable.
The paid requirements phase was not a discovery gate. It was the purchase of the right to know what we were pricing.

Which is the correction to the popular story, and it is more useful than it looks. AI did not make fixed price possible. It made the census cheap — and the census is what fixed price had always been waiting on. Every serious fixed-price practitioner I have ever met already knew they needed to measure before they promised. What they could not do was afford it at any useful breadth.

The regulation has said this for decades

If my own record is a poor witness in its own defence — and it is — the procurement law of the largest buyer on earth is not.

United States federal acquisition regulation defines a firm-fixed-price contract as one that "places upon the contractor maximum risk and full responsibility for all costs and resulting profit or loss." 1

Sit with that for a moment, because it is not how anybody in professional services describes a fixed price to a client. Buyers experience fixed price as certainty. The regulation defines it as a transfer. Same instrument, and the two parties are describing opposite properties of it — which is fine, and is exactly how insurance works, and is the reason the fee is not simply a discounted forecast of effort.

Then the eligibility test, which is the part I would put on a wall. Fixed price is suitable when "Performance uncertainties can be identified and reasonable estimates of their cost impact can be made, and the contractor is willing to accept a firm fixed price representing assumption of the risks involved." 1

Three tests, written into procurement law

Identifiable. You can name the uncertainties. That is a census.

Estimable. You can put a cost impact on each of them. That is a band.

Willing. The supplier accepts the residue, knowingly. That is an underwriting decision, and it is the only one of the three that is a choice.

There is one more detail in the same part of the regulation worth noticing, because it anticipates an instrument this book spends two chapters on. Even the regulation's own "fixed" price has a typed reopening built into it: a fixed-price contract with economic price adjustment "provides for upward and downward revision of the stated contract price upon the occurrence of specified contingencies." 1 Named triggers, agreed in advance, rather than an argument at the point of pain. The instinct that a fixed price should reopen only against pre-declared events is not ours and it is not new.

The incumbent says it too

The most useful sentence I found while researching this book came from the firm with the most to gain from the opposite claim. Kate Smaje, McKinsey's global leader of technology and AI: "Outcomes-based pricing didn't start because of AI, but the type of work AI transformation demands suits it." 2

The same reporting puts roughly a quarter of the firm's global fees through performance-based arrangements, developed over several years of multi-year transformation work rather than arriving with the current model generation. I use the smaller quote rather than the bigger number deliberately: the quarter establishes that a shift is happening; the sentence establishes what caused it, and only one of those is contested.

Why now — the pressure, not the opportunity

Everything above argues that this is not new. Here is why it is nonetheless urgent, and the reason is not technological.

In May 2026 a United States executive order made fixed price the procurement default. Any non-fixed-price contract — cost-reimbursement, time-and-materials, labour-hour — "must be justified in writing by the contracting officer to the agency head." 3 The order names the number that motivated it: approximately $120 billion obligated on cost-reimbursement consulting contracts in a single fiscal year.

Read that as an inversion of the burden of proof. It used to be that a supplier proposing a fixed price had to explain why they were confident. Now, in the largest procurement market in the world, a supplier proposing anything else has to explain why they are not.

The consequence for the supply side is uncomfortable and worth stating flatly: the market will get fixed prices whether or not the firms quoting them have the machinery to hold them. Commercial pressure does not wait for capability. It produces the form and lets the substance catch up, or not.

One nuance keeps that observation honest. The same order concedes legitimate exceptions — research, and the pre-production developmental phase of major systems acquisition. Even a fixed-price mandate carries an eligibility boundary, written by the buyer, in advance. That is an external precedent for the argument I make in Chapter 9, arriving from the direction you would least expect it.

The counterweight, placed here rather than hidden

A book arguing for a fee shape has an obligation to say what the evidence on that fee shape actually looks like, and it does not flatter the argument.

Work by Jørgensen, Mohagheghi and Grimstad, published in the International Journal of Project Management, found the use of fixed-price contracts connected with a higher risk of project failure than time-and-materials in software projects. 4 A widely repeated pair of success-rate percentages circulates alongside that finding; I could not confirm them against the paper itself, so they do not appear in this book. The direction is well-sourced. The numbers are not mine to lend you.

Now take the ground, because that finding is my antagonist rather than my problem.

Fixed price as normally practised — quoted against an unmeasured estate, with no census, no bands, no operational exclusions and no typed reserve — deserves that record entirely. It is bravado with an invoice schedule. It is the exact thing I described in Buy Certainty First as gambling with better stationery: a risk transfer accepted without a theory of the risk.

That is the only time I will use that line in this book. It is a good line and I have leaned on it enough; repeating it would be the rhetorical equivalent of silent absorption.

The precise delta

So what did AI change? Three things, and being vague about this is exactly how "AI makes fixed price easy" got into circulation.

Not the fee shape

It predates the current model generation by decades in regulation and by twenty-two years in my own filing cabinet.

The cost structure

Production got cheaper. Verification got dearer. Both directions are real and only the first one gets discussed, which is why Chapter 13 exists.

The correlation structure

Everybody is running the same small number of models on the same shared layer. That changes what your portfolio does when something breaks, and it is the subject of Part III.

Which leaves one question standing, and it is the one the rest of Part I answers. If a fixed price is a risk transfer — and it has been, in writing, since long before any of us — then the discipline for holding transferred risk is the thing a firm actually needs.

That discipline exists. It is written down. It is roughly three hundred years old, and it belongs to a profession that is not ours.

03
Part I: The Business You Are Actually In

The Underwriter's Table

Eight rows. Each one is a job that exists because the commercial promise cannot be kept without it — and most firms are doing three of them.

The best sentence I found while writing this book is not mine. It was written by actuaries, for actuaries, in a document about risk classification, and it states my thesis better than I have ever managed to:

Though an individual exchanges the uncertainty of occurrence, timing and magnitude of a particular event for the certainty of a fixed price, that exchange in no way makes the uncertain known. Nor need it. The insurance program assuming the financial uncertainty is not able to fix the occurrence or, often, the magnitude of a specific risk merely because it assumes that risk. But it should find a way of establishing a fair price for assuming it. — American Academy of Actuaries, Committee on Risk Classification5

Read it as a description of your last fixed-price engagement, because it is one.

The client's certainty and your certainty are different objects. The client gets a number that will not move. You get a distribution. Nothing about signing the contract collapsed that distribution — it did not become narrower, or better understood, or less likely to surprise you. All that happened is that you agreed to stand in front of it.

So the fee is not a discounted forecast of effort. The fee is the price of converting one of those objects into the other, and if your pricing conversation does not have a term in it for that conversion, you are performing the service for free and calling the remainder margin.

I got to the same place from the other direction, and less elegantly: a fixed price for an uncertain activity is a risk transfer from client to supplier, and that sentence should make a serious firm nervous.

Selling certainty does not require having certainty. It requires a defensible method for pricing the uncertainty you agreed to hold.

The table

Here is the mapping. It is not decoration and it is not a metaphor for being careful. Each row on the right is a job that exists because the row on the left cannot be done without it.

Mapping between AI-native engagement machinery and underwriting jobs
Your AI-native engagement The underwriting job
Preflight censusRisk assessment
Eligibility rules and bandsRisk classification
Fixed feePremium for expected cost and retained risk
Included human dispositionsCovered consumption
Typed reserveExplicit loss allowance
Unsupported-source rulesCoverage exclusions
Re-band or declineReprice or refuse the risk
Engagement actualsClaims and loss-ratio data

Tables like that are easy to nod at and hard to use, so walk it. For each row: what the job is for, and what specifically breaks when it is missing. The second half is the part worth reading.

1. Census — risk assessment

Inspect the exposure before accepting it, with something that measures rather than estimates. Without it, you have priced a picture rather than an estate, and every discovery during delivery is a dispute, because there is no recorded value for the discovery to be measured against.

2. Bands — risk classes

Group engagements whose expected cost genuinely behaves alike, so that experience from one can price another. Without it, you get adverse selection — and the reason that is dangerous rather than merely suboptimal is that it is invisible in any individual deal. Every engagement looks defensible; the composition of the book is the problem. Chapter 6.

3. Fee — premium

Expected cost, plus the cost of settling surprise, plus the cost of the capital standing behind the promise. Without it, you are pricing delivery and donating the risk-bearing, which is the most expensive service most firms give away. Chapter 5.

4. Included dispositions — covered consumption

A named quantity of the scarce resource sits inside the price, and the rest does not. Without it, the fixed price means unlimited senior anxiety for one number — and the meter that everybody in delivery can feel is the one nobody in commercial can see.

5. Typed reserve — loss allowance

A named absorption layer for pre-declared classes of surprise, with a counter on it. Without it, the margin silently does the reserve's job. Which works, right up until it does not, and by then nobody can name the moment it stopped working or the events that consumed it.

6. Unsupported-source rules — exclusions

Operational limits stated at signature, and priced. Without them, "just have a look at this one as well" is how a fixed price dies — not in a single decision, but in eleven small ones nobody logged. Chapter 8.

7. Re-band or decline — reprice or refuse

The ability to say not at that price or not at all. Without it, the band stops predicting anything, because everything gets forced into it — and a product that cannot say no has quietly become a project with a product's price. Chapter 9.

8. Engagement actuals — claims data

The record of what happened, in a form that changes the next price. Without it, engagement twenty is exactly as risky as engagement two, and every hard-won lesson lives in the head of whoever was on the engagement. Part VI.

Which of the eight do you actually do?

Not aspire to. Not have a slide about. Do, with an artefact behind it that somebody outside the delivery team could find without asking.

I will state my prediction as a prediction, so you can falsify it on yourself in an hour. Most firms do rows one, two and three. Those are the visible, saleable parts — they are what makes a proposal look productised, and they are what a buyer asks about. The neglect runs in a reliable order: exclusions, then decline, then loss history.

And the order is not accidental. Exclusions feel adversarial at the moment of sale, so they get written last and softened. Decline feels like leaving money on the table, and nobody is promoted for it. Loss history has no customer at all — no client asks for it, no partner is measured on it, and it produces nothing this quarter. So it is always the thing that will be started properly next year.

The afternoon audit

Three columns, eight rows. Fill it in for one offer.

The job The artefact that proves we do it Who owns it
Risk assessment
Risk classification
Premium
Covered consumption
Loss allowance
Exclusions
Reprice or refuse
Claims and loss-ratio data

The blank rows are the finding. Not the filled ones.

The organ that separates underwriting from risk management

An obvious objection at this point is that all of this is just good project risk management wearing a costume, and if that were true the costume would be a waste of everybody's time.

The distinguishing organ is decline.

A risk register manages a project you have already agreed to do. It identifies, it rates, it assigns owners, it tracks mitigation. It has never in the history of the discipline refused anything, because by the time it exists the refusal decision is months in the past. Underwriting sits before that: whether to write the risk at all, at what price, in which class, with what reserve, and what will not be covered.

If there is no artefact in your firm that can say no before a proposal is written, you do not have an underwriting function. You have a sales function with a legal review. That is a whole chapter — Chapter 9 — but the flag belongs here, because it is the row that most sharply distinguishes this frame from the version of it that would be a waste of your time.

Borrow the discipline, not the licence

Now the boundary, and it belongs at the moment the analogy is introduced rather than in a disclaimer at the back.

This is a mechanism transfer. It is not a regulated product, and I am not arguing that a professional-services firm should sell insurance. I have already written the chapter on why that boundary is expensive to cross: compensation for a customer's project loss is a risk-transfer instrument, and depending on how it is structured, arrangements of that kind can amount to financial products or contracts of insurance under Australian law. The product design consequence is blunt — do not brand the offer as a policy, and get counsel before any indemnity.

The line I used there holds here without modification: customers who need insurance products should buy insurance products from parties authorised to sell them.

What the analogy carries — and what it does not

It carries: the questions an underwriter must answer before accepting an exposure, in the order they have to be answered, with the vocabulary that makes each one specific.

It does not carry: statutory practice, licensing posture, reserving standards you are obliged to meet, or any permission whatsoever to write an indemnity against somebody else's loss.

Borrow the discipline, not the licence.

The market got here first

One reason to believe this is a mechanism transfer rather than analogical drift: reinsurers have already built the product, without any reference to consulting.

Munich Re sells cover "for AI systems designed to address a wide area of AI-related risks for AI providers and corporate adopters caused by AI performance errors, including: Contractual liabilities; Own damages/financial losses; and Legal Liabilities." 6 The same product backs performance warranties that let a supplier indemnify its clients for losses directly related to AI errors. Armilla places affirmative AI liability cover at Lloyd's as a coverholder. 7

Be careful about what that proves, because it is tempting to over-read it. Their rating methodology is not public. I do not know how they price the exposure, and I am not going to imply that I do. What their existence establishes is the shape: people whose entire business is pricing variance have looked at AI performance variance and concluded it can be priced.

Which retires one objection permanently. "AI variance is unpriceable" is no longer a defensible position. The only live question is whether you can price yours — and that question has an answer, and the answer is the rest of this book.

Where the eight rows go from here

  • Rows 3, 4 and 5 — premium, covered consumption, reserve — are Part II, along with the classification problem in row 2.
  • Rows 6 and 7 — exclusions and refusal — close Part II, with the case they cannot handle in Part III.
  • Row 1 and the whole table run end to end on one engagement in Part V, including the events that were argued about and the one the schedule got wrong.
  • Row 8 — the actuals — is Part VI, and it is the row that decides whether any of this compounds.

Key takeaways

  • The client's certainty and yours are different objects; the fee is the price of the conversion.
  • Eight jobs, and each exists because the commercial promise cannot be kept without it.
  • Most firms do census, bands and fee — and neglect exclusions, decline and loss history, in that order.
  • Decline is the organ that separates underwriting from risk management. A risk register has never refused anything.
  • Borrow the discipline, not the licence: this transfers mechanism, never statutory practice or the right to indemnify.

The table is a list of jobs. Before any of them can be done, though, there is a sorting problem — because two things that both look like "we don't know yet" turn out to need completely different machinery, and treating them alike is how the first three rows quietly fail.

04
Part II: Pricing the Exposure

Two Kinds of Not Knowing

One kind gets smaller when you look at it. The other does not, and no amount of inspection will change that. Sorting them is what decides which of the eight jobs applies.

Open any proposal you have written in the last year and find the assumptions section. It will contain a list of things you do not know. What it will not contain is any indication that the items on that list are two completely different species of object, requiring two completely different responses, and that treating them as one list is why the section is usually decoration.

Epistemic uncertainty is complexity that already exists and has not been observed. How large the estate is. How many configurations are in use. Which records are missing. What the dependency pattern looks like. How many source classes are actually present rather than declared. It is a measurement problem, and looking harder makes it smaller.

Residual uncertainty is everything else. Client behaviour. Novel exceptions that have not happened yet. Model failures. Vendor outages. A regulator's decision. Conditions that are genuinely without precedent. Looking harder does not shrink it by a single unit, because there is nothing there yet to look at. It has to be capped, reserved for, excluded, or separately priced.

The diagnostic

One question sorts the list, and it is the artefact of this chapter:

Ask, of every unknown

Would another week of inspection change my estimate of this?

Yes → epistemic. Census it. Money spent on inspection here has a return.

No → residual. Price it. Money spent on inspection here is theatre.

The reason that question works — and the reason the more natural questions do not — is that it keys on reducibility rather than on size or on how frightening the item is. Most firms sort their unknowns by anxiety, which sorts nothing, because the scariest item on the list is frequently the one a two-day census would have retired and the cheapest-looking item is frequently the one that eats the engagement.

Big-and-reducible is a census job. Small-and-irreducible is a reserve job. They can appear on the same page in the same font and require opposite responses.

Run it on a real list. Here is a set of lines of the kind that turn up in an assessment proposal, sorted:

Proposal assumptions sorted into epistemic and residual uncertainty
Epistemic — census it Residual — price it
"We assume the estate is broadly as described.""We assume key stakeholders will be available."
"We assume all sources are in supported formats.""We assume no significant regulatory change during the engagement."
"We assume documentation is reasonably current.""We assume the client's other vendors will cooperate."
"We assume a single authoritative system of record.""We assume no material change in requirements."

Two of those are genuinely hard to place, and the difficulty is instructive rather than annoying.

"Documentation is reasonably current" looks epistemic — you can go and read the documentation. But "current" is a claim about the relationship between two things, one of which is the live system, and if you cannot observe the live system then the currency of the documentation is not measurable within your boundary. The honest sort depends on your access, which means the same assumption is epistemic for one engagement and residual for another. That is a feature. It is what makes the sort worth doing per engagement rather than once per offer.

"No material change in requirements" looks residual, and mostly is. But a proportion of requirement change is not change at all — it is the client discovering what they already meant, which a good census surfaces at the door as an ambiguity rate rather than discovering in week five as a change request. Part of that item is measurable now. Splitting it is worth ten minutes.

A census cannot reduce an unknown that has not happened yet.

Why the split is load-bearing

Because it tells you which lever to pull, and pulling the wrong one is expensive in both directions.

Every hour spent censusing a residual uncertainty is wasted money that also produces false confidence — which is the worse half. The inspection comes back clean, the finding is recorded as "no issues identified", and the risk is exactly where it was, now with a document suggesting otherwise.

Every residual treated as epistemic becomes a promise you cannot keep. It arrives in proposals as a sentence that sounds reassuring and functions as a liability: we'll clarify that during delivery.

Pitfall: "we'll clarify unknowns during delivery"

That sentence is how fixed-price certainty products become ordinary projects with better marketing. If the unknown is material, it is one of exactly three things: a typed deliverable of this purchase, a named reserve trigger, or an explicit exclusion. It is never a smile in the steering committee.

The asymmetry that makes this urgent

Here is the part that is specific to now, and it is the reason this book exists in 2026 rather than in 2014.

AI collapses the cost of reducing epistemic uncertainty. Reading the whole estate instead of a sample. Parsing every declared requirement. Reconciling every source against every other. Classifying complexity before commitment rather than discovering it during delivery. All of that used to be a paid discovery phase with a consultant-week price tag per field, and it is now close to a byproduct of the machinery that will deliver the work anyway.

AI does almost nothing to residual uncertainty. The client's decision-maker is still on leave. The regulator still has not ruled. The genuinely novel case is still novel, and no amount of breadth turns an unprecedented thing into a precedented one.

With one exception, running the wrong way. Shared models, shared parsers, shared kernels — a species of residual risk that barely existed when every engagement was carried by different humans making different mistakes on different days. AI made that worse, and it is the subject of Part III.

The correction to the sales meme

The boundary of fixed-priceability moves outward exactly as far as the epistemic share of a given offer's uncertainty. Not a general expansion. An offer-by-offer one, measurable in advance, and different for two offers sold by the same firm in the same month.

Which is why "AI makes fixed price easy" is not an exaggeration of a true claim. It is a claim about the wrong variable.

The vocabulary this runs on

Sorting an unknown is not the same as resolving it, and a large proportion of epistemic unknowns will not close inside the engagement even with a good census. They still have to leave the engagement somehow, and the way they leave is published doctrine that this book uses rather than re-derives.

Typed uncertainty: every unresolved item becomes a bounded terminal state with a precise assertion and a next consumer. Directly mapped. Partially mapped. Not observed. Insufficient evidence. Inaccessible within the audit boundary. Unsupported source type. Ambiguous — human decision required. Excluded from this phase.

One clause in that catalogue carries more weight than the rest combined, and I am going to repeat it here because this book cannot function without it: not observed does not mean does not exist. A bounded audit almost never warrants absence. Clients and lawyers will hear "we did not find X" as "X is not there", and if the product permits that slide you have manufactured both a liability and a false map in a single sentence.

Why the vocabulary belongs in this chapter rather than being deferred to the machinery chapters: the typed states are how an epistemic unknown exits the engagement when the census could not close it. Without them, every unclosed epistemic item silently becomes residual and lands in the reserve — which is how reserves get exhausted by things that were measurable all along, and how a well-designed reserve gets blamed for a census that was not run properly.

Buyers hear typed unknowns as incompleteness — once

Clients tend to read a typed unknown as a gap in your work rather than as a property of the world. That lasts exactly until they have been burned once by an over-confident report, at which point they become the strongest advocates for typing you will ever meet.

The reframe that works in the room is short: incomplete false confidence is the risk; typed incomplete knowledge is the asset. Show them a sample decision pack where not observed and human decision required appear as visible, valuable rows rather than as apologies.

One operational note worth carrying, from the pricing side: track the rate of ambiguous — human decision required by band. If too many items land there, either your sensors are weak, the requirements language is broken, or the product boundary is in the wrong place. A spike in that state is a product signal, not a badge of professional complexity.

Where each half goes

  • The epistemic share drives the census and the band — Chapters 5 and 6.
  • The residual share drives the reserve and the exclusions — Chapters 5 and 8 — and where neither can hold it, the decline in Chapter 9.
  • The correlated slice of residual gets its own treatment, because it fits none of the above — Chapters 10 and 11.

Notice what that list implies. Every unknown has exactly one route, and there is nothing left over. If you finish the sort and an item fits none of them, that is not a risk you should carry carefully. It is a design failure, and the correct response is to shrink the promise until the item is out of it.

Key takeaways

  • Sort unknowns by reducibility, not by size or by how alarming they are.
  • The diagnostic is one question: would another week of inspection change my estimate of this?
  • Censusing a residual is waste that also manufactures false confidence.
  • AI collapsed the cost of the epistemic half and left the residual half alone — except for one class, which it made worse.
  • Every unknown routes somewhere. A leftover is a signal to shrink the promise.

With the uncertainty sorted, the next question is arithmetic. If the fee is a premium rather than a forecast, what exactly is it made of — and why does an eighty-year-old actuarial document insist on two separate charges where most firms carry one?

05
Part II: Pricing the Exposure

What the Fee Is Made Of

Most firms can name two components. The profession that has priced risk for three centuries names four — and insists on two separate charges where we carry one.

Ask a partner what their fixed fee is made of and you will reliably get two terms: expected delivery cost, and margin. It is not a foolish answer. It is the answer the accounting system produces, and for a labour-priced unit it was very nearly complete.

Then ask a follow-up: what is the margin for?

If the answer is "profit", the fee has no reserve in it, and the margin is quietly doing a reserve's job. That is the whole diagnostic, it takes eleven seconds, and it is the most common finding in this book.

Four terms, not two

The actuarial statement of the same object is more useful, and it is worth quoting precisely because the list is longer than ours by exactly the items that hurt.

"A rate is an estimate of the expected value of future costs," and a rate "provides for all costs associated with the transfer of risk" — where those costs are enumerated as "claims, claim settlement expenses, operational and administrative expenses, and the cost of capital." 8

Translate each into a professional-services fee.

1. Expected delivery cost

The one everybody has. Worth noting immediately what dominates it in an AI-native offer: not machine breadth, which is nearly free, but the number of consequential judgements a named human must own. That is the metered resource and it is the term that actually moves.

2. Settlement cost

The expected cost of handling surprise: reserve draws, remediation after a failed acceptance, verification, escalation, the re-run. Note carefully that this is not the reserve — it is the expected value of the reserve's use, which is a term in the price rather than a pot of money. Almost nobody carries it, and it is not small.

3. Operating cost

The machinery: the harness, the evaluation set, the adapter library, the kernel maintenance, the register review. Ordinary, and it belongs in the fee rather than in overhead, because it scales with the size of the book rather than with the size of the firm.

4. Cost of capital

The one that will be new. Your working capital, your ability to absorb one bad engagement without changing behaviour, and your reputation are all posted as collateral the moment you sign. That collateral has a cost in every engagement, including all the ones where nothing goes wrong. A firm with no answer here is financing its clients' risk from its own balance sheet at a rate of zero.

Your fee is not delivery cost plus margin. It is expected cost, plus the cost of settling surprises, plus operating cost, plus the cost of the capital you just posted.

Prospective, which is the whole difference

One more line from the same document, and it explains why the census has to happen before the number rather than during the engagement: "Ratemaking is prospective because the property and casualty insurance rate must be developed prior to the transfer of risk."

You do not get to adjust after learning. That is the entire difference between a price and an invoice, and it is the reason every hour of measurement moved to the front of the engagement is worth more than the same hour spent later.

The sentence that carries the chapter

Under Considerations, between credibility and catastrophes, the same standard says something that took me a while to appreciate and then reorganised how I think about a fixed fee:

The rate should include a charge for the risk of random variation from the expected costs. This risk charge should be reflected in the determination of the appropriate total return consistent with the cost of capital and, therefore, influences the underwriting profit provision. The rate should also include a charge for any systematic variation of the estimated costs from the expected costs. This charge should be reflected in the determination of the contingency provision. — Casualty Actuarial Society, Statement of Principles Regarding P&C Insurance Ratemaking8

Two charges. Two different kinds of wrongness. Read them into our vocabulary slowly, because this is the most important external quotation in the book.

Random variation is this engagement is harder than average, in the ordinary way. The estate is messier than the median estate of its band; two sources disagree; an approach fails and another is generated. That is priced into the margin — and this is exactly what it means, financially, for the supplier to absorb interior variation. It is not generosity and it is not absorption in the moral sense. It is a charge that was already in the number.

Systematic variation is the picture the estimate was built on was wrong. Not messier than expected — differently shaped than expected. That gets its own separate charge, and in the profession's language the charge lands in the contingency provision.

Random variation → margin

The engagement is harder than the band's average, in a way the band already anticipated could happen. Absorbed silently. Priced once, for all engagements in the class.

Our name for it: interior variation.

Systematic variation → separate charge

The assumption behind the estimate was wrong for this engagement. Drawn against a named trigger, visibly, with a counter. Priced as its own provision.

Our name for it: the typed reserve.

Now the consequence that should change behaviour, and it is the reason this chapter is here rather than in an appendix.

A firm holding one number for both is holding a margin that is silently doing a reserve's job. When systematic variation arrives — and it does, several times a year — the margin is what disappears. And because it disappears gradually, in increments nobody logged, against events nobody classified, there is no moment anybody can point to and no list of causes anybody can review. The firm ends the year with worse margin than it modelled and a plausible story about a difficult client.

Myth vs reality

The myth: the typed reserve is a clever new device for the AI era — a way of making fixed price work now that delivery is cheap.

The reality: it is the contingency provision. It has been a stated professional obligation since 1988, it exists because random and systematic variation are different objects, and every insurer on earth carries both charges because carrying one of them does not work.

The reserve was never mine to invent

The credibility of this whole frame depends on not overclaiming, so let me be exact about how much of this is new. Almost none of it.

Project management has had the same split for decades. Contingency reserve sits inside the baseline for known risks with active response strategies — the known-unknowns. Management reserve sits outside the performance measurement baseline for unforeseen in-scope work, and drawing on it goes through the change-control process. 9

Construction has the same pair under different words, and the distinction is sharper than ours: an allowance covers planned scope that has not been fully specified yet — the fixtures the owner has not chosen. A contingency covers unanticipated cost that nobody had at budget time. Change orders are the primary driver of contingency depletion. 10 A conventional contingency percentage circulates in that industry; it comes from vendor guidance rather than a standards body, so I will call it convention and not lend you a number.

Which is a nearly exact match for the split between included findings and the reserve. Every lump-sum builder already holds a reserve. What is different in what I am proposing is narrow, and I want it stated narrowly rather than allowing anyone to infer a generosity that is not there:

  • typing by event class rather than by risk-register entry;
  • a published consumption rule, so a draw is a lookup rather than a negotiation;
  • and an exhaustion decision with an owner, a window and a stated default.

The protocol for all three is published and owned in its own treatment; this book uses it and does not re-teach it. What this book adds is a different reading of the same event — exhaustion as loss development — which arrives in Chapter 16 with a worked engagement attached.

The census is the arithmetic input

None of the four terms can be filled in without measuring the thing first, which is why the census is row one of the table rather than a nice-to-have.

Measure the input surface before quoting, using the same deterministic sensors that will deliver the engagement — so the measurement is a free byproduct rather than a paid discovery phase. The single most important property of that arrangement is easy to skim past: complexity is measured before the promise is made, by the thing that will have to keep the promise. A census run by a different instrument than the delivery machinery is a survey, and surveys are optimistic.

I have put the compressed version of this chapter in print twice before. Once as a description: that is not old-fashioned time-and-materials estimation; it is machine-measured product configuration. And once as a distinction that makes it operable: estimation asks a person to invent a number under incomplete information and social pressure; configuration asks a system to measure inputs against a published product definition.

Humans still design the bands, the materiality rules and the reserve triggers. What humans stop doing is acting as the unreviewable join algorithm for the price.

The cost curve, stated honestly

One correction that keeps the arithmetic from being oversold, and I would rather make it myself than have a reader make it for me. AI does not make the cost curve flat. It makes it flatter, and it changes which variable drives it. The variance that historically killed fixed pricing — how many documents, how messy, how many undocumented systems — now lands mostly on cheap parallel machine work. The residual variance is how many findings need a senior human's call, and that is precisely what the included-disposition count meters.

So the envelope is not padding for total ignorance the way old fixed prices were. It prices a measured residual. The padding shrinks to a reserve, the reserve is typed, and the types are visible on a counter both parties can read.

And where the sources give me no number — band prices, reserve percentages, a cost-of-capital rate — I will give you the shape and refuse to invent the figure. Copy the mechanism. Do not copy imaginary numbers as if they were market data.

Key takeaways

  • Four terms: expected cost, settlement cost, operating cost, cost of capital. Most firms carry two.
  • Ask what the margin is for. If the answer is "profit", it is doing a reserve's job.
  • Random variation belongs in the margin; systematic variation gets its own charge. Two kinds of wrongness, two charges.
  • The reserve is the contingency provision. Not new, not ours, and a professional obligation for four decades.
  • Configuration, not estimation — and where there is no number, the shape.

Which leaves the question of what population the premium is a premium for. That is a classification problem, and it is the one most firms have quietly decided they can skip because their clients are all different.

06
Part II: Pricing the Exposure

Bands Are Risk Classes

One sensible price is the most dangerous number a firm can publish. Not because it is wrong — because of who it attracts.

The objection arrives in the first ten minutes of every conversation about fixed-price productised services, and it is a reasonable objection asked in good faith:

How can you possibly fix-price this? Our clients differ by an order of magnitude.

I opened a chapter on exactly that question once before, and I gave a mechanical answer — the envelope, the census, the bands. What I did not do was explain why the question is not an argument against fixed price at all. It is the argument for risk classes, and until you see that, bands look like packaging.

What a blended price actually does

Do not warn about this. Walk it, one buyer at a time, because the mechanism is invisible from inside and obvious from outside.

You publish one number for a class of engagement. Four buyers look at it in the same quarter.

Buyer A — true cost well below your number. They compare, find somebody cheaper, and go. You never learn why. They do not send feedback; they send nothing.

Buyer B — true cost at your number. They buy. This is the engagement you had in mind when you set the price, and it is now a minority of what you sell.

Buyer C — true cost above your number. They buy quickly, because you are cheap. They are delighted with you. They refer you to people whose estates look like theirs.

Buyer D — true cost far above your number. They buy fastest of all, and somewhere in the sales process they will tell you their estate is fairly straightforward. They believe it.

Nobody in that sequence has behaved badly. Every one of them made a rational purchase with the information available. And within a year your book is composed disproportionately of the engagements your price was wrong about.

The part that makes it lethal rather than merely unfortunate: it is invisible in any individual deal. Deal review looks at engagements one at a time, and each one is defensible on its own terms. The composition is the problem, and composition is not a thing anybody in a professional-services firm is assigned to look at.

You do not get a random sample of the market. You get the part of it that agrees with your mistake.

The name for it

"Risk classification is one means of minimizing the potential for adverse selection. It reduces adverse selection by balancing the economic forces governing buyer and seller." 5 And the mechanism, from the same document: adverse selection occurs when prices are not reflective of expected costs.

Naming it is not academic. An unnamed pattern gets explained away one deal at a time — that client was unusual, that estate was a mess, that one is on us. A named pattern gets measured, and the measurement is Chapter 20's row three.

Why classes work when the individual engagement does not

There is an objection behind the objection, and it deserves a real answer rather than a reassurance: every engagement is genuinely unique, so how can a class mean anything?

The answer does not require pretending engagements are alike. "While any individual risk in a given class is no more predictable than it was before the transferring or pooling of the risk occurred, a reasonable price may be established by observing the losses of the class and relating the price to the average experience of the class." 5

The ratemaking side says the same: where an individual risk's own experience is not credible, it is appropriate to use the aggregate experience of similar risks, and a rate estimated that way is an estimate of the cost of the transfer for each individual in the class.

That is the actuarial licence for eligibility bands. Uniqueness at the item level and predictability at the class level are perfectly compatible — it is the entire basis on which insurance functions, and nobody argues that houses are interchangeable.

Worth noticing, too, is what that document says about the alternative. Pricing by "wisdom, insight and good judgment concerning the nature of the particular hazard involved… usually is not the best method but sometimes is the only one available." Senior estimation theatre has a formal position in the actuarial literature, and the position is last resort.

The design constraint most band schemes get wrong

Having established that classes work, the interesting problem is how many to have — and this is where band design stops being packaging and becomes a genuine trade with no clean answer.

"There is a point at which partitioning divides data into groups too small to provide credible patterns. Each situation requires balancing homogeneity and the volume of data." 8 Classes may improve homogeneity, at the expense of credibility.

Too few bands

Each band contains engagements that are not really alike. The band average predicts nothing about the engagement in front of you, and the band's actuals are a blur of two or three different populations.

Too many bands

Each band is homogeneous and empty. You never accumulate enough history in any one of them to learn anything, and every engagement is effectively priced from first principles again.

There is no correct number. There is a defensible position that you revisit with data. And the observation that lands with practice leads: a firm that has never felt that tension has not been banding seriously. If your bands have never been uncomfortable in either direction, they are packaging with thresholds.

Practical guidance rather than a figure, since I do not have anybody else's population: start with fewer bands than feel right. A band with history beats a band with precision, because the history is what lets you tell whether the precision was real. Split only when a band's actuals visibly contain two populations rather than one — bimodal effort, or two clusters of exception classes that never co-occur.

Classes have to be declared before the data

A discipline that most firms will fail on the first attempt, and it costs a year when they do.

"A major difficulty with this approach is the need to choose the relevant similar risk characteristics and related classes before the observation period. There often is not a clear-cut optimal set of characteristics."

You cannot accumulate loss data against a band that did not exist when the work was done. Which means two things operationally. Census fields must be fixed in advance — you can only learn about drivers you measured. And re-banding must be a governed act with a date, not drift.

The failure that prevents: a firm that quietly redefines its bands twice a year has a loss history that cannot be aggregated at all, and it will discover this at exactly the point it wanted to use it — usually the week it decides to scale the offer.

Five principles, and the one that transfers least obviously

"The system should reflect expected cost differences. The system should distinguish among risks on the basis of relevant cost-related factors. The system should be applied objectively. The system should be practical and cost-effective. The system should be acceptable to the public." 5

Four of those translate directly. The fifth needs work, and it turns out to be the commercially important one.

Acceptable to the public, in our world, means the buyer must be able to see why they are in their band. You do not publish your fully loaded cost model. You do publish enough drivers that a buyer can look at their own numbers against your thresholds and understand the assignment. Opacity about drivers recreates the budget dance and the distrust that comes with it; transparency about drivers creates configuration trust, which is a different and more durable thing.

Applied objectively also does more work than it looks. It means the band comes from the census output, not from who happened to be in the room when the number was set.

What a good driver looks like

The test for any candidate driver is one question: does it move with expected cost, or with apparent size? Those two feel the same on a proposal and diverge in delivery.

Volume counts — workbooks, documents, contracts, tickets — move with apparent size. They are easy to census, easy for a buyer to accept, and in an AI-native offer they mostly drive machine work, which is nearly free. They feel like drivers because they used to be.

The one worth naming, because most firms already collect it and none of them use it, is the early ambiguity rate: the proportion of declared requirements or obligations that cannot be parsed into testable claims. It predicts disposition load better than raw counts, because disposition load is the actual cost driver and volume is a proxy for it that stopped working when breadth got cheap.

Chapter 18 is the engagement where a driver turned out to be decoration, what it cost, and what replaced it. It is worth knowing in advance that this happens, because the first time a driver fails the instinct is to blame the engagement.

Pitfall: the silent band override

Sales drops a band to win a logo, in a proposal footnote, for a good reason. If anyone can redefine a band that way, you are back to authored pricing with extra steps and a product's name on it — and, worse, the engagement will be recorded against a band it was never in, which corrupts the only data that would have told you.

The light control: a monthly change board that reviews proposed threshold edits against review-economics data. Product configuration is a managed object. Overrides are logged as exceptions with a name attached, or they are not permitted.

And the internal review most firms skip: if a band consistently loses money, the band is wrong. Fix the product definition. Do not save it with unpaid heroics — that is the silent absorption of the next chapter, moved one level up and made structural.

The same client, priced twice

One more principle to plant here and collect later. "When an individual risk's experience is sufficiently credible, the premium for that risk should be modified to reflect the individual experience." 8

Engagement two with the same client is not priced from the class. It is priced from that client's own loss history — how fast they provisioned access, whether their decision-makers appeared inside the stated window, how stable their requirements turned out to be. That is a real and slightly awkward commercial move, and Chapter 20 handles it properly, including how to present it without it reading as a punishment.

Key takeaways

  • A single blended price does not get you a random sample of the market. It selects for the engagements you underpriced.
  • Adverse selection is invisible in deal review, because every individual engagement is defensible.
  • Uniqueness at the engagement level and predictability at the class level are compatible — that is what classes are.
  • Homogeneity trades against credibility. Start with fewer bands; split only when a band's actuals show two populations.
  • Publish the drivers, not the cost model. And a band that always loses money is a product defect, not a run of bad luck.

The band classifies the engagement before it starts. What it cannot do is classify what happens inside it — and there the trouble is a single word that has been doing four jobs at once for as long as any of us have been selling work.

07
Part II: Pricing the Exposure

The Variance Schedule: Four Classes

"Scope" is the least useful word in professional services, because it hides four unrelated things behind one noun — and every argument about it is therefore an argument about the wrong question.

Watch a scope argument for five minutes and you will notice something odd about it. Neither side is being unreasonable, both are describing real events accurately, and they are talking past each other completely. That is not a communication failure. It is what happens when one word is carrying four different objects with four different owners, four different commercial consequences and four different speeds.

Here is the schedule that separates them. This is the artefact this book is built around.

The four-class variance schedule
Class What it is Who owns it What the buyer pays Speed
1. Provider-owned interior variation Production movement inside an unmoved perimeter Supplier, silently Nothing Instant — no decision by anyone
2. Metered residual scarcity Consequential human judgement, sign-off, verification Priced in the band; overflow draws reserve Included quantity, then a visible draw Logged, not negotiated
3. Boundary mutation A named perimeter field moved against its recorded value Both parties, by re-contract A conversation with four options Days, named owner each side
4. External and correlated tail A cause outside this engagement, often shared across many Nobody, by default — which is the problem Entirely dependent on what was written Simultaneous, portfolio-wide

Class 1 — provider-owned interior variation

An approach fails its test and the system generates another. An integration needs three regeneration attempts rather than one. The evidence base is rebuilt because the first assembly used a source that turned out not to be authoritative. Fifteen thousand pages instead of three thousand. Two authoritative sources disagree and reconciliation takes four passes. A source arrives in a form nobody anticipated and an adapter gets written on a Tuesday afternoon.

None of that is a change request. The buyer bought a state; the search that produced it is my business. If they paid for candidate mortality they would be funding the search they explicitly did not buy — and funding it at exactly the moment it was working, which is the wrong incentive pointed at the wrong party.

My own less formal version of the same observation, from the conversation this book came out of: AI is just gobbling up all those middle bits to keep the project bounded in time and resources and how it operates in the real world. The middle bits are gone. The judgements are not.

The rule that makes it survive a delivery organisation

Everything above is a description. This is the mechanism, and without it the description is a preference that will lose an argument on a Thursday night.

Interior is the default, and it is a residual rather than a judgement. A supplier cannot classify something as interior by deciding to. It can only observe that no named perimeter field moved. Interior is what remains when eight recorded values have each failed to change.

Nobody has to be brave

Under the old clause, absorbing something is an act. Somebody chooses not to raise a request. That choice is invisible, unrepeatable, and becomes an unacknowledged precedent the next engagement inherits.

Under this rule, absorbing is what happens when nobody does anything. It removes the delivery lead's discretion in the direction where discretion is most expensive — at ten o'clock at night, with a client on the phone and a relationship in the room.

And the financial reading, connecting back to Chapter 5: Class 1 is random variation, and it is priced into the margin. That is what "the supplier absorbs it" means on a profit-and-loss statement rather than in a contract clause. It is not a concession. It is a charge that was already in the number.

Class 2 — metered residual scarcity

Material human dispositions. Licensed sign-off. High-consequence ambiguity. Exceptional technical review. Verification and remediation. Genuine novelty that nobody has a route for yet.

None of that became free when generation became cheap. The machine may investigate, synthesise and nominate every one of those calls. It may not become a disposition, and it may not mutate an authoritative state.

Which is why this is the metered resource the band actually prices, and why the interior can be free without the engagement being free. AI makes the cost curve flatter, not flat, and what remains scarce is the count of consequential judgements a named human must own.

The commercial mechanics are three things. An included quantity per band. Written materiality rules — what counts as a finding, what is context-only, what merges, decided in advance rather than argued in week six. And beyond the included count, exactly two routes: a band uplift if the census was wrong, or a reserve draw if exceptions appeared inside a correctly assigned band.

Showing that meter to the buyer is honest rather than awkward, for a reason worth stating: the meter is the thing the band was priced against. Hiding it does not make the price feel more generous; it makes the price unexplainable.

Class 3 — boundary mutation

A named perimeter field has moved against its recorded value. The buyer changes the promised outcome. A new source class appears. The input estate crosses the agreed band. The required authority changes. A regulatory obligation appears. The deadline moves. The reserve exhausts.

The compact rule is four words long and it does more work than any clause I have written: implementation movement is not contract movement. Internal mess is the supplier's variance; changed commercial intent is a boundary mutation.

The financial reading again: Class 3 is systematic variation — the picture the estimate was built on has changed. Chapter 5 established that this gets its own charge, which is why the reserve exists and why exhaustion is a commercial event rather than an accounting one.

What I am deliberately not doing here is re-deriving the instrument. The eight perimeter fields, the recorded-value discipline, the trigger design and the re-contract routine are published, and re-teaching them would cost this book the chapters it needs for its own argument. One paragraph, one citation, move on.

Class 4 — external and correlated tails

Named here, developed in Part III — but named with enough detail that the gap is visible.

The familiar members are old news: a third party delays an integration; a regulator changes an admissible process; the client cannot produce an authorised decision-maker; organisational adoption fails. Every experienced practitioner has a clause for those.

The two that barely existed five years ago do not have a clause anywhere:

  • a shared model upgrade changing behaviour across every live engagement in the same week;
  • one common parser, connector or kernel defect affecting the whole portfolio at once.

Look at why these fit none of the first three classes, because that is the diagnosis rather than the complaint. Class 1 assumes the cost is yours and bounded by this engagement. Class 2 assumes it is metered per engagement. Class 3 assumes a perimeter field moved on this engagement. A correlated tail moves nothing on any single perimeter and draws every reserve simultaneously.

It is the class AI genuinely made worse, and it is the reason this book has a Part III.

Three dispositions, three Thursdays

A classification is only useful if each class has a different consequence attached automatically. Otherwise it is a relabelling with better vocabulary. So: what does each one actually look like on a Thursday?

Thursday one — interior

Nothing happens. The work is harder than expected, the team does more of it, and the buyer never learns. No request, no note, no log entry against them. The default is silence.

Thursday two — reserve draw

A line moves on a counter the buyer can already see, and somebody sends a one-paragraph note naming which trigger fired. The delivery lead decides, against a published rule. Logged, not negotiated.

Thursday three — boundary mutation

A named person on each side has a scheduled conversation inside five days, with a delta and four options in front of them: uplift the band, extend the reserve, narrow the boundary, or stop.

Three classes, three decision-makers, three speeds. The design principle underneath is worth keeping: a taxonomy whose classes share a decision-maker is a relabelling. One whose classes have different decision rights is a governance design.

Class 4 fits none of those three, which is precisely its problem, and precisely why it needs its own part of this book rather than a fourth row on a form.

Silent absorption is unpriced risk capital

Now the part of this chapter I most want a reader to carry, and I want it to stay commercial rather than drifting into a lecture about discipline.

Absorbing everything is not generosity. It is an unpriced promise — and an unpriced promise is what destroys suppliers, quietly, one engagement at a time.

The sentence that produces it is always the same, always said kindly, and almost always by somebody senior: we'll absorb it, they're a good client. Name what that actually is. An undocumented, unpriced, unrepeatable concession that immediately becomes the baseline for the next engagement, and cannot be withdrawn later without looking like a downgrade. Generosity with the receipt thrown away — and it compounds, because next time the starting position is the concession rather than the contract.

The term this book adds on top of that: unpriced risk capital. Because that is what it is. Your firm made a genuine capital contribution to your client's project, recorded nowhere, recoverable never, and invisible to every instrument either party owns.

Why visible absorption is worth more than silent absorption

A supplier who can show four absorbed exceptions and one governed decision has evidence of absorption. It can be put in front of a buyer at renewal. It changes what they believe the fixed price was buying.

Silent absorption produces no evidence, buys no credit, and cannot be sold. Same money. Same events. One of them is an asset and one of them is a donation.

The moral version of this argument is easy and unpersuasive, and nobody with a pipeline has ever been moved by it. The commercial version is stronger, so use that one.

Run it on your own last engagement

Four steps, one afternoon, and you already have the raw material.

  1. Take the last engagement that closed. Not the interesting one — the last one, because selecting for interest is how this exercise gets ruined.
  2. List every surprise: everything that made somebody say hang on. Emails count. The list will be longer than your change log, and that gap is the first finding.
  3. Put each item in exactly one class. Where two people disagree, mark it — those are your underspecified rules, and Chapter 15 supplies the tie-breakers for the common ones.
  4. Price the Class 1 column at your own internal cost.

That number is what your firm contributed as unrecorded risk capital on one engagement. Multiply by your engagement count for the year. Almost nobody has this number, and the reason to compute it is not that it will be shocking — it is that until it exists, every conversation about absorption is a conversation about feelings.

Key takeaways

  • "Scope" hides four objects with four owners, four consequences and four speeds.
  • Interior variation is the default and a residual — you cannot classify something as interior by deciding to.
  • The default is what makes it survivable: absorbing happens when nobody does anything, so nobody has to be brave.
  • Classes 1–3 have automatic consequences. Class 4 has none, which is Part III's subject.
  • Silent absorption is unpriced risk capital. Visible absorption is an asset you can show a buyer.

Three of the four classes now have a home. The fourth needs something written down before it needs anything else — and it turns out that the thing everybody writes last, in the final hour before a proposal goes out, is a pricing instrument rather than legal small print.

08
Part II: Pricing the Exposure

Exclusions Are Part of the Product

They get written in the last hour before a proposal goes out, from a template, by whoever is still awake. It is the most expensive hour in the document.

You know how the exclusions list gets written. The pricing is settled, the deliverables are agreed, somebody has done a final read for tone, and then there is a section near the back that says Assumptions and Exclusions. It gets populated from the last proposal, adjusted for the obvious differences, and sent.

Nobody is being lazy. It is genuinely the least interesting part of the document to write, and it has never in anybody's memory changed whether a deal was won.

It is also the only part of the document that changes what the number means.

Exclusions are a pricing input

In the actuarial standard, exclusions do not sit in a legal annexe. They sit in the same list of ratemaking Considerations as credibility, exposure units and catastrophes: "Consideration should be given to the effect of salvage and subrogation, coinsurance, coverage limits, deductibles, coordination of benefits, second injury fund recoveries and other policy provisions." 8

They are inputs to the rate. They change the price. Which gives the test this chapter turns on, and it is one you can run on your own last two proposals in five minutes:

If your exclusions do not change your price, they are decoration.

Two proposals with identical numbers and materially different exclusion lists cannot both be correctly priced. One of them is carrying an exposure it is not being paid for, and it is not the one with the longer list.

Underneath the test sits the structural reason exclusions exist at all, and it is not "we would rather not do that work". You may only absorb what you can measure, price and reserve for. Everything else is an unpriced promise, and Chapter 7 has already established what unpriced promises do to a supplier.

What must not be absorbed

Six fences, published in their own treatment and used here rather than re-derived: physical work; third-party decisions and timetables; buyer delay; unlimited exception tails; liability the supplier cannot control; and open-ended intent.

One line out of that list is worth carrying on its own, because it is the most expensive misreading of the entire AI-native architecture: better analysis lowers the probability of an error and does nothing whatsoever to the cost of the one you make. Machine breadth reduces the chance of missing something. It does not touch the magnitude of what happens when you do, and a supplier who accepts an uncapped obligation because the analysis got cheap has confused two different quantities and will discover the difference exactly once.

And one fence is worth naming as the one most suppliers breach, so you can check yourself: buyer delay. It is breached in the generous direction, always for a good reason, and it does two things at once. It teaches the buyer that dates are decorative, and it consumes calendar — which is where your scarce humans actually get spent.

The four refusals, translated

This chapter's own contribution starts here. I designed an equipment-continuity product some years ago whose entire commercial integrity rested on refusing four specific promises, and the discipline generalises out of its domain more cleanly than I expected.

The original rule: sell concrete service acts with clocks and cut-offs; refuse the promises whose physics you do not control. The four refused were guaranteed uptime, zero downtime, delay compensation or project-loss indemnity, and universal 24-hour repair.

Every one of them has a twin in professional services, and the twins are sold constantly.

The four refusals translated into professional services
Refused in equipment continuity The professional-services twin
Guaranteed uptime"Guaranteed adoption" — the recommendation will be implemented
Zero downtime"No findings missed" — complete coverage of your estate
Delay compensation, project-loss indemnity"We'll cover your project loss if we're late"
Universal 24-hour repair"Every exception resolved within a week"

A list of refusals persuades nobody, so walk one of them as a mechanism.

Why a universal clock cannot work

Repair completion time is a heavy-tailed distribution. The head is a stocked wear item and a competent technician. The tail is diagnosis ambiguity, a long-lead part, a remote site, crane time, or a failure that is not the component anybody photographed.

So a single universal clock is set to one of two places. Set it to the head, and it systematically fails the tail — which trains customers to treat exceptional cases as service failures, and trains staff to take unsafe shortcuts to hit a clock that was a lie for part of the distribution. Set it to the tail, and it is commercially meaningless because it promises nothing anybody wanted.

Now apply that, unchanged, to "every finding dispositioned within five days".

Disposition time is also heavy-tailed. Most findings are a stocked answer and a competent reviewer. Some need an authority who is not available this week, or a legal opinion that has not been formed, or a judgement that genuinely requires two people in a room. The universal clock fails for precisely the same structural reason — and, worse, it fails on the exact findings that mattered most, because difficulty and consequence are correlated.

The mechanism of refusal transfers as cleanly as the failure does. Commit to clocks on acts you control: acknowledgement, classification, escalation to a named qualified person, delivery of the evidence pack by a date. Then, if you offer resolution times at all, bound them to defined classes with explicit eligibility rather than across the whole distribution.

The consulting translation

Every one of these is a bounded state a supplier can keep and evidence. Every one of the things they replace is an outcome somebody else controls.

  • Promise an assessed estate, not perfect enterprise knowledge.
  • Promise a verified decision, not improved client profit.
  • Promise an evidence pack by a date, not organisational adoption.
  • Promise a tested increment, not transformation success.
  • Promise a response commitment, not control over every dependency.

The fence around this book's own enthusiasm

An underwriting frame is seductive in a specific way: once you have the vocabulary, it starts to feel as though any risk can be priced if you are clever enough about the trigger. It cannot, and my own corpus says so, and this book is going to stay consistent with it.

A promise — and therefore a falsifier attached to it — may attach only to a state the supplier can (a) keep and (b) evidence. Both tests, not either.

The boring case fails both and everybody spots it. The interesting judgement lives in the two cases that pass one test and fail the other.

Keepable, not evidenceable

"The client's team now understands the architecture."

You can genuinely cause it. Nobody can produce an observation showing it did not happen. Repair: attach to an artefact and a demonstrated use of it.

Evidenceable, not keepable

"Backlog reduced by thirty per cent this quarter."

Beautifully measurable, and dependent on volumes, staffing and steps you do not control. Repair: attach to your contribution — the assessed portion, the demonstrated throughput of the mechanism under stated conditions — and let the buyer own the aggregate.

One test to carry into a drafting session, and it settles most arguments in under a minute: if it fires, can we tell whether we caused it? If the answer is no, the promise is attached to the wrong thing, and no amount of measurement precision will rescue it.

Outcome contracting is not the escape hatch

The obvious response to all of this is that "outcome-based" contracting solves it — that if you price against results rather than activities, the exclusions problem takes care of itself. The public sector has decades of evidence on that and it is not ambiguous.

The gaming taxonomy is documented and has names for each behaviour: contracts risk "'cherry picking', where eligible individuals are not referred or accepted onto a service if they seem unlikely to achieve payable outcomes, 'creaming', in which providers focus their efforts on those individuals who are easiest to help, and 'parking', which is neglect of those who may be more difficult to achieve outcomes with." And underneath that, the harder problem: "it can be difficult to set simple, measurable outcomes that align effectively with complex social problems." 11

Then the finding that contains three separate failures in one paragraph. The US Government Accountability Office examined performance-based logistics arrangements across fifteen programme offices. Only one had updated its business case with actual cost and performance data — and in that single case "it determined that the performance-based logistics contract did not result in expected cost savings and the weapon system did not meet established performance requirements." Meanwhile programme officials "typically relied on cost and performance data generated by the contractors' information systems", and the offices "had not determined whether contractor-provided data were sufficiently reliable." 12

Read all three. Fourteen of fifteen never re-tested the premise. The one that did test it, fired — which is what a test is for, and exactly why nobody volunteers to run one. And the outcome evidence was authored by the party being assessed, which is the whole of Chapter 12's argument, breached in contract form twenty years before anyone said the words "LLM as judge".

So "outcome-based" is not a solution. It is a different set of failure modes. The first three are prevented at the perimeter — by eligibility and exclusions, not by any instrument downstream. If your intake admits selective participation, nothing later will fix it. The fourth is prevented by the keep-and-evidence fence above.

The clause without which the exclusions are fiction

If sales can reintroduce a refused promise in a deal footnote, the product definition is fiction. That is not a governance nicety; it is the difference between a product and a set of intentions.

Treat out-of-menu promises the way you would treat giving away professional indemnity: exceptional, logged, reviewed by somebody who does not carry the deal — or refused. The commitment menu is a cultural artefact as much as a contract schedule, and it needs a logged exception path rather than an honour system, because under competitive pressure honour systems produce exactly one outcome.

The sequencing rule that gives exclusions credibility

Everything above is worthless if the exclusions arrive at the wrong moment. They are stated by the supplier, in the contract, before anyone needs them.

A fence disclosed at the moment it is needed is not a fence. It is an excuse with a clause number.

Which is also the argument for putting them near the front of a proposal rather than the back. A document that hides its limits at the end has reproduced the exact pathology it claims to replace, and the buyer will read it that way the first time one of the limits becomes relevant.

And there is a commercial return on the sequencing that surprises people. Buyers who have been burned once by an over-confident supplier recognise a real boundary immediately, and they pay for it. Predictability is a premium attribute when the supplier possesses machinery capable of holding the risk — it is not a discount, and pricing it as one teaches the market the wrong category.

Key takeaways

  • Exclusions are a ratemaking input. If yours do not change the price, they are decoration.
  • Better analysis lowers the probability of an error and nothing at all about the cost of the one you make.
  • A universal clock across a heavy-tailed distribution fails the tail or means nothing. Commit to acts you control.
  • A promise may attach only to a state you can keep and evidence. Both tests.
  • Stated by the supplier, in the contract, before anyone needs them — or it is an excuse with a clause number.

Exclusions handle the risks you can name and refuse in writing. What they cannot handle is the opportunity where the problem is not one risk but the perimeter itself — where the honest answer is not a clause but a different engagement, or none.

09
Part II: Pricing the Exposure

Decline Is a Mechanism, Not a Mood

The dangerous opportunity is the good one. Well funded, real relationship, real problem, real budget — and a perimeter that will not hold.

Chapters about saying no usually open on a bad deal, which teaches nothing, because nobody needs help refusing a bad deal. The engagements that damage a practice are almost never the ones that were badly delivered. They are the ones that were accepted while somebody in the room already knew the promise could not be bounded.

So picture the other kind. A well-funded group, a real relationship, a genuine problem, a budget that exists, and a buyer who would like to work with you specifically. Everything about it is attractive except the perimeter.

That is the opportunity this chapter is about, and the reason decline has to be a mechanism rather than a mood is that in this room, mood will lose.

The organ, restated once

A risk register manages a project you have already agreed to do. It identifies, rates, assigns owners and tracks mitigation, and it has never in the history of the discipline refused anything — because by the time it exists, the refusal decision is months in the past.

Which produces a diagnostic worth applying to your own firm before you read further: if there is no artefact in your business that can say no before a proposal is written, you do not have an underwriting function. You have a sales function with a legal review.

The quotability self-test

Eight rows. Run it before anybody writes a number. For each field, one question: can I write a recorded starting value and an observable trigger for this?

The eight-row quotability self-test
# Perimeter field The question
1Promised stateIs there a sentence to record — one that has held still through two meetings?
2Authoritative input estateWho warrants it, and can they warrant it today?
3Volume / bandCan it be censused this week?
4Buyer-controlled dependenciesNameable and datable?
5Authority and accessDoes the approving body exist, with known composition?
6Consequence / liability classCan the tail be described — and capped, in writing?
7Acceptance ruleCan it be written, and could it come back negative?
8Fixed time boundaryIs there a date, and does it mean something commercially?

Mark each one ✓, ~ or ✗. The marks are the decision — not a discussion input, the decision.

Take the opportunity from the opening. Field 1 fails: the decision has been redefined twice in three meetings, so there is no sentence to record. Field 2 fails: which entity's records govern depends on a restructure that has not completed, so nobody can warrant the estate today — the warrantor is one of the things being decided. Field 3 passes, as it happens; the estate is censusable this week. Field 4 is marginal: the people who owe inputs are the people whose roles are being restructured, so they are nameable but not datable. Field 5 fails: the approving body will exist after the restructure and its composition is unknown. Field 6 fails: which regulator applies is one of the things the restructure decides, so the consequence tail cannot be described and therefore cannot be capped. Field 7 fails, entirely downstream of field 1. Field 8 passes and is meaningless without the rest.

Five of eight cannot carry a recorded value, and three of those five sit outside both parties' control — which is the distinguishing feature, not the count. A field the buyer could fix is a precondition you can write into the engagement. A field neither party can fix is a boundary.

Note the dependency before counting failures, because it changes the remedy. Field 7 is downstream of field 1: repairing one field repairs two rows. A reader running this on their own live opportunity should map the dependencies first, or they will over-diagnose.

The chain that terminates

No recorded value, therefore no delta. No delta, therefore no trigger. No trigger, therefore every surprise resolves into an argument.

An instrument that cannot distinguish interior variation from boundary mutation is not an instrument, and a fixed price laid over it is not a commitment. It is a wager with a schedule attached.

Hard is what the machinery is for. Unclassifiable is what the machinery refuses.

That distinction is what stops this chapter from being timid. Hard engagements are the entire point of everything in Parts I and II. The opportunity above is not hard. It is unclassifiable, which is a different property, and no amount of care compensates for it.

What happens if you ignore the verdict

Not a warning — a sequence, because it is the same sequence every time and its predictability is the argument.

You absorb the first two redefinitions as goodwill. Each one is individually small, and each one is genuinely the right call in the room. In month two you discover the approving body has changed composition, which quietly invalidates the acceptance design you never quite finished writing. In month three you reopen the price — arriving at exactly the failure the perimeter was built to prevent, except that you have now also spent the credibility of a fixed-price promise on the way there.

And the compounding cost, which does not show up in the post-mortem: the next fixed price you quote that buyer will be discounted in their head by their memory of this one.

The shrink, which is the actual answer

Refusal without a shrink is a lecture, and nobody with a pipeline has ever been moved by a lecture. So here is the move, applied rather than listed.

The four-step shrink

  1. Identify which perimeter field is unstable. Not "this feels risky" — which of the eight cannot be answered, or cannot be held.
  2. Remove precisely the commitment that depends on it. That one. Not a defensive haircut across the whole proposal. The censusable component stays exactly where it is.
  3. Re-home the removed part as one of three things: a client responsibility with a named owner, a separate commercial object, or a typed terminal state.
  4. Re-check that what remains is still worth buying. The step people skip — and the one that decides whether this is a shrink or a decline wearing a shrink's clothes.

Applied to the opportunity above: fields 1, 2, 5 and 6 are unstable, with 7 downstream of 1. So sell the bounding. A smaller fixed commitment whose entire deliverable is a filled perimeter — which entity governs, which regulator applies, who approves, and what decision is actually being asked.

Notice what that product is. It is the missing rows, produced as a deliverable. And a filled perimeter is precisely what makes the larger engagement quotable afterwards — by me, or by anybody else, which is the test of whether the bounding is a real product or a hold on the account.

The buyer is not being sold a smaller version of what they wanted. They are being sold the thing that has to exist before what they wanted can be bought at all.

And say the uncomfortable half plainly: it is also less revenue this quarter, from the same budget. Both are true, and a chapter that mentioned only the first one would be asking you to believe something you know is false.

Re-band, the middle option

Between "quote it" and "refuse it" sits the option most firms skip. When the census turns out to have been wrong, that is a measurement bug or an access lie — not a reason to return to author-luck pricing. Re-census, uplift the band, or stop. Do not absorb it as culture.

The distinction that tells you which is which: did the estate turn out to be different from what was measured, or did the driver fail to predict cost for an estate that was measured correctly? The first is a re-band. The second is a product-definition problem, and it gets its own worked case in Chapter 18.

Why refusal is commercial rather than moral

The moral version of this argument is easy and unpersuasive. The commercial version is stronger, so here it is.

Forcing out-of-band work into a fixed price does three expensive things at once. It destroys the band's meaning for every future buyer, because the band no longer predicts anything. It teaches your own sales system that drivers are negotiable, which reintroduces private-judgement pricing under a product's name. And it converts a product back into a bespoke project with a product's price and a project's cost — which is the worst combination available anywhere in professional services.

There is a fourth cost, quieter and worse than the other three. A lane full of forced exceptions cannot teach you anything. A half-comparable population produces numbers nobody can act on, and a firm that cannot measure its own product cannot improve it. That cost connects directly to Part VI, and it is the one that turns a bad quarter into a permanently blind business.

Now the objection, and it deserves an honest answer rather than a pious one. Our competitors will just say yes.

Some will. Some of them will win the deal and lose the money, and that is a market you can wait out — particularly since a buyer who has been over-promised to is a buyer looking for somebody credible in eighteen months. But refusing costs revenue now, and a firm without a pipeline cannot afford principles. Which is exactly why the realistic move is almost always the shrink rather than the walk-away, and why shrinking is the harder skill and the one nobody teaches.

Fixed price was never absolute

My own doctrine is a poor witness in its own defence, and construction lawyers have been saying the quiet part for decades.

Lump-sum contractors "often include significant contingencies in their pricing"; they are "naturally incentivized to seek opportunities to reopen the fixed price" where those contingencies prove insufficient; and then the sentence worth carrying into every negotiation: "In truth, there is no such thing as an absolute fixed price contract." 13

The same source states the eligibility test from the other end, which is the more useful half: lump sum "may still be preferable for well-defined, low-risk projects where scope and owner requirements are clear from the outset."

Legal services arrives at the identical boundary by a different route. Hourly billing persists in "complex, high-stakes, or open-ended" matters because "the scope keeps evolving… and the risk is asymmetric and dynamic", so "pricing these engagements upfront requires embedding significant risk premiums, which often makes fixed-fee structures impractical." 14

Two industries, opposite ends of the market, same boundary. The honest reading is that the perimeter does not abolish the pressure to reopen — nothing does. It governs when and how the reopening happens: early, typed and evidenced instead of late, adversarial and improvised.

What survives a refusal

Not nothing. What is preserved when the fee shape is declined is the stable unit of commitment — a per-project activation, a verified decision, an assessed estate, a protected period, a guaranteed response commitment, a capacity tier. Fixed price is a strong signal, not a law. What the unit must increasingly not be is "however many hours our internal process happens to consume".

And the honest hybrid is a legitimate answer rather than a failure: fixed onboarding and core outcome, metered consumption, paid dispositions, explicit capacity reservation, capped exceptional work. Chapter 18 walks one that was built under exactly that pressure.

When I first said all of this out loud, I put a hedge on it — that this is probably a pretty good reusable construct for AI-native successor products and businesses, a pattern or a template, not the only one. By now that reads as confidence rather than doubt. A rule that could not decline anything would not be a template. It would be a slogan.

Key takeaways

  • Run the eight rows before anybody writes a number. The marks are the decision.
  • Fields outside both parties' control are the distinguishing feature, not the failure count.
  • Hard is what the machinery is for; unclassifiable is what it refuses.
  • The shrink, not the walk-away — and re-check that what remains is worth buying, which is the step people skip.
  • Refusal is commercial: forcing work into a band destroys the band, teaches sales the drivers are soft, and blinds your measurement.

Part II has now handled every risk you can measure, meter, exclude or refuse one engagement at a time. Which leaves one that does not arrive one engagement at a time — and it arrives on the same Tuesday for all of them.

10
Part III: The Tail Nobody Reserved For

Your Own Machinery Correlates Your Book

Every fixed-price portfolio relies on one assumption that nobody states out loud — and the flywheel is systematically destroying it.

Here is a sentence every practice lead has said, and it is true: we can absorb a bad engagement, because the other nineteen were fine.

That is not luck and it is not resilience. It is pooling, and pooling has a precondition that is doing all the work in that sentence while being invisible in it.

The precondition nobody states

"Both types involve the transfer of financial uncertainty from one party to another and the subsequent pooling of risks. In both cases, the exposure to loss by the sharing mechanism should be broad enough to assure reasonable predictability of the total losses." 5

And underneath it, the mathematics: the law of large numbers is a statement about independent random variables. That word is doing all the work, and nobody notices it — because in a human-delivered practice, independence was free. Every engagement was carried by different people, making different mistakes, on different days, using different judgement. You did not have to buy independence. You could not have avoided it.

A portfolio does not protect you because it is large. It protects you because its members fail for unrelated reasons.

Where the analogy is strongest — catastrophe

The profession that prices variance for a living does not treat correlated loss as a worse version of ordinary loss. It treats it as a structurally different object with its own line in the rate: "Consideration should be given to the impact of catastrophes on the experience and procedures should be developed to include an allowance for the catastrophe exposure in the rate." 8

Which yields the transferable rule, and it is short enough to keep:

Independent variance is poolable, and therefore priceable as margin. Correlated variance is not poolable, and must be controlled, excluded, or capitalised.

There is no fourth option, and "we hold a reserve" is not one of the three.

The flywheel correlates the book

Now the uncomfortable part, and it is this chapter's actual contribution.

Consider the list of things every compounding argument — including all of mine — tells you to build. One shared kernel. Shared parsers. A shared evaluation harness. One pinned model. One adapter library. Reusable acceptance tests. A common exception taxonomy loaded as a starting condition.

Every single one of those is a deliberate act of correlating your own book.

The machinery that makes engagement twenty cheaper than engagement two is the same machinery that makes all twenty fail together. That is not an unfortunate side effect of doing it badly. It is what shared machinery is.

Why nobody notices: the two goals are pursued by different people in different meetings. The delivery lead is optimising for reuse, correctly. The commercial lead is optimising for portfolio spread, correctly. Neither of them is wrong, and neither of them is looking at the other's variable.

The inversion

The more engagements you run on the same machinery, the worse your concentration gets, not the better.

Which is the exact inverse of how a portfolio is supposed to behave, and it is the price of the flywheel. Growth on a shared kernel increases the exposure that growth was supposed to spread.

Be precise about the trade rather than moralising about it, because I am not arguing against shared machinery. Part VI argues for it, at length, and the economics only work with it. What I am arguing is that the shared kernel has a cost line nobody has been writing down — and the controls in Chapter 11 are that line.

The defence firms reach for does not survive contact: portfolio pricing diversifies only independent exceptions. It does not protect against a common failure in the production system, and a common failure in the production system is now the most likely large loss a well-machined firm will experience.

The regulators, who were not writing about us

My own doctrine is a poor witness here, so use bodies that had no interest whatsoever in this argument when they wrote it down.

The Bank of England, on concentration: "A reliance on a small number of providers for a given service could also generate systemic risks in the event of disruptions to them, especially if is not feasible to migrate rapidly to alternative providers." And on the shape of the exposure: "under a scenario in which customer-facing functions have become heavily reliant on vendor-provided AI models, a widespread outage of one or several key models could leave many firms unable to deliver vital services." And most directly of all: "From a systemic risk perspective, the potential for AI-based participants to take increasingly correlated positions is an important consideration." 15

The Financial Stability Board, on the mechanism itself, which is the sentence I would put in front of any delivery lead who thinks this is speculative: "Most LLMs are trained using the same underlying architecture and many are trained, at least in part, on common sources of web crawl data. The homogenisation in training data and model architecture can lead to correlated outputs." 16 The same report names third-party dependencies and service-provider concentration as a distinct vulnerability.

Translate that for a five-person practice, because the scale difference conceals the identity of the structure. These are financial-stability regulators describing your delivery stack. They are worried about it at national scale. The same structure exists in your book of twelve engagements — without the capital buffers, without the reporting, and without anybody whose job is to look at it.

The demonstrations

Two events, walked rather than listed, because the transferable part is not the size. It is the shape.

CrowdStrike, 19 July 2024

Microsoft estimated that the faulty update affected approximately 8.5 million Windows devices. 17 One analysis put direct financial losses to the Fortune 500 alone, excluding Microsoft, at "at least $5.4 billion", with cyber insurance covering "10% to 20% of the losses" — and estimated preliminary insured losses of "between $400 million and $1.5 billion, potentially the single worst loss in the cyber insurance sector over 20 years." 18

Note the vocabulary the insurance trade press reached for. Not "a big outage". An accumulation loss event — a loss that accumulated across an entire book from one cause. That is a category, it has its own pricing treatment, and it is the category a professional services firm has just quietly entered.

AWS us-east-1, October 2025

From Amazon's own summary: "The incident was triggered by a latent defect within the service's automated DNS management system that caused endpoint resolution failures for DynamoDB." 19

It was not an attack. It was a Tuesday.

Both root causes share a shape, and the shape is the thing to carry: a latent defect in automation, triggered by a routine change. Nobody attacked anything. Nobody was negligent in any way a review would have caught. An ordinary release, in a shared dependency, producing simultaneous failure across thousands of organisations that had no relationship with each other.

Which is precisely the class of event a per-engagement interior-variation margin cannot absorb. Not because the margin is too thin — because every margin in the portfolio is drawn down in the same week, and margins were sized on the assumption that they would not be.

What the insurance market does about it

Instructive, because the answer is not "price it higher".

Lloyd's excluded catastrophic state-backed cyber attacks, and the reasoning is worth reading exactly: "The ability of hostile actors to easily disseminate an attack, the ability for harmful code to spread, and the critical dependency that societies have on their IT infrastructure, including to operate physical assets, means that losses have the potential to greatly exceed what the insurance market is able to absorb." 20

And required that the exclusion be legible rather than buried: "we wish to reiterate that policy language should be clear so that the scope of cover is understood by all the parties and the exposure is properly assessed and monitored by syndicates." 21

There is a nuance in that decision which must not be lost, because it is the difference between the lesson and its opposite. Lloyd's did not exclude because the risk was bad. Bad risks are what insurance is for. It excluded because the risk was non-diversifiable — because pooling, which is the entire mechanism, does not work on it.

And note what they did not do. They did not stop writing cyber. They wrote a boundary, made it legible to both parties, and monitored what remained. Three moves, and Chapter 11 transfers all three.

The gap, stated as a gap

I would like to give you a number here and I am not going to, because there isn't one.

Nobody has published a correlation coefficient for delivery failures across engagements sharing a model vendor. The regulators establish the mechanism; nobody has the magnitude. Any figure I offered would be mine, would have to be framed as mine, and would be exactly the fabrication this book spends a chapter warning against.

What a reader can do in the absence of a published number is better than what the number would have given them anyway: measure their own. The register in the next chapter produces a firm-specific exposure map, which beats an industry average that would not have described your stack.

Three homes, and the reserve is not one of them

If the class cannot be diversified, there are exactly three remaining places for it: controls, priced into the band; exclusions, stated at signature with a typed terminal state; and capital, acknowledged on the balance sheet rather than assumed away.

What it must not be routed into is the reserve. A reserve sized for independent exceptions cannot absorb a simultaneous draw, and Chapter 17 does that arithmetic on a specimen so it is a demonstration rather than an assertion.

Key takeaways

  • Pooling requires independence. In a human-delivered practice independence was free; in a machined one it has to be bought.
  • Every element of the flywheel is a deliberate act of correlating your own book.
  • Growth on a shared kernel increases the exposure it was supposed to spread — the inverse of a portfolio.
  • Regulators name the mechanism: homogeneous models and vendor concentration produce correlated outputs and correlated failure.
  • Non-diversifiable is not the same as bad. Lloyd's excluded, made it legible, and monitored the residue — the three-part template.

Which leaves the practical question. If you cannot pool it and you cannot reserve for it, you have to control it — and the controls turn out to be four things your engineering team has already asked for and lost the budget argument on.

11
Part III: The Tail Nobody Reserved For

Underwriting Controls That Look Like Engineering

Version pinning, independent evaluations, rollback and shared-component regression are commercial instruments with a line in the price. The reason they keep losing budget is that nobody has ever computed the other side of the comparison.

You have sat in this meeting. The engineering lead asks for two weeks to build a regression suite across the shared components — the parsers, the adapters, the harness that every live engagement runs through. It is a good request, made with a straight face, supported by an example from last quarter.

It loses to a client deliverable. Not because anybody in the room is short-sighted, and not because the engineering lead argued it badly. It loses because nobody in the room can say what it is worth, and a request with no number always loses to a request with a client name on it.

That is a pricing failure wearing an engineering costume, and this chapter is about supplying the missing number.

The comparison being made is the wrong one

Those four controls are not hygiene. They are how you cap a non-diversifiable exposure that you have already accepted on behalf of every engagement in the book — accepted, note, the day you standardised the machinery, in a decision nobody framed as an underwriting decision.

So the comparison in that meeting is not "two weeks of engineering versus a client deliverable". It is "two weeks of engineering versus the simultaneous drawdown of every reserve you hold".

Which is why the decision keeps going the wrong way: one side of the comparison has a number and the other has never been computed. Fix that first, and the argument stops needing to be won.

The accumulation register

Borrowed in name from how insurers track aggregation — the practice of knowing, at any moment, how much of your book is exposed to one event. Specified here so it can be built this week, with a spreadsheet, by one person.

One row per shared component.

The accumulation register: columns and what belongs in each
Column What goes in it
ComponentThe model version, parser, connector, adapter, harness, kernel module, prompt library, eval set
Live engagements depending on itCount and names. This is the blast radius, and it is usually larger than expected
What changes it without asking usVendor push, silent model update inside a pin window, upstream library release, API deprecation, provider policy change
What we would observeThe concrete signal: eval score drop, schema mismatch, throughput change, a class of output going quiet
Detection timeFrom change to our knowing. The single most important number on the row
Rollback available?Yes or no, and how long it takes in practice rather than in principle
Exposure if it fails across all rows at onceThe sum, stated as a shape if you cannot yet state it as a figure

Two rows worth walking, because the value of the register is not in the easy ones.

The comfortable row. A parser you wrote, version-locked, covered by continuous integration. Six engagements depend on it. Nothing changes it but you. What you would observe is a failing test, immediately. Detection: instant. Rollback: yes, minutes. That row takes ninety seconds to fill in and is worth filling in anyway, because it establishes what "good" looks like on the same page as the alternative.

The uncomfortable row. A vendor-hosted model whose behaviour can shift without a version bump. Nine engagements depend on it. What changes it without asking you: the vendor, on their schedule, sometimes without a changelog entry that means anything to you. What you would observe: nothing, unless you are running an evaluation set — the output stays well-formed. Detection time: the honest answer is "a client would tell us". Rollback: only if you pinned, and only if the pinned version is still served.

The finding is almost always the detection column

Most firms can list their dependencies in twenty minutes. Very few can fill in detection time at all — and the rows where the honest answer is "a client would tell us" are the unpriced catastrophe rows.

The question that makes the register pay for itself, and it takes an afternoon to answer for a whole book: if this changed badly on a Monday, when would we know?

The four controls, priced as controls

Each one: what it does, what it costs, and — this is the part that has been missing from the budget conversation — what it buys, expressed in the variable that actually moves.

Version pinning

Does: converts a vendor's release calendar into your own.
Costs: model currency, some capability lag, and the discipline of running a migration project you chose rather than one you inherited.
Buys: the tail becomes a scheduled event. You migrate when your harness says the new version passes, not when the vendor ships.

Independent evaluation sets

Does: detects behaviour change before a client does.
Costs: building a set that stays honest — and resisting the constant pull to write cases that match what the system already does.
Buys: detection time. Which is the variable that decides whether one engagement is affected or all of them.

Rollback

Does: returns you to a known-good configuration.
Costs: architectural discipline, previous versions kept warm, and refusing to let state drift make the old configuration unrunnable.
Buys: loss duration — the difference between an incident and a re-delivery.

Shared-component regression

Does: runs across every engagement's dependency path, not just the one currently being worked on.
Costs: the suite, and the CI time.
Buys: it is the only control that tests the correlation itself rather than any individual instance. The other three protect engagements. This one protects the portfolio.

That last distinction explains why the regression suite is always the one that loses the budget argument. Its benefit is the thing that did not happen to eight clients nobody in the room was thinking about. It is structurally unpersuasive, and it is the most valuable of the four.

Lloyd's three-part template, transferred

Exclude what you cannot absorb. Make the wording clear enough that all parties understand the scope. Monitor the residual exposure.

Applied to a professional-services contract, that becomes three concrete things:

  • A named external-dependency exclusion, with the state it produces typed as a valid, deliverable terminal state rather than as a failure — blocked pending third party, model behaviour changed outside pinned version, upstream schema withdrawn.
  • Stated at signature, in the same vocabulary as everything else in the schedule, so it reads as part of the product rather than as a disclaimer.
  • And an internal register reviewed on a cadence — because the third part of the template is the one everybody drops, and it is the one that makes the first two honest.

The typed-state move matters commercially and is easy to skim past. An excluded event that produces a typed deliverable is still a deliverable. The engagement does not become undefined; it becomes explicit about what it could not reach. A decision pack that says "this portion was blocked pending an external determination, and here is what would change if it resolves the other way" is a better instrument than one that quietly omits it.

Where the exclusion has to stop

This is the chapter's honesty clause, and burying it would make everything above dangerous.

You may exclude what you do not control. You may not exclude your own production system's failures.

A model regression inside a version you chose to run is interior variation with an unusually wide blast radius. It is not an external tail. It is your architecture choice producing your cost.

The test, stated so it can be applied under pressure at the moment somebody wants to reach for the clause: did we choose to depend on this, and could we have detected it? Two yeses and it is yours.

Pitfall: "the model changed" as a universal excuse

It is only an external event if the change crossed a boundary you had pinned, declared and monitored — which is to say, it is only external if you had already done the work. A supplier who classifies its own dependency choices as external tails has invented a way to make buyers fund its architecture. The clause would be unarguable, indefensible, and it would work exactly once.

Independence, arriving early

One connection worth planting here, because most readers will not make it on their own and it changes how the eval set gets built.

An evaluation set only detects drift if it was not authored to match what the producer already believes. A harness written by the team whose work it checks converges on that team's assumptions and stops seeing them — not through carelessness, but because the same people wrote both the expectation and the check.

So independence shows up twice in this book, doing two commercial jobs with one property: as an acceptance instrument in Chapters 12 and 13, and as a correlation control here. Same requirement, and if you build it once you get both.

What the market's existence establishes

Munich Re and Armilla are already selling cover against AI performance variance with defined triggers and measured baselines. I said in Chapter 3 that their rating methodology is not public and I am not going to pretend otherwise here either.

What their existence establishes is narrow and sufficient: people whose entire business is pricing variance have decided that AI performance variance can be priced. "It is unpriceable" is off the table. The only live question is whether you can price yours — which requires the register, which requires the ninety minutes.

Build it this week

Ninety minutes, one spreadsheet

  1. List every shared component your live book depends on. Include the ones you did not build, and the ones you have forgotten you depend on — that second category is where the surprises are.
  2. For each, name the engagements that depend on it. Count them.
  3. For each, answer four questions: what changes this without asking us? What would we observe? How long until we knew? Can we roll back, and how fast?
  4. Look only at the rows where detection is "a client would tell us". Those are your unpriced catastrophe rows.
  5. Decide, per row: control it, exclude it, or hold capital against it. There is no fourth option — and "reserve for it" is the wrong answer, which the next chapter demonstrates rather than asserts.

The row that makes you uncomfortable is the point of the exercise. If every row is comfortable, you have not listed the vendor-hosted ones.

Key takeaways

  • The regression suite loses the budget argument because only one side of the comparison has ever been computed.
  • The register's value is the detection column, and most firms cannot fill it in.
  • Three controls protect engagements; only shared-component regression protects the portfolio.
  • Exclude, make it legible, monitor the residue — and type the excluded state as a deliverable rather than a failure.
  • You may not exclude your own architecture choices. Did we choose it, and could we have detected it? Two yeses and it is yours.

Part III has dealt with the risk that arrives from outside the engagement. Part IV deals with the one that arrives from inside it — from the incentive a fixed price creates in the party doing the work, which is not a character problem and cannot be fixed by hiring better people.

12
Part IV: Who Is Allowed to Say It Is Done

The Producer Cannot Be the Verifier

Under time-and-materials, effort and payment move together. Under fixed price they move in opposite directions — and nothing else about the engagement has to change for that inversion to start producing consequences.

Three observations, none of them controversial on its own.

The supplier observes the actual production path and what it cost. The buyer largely observes the supplier's claim that completion occurred. And under a fixed fee, the supplier benefits financially from reducing internal effort.

Put together, that is money sitting on one side of an information asymmetry, with the party holding the information also holding the scoring pen.

The framing that makes this usable

Before going further I want to fix the register, because this argument is routinely made badly and the bad version cannot be said in front of a client.

None of this requires dishonesty. It is the ordinary consequence of optimisation. A team that knows which artefact determines acceptance will produce that artefact well, and will produce the things nobody is scoring to whatever standard time allows. That is not a moral failing. It is what competent people do under a deadline.

The version I have used for years converts the character question into a structural one, and it is the only form of this argument I have ever seen actually land in a room: it doesn't matter how honest your developer is if you have no independent way to verify the honesty.

That sentence comes out of an older argument about why custom software lost to packaged software, and the argument is worth thirty seconds because it is the same shape one level down. The standard story is that custom was too expensive and too slow. That story is incomplete, and being incomplete is why it cannot explain anything useful. Custom software did not lose on build cost alone. It lost on verification cost — a non-technical buyer could not inspect the work, could not tell a good developer from a confident one, and could not tell six months of progress from six months of invoices. Every status update was a claim they had no way to check.

An AI-native fixed-price engagement reproduces that buyer position exactly, one level up. The buyer cannot inspect the production path, cannot tell a well-verified result from a confidently presented one, and cannot tell coverage from the appearance of coverage.

The lineage, attributed properly

Almost everybody gets this wrong, and getting it right costs one sentence and buys a small amount of credibility.

Charles Goodhart, 1975, watching the Bank of England target monetary aggregates that promptly lost their meaning: "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."

The famous compression — "When a measure becomes a target, it ceases to be a good measure" — is Marilyn Strathern's, from a 1997 paper on audit in the British university system, and it is routinely hung on Goodhart. 22 Donald Campbell adds the stakes clause: the more a quantitative indicator is used for decision-making, the more subject it is to corruption pressures, and the more apt to distort the process it was meant to monitor.

I have assembled all three before with their references, so I will not re-do the history here. What matters is the operative reading for delivery, which is sharper than the general statement: a quality gate is a measure — a proxy standing in for "the work is actually good". Hand that measure to the worker as its target and it stops measuring. The worker does not need to do the job; it needs to clear the gate, and those are only the same thing as long as it does not know where the gate is.

Capability makes it worse

Here is where intuition fails, and the failure is expensive because it runs in the reassuring direction.

The natural assumption is that better people and better tools protect you: a sharper team is less likely to satisfy a proxy hollowly, and a stronger model is less likely to produce something that passes and does not work.

The research says the opposite. From DeepMind's work on specification gaming — which they define as "a behaviour that satisfies the literal specification of an objective without achieving the intended outcome" — "Even for a slight misspecification, a very good RL algorithm might be able to find an intricate solution that is quite different from the intended solution, even if a poorer algorithm would not be able to find this solution… This means that correctly specifying intent can become more important for achieving the desired outcome as RL algorithms improve." 23

Myth vs reality

Myth: our premium work is protected, because it is done by our best people with our best tools.

Reality: a better producer finds the loophole a worse one would have missed. Capability is an aggravating factor, not a mitigating one, and the highest-stakes work is therefore the most exposed rather than the least.

And the sentence that explains the timing of this entire book, which is neither Goodhart's nor Strathern's: "the importance of Goodhart effects depends on the amount of power directed towards optimizing the proxy, and so the increased optimization power offered by artificial intelligence makes it especially critical for that field." 24

Read that against what an AI-native service actually is. It is, by construction, a large amount of optimisation power aimed at whatever you wrote down as the test.

The 2026 demonstration

Until recently this was an argument by analogy, and analogies lose arguments with sceptical partners. It is not an analogy any more.

A controlled study put two production coding agents to work re-implementing a component library in a different framework, under a hidden 222-test oracle, across eighteen runs and — this is the design choice that makes it valuable — three different oracle-availability conditions.

Without the oracle, the library is present but unfinished, revealed by scores. With the oracle in the loop, the score reaches near-perfect, but from a demo holding the tested behavior directly, the library left dead or absent. We call this building to the test… The agent does not, on its own, validate what it ships as a user would. — Ma, Kereopa-Yorke and Schultz, “Building to the Test”, arXiv:2606.2843025

The score is not the outcome. Near-perfect against the oracle, with the artefact that was actually requested dead or absent. That is an acceptance meeting where everything passed and nothing was built.

Withholding the oracle changed what got built. The availability of the answer key is not an administrative detail. It is a design variable that determines the shape of the work.

The producer cannot be the verifier. Stated by the researchers, about systems with no financial incentive whatsoever. Add a fixed fee and the pressure only increases.

The disposition does not stay where you put it

One further finding, and it is the part that should genuinely worry a practice lead rather than merely interest them.

Anthropic's work on production reinforcement learning found reward hacking generalising well beyond the objective that was gamed: "when large language models learn to reward hack on production RL environments, this can result in egregious emergent misalignment… the model generalizes to alignment faking, cooperation with malicious actors, reasoning about malicious goals, and attempting sabotage." 26

Translate that carefully and without over-reading it, because the over-read is available and it is not what I am claiming. I am not saying your delivery team will sabotage anything. I am saying that a system — human or machine — which learns that clearing the gate is the job carries that disposition into the parts of the engagement nobody is scoring. The learned behaviour is not "satisfy this test". It is "find what is being checked".

Which is the practical reason the gate must be held by somebody who does not carry the delivery number. Not because they are more honest. Because they are optimising for a different thing.

The failure image

Every partner over a certain age recognises this instantly, which is why I keep returning to it: an engagement everybody was happy with, delivered on time, praised in the steering committee — and no artefact anywhere that could have said no.

Ask the awkward question that follows the invoice: what, exactly, was accepted? In most engagements the honest answer is a feeling, held by several senior people at roughly the same time.

That is not a defect of those people. It is a missing instrument — and three things become possible the moment it exists. The buyer can tell completion from exhaustion, rather than an engagement ending because the calendar ended. The supplier can defend the perimeter at the exact moment it matters most, which is when the money changes hands. And disputes about whether the promise was kept stop resolving to whoever is more senior or more determined, and start resolving by looking at something.

What this chapter is not claiming

One boundary, so the seam with the falsifier work stays clean. The semantics of falsification — what makes an engagement wrong to start versus wrong to accept, and why those two fail independently — is a separate treatment with its own book, and I am not re-opening it here.

This chapter argues exactly one thing: the independence of the oracle.

Key takeaways

  • Fixed price inverts the relationship between effort and payment, which arms Goodhart without anybody deciding anything.
  • None of it requires dishonesty. Make the structural argument; the character argument cannot be said in front of a client and is also wrong.
  • Capability aggravates specification gaming. Your best work is the most exposed, not the least.
  • "The agent does not, on its own, validate what it ships as a user would" — and withholding the oracle changed what got built.
  • The gate belongs with somebody who does not carry the delivery number.

If the producer cannot be the verifier, then the contract needs more objects in it than most contracts have — and the third one has a price that has to go somewhere.

13
Part IV: Who Is Allowed to Say It Is Done

Three Objects, and What the Third Costs

Independence sounds expensive until you see the four cheapest forms of it — and the ratio that turns "we can't afford that" into a number somebody can manage.

Here is a sequence no supplier should be permitted to run end to end:

define the proxy  →  produce the artefact  →  score the proxy  →  declare completion

Every fixed-price engagement without an independent oracle is running exactly that sequence, and most of them do not know it — because the four steps are owned by four different people inside the same firm. Different desks, different names on the org chart, one objective function.

Three separate objects

The contract needs three things where most contracts have two.

1. The promise

The state, decision, artefact or commitment being purchased. What the buyer will have that they did not have before.

2. The declared acceptance surface

The evidence both parties know must exist. Published, legible, agreed at signature. This is not a secret and should not be.

3. The independent oracle

Tests, samples, authoritative reconciliations or live receipts that the producer does not unilaterally control. The thing that can return a result the producer did not choose.

Compressed to a rule: intent visible, acceptance legible, verification independent. The agent may propose completion. The harness disposes it.

Most contracts collapse objects two and three into a single "acceptance criteria" clause. Which means the document that defines completion and the instrument that proves it are the same artefact, held by the same party — and no amount of careful drafting inside that clause fixes a structural problem.

The distinction that resolves the apparent conflict with hidden-gate delegation is worth being precise about, because a reader who knows that work will spot the tension. Not every criterion can or should be hidden; commercial fairness requires clarity about what is being bought. But the producer need not possess the complete answer key. Intent stays visible because intent cannot be gamed, only pursued. The answer key is a different object and does not have to travel with it.

Four cheap forms of independence

The objection this pre-empts is that independence means standing up a second delivery team. It does not, and pretending it does is how the requirement gets dismissed in the first meeting.

Four cheap forms of independent verification
Form What it is What it costs When it fits
Held-out cases A slice of the estate the producer never sees until scoring, selected by census rule rather than by the delivery team The discipline of carving it out before work starts Work repetitive enough that a sample generalises
Buyer-selected samples The buyer picks after generation, from a population the supplier cannot curate Almost nothing — and it is the most persuasive of the four in a sales conversation The buyer has enough domain knowledge to pick meaningfully
Authoritative reconciliation Check output against a system of record neither party authored, within a stated tolerance Access, and agreeing the tolerance in advance An external source of truth exists at all
Live receipts Production behaviour observed over a period rather than demonstrated at a point Time, and a commitment extending past invoicing The promise is a maintained state rather than a delivered artefact

Walk the cheapest one properly, because it is also the most under-used.

Buyer-selected sampling works like this. At signature, the supplier declares the population and its boundary — every contract in these three repositories, every finding in this category, every record processed under this rule. After generation is complete, the buyer draws n items from that population. The supplier scores them against the declared acceptance surface, in front of them, in one session.

Total cost: an afternoon, and the willingness to be wrong in public. Total value: an acceptance event that means something, because the producer did not choose the evidence.

Which is what all four forms have in common, and it is the entire property. The producer does not choose the evidence. Everything else is implementation detail.

If you need the thirty-second version for a procurement lead, they already hold the distinction in another form. A SOC 2 Type 1 report "is as of a point in time… It only covers the design effectiveness of the internal controls." A Type 2 report "covers a period of time… the operating effectiveness of the internal controls over time." 27 A demo is Type 1. Live receipts are Type 2. Everybody in the room understands immediately.

Independence has to be mechanical

Not a second opinion. A different question, checked a different way, against something outside the model.

The rule, which costs suppliers money and therefore needs grounding outside my own corpus: no single-model self-verification may be presented as independent proof.

Be precise, because that rule is both over- and under-applied. It does not forbid using a model to check work — models catch real defects and refusing to use them would be silly. It forbids presenting that check as the independent evidence on which completion turns, and it forbids treating agreement between two runs, two prompts, or two models of the same family as verification.

The grounding. Across an evaluation of more than 350 large language models, "models agree 60% of the time when both models err" — and, closing the obvious escape route, "larger and more accurate models have highly correlated errors, even with distinct architectures and providers." 28 And the trend, which closes "this will fix itself as models improve": "model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures." 29

Now hold the counterweight in the same breath, because the literature is not one-sided and presenting only the alarming half would be exactly the selective citation this book criticises two sections from now. Strong model judges do match human preference well — "achieving over 80% agreement, the same level of agreement between humans." 30

So the honest position is narrower than a prohibition and sharper than a preference:

Model judgment is a legitimate link in an evidence chain. It is never the closer. Ranking candidates is not the same job as closing a commercial promise.

One connection most readers will not make on their own, and it is worth thirty seconds: correlated model error is the same phenomenon as the correlated portfolio failure in Chapter 10, appearing at the level of the check rather than the delivery. One underlying cause — homogeneous models — producing two commercial consequences at two different altitudes. If you built an independent evaluation set as a correlation control, you have most of an acceptance oracle already.

The structural precedent

Whole industries settled this decades ago and made it a matter of structure rather than process quality or good intentions.

Under ISO/IEC 17020, third-party inspection requires a Type A body — independent of the entities involved in design, production, ownership or maintenance of the thing being inspected. As an accreditation body puts it: "Rock solid demonstrations of impartiality require the IB to ensure that its own staff are not involved in any aspect of ownership, design, manufacture, or other relationship as regards the object of inspection or its manufacturer / supplier." 31

Note the shape of that requirement, because it is the anti-washing test applied to verification: independence is declared structurally, and then the accreditor tests the declaration against how you actually operate. Saying you are independent is the beginning of the process rather than the end of it.

The objection that decides whether any of this happens

It costs money out of a fixed fee. True. Three answers, in increasing order of usefulness.

The cost exists either way. Unpriced, it arrives later as rework, dispute, remediation, and a quiet discount on the next proposal to the same buyer. You are not choosing whether to pay it. You are choosing whether it is in the number.

The cheap forms cost a fraction of a second delivery team. See the table above. Buyer-selected sampling costs an afternoon and produces the strongest acceptance evidence available to a small firm.

It is measurable — which converts an argument into a number a partner group can actually manage.

The verification ratio

Verification effort divided by production effort removed, tracked per engagement type. One timesheet field and one column in the delivery record.

The phenomenon is well evidenced, and it has a name. DORA: "The verification tax: Time saved writing is often re-spent auditing." And, more precisely: "While AI successfully accelerates initial code generation and reduces the friction of starting new tasks, the time saved in creation is frequently re-allocated to auditing and verification." 32

Two observations from the same source matter more to an underwriter than to an engineering manager. "Verification is a fundamentally different cognitive task than creation" — so the effort does not simply move, it moves to different and scarcer people. And: "higher AI adoption is associated with an increase in both software delivery throughput and software delivery instability."

Read that second one as a change in the shape of the loss distribution rather than as a productivity note. More claims, arriving faster, with more variance. That is a different pricing problem from "delivery got cheaper", and it is the problem a firm actually has.

The corollary that inverts the usual assumption

Verification cost rises with stakes faster than production cost falls. So the higher the stakes, the less of the compression you keep — which is the opposite of what most partnerships assume when they reason that their premium work is protected.

A firm that removes twenty junior hours of analysis and adds eight senior hours of review has captured a materially smaller gain than its timesheet suggests. In high-consequence work it may have captured none.

That does not kill outcome pricing. It bounds it, and the bound is measurable — which is the entire reason to track the ratio. A firm that has never computed it does not know how much of its productivity gain is real.

What this book owes its own thesis

There is a famous number in this territory and I am going to handle it the way I would want a supplier to handle their own loss data.

METR ran a randomised controlled trial with experienced open-source developers working on their own repositories. Using early-2025 tooling they took 19% longer than without it — having forecast a 24% speed-up beforehand, and believing afterwards that they had been sped up by 20%. 33

It is a wonderful result and it is doing a lot of work in a lot of arguments. It also comes with a statement from its own authors, in February 2026, about their follow-up experiment: "we believe that the data from our new experiment gives us an unreliable signal." They attribute the unreliability to selection effects they cannot control for, and say the original figure likely does not reflect current conditions. 34

So the safe claim is the shift of effort from creation to verification, which is consistently observed, rather than a slowdown number that its own researchers have walked back. Cite the 19% as history, always paired with the retraction, or do not cite it at all.

I am labouring this for a reason that is on-thesis rather than pedantic. An underwriter who mis-states their own loss data is not underwriting. A book making that argument cannot then reach for a convenient statistic whose authors have withdrawn the signal, and the discipline it is asking of a reader has to be visible in its own citations first.

Key takeaways

  • Three objects, not two. Collapsing the acceptance surface and the oracle is the structural error.
  • The property that matters is one thing: the producer does not choose the evidence.
  • Model agreement is not independence — and it gets less independent as models improve, not more.
  • Independence costs money either way. Priced, it is a line item; unpriced, it is rework and a discount on your next proposal.
  • Track the verification ratio. The higher the stakes, the less of the compression you keep — and you cannot manage what you have never computed.

Four parts of doctrine are now on the table. None of it is worth anything until it has been run on a real engagement — including the events that were argued about, the reserve that ran out, and the one the schedule got wrong.

14
Part V: One Engagement, Underwritten

The Census, the Band and the Fee

The underwriting decision itself, made in public: what was measured, what the measurement bought, and how the fee was composed.

Status of what follows

This is a composite specimen, assembled from delivery patterns in my own material. It is a worked model, n=1. It is not a client case study and no client is described in it.

Where a figure would be market data I do not have, I give the shape. Copy the mechanism. Do not copy the numbers as if they were portfolio evidence. That is the same discipline my earlier books applied to their own specimens, and it is in the first paragraph rather than a footnote for a reason: a book that asks for honesty about loss data cannot blur its own.

The offer. An Obligation Estate Assessment. A bounded assessment of an organisation's third-party contract estate, ending in a decision pack: an obligations inventory with typed coverage, an exposure map, named gaps, and a recommended disposition per material obligation. Fixed fee, eight weeks, decision-complete. Remediation and construction are separate purchases.

The census, field by field

Eight fields. What matters is not the list but the second column — what each field was for, and what cost driver it was a candidate for. A census field that is not a candidate driver is a statistic.

The preflight census fields and what each was measuring for
Census field What it measures Why we thought it drove cost
Repositories in scope, with access statusWhere the estate lives and whether we can reach itAccess is the largest single source of calendar loss
Contract count by classMaster agreements, SOWs, amendments, NDAs, purchase ordersThe obvious volume proxy — and the one Chapter 18 finds wanting
Format distributionNative text, structured export, scanned image, legacy systemDecides whether extraction is free or needs an adapter
Declared obligation categoriesThe client's own register of what it believes it owesThe comparison surface for coverage
Early ambiguity rateShare of declared obligations that cannot be parsed into testable claimsPredicts disposition load better than raw counts
Unsupported source types at the doorFormats outside the sensor set for this phaseBecomes an exclusion or a reserve trigger — never a silent absorption
Named counterpartiesConcentration and jurisdiction spreadDrives which legal authority must be involved
Buyer-side dependency ownersWho owes us access, decisions and sign-off — by name and roleConverts "the client is slow" into a nameable, datable field

The property that matters most is easy to skim past: complexity is measured before the promise is made, by the thing that will have to keep the promise. The census runs on the same deterministic sensors that will do the extraction, so the measurement is a byproduct rather than a paid discovery phase — and, more importantly, it is not optimistic in the way a survey is.

The discipline that goes with it: if you skip the census and just know the band from a sales call, you have reintroduced the unreviewable join algorithm at the front of the product you built to kill it.

The epistemic and residual sort

Chapter 4's diagnostic, run against this engagement's actual unknowns. This is the page I would put in front of a sceptical partner first, because it is where the abstraction turns into routing.

Epistemic, closed by census

Estate size. Format mix. Counterparty spread. Coverage of the declared obligation categories. Which unsupported formats are present at the door. All of it measured in the week before the number was written.

Epistemic, not closed

How many obligations would prove genuinely ambiguous once read. The census sees the shape — the ambiguity rate — but not the content. Routed to the included disposition count, with a reserve trigger sitting above it.

Residual → reserve trigger

Access delayed or narrowed after contract. Buyer decision-maker unavailable beyond a stated window. Obligation density above the band's assumption. Requirement influx inside the declared package.

Residual → exclusion, typed

A regulator changing an admissible interpretation mid-engagement → blocked pending external determination. A counterparty refusing to confirm a disputed amendment → insufficient evidence. Both deliverable states, not failures.

Residual → buyer responsibility, named

Legal sign-off on material obligations. Access provisioning. Warranty of which repository is authoritative. Each with a named owner and a role, in the schedule.

Residual → correlated, routed to controls and exclusion

Shared model, shared parser, shared extraction harness. Not routed to the reserve — to register rows, controls in the band, and a named exclusion. Chapter 17 shows the arithmetic behind that refusal.

What this page makes visible is the thing worth taking away from it: every unknown has exactly one route, and there is nothing left over. A leftover is not a risk to be carried carefully. It is a design failure, and it is the moment to shrink rather than to be brave.

Band assignment, with drivers published

Three bands. This estate lands in M on contract count and format mix — with the ambiguity rate sitting near the top of M's range.

That last detail was flagged to the buyer at assignment rather than discovered later, and the buyer was shown the drivers, the thresholds and their own numbers against them. Not the internal cost model; the drivers.

And a judgement call was recorded at the time, because a specimen that only shows the clean parts is not a specimen. The ambiguity rate near M's ceiling could have justified uplifting to L. It did not, on the reasoning that M's included disposition count plus the reserve covered the modelled overrun.

Write down that it was a call. Chapter 15 shows what it cost. Chapter 18 shows what it eventually changed. Neither of those is available if the decision was never recorded as a decision.

The fee build

Chapter 5's four terms, filled in — as composition, not as figures.

Expected delivery cost. Dominated by dispositions on material obligations, not by extraction breadth. The machine reading ten thousand contracts is nearly free. The calls a named lawyer has to own are not, and there are considerably fewer of them than there are contracts.

Settlement cost. The expected value of the reserve's use — not the reserve itself. Plus remediation after the held-out sample, plus escalation on ambiguous obligations.

Operating cost. Harness maintenance, evaluation-set upkeep, register review, the adapter library that made three of this engagement's formats free.

Cost of capital. Working capital across eight weeks, plus the capacity to absorb one bad engagement in this band without the firm changing behaviour. Stated as a term with a rationale. The percentage is local and I do not have anybody else's.

Deliberately absent, and worth naming so the absence is not read as an oversight: no rate card, no reserve percentage, no band price. Those are local, and lending you mine would be exactly the failure this book spends Chapter 22 warning against.

The variance schedule, as signed

All four classes with this engagement's specifics. This is the artefact Chapter 15 classifies against, so it has to exist before the events do.

Class 1 — interior, absorbed

Extraction retries. Adapter writing for anticipated format classes. Regeneration after a failed internal check. Reconciliation passes. Volume movement within the declared band. Rebuilds caused by our own errors.

Class 2 — metered

Material obligations through human disposition, up to the included count. Materiality rules written down at signature: what counts as a material obligation, what is context-only, what merges. Beyond the count: band uplift if the census was wrong, reserve draw if it was not.

Class 3 — perimeter fields

Promised state; authoritative repository set; band ceiling; buyer-controlled dependency set; approving authority; consequence class; acceptance rule; end date. Each with a recorded value, dated and agreed. The field discipline is published elsewhere and used here rather than re-derived.

Class 4 — the row most engagements do not have

Named external and shared-dependency exclusions, each producing a typed terminal state rather than a failure: blocked pending external determination; model behaviour changed outside pinned version; upstream schema withdrawn. Backed by named register rows, not by a clause alone.

The reserve, sized and typed

Five units. Five admitted trigger classes, each with a published consumption rule:

  1. access delayed or narrowed after contract;
  2. an unsupported source type the client still needs interpreted;
  3. obligation density materially above the band's assumption;
  4. security or privacy review cycles stalling the sensors;
  5. a buyer decision-maker unavailable beyond a stated window.

The counter sits on the same surface as disposition consumption, visible to the buyer from week one — not in a supplier-side spreadsheet that gets shared when it becomes convenient.

Exhaustion behaviour agreed at signature: a named owner on each side, a five-business-day window, a delta with evidence, four fixed options, and a stated default if nobody decides. Chapter 16 runs it.

The acceptance design, with the oracle named at signature

This is where Chapter 13 stops being doctrine.

The promise: an obligations inventory with typed coverage against a declared repository boundary, an exposure map, and a recommended disposition per material obligation.

The declared acceptance surface: coverage percentage against the declared boundary with every gap typed; reconciliation of extracted obligations against the client's own register within a stated tolerance; named human disposition on every material obligation.

The independent oracle — and two of the four cheap forms were used deliberately, because they fail differently:

  • Buyer-selected sample. Thirty contracts drawn by the client's counsel after generation, from the declared population, scored in front of them in one session.
  • Authoritative reconciliation. Extracted renewal and termination dates checked against the client's finance system, which neither party authored.

And what is deliberately not the oracle, stated in the schedule so it cannot drift: our own extraction confidence scores, and a second model agreeing with the first.

The self-test, run

Chapter 9's eight rows, marked for this opportunity. Six clean, two marginal — and the two marginal ones are the interesting part.

Field 4 — buyer-controlled dependencies

Nameable, and datable for only two of the three owners. The third — legal sign-off — could be named but not committed to a date.

Remedy: it became a reserve trigger with a stated window rather than an assumption. That decision is why Event 7 in the next chapter has a rule to point at.

Field 6 — consequence class

The client wanted advice they could rely on for a regulatory filing. That tail could not be described in a sentence, and therefore could not be capped in a way either party would sign.

Remedy: the promise was shrunk. We assess and type; we do not opine on regulatory adequacy. That exclusion sits in the front of the schedule, not the back.

Record what the shrink cost, because a specimen that hides that is not honest: it removed the most valuable-sounding thing in the original conversation, and the fee reflected it. It also improved the engagement — but that only becomes visible in Chapter 16, and claiming it here would be hindsight wearing a lesson's clothes.

The engagement is now underwritten. Everything above was written before anybody did any work.

Then it ran, and sixteen things happened.

15
Part V: One Engagement, Underwritten

The Event Log, Classified

Sixteen events, eight weeks, every one classified — including the four that were genuinely arguable, and the one the schedule got wrong.

Here is the whole log, in the order things happened, with no editorial sorting.

Before the table, one thing worth saying about what makes it worth reading. The interesting part is not the twelve events that were obvious — those are just delivery. It is the four that could credibly have gone two ways, and the one that was classified correctly by the letter of the rule and wrongly by the system.

The complete engagement event log, classified into four variance classes
# Wk Event Class Rule applied Consequence
11Repository access granted three days after the agreed date2 → reserveTrigger 1: access delayed after contractDraw 1 of 5
21A contract family arrives in a format the sensor set did not anticipate; adapter written1Shape is methodNone
31First extraction pass mis-types renewal clauses; whole pass regenerated1Our own error is never a customer changeNone
42Contract count lands well above the census midpoint, inside the band ceiling1Volume within a declared class, inside the bandNone — the miss
52A second repository is discovered, holding the same document classesdisputed → 3The band ceiling is a recorded perimeter valueBand uplift, re-contract
63Two authoritative sources disagree on the current amendment; four reconciliation passes1 + 2Passes are interior; the call is a dispositionDisposition consumed
73Client counsel unavailable for nine daysdisputed → 2Trigger 5, not generosityDraw 2 of 5
84A scanned-fax contract family — unsupported source — that the client still needs interpreted2 → reserveTrigger 2Draw 3 of 5
94Clause-classification behaviour shifts on one family after a minor model update inside the pinned windowdisputed → 1We chose the dependency and could detect itNone — register row updated
105Material obligations running well above the band's density assumption2 → reserveTrigger 3Draw 4 of 5
115Client asks to add two subsidiaries' contracts3Authoritative input estate movedRe-contract, band uplift
126Buyer-selected sample: three of thirty fail reconciliation on effective-date logic1Our defect; the oracle workedRemediation absorbed
136A regulator changes an admissible interpretation affecting one counterparty family4Named exclusion; typed terminal stateBlocked pending external determination
147The client's risk committee sends a new obligation category to be assesseddisputed → 3A category is the promised state, not volumeRe-contract
157One repository restricted to metadata only after a security review2 → reserveTrigger 4Draw 5 of 5 — band spent
167A second security review stalls the sensors on the final repositoryexhaustionReserve spentProtocol runs — Chapter 16

The composition, and why it is the right result

Nine of sixteen events resolved to Class 1 and produced nothing. No request, no note, no log entry against the buyer.

A reader's instinct at this point is that nine invisible events sounds like a supplier absorbing too much. It is the opposite. A hundred-row change log is not governance; it is a record of a hundred arguments, and its length is a symptom rather than an achievement.

What the buyer actually experienced across eight weeks: five reserve draws they could watch on a counter, two re-contracts with a delta and four options each, and one exclusion that became a named gap in the decision pack. Everything else was invisible to them — which is precisely what the fixed price bought.

The disputed calls

Four events could credibly have gone the other way. Each one gets both arguments before the rule, because a rule that has only ever been shown settling easy cases has not been shown doing anything.

Dispute 1 — Event 5: the second repository

The case for interior. It holds the same document classes we already census. It is volume, and volume within a declared class is ours. That was the delivery lead's position and it is not a silly one — it is a direct application of a rule we had written down.

The case for the perimeter. The recorded value on the authoritative-input-estate field named the repositories. Volume is a band question until it crosses the band ceiling — and the band ceiling is itself a recorded perimeter value, no different in kind from the promised state or the end date.

The rule that settles it. Shape is method; class is perimeter; and the band ceiling is a recorded value like any other. An adapter is a Tuesday. A new class of source is a conversation. A repository of the same class that breaks the ceiling is a band conversation — which is a re-contract taking the smallest of the four options.

What made it feel wrong. The client had not hidden the repository. Nobody had behaved badly. The instinct in the room was that re-contracting over an honest omission was punitive — which is exactly the instinct a recorded value exists to overrule, and exactly why the value is recorded before anybody has a position to defend.

Dispute 2 — Event 7: nine days of counsel unavailability

The case for absorbing it. Nine days is not much. The relationship is good. Raising it feels petty, and there is a real chance it damages goodwill worth more than the draw.

The case for the reserve. Buyer delay is a named trigger with a stated window, and it consumes calendar — which is where the scarce humans actually get spent. Absorbing it teaches the buyer that dates are decorative, and this is the fence most suppliers breach in the generous direction.

The rule. Buyer delay draws the reserve first, then the perimeter. It is a published rule rather than a judgement call precisely so that it does not have to be made at ten o'clock at night by whoever is least able to refuse.

What the draw actually bought. A one-paragraph note naming the trigger. And then something nobody predicted: the client's own operations lead used that note internally to get counsel's time released faster on the next two requests. Visible absorption produced a behaviour change on the buyer's side. Silent absorption would have produced nothing except a slightly later delivery and a slightly thinner margin.

Dispute 3 — Event 9: the model that moved under us

The most important dispute in the chapter, because it is where Part III's argument meets its own honesty clause.

The case for Class 4. The vendor changed behaviour. We did not ask for it, we did not cause it, and a model update is exactly the shared-dependency event the last two chapters described at length. There is a clause available and it would work.

The case for Class 1. We chose to depend on that model. The change happened inside a window we had pinned. Our own evaluation set detected it in two days — which is proof that it was detectable. Chapter 11's test: did we choose to depend on this, and could we have detected it? Two yeses.

The rule. A model regression inside a version we chose to run is interior variation with an unusually wide blast radius. It is not an external tail. A supplier who blurs that line has found a way to make the buyer pay for its own architecture choices.

The uncomfortable part, recorded. That call cost two days of re-running, absorbed, on an engagement whose margin was already carrying a band decision made at the ceiling. It is still the right call. The register row was updated — pin window narrowed, evaluation set extended with the clause family that moved — which is the cost being converted into a control rather than into an invoice.

Dispute 4 — Event 14: the new obligation category

The case for the reserve. Requirement influx beyond the declared package is a named trigger in plenty of envelopes, and the volume here is small — one category, a few dozen obligations.

The case for the perimeter. A category is not volume. Field 1 records the promised state — the obligation categories the assessment covers. Adding one moves the recorded value. And because the acceptance rule in field 7 is downstream of field 1, two rows move at once.

The rule. Influx within a declared category is a reserve question. A new category is the promised state. Volume moves inside a boundary; categories move the boundary.

Why it mattered commercially. The reserve had one unit left. Routing a perimeter change into it would have exhausted the reserve on the wrong kind of event and left nothing for Events 15 and 16 — which is how a correctly designed reserve gets blamed for what was actually a classification error.

The one the schedule got wrong

A framework that has only ever been applied to clean cases has demonstrated its author's imagination and nothing else. So here is the miss, and it is more instructive than any of the four disputes.

Event 4. Contract count landed well above the census midpoint, inside the band ceiling. Classified interior. Correct, by the letter of the rule, and I would classify it the same way again in isolation.

What the loss review found three weeks later: the band's obligation-density assumption was keyed to contract count. So the same volume movement that was correctly absorbed under one rule silently invalidated the assumption behind another — and surfaced in week five as Event 10, which drew the reserve.

The trap, named precisely

The schedule treated two dependent variables as independent. Volume was absorbed as interior. Density was reserved as a trigger. Nobody had written down that the second was a function of the first.

Check your triggers for dependency. A schedule whose classes are individually correct can still be wrong as a system, and it will fail silently, three weeks downstream of the event that caused it.

What it cost. One reserve unit that should never have been a reserve unit. It should have been a band-uplift conversation in week two — when it would have been cheap, unsurprising and entirely uncontroversial, because the census delta was sitting right there — rather than a density trigger in week five, by which point the reserve was already half spent and the conversation had a different temperature.

The fix, and note that it is a product change rather than a delivery change: the density trigger is now expressed per hundred contracts rather than in absolute terms, so volume movement inside the band cannot silently consume the reserve. That is the first entry in the variance model, and it arrived from a mistake rather than from a workshop.

The classification card

Seven rules, extracted from the log. This is the artefact a delivery lead puts on a wall — and the test of whether it is any good is whether somebody who was not on this engagement could apply it to theirs.

  1. Default to interior. Nobody has to be brave.
  2. Interior is a residual, not a judgement. It is what remains when every named field has failed to move. You cannot classify something as interior by deciding to.
  3. A reserve draw requires a named trigger class, not a feeling. If no trigger fires, it is interior.
  4. A boundary mutation requires a named field with a recorded starting value and a measurable delta. Name the field out loud, or it is not one.
  5. Class 4 requires a shared cause with a blast radius beyond this engagement — and it is only external if you had pinned, declared and monitored the dependency first.
  6. Shape is method; class is perimeter. Volume is a band question; the band ceiling is a perimeter value.
  7. Categories move the boundary; volume moves inside it.

The two-person test

Hand the log to two people who did not deliver the engagement, and have them classify independently.

Agreement means the rules are specified. Disagreement names the row that is underspecified — which is a finding rather than a failure, and it is the cheapest way to find a gap in a variance schedule that exists.

On this specimen, two people would split on Events 5, 7, 9 and 14. Which is precisely why those four have written rules and the other twelve do not need them. A variance schedule does not need a rule for everything. It needs a rule everywhere two reasonable people would disagree, and the log is how you find out where those places are.

The absorption total

What Class 1 consumed on this engagement, valued at our own internal cost: the regenerated extraction pass, the adapter, four reconciliation passes, the two-day model re-run, the effective-date remediation after the buyer's sample, and the volume overrun.

I am going to give you the method rather than my figure, and the reason is not coyness. My number describes my cost base on one composite engagement; yours describes your business. Sum the internal cost of every Class 1 event. That is what your firm contributed as unrecorded risk capital on one engagement. Multiply by your engagement count for the year.

The point is not that the number will be shocking. It might be entirely reasonable — absorbing interior variation is what the margin is for, and a healthy number here is evidence the model is working. The point is that almost no firm has ever computed it, and until it exists, every conversation about absorption is a conversation about feelings.

Silent absorption produces no evidence, buys no credit, and cannot be sold.

What the log made visible that no timesheet would have

One finding, and it is the most valuable thing this engagement produced.

Contract count predicted extraction effort — which was nearly free. It did not predict disposition load, which was the entire cost. The ambiguity rate did better, and it had been sitting near its band ceiling from day one, flagged and then not acted on.

A timesheet would have shown the hours. It would have shown that the engagement was tighter than modelled, and somebody would have written a sentence about a demanding client. Only the classified log shows which assumption was wrong — and that is a different kind of information, because it changes the next price rather than the next debrief.

That observation is the first substantive entry in the variance model. It is also, as Chapter 18 will show, an observation we recorded and did not act on — which turns out to have a cost of its own.

First, though: the counter reached five, and the most predictable event in a bounded engagement arrived exactly on schedule.

16
Part V: One Engagement, Underwritten

The Reserve, Drawn

Four of five, in week seven, on a surface the buyer had been able to see since week one. Nobody was surprised by the arithmetic, because the arithmetic had been public for two months.

Reserve exhaustion is the most predictable event in a bounded engagement. It has a counter on it. Everybody watched it approach for three weeks.

Which makes the way most engagements handle it genuinely strange. It is the one commercial event in delivery that arrives with a public countdown attached, and it is routinely handled as though it were a shock.

The draws

Each one: which trigger fired, in which week, what it consumed under the published rule, and — the column that decides everything downstream — what the buyer received at the time.

The complete reserve consumption history for the specimen engagement
Draw Wk Event Trigger What the buyer received
1 of 51Repository access three days lateAccess delayed after contractOne-paragraph note naming the trigger; counter moves
2 of 53Counsel unavailable nine daysDecision-maker beyond stated windowNote; counter; a revised disposition schedule
3 of 54Scanned-fax family, unsupported, still neededUnsupported source the client needs interpretedNote; counter; the alternative offered — and declined
4 of 55Obligation density above band assumptionDensity above band assumptionNote; counter; explicit warning that one unit would remain
5 of 57Repository restricted to metadata after security reviewSecurity review stalling sensorsNote; counter reads spent; exhaustion protocol armed

Two details in that table are doing more work than the rest.

Draw 3. The buyer was offered the exclusion — we can type this family as unsupported and exclude it from this phase — and chose the draw. That is what a priced option looks like in practice. A buyer who can choose between "we exclude this" and "this costs you a unit of your reserve" is participating in the underwriting rather than receiving it, and the relationship afterwards is different in a way that survives into the next engagement.

Draw 4. The note included a sentence that made exhaustion survivable three weeks later: one unit remains, and here is what would consume it. Nine words of forecasting, sent at a moment when there was no pressure, that converted the week-seven conversation from an announcement into a continuation.

The counter, week by week

1 / 5 Week 1 — access three days late

2 / 5 Week 3 — counsel unavailable nine days

3 / 5 Week 4 — unsupported source, client's choice

4 / 5 Week 5 — density above assumption

Spent Week 7 — access narrowed after security review

Why the public counter is structural rather than polite

Two protections come out of one artefact, and neither of them depends on anybody's good faith.

A buyer who has watched the counter move cannot be surprised by exhaustion. A supplier who has shown the counter cannot credibly be accused of manufacturing it.

The general form of the argument is already published and it is worth restating in one line: showing the meter is honest rather than awkward, because the meter is the thing the band was priced against.

And the failure a hidden counter guarantees is worth being precise about, because it is not the one people expect. It is not that the buyer objects to the money. It is that they object to learning about the mechanism and its consumption at the same moment — which reads, in every case, as a mechanism invented to justify the consumption.

The two failure modes, walked on this engagement

Silent absorption

We eat Draw 3's scanned-fax family. Then Draw 4's density overrun. Margin erodes where nobody is looking, and each individual decision is defensible and kind.

The buyer never learns the reserve meant anything, because its consumption was never visible. And when we finally have to raise Event 16 — because eventually we must — there are three weeks of precedent demonstrating that we did not need to. Every earlier act of generosity has become evidence against us.

The ambush

A change request lands in week seven with no warning, citing a schedule the buyer has not read since signature. The commercial position may be entirely correct — the clause is there, the trigger fired, the arithmetic is right.

The relationship damage happens anyway, because from the buyer's side an unannounced invoice and an opportunistic one are indistinguishable.

The fix for both is neither goodwill nor better relationship management. It is a protocol agreed at signature, when nobody is under pressure and neither party has a position to defend.

Exhaustion, handled

The protocol is published and owned elsewhere, so: counter published from day one; a named owner on each side; a five-business-day window; a delta with evidence; four fixed options — uplift the band, extend the reserve at a published rate, narrow the boundary, or stop with work to date delivered in its typed state; and a stated default if nobody decides.

What ran here. Event 16 arrived on a Wednesday in week seven. The named owners met on the Monday. The delta showed what was recorded at signature, what was now true, what remained of the schedule under each of the four options, and — the part that made the meeting short — what each option would do to the usefulness of the decision pack itself.

What this chapter adds: exhaustion is loss development

That is a reading rather than a repeat, and it changes what the four options are.

In an insurer's language, this is the moment incurred experience is compared against the provision and the provision is found short. "The true value of the liability for losses or loss adjustment expenses at any accounting date can be known only when all attendant claims have been settled." 35

Which means the decision exhaustion forces is a reserve-strengthening decision, and it has exactly the same three answers an insurer has: put more capital behind it, narrow what is covered, or stop writing it. The fourth option — extend the reserve at a published rate — is the professional-services version of charging for the strengthening rather than absorbing it.

Reading the four options that way does something useful in the room. It stops the conversation being about whether the supplier is being reasonable, and makes it about which of three structurally different commercial positions both parties want to be in for the remaining fortnight.

What was decided

Narrow the boundary. The final repository was typed inaccessible within audit boundary and excluded from this phase.

Then the detail that is easy to miss and is the most interesting thing in this chapter.

The excluded repository appeared in the decision pack as a named gap with a named consequence for one of the three recommended dispositions. Not a footnote and not a caveat — a row that said, in effect: this option depends on obligations we did not see, and here is what would change if that population contradicted our assumption.

The exclusion improved the deliverable

A decision pack that says what it could not see is a better instrument than one implying it saw everything. The board reading it now knows which of its three options is sensitive to an unexamined population, and can decide whether to fund looking or accept the exposure.

This is what typed uncertainty is for. Not observed is a deliverable state. It is only embarrassing in a product that promised omniscience.

What each draw left behind

This is the row that connects Part V to Part VI, and it is the difference between a reserve that is a leak and a reserve that is an investment.

Every draw should deposit four things: a typed exception class with a detection condition; a deterministic test that flags the same signature automatically next time; a decision rule stating how it resolves; and a routing trigger naming who looks at it and at what seniority.

Without those four, a day of senior time bought one answer. With them, it bought a class.

What this engagement's five draws actually deposited:

  • Draw 1 → access provisioning became a dated buyer-side dependency with a named owner in the schedule, rather than an assumption in a paragraph.
  • Draw 2 → decision-maker availability became a stated window with a defined consequence, which is why Event 7 had a rule rather than an argument.
  • Draw 3 → the scanned-fax family became a detection condition at census time. On the next engagement it is a flag at the door, not a surprise in week four.
  • Draw 4 → the per-hundred-contracts density trigger from Chapter 15's miss.
  • Draw 5 → security review became a named class with an escalation route and a seniority attached, because it had now happened twice on one engagement.

And the commercial consequence this book adds on top: a class can become a band driver in the next contract. That is how a reserve shrinks across engagements rather than being re-drawn forever — and it is the difference between a firm that gets better at bounding work and one that simply gets more practised at absorbing it.

"Won't clients demand the unconsumed reserve back?"

Answer with the published rule rather than a negotiation: what is not drawn is not consumed.

It survives contact because the reserve is not a contingency held against the buyer's money. It is the priced home for an absorption obligation the supplier accepted, and its unconsumed portion is what the buyer paid for and did not need — in the same way an unclaimed policy is not refunded at the end of the year.

But there is an honest limit on that answer, and a supplier who ignores it is being lazy rather than principled. If the reserve is never drawn across many engagements, it is mispriced. The correct response is to reduce it in the product definition, not to defend it deal by deal against buyers who have noticed. That is a loss-ratio finding, and it belongs in Chapter 20 rather than in a negotiation.

Why visible absorption beats silent absorption

Make this argument without any appeal to virtue, because the commercial version is much stronger.

The buyer who watched the counter move from one to five experienced the week-seven decision as governance. A buyer seeing the counter for the first time at exhaustion experiences it as a bill. Same money. Same events. Entirely different relationship — because surprise had a language and a schedule.

And the second-order return, which is the one most suppliers throw away without noticing: a supplier who can show five absorbed exceptions and one governed decision has evidence of absorption. It goes in front of the buyer at renewal. It changes what they believe the fixed price was buying, and it does so with a document rather than an assertion.

Same money, same events, entirely different relationship — because surprise had a language and a schedule.

Key takeaways

  • Exhaustion is the most predictable event in a bounded engagement. Handle it as a scheduled event, because it is one.
  • The public counter protects both parties, and neither protection depends on good faith.
  • Read exhaustion as loss development: it forces a reserve-strengthening decision with three real answers.
  • Every draw deposits a class, a test, a rule and a route — or it bought one answer instead of a capability.
  • A reserve never drawn is mispriced. Fix the product definition, do not defend it deal by deal.

Five draws, one exhaustion, one narrowed boundary — all of it inside a single engagement, all of it absorbable by a single reserve. Which raises a question that reserve was never designed to answer: what happens when the event is not inside any engagement at all?

17
Part V: One Engagement, Underwritten

The Week the Model Changed

Three ordinary things happen in the same week, and every reserve in the book is drawn at once. The arithmetic is the point.

A Tuesday. Three events, none of them dramatic, none of them anybody's fault.

  • The successor to a pinned model is released, and the pinned version enters a deprecation window.
  • A document-source connector is deprecated on ninety days' notice.
  • A shared parser starts handling a common clause family differently after an upstream library update.

This is a simulation, and it is labelled as one

I have no correlated-tail loss event of my own to report. The regulators supply the mechanism; nobody supplies the magnitude, and Chapter 10 said so.

The purpose of what follows is to expose an arithmetic relationship, not to assert a frequency. If it reads as a prediction, it has been written badly.

What makes it worth running rather than speculative: the inputs are the register from Chapter 11 and the reserve design from Chapter 14, both of which are real artefacts you can build. The scenario is constructed. The structure it runs against is not.

And one thing the simulation cannot do, stated before it starts: it cannot tell you how often this happens. That number does not exist and I am not going to invent one. What it can tell you is what happens to your book when it does — which is a question you can answer this week, about your own stack, and should.

The setup

Nine live engagements of the same offer, at different weeks of their eight-week clocks. Each one carries its own reserve, sized against independent exception classes, following every rule in Part II.

All nine run on the same kernel: one extraction harness, one clause-classification model, one connector library, one evaluation set, one adapter catalogue.

Every one of those nine reserves was sized correctly for the engagement it sits in. That is the point of the exercise. The design is not sloppy. It is sound at the wrong altitude.

The arithmetic, in stages

1. What a per-engagement reserve is sized for

The exception rate within an engagement, over its own duration. Five units against five trigger classes, calibrated on the population of one engagement's surprises. Everything about that sizing assumes the engagement is the unit of exposure.

2. What a correlated event does

It draws against every reserve in the portfolio in the same week — not because the events are similar, but because they are the same event arriving nine times.

3. Why the sum is the number nobody has

The exposure is not the largest single reserve. It is the sum across the book, arriving at once. Every firm knows its largest reserve. Almost none has ever added them up and asked what event could draw all of them.

And three things make it materially worse than a single large loss of the same total size.

No time to re-band. Re-banding is a per-engagement conversation with a five-business-day window and a named owner on each side. Nine of them simultaneously is not a process. It is an outage of your commercial function.

No time to re-contract. The four options each require a delta with evidence, prepared by somebody senior enough to defend it. Nine parallel re-contracts consume exactly the senior attention that is also required to fix the technical cause — and those are the same three people.

Every delivery lead discovers it at the same moment. So the firm's escalation path saturates at precisely the point it is load-bearing, and the queue is served in the order people shouted rather than in the order of exposure.

The shape of the total, stated as a shape because that is what I have: reserve exposure scales with engagement count, while reserve capacity was set per engagement. So the ratio between what you can absorb and what can be drawn degrades linearly as the book grows on shared machinery.

The exposure is not your largest reserve. It is all of them, on the same Tuesday.

Why the standard defences fail

"We hold a reserve." The reserve was sized against a distribution that assumed independence. It is not that it is too small; it is that it is the wrong instrument, in the same way that a fire extinguisher is the wrong instrument for a flood rather than an undersized one.

"We have a diversified portfolio." Worse than not an answer. Portfolio pricing diversifies only independent exceptions. Nine engagements on one kernel are, for this class of event, one engagement with nine invoices.

Which is the inversion from Chapter 10, now with arithmetic attached: the more engagements you run on the same machinery, the worse the concentration gets. Growth increases the exposure it was supposed to spread.

And, once more, because a reader could reasonably conclude the wrong thing here: this is not an argument against shared machinery. The shared kernel is why engagement twenty is cheaper than engagement two, and Part VI argues for it at length. It is an argument that the kernel has a cost line, and that the controls are that line.

What each control would have changed

What each underwriting control changes on the simulated Tuesday
Control What it changes on this Tuesday The variable it moves
Version pinningThe successor release is a scheduled migration rather than an arrival. The deprecation window becomes our project plan.Converts an arrival into a date we chose
Independent evalsThe parser's clause-family shift is caught on the eval run, before it reaches any deliverable.Detection time
RollbackThe parser reverts to the known-good version while the fix is built; delivery continues on all nine.Loss duration
Shared-component regressionThe connector deprecation surfaces as a failing test across all nine dependency paths at once, rather than as nine separate discoveries over three weeks.Tests the correlation itself

Note what the table shows about budget politics: three of the four controls protect engagements. Only the fourth protects the portfolio — and it is the one that always loses the budget argument, because its benefit is the thing that did not happen to eight clients nobody in the room was thinking about.

Detection time is the master variable

Walk it, because this is the sentence that should change a spending decision.

At two days' detection — an evaluation set running on a cadence — the parser change affects work in progress on two engagements and is repaired before anything is delivered. The cost is engineering time and a small amount of rework. Nobody outside the firm ever knows.

At three weeks' detection — which is what "a client would tell us" means in practice — it has reached deliverables on all nine. And now the cost is not repair. It is re-delivery plus credibility, across your entire live book, in the same fortnight, with nine buyers comparing notes at whatever industry event happens next.

Same defect. Same fix. Two orders of consequence apart, and the only variable that moved was how long it took to notice.

The register, filled in

What a reader's own first pass looks like — including the rows that are uncomfortable, because the comfortable rows are not why you build it.

The accumulation register, filled in for the specimen portfolio
Component Live engagements Changes without asking us What we would observe Detection Rollback
Clause-classification model (pinned minor)9Vendor minor updates inside the pin windowEval score drop on specific clause families2 daysYes — prior pin kept warm
Extraction harness (ours)9Only usCI failureImmediateYes
Document-source connector (vendor)7Deprecation notice; auth changesIngest failures, partial estatesHours if monitored; weeks if notPartial — no supported prior version
Clause parser (upstream library)9Upstream release, transitivelySilent reclassification — nothing failsOnly via eval setYes, if version-locked
Client finance systems (reconciliation source)4Client change windowsReconciliation tolerance breachesDaysN/A — buyer-side dependency

The parser row is the one that should worry you. Nothing fails. No error appears. No monitor goes red. The output is well-formed, plausible, and wrong — on every engagement — until something outside the system notices. That is the shape of the expensive version of this event, and it is invisible to every instrument a firm normally runs.

And note where the register does its real work. Not in the "changes without asking us" column, which most teams can fill in from memory in twenty minutes. In detection — which is the column most firms cannot complete at all, and completing it is the finding.

The external anchor for the shape

Both public demonstrations share the same root-cause shape, and it is the shape rather than the scale that transfers. From Amazon's own summary of the October 2025 event: "The incident was triggered by a latent defect within the service's automated DNS management system." 19 And CrowdStrike reached approximately 8.5 million devices from a single update, which the insurance market classified as an accumulation loss event rather than an outage.

Routine change. Shared dependency. Simultaneous downstream failure. A detection path that ran through the customers. Change the nouns and it is the parser row above.

The commercial conclusion

Class 4 does not belong in the reserve

A reserve sized for independent exceptions cannot absorb a simultaneous draw. Putting correlated risk there means the reserve is doing a job it was never priced for, and it will fail on the one occasion it is most needed.

It belongs in exactly three places:

  • Controls, priced into the band. Visible, recoverable, and the subject of Chapter 11.
  • Exclusions, stated at signature, producing a typed terminal state rather than a failure.
  • Capital, acknowledged on the balance sheet rather than assumed away — which is why Chapter 5 has a fourth fee term.

A firm that routes correlated tails into the reserve has mispriced its whole book, and will find out on one Tuesday.

Where the line fell, and why it was uncomfortable

Run Chapter 11's test across the three events in the simulated week, and the result is not the one a supplier would choose.

The connector deprecation is genuinely external. A third-party decision on a third-party timetable, declared as an exclusion at signature, producing a typed state. One of three.

The model successor is not. We chose the dependency, we pinned it, and the deprecation window was published. That is a migration we owe ourselves and it belongs on our plan and in our operating cost.

The parser change is not either, by exactly the same test — and it is the one that would cost the most. Which is the discomfort in one line: the most expensive event in the simulated week is one we own.

Say plainly why the line is drawn there anyway. A supplier who classifies its own dependency choices as external tails has invented a mechanism for making buyers fund its architecture. The clause would be unarguable, indefensible, and it would work exactly once — after which every fixed price that firm quotes carries a discount for the possibility of the same manoeuvre.

Key takeaways

  • Nine correctly sized reserves are still the wrong instrument for one event arriving nine times.
  • Reserve exposure scales with engagement count; reserve capacity was set per engagement. The ratio degrades as you grow.
  • Detection time is the master variable — two days versus three weeks is repair versus re-delivery across the whole book.
  • The row where nothing fails is the expensive one. Well-formed and wrong beats an error every time.
  • Class 4 has three homes: controls, exclusions and capital. The reserve is not one of them.

The machinery has now been shown working and shown under stress. What it has not been shown doing is refusing — and a rule that has never refused anything has not yet been tested.

18
Part V: One Engagement, Underwritten

The One We Did Not Write

A second engagement that looked like the first one and was not — the census variable that turned out to be decoration, and what replaced the fixed price.

Same offer. Similar-looking estate. Band assigned by the same drivers, from the same census, at the same point in the sales cycle. Everything that had worked once was applied again, correctly.

Then the effort departed from the band in a direction the census could not see.

The departure

What the census said: contract count mid-band. Format mix clean — almost entirely native text, so extraction was going to be nearly free. Ambiguity rate comfortably inside M's range, and noticeably better than the first engagement's. Two repositories, three named buyer-side owners with dates against all three.

On paper, an easier engagement than the one in Chapter 15.

What happened: disposition load ran far above the included count from week two, and kept running. Not a spike — a level shift. Every week, consistently, more material obligations requiring a named human call than the band had priced.

And the reserve draws did not fit any trigger cleanly, because nothing had gone wrong. Access was fine. Nobody was late. No unsupported format appeared. No security review stalled anything. The work was simply costing more per unit than the band assumed.

Which is the most dangerous shape a departure can take, and it deserves naming: no event fired. There was nothing to classify, nothing to escalate, and nothing to tell the buyer. A variance schedule handles surprises. It does not handle a wrong price — and a delivery team six weeks into an engagement is structurally the last group in the firm to notice the difference.

The diagnosis, run properly

Four possible causes. Each has a different remedy, and only one of them falsifies the offer. Run them in order rather than jumping to the interesting answer, because the interesting answer is the one you want it to be.

Four diagnoses for a band that failed to predict cost
Diagnosis What it would look like Remedy
Measurement bugThe census measured something other than what it claimed — a counting error, a mis-scoped repositoryFix the sensor. Re-census. Not a pricing problem
Access lieThe estate was not what was warranted at signatureA perimeter field moved; re-contract. A Class 3 event with late detection
Band design faultThe driver does not predict cost for this populationProduct change: the driver set is wrong
Novel consequential judgementThe cost is in dispositions no machinery can reduce, for this class of workThe offer's territory is smaller than we thought. This is the one that falsifies

The eliminations, briefly, because the elimination is the method.

Measurement bug — ruled out. A recount matched the census within tolerance. The instrument measured what it said it measured.

Access lie — ruled out. The repositories were exactly as warranted. Nobody had concealed anything, and the estate was genuinely what it appeared to be.

Novel consequential judgementpartially true, and this is where an honest diagnosis gets uncomfortable rather than tidy. Some of the excess was irreducible legal judgement of a kind no census would ever have predicted or any machinery reduced. But not the level shift. A level shift is systematic, and irreducible novelty is not.

Band design fault — the answer.

What the driver missed

The true cost driver was not volume, and it was not the ambiguity rate either. It was clause heterogeneity — the proportion of contracts using bespoke drafting rather than a house template.

The mechanism, stated so it generalises past contracts: a bespoke clause cannot be classified by pattern. Every one becomes a disposition. Which means a house-template estate of ten thousand contracts is cheaper to assess than a bespoke estate of two thousand — and the census had been counting the wrong noun the entire time.

The census was counting the wrong noun.

And here Chapter 15's finding arrives as a prediction that came true. The classified log on the first engagement had already shown, in writing, that contract count predicted extraction effort — which was nearly free — and did not predict disposition load, which was the whole cost. That observation sat in the variance model for one engagement before it cost money.

Which produces the admission that makes this chapter worth including at all: we had the signal and did not act on it.

Not because anybody was careless. The finding was recorded, it was correct, and it was one line in a post-engagement review at a moment when the engagement had gone well and the next proposal was already in draft. A finding recorded is not a finding used, and the gap between those two verbs is where a variance model actually fails. Chapter 20's design rule — every row must change a decision — exists because of this.

Honouring the falsifier

This is the section that decides whether this book means anything.

What was not done: absorb it as culture. That option was available. It is what most firms do, it is defensible in the moment, and it produces one more half-comparable engagement that teaches nothing — while quietly training everybody in the firm that bands are aspirational.

What was done, in order:

  1. Re-census against a new candidate driver — bespoke-drafting share, measurable at the door from a sample rather than requiring the full estate.
  2. Re-price the driver set in the product definition, not in this proposal's footnote. Band drivers changed for everyone, dated, through the change board. That is the difference between fixing a product and rescuing a deal.
  3. Hybrid the engagement in front of us, because the client still had a real problem and the fixed price on the old drivers was no longer honest to hold.

The hybrid, drawn rather than described

"Use a hybrid" is useless advice unless the shape is on the page. Three commercial objects, each with its own acceptance:

Fixed price on the censusable component

Extraction, coverage, and a typed inventory against the declared boundary. Everything the census genuinely predicted stayed exactly where it was — which is step two of the shrink, applied rather than quoted.

Metered dispositions above a named count

At a published rate, with the counter visible from week one — the same surface as the reserve counter, for the same reason. The buyer can see the meter that the band was priced against.

The bespoke population, sold as its own bounded exercise

With its own acceptance. Which is selling the bounding of the part that could not be bounded — the same move as a decline, applied to a component rather than to an engagement.

What survives the change of fee shape is the thing that mattered: the stable unit of commitment. Fixed price is a strong signal, not a law. The buyer is still purchasing a bounded state with independent acceptance. Only the metering changed.

And the alternative that was rejected deserves a paragraph rather than a clause, because it is the one most firms take. Forcing the engagement into the existing band would have destroyed the band's meaning for every future buyer, taught our own sales system that drivers are negotiable, and converted a product back into a bespoke project with a product's price and a project's cost.

The decline that followed

Six weeks later, a third opportunity. Run Chapter 9's self-test live, and show the marks.

The quotability self-test run on a declined opportunity
# Field Mark Why
1Promised stateRedefined twice in three meetings: "are we exposed", then "what is our posture", then "what do we tell the board"
2Authoritative input estateWhich entity's contracts govern depends on a restructure that has not completed
3Volume / bandCensusable this week
4Buyer-controlled dependencies~Nameable, not datable — the owners are the people being restructured
5Authority and accessThe approving body will exist after the restructure; composition unknown
6Consequence / liability classWhich regulator applies is one of the things the restructure decides
7Acceptance ruleEntirely downstream of field 1
8Fixed time boundaryAvailable, and meaningless without the rest

Five of eight cannot carry a recorded value, and three of those five sit outside both parties' control. Note the dependency before counting: field 7 is downstream of field 1, so repairing one field repairs two rows — which is why the count alone is not the decision.

The chain terminates: no recorded value → no delta → no trigger → every surprise resolves into an argument. Not a hard engagement. An unclassifiable one.

What the buyer actually bought

The shrink: a smaller fixed commitment whose entire deliverable is a filled perimeter — which entity governs, which regulator applies, who approves, and what decision is actually being asked. Four rows, produced as a deliverable, with an acceptance test attached to each.

Why that was worth buying rather than a consolation prize: a group mid-restructure genuinely needs to know which entity governs before it can decide anything else. And the test of whether the bounding is a real product or a hold on the account is uncomfortable and worth applying — the filled perimeter makes the larger engagement quotable afterwards by anyone, including a competitor. If your bounding product only makes the work quotable by you, it is a hold on the account with a deliverable stapled to it.

The commercial outcome, recorded including the part that hurts. The shrink was a fraction of the revenue of the engagement that was asked for, in the same quarter, out of the same budget. Say so. A book that pretends refusal is costless is asking a reader to believe something they know is untrue, and they will discount everything else accordingly.

And the offsetting fact, stated as a fact rather than as consolation: the larger engagement came back seven weeks later with four of the five failed rows carrying recorded values — because the shrink had produced them.

What a decline is worth as data

Here is the structural contribution of this chapter, and it is the cheapest thing in the entire book to adopt.

A declination is not a lost deal. It is a data point about where the offer's territory ends, and it belongs in the variance model with the same status as a completed engagement.

Firms record wins. Firms record losses. Almost nobody records refusals with reasons — which is remarkable, because it costs one paragraph at a moment when the decision is already being made and the reasoning is already fully formed in somebody's head. The marginal cost is approximately zero and the marginal value compounds.

The declination entry — five fields

  1. The opportunity, described in the offer's own census vocabulary — so it is comparable to the ones you accepted.
  2. Which perimeter fields could not carry a recorded value.
  3. Which of those sat outside both parties' control.
  4. What was offered instead, and whether it was bought.
  5. The reopen condition — what would have to become true for this to be quotable.

Field 5 matters more than it looks. Without it, a declined opportunity gets re-litigated every time a partner remembers it. With it, the answer is a lookup.

And the aggregate value, which is the reason to keep the log rather than the individual entries. After a dozen entries, the declination log tells you the shape of the boundary of your own offer. That is information you cannot get from the engagements you won, because all of those are inside the boundary and none of them touch its edge.

The honest register to close Part V

What Part V has produced: one worked engagement, one simulation, one re-priced offer, and one refusal.

That is a specimen. It is not a portfolio, and the difference is not rhetorical.

It demonstrates that the machinery runs. That it classifies contested events with rules two independent people would apply the same way. That it survives its own miss and converts the miss into a product change. That it can be stressed at portfolio scale and produce a number. And that it can say no, at a cost, with the cost recorded.

It does not demonstrate band economics across a cohort, loss ratios by band, or that engagement twenty is safer than engagement two. Those require a population I do not have.

Chapter 22 states exactly what that does and does not prove, before anybody else does — which is the only defensible order to do it in.

First, though: the question Part V has been quietly begging. All of that machinery, all of that recording, all of that refusing — what is it for? The answer is an asset, and it only exists if somebody writes it down.

19
Part VI: The Loss History

Fixed Price Makes You the Residual Claimant

A delivery improvement takes an engagement from a hundred units of effort to eighty. Everybody agrees this is good. Almost nobody in the room can say who gets the twenty.

The answer is not determined by how good the improvement is, or by who built it, or by how the benefit is described in the internal announcement. It is determined entirely by a fee shape signed months earlier by somebody who was not thinking about this.

The two answers

Under time-and-materials

Work drops from a hundred to eighty. Billings drop by twenty. The client captures the gain — automatically, without asking, and usually without noticing.

The supplier funded the improvement and handed it over. Do that repeatedly and you have an investment programme whose entire return accrues to somebody else's balance sheet.

Under a fixed commitment

The client pays for the bounded promise. Machinery that lowers the next delivery's cost accrues to the supplier — as margin, as capacity, or as delivery speed.

Which turns the improvement from an act of generosity into an investment with a return, and changes what a rational firm should spend on it.

The causal chain, stated once and cleanly:

fixed customer commitment  →  supplier owns production variance  →  write-back lowers future cost or risk  →  supplier captures part of the improvement  →  investment in reusable machinery becomes rational
Fixed price does more than transfer variance. It gives the supplier an economic reason to turn delivery learning into infrastructure.

Why this makes it doctrine rather than observation

This is the reason the square and the flywheel are one business model rather than two compatible shapes that happen to appear in the same books.

The square assigns ownership of variance. The flywheel determines whether that ownership becomes profit or pain.

Own the variance without the machinery, and you have taken a position — a leveraged one, held through your own margin, with no mechanism for it to improve. Build the machinery without owning the variance, and you have improved somebody else's economics at your own expense.

Which reframes the investment argument entirely. Reusable machinery is not an efficiency programme and should stop being justified as one. It is the thing that makes your retained variance shrink over time — and under a fixed price, the shrinkage is yours.

Six conditions, each of which can fail

The effect is conditional, and the conditions are strict enough that most firms will fail at least one. State them as things a reader checks rather than as caveats.

  1. A credible cohort of sufficiently similar future engagements. Without it, the machinery is bespoke tooling with a reusable name.
  2. Recurring machinery outweighing client-specific variation. If every engagement forks the system, you have a codebase per client and a shared repository by accident.
  3. Verification costs not growing as fast as production breadth. Chapter 13's ratio, arriving here as a condition rather than a metric. If verification scales with output, the compression is consumed by checking it.
  4. Available reuse rights. Learning that cannot lawfully cross the client boundary is not an asset, however good it is.
  5. The resulting capability actually reaching future work. Built and shipped are different verbs. A kernel nobody loads is a library.
  6. Competitive repricing not immediately returning every gain to the market.

Note the asymmetry, because it decides how much of this you can manage. Conditions one to five are inside your control. Condition six is not — and it is the one most likely to be assumed away in a business case.

The denominator nobody computes

Not headcount. The right question is narrower and much less flattering: how many eligible future commercial units can actually consume this machinery?

A hundred-and-fifty-person bench creates potential distribution. It does not prove amortisation. If six people ever encounter the problem, or every engagement forks the system, the effective reuse population is six — or zero.

Make it operable by writing it down before the investment, as a named cohort with a named qualification rule. "Engagements in bands M and L of this offer over the next four quarters" is a denominator: it has a number, a date and a test. "Our practice" is not a denominator; it is a hope with a headcount.

The failure this prevents is common and expensive in a quiet way: a firm builds a genuinely excellent piece of machinery for a problem that occurs twice a year, and books the cost against a bench that will never touch it. Nothing about that is visible in a project review, because the machinery works.

The four ways the money moves

Capturing a gain and keeping it are different problems, and the second one is decided by behaviour outside your firm. A sibling book works this in full; the summary matters here because it tells you what fixed price is actually buying you.

The four capture conditions for a delivery compression
Outcome Who captures the saving Verdict
Unsold capacityThe customer, entirelyThe passive case — what happens when nobody decides anything
Fully sold capacityNobody — absorbed by volumeOld economics defended; unstable at equilibrium
Outcome pricingThe firm, conditionallyMargin expands if competitors hold and verification cost stays contained
Demand expansionThe firm, via a larger poolThe only outcome that creates a new reason to pay

The sharpest way to state this chapter's relationship to that argument: fixed price is how you get into the third column. The residual-claimant mechanism is precisely what moves a firm out of the passive default and into the branch where the gain is captured rather than donated.

And Chapter 13's verification ratio is what decides whether you stay there. Outcome pricing has two failure conditions, and verification cost is the one firms forget to model — which is why the ratio is a condition in the list above rather than a metric in an appendix.

One further bound, stated because it is honest rather than because it helps: even the third column only holds while competitors hold. It takes one firm with a thinner book to reprice the category, and no amount of internal discipline prevents that.

The adjacency, named and left

How first-engagement learning is funded, capped, ledgered and killed if it fails — as a pre-approved, expiring subsidy on a separate capability ledger, proved only when engagement two runs on ordinary staff — is a published treatment of its own, and this book does not re-open it.

The stricter bar, honoured rather than softened

Everything above makes an investment case. Before Chapter 20 turns that into an asset claim, there is a constraint I have to impose on myself, and it comes from my own corpus rather than from a critic.

"Compounding is not 'we captured some learnings'. It is a measurable change in the starting position of the next engagement — and one engagement cannot evidence it."

Which means the claim I am allowed to make in the next chapter is narrower than "our loss history is our moat". It has to say what engagement two starts with that engagement one did not have — named artefacts, in locations the next run actually reads.

And the falsifiers come with it, from the same source. Escalation returning at the same density with the same people. Exception classes that are still tribal knowledge rather than typed routes. Pricing rules that bend whenever a principal is in the room. A "playbook" that turns out to be a slide deck. Any two of those and the flywheel is a story.

The one it is easiest to fool yourself about deserves its own sentence: faster delivery caused by individual familiarity is not flywheel evidence. The same people getting better at the same work is a good thing, it feels exactly like compounding from inside the engagement, and it is not an asset — because it walks out of the building when they do.

Same P&L line, different asset position

A first engagement at break-even is rational if the loss is purchasing reusable capability, and irrational if it is purchasing heroics.

Two very good people working very hard, solving everything by hand, brilliantly, with a delighted client and the lessons shared over drinks — identical on the profit-and-loss statement to an engagement that left a delivery vessel, typed exception classes, corrected band drivers and a reusable acceptance harness. The difference is invisible to finance and total in every other respect.

Key takeaways

  • Under T&M the client captures your delivery improvement; under fixed price you do. That is the whole economic joint.
  • The square assigns ownership of variance; the flywheel decides whether that ownership becomes profit or pain.
  • Six conditions, five of them yours. Condition six — competitive repricing — is the one that gets assumed away.
  • Write the reuse cohort down as a named population with a qualification rule. A bench is not a denominator.
  • Faster delivery from familiarity is not compounding. Same P&L line, entirely different asset position.

The asset now has a name and a set of conditions. So what is actually inside it? A competitor can rent the same models and copy the offer by Tuesday. The next chapter is about what they cannot.

20
Part VI: The Loss History

The Variance Model

What a competitor can copy by Tuesday, what they cannot — and the professional obligation that turns a record into an asset.

Your offer. Your language. Your report format. Your deck. Your prompts. Your model — they rent it from the same three suppliers you do. A competent competitor can have a convincing version of your proposal by the end of the week, and if their designer is better than yours it will look nicer than the original.

What they cannot copy is a record of what happened when your machinery met these estates. That record has a name in another profession, and the name is not "lessons learned".

Seven rows

Each row: what it records, where the data comes from in an ordinary delivery record, and the decision it changes. That third column is the one that separates an instrument from a table.

The seven-row variance model
# What it records Where the data is Decision it changes
1Which census variables predicted effort — and which were decorationCensus output against actual disposition loadThe driver set, and therefore every future band assignment
2Which configurations produce exceptionsThe classified event logEligibility rules — what gets declined at the door
3Which bands stayed profitable after review, remediation and exception costBand economics reviewBand thresholds, and whether a band should exist at all
4Which acceptance tests caught real failures rather than ratifying themOracle results against post-acceptance findingsThe acceptance design; which oracle form fits this offer
5Which unknowns can safely terminate without human resolutionTyped-state frequencies by classThe included disposition count — the largest single cost lever
6Which client behaviours consume reserveReserve draw log, by client and triggerIndividual risk rating on repeat clients
7Which principal judgements can be compiled into the machineryEscalation log and its resolutionsThe transfer position — what engagement two does without you

Row 1 already has evidence attached in this book, which is why it is first. Contract count predicted extraction effort — nearly free — and did not predict disposition load, which was the whole cost. Chapter 15 found it in a classified log. Chapter 18 paid for not acting on it. One row, two engagements, one product change.

Row 5 is the one that is systematically under-appreciated, so walk it. Every unknown that can terminate in a typed state without a human is a disposition removed from the band's cost — and disposition count is the dominant term in the fee. Row 5 is also where a firm discovers that ambiguous — human decision required has quietly become a dumping ground for anything difficult, which is a product signal rather than a badge of professional complexity.

And the design rule for the whole model, which is the thing to hold onto when somebody proposes an eighth row: every row must change a decision. A row that records something interesting and changes nothing is a report. Reports do not compound.

The professional obligation

This is stated better in the actuarial literature than in any consulting methodology I have read, including my own.

Incurred results should be measured for reasonableness against relevant indicators "and expressed wherever possible in terms of frequencies, severities, and loss ratios. No material departure from expected results should be accepted without attempting to find an explanation for the variation." 35

Read that as a delivery discipline rather than an accounting one. Every engagement where actual effort departed materially from the band gets an explanation before it closes. Not a shrug. Not "that client was difficult". An explanation that changes a rule, a test, a driver or a band.

What makes this hard is not the analysis. It is the timing. The explanation has to be produced at the moment everybody wants to move on to the next engagement, by the people who are already late for it — which is why it belongs in a closing checklist with a name against it rather than in a culture.

Three borrowed words worth glossing once, because they are more precise than anything in the consulting vocabulary and most firms argue about all three without measuring any of them. Frequency: how often a class of surprise occurs. Severity: what it costs when it does. Loss ratio: what the class cost against what was priced for it.

Individual risk rating

"When an individual risk's experience is sufficiently credible, the premium for that risk should be modified to reflect the individual experience." 8

Applied: engagement two with the same client is priced from that client's own loss history rather than purely from the class. That is row six doing commercial work rather than sitting in a report.

Concretely, the behaviours that move a repeat client's band: access provisioning speed; decision-maker availability against the stated window; requirement stability; how often their own register turned out to disagree with their systems; and how many of their obligations landed in ambiguous — human decision required. All five are already in the draw log, and none of them requires a new instrument.

How to present it without it reading as punishment, because this is where most firms flinch and then quietly absorb the difference instead. Show it as configuration, exactly the way the band drivers were shown at first assignment: here is what we observed last time, here is the driver it moved, here is what it would take to move it back.

A client who can see the lever will often pull it — which is cheaper for both parties than a premium, and produces a better engagement for them as well as for you.

And the reciprocal, which is what makes it credible rather than extractive: a client whose behaviours were better than the class gets the benefit. If individual rating only ever moves in one direction, it is not a rating system. It is a price rise with a methodology attached, and buyers work that out inside two cycles.

Why a closed engagement is not a settled one

"The true value of the liability for losses or loss adjustment expenses at any accounting date can be known only when all attendant claims have been settled."

In delivery terms: an engagement's real cost is not known at the acceptance meeting. Remediation requests, warranty questions, the escalation that arrives in month three, the finding a client disputes in month five — that is loss development, and it lands after the file is closed and the team has moved on.

A systematic error, not a nuisance

A firm that closes its book at acceptance systematically understates its own costs — and does so consistently in the same direction, which means the bias never averages out across a population. Every band looks slightly more profitable than it is. The error compounds with volume.

The remedy is a review at a fixed lag, not at close. Ninety days is a defensible starting point: long enough to capture the remediation tail, short enough that the evidence is still recoverable. And the review is short — did anything arrive after acceptance, what did it cost, which row of the model does it change?

Note the connection to Chapter 16, because it makes the pattern visible. Reserve exhaustion was the first loss-development event, and it happened inside the engagement with a counter on it. This is the same event happening outside the engagement, without a counter, which is why it needs a scheduled date instead.

What engagement two starts with

The stricter bar from Chapter 19, answered specifically from the specimen rather than asserted. Named artefacts, in locations the next run actually reads:

  • Typed exception classes with detection conditions. The scanned-fax family is now detected at census rather than discovered in week four.
  • Deterministic tests for those signatures, so the same shape flags automatically rather than being recognised by somebody who happened to be on the last one.
  • Corrected band drivers — the per-hundred-contracts density trigger from Chapter 15's miss, and the bespoke-drafting share from Chapter 18.
  • A reusable acceptance harness with local cases swapped in rather than rebuilt.
  • Escalation classes encoded with owners and seniority, so a repeat escalation routes without a conversation about who should look at it.
  • A declination entry with a reopen condition, so a bad-fit opportunity is not re-litigated every time somebody remembers it.

That is a list you can audit. Which is the point — "we learned a lot" is not, and the difference between those two sentences is the entire subject of this chapter.

A library is a cost centre. A substrate is a cost driver.

A document dump does not compound. It may increase retrieval volume while faithfully preserving contradictions and obsolete practice. Compounding requires governed write-back:

engagement produces a correction  →  correction receives provenance and rights clearance  →  correction changes a rule, test, driver or reusable method  →  a later engagement invokes it  →  delivery cost, risk or expert intervention measurably falls

And the honest test at the end of every engagement, which takes one minute and is rarely asked: did we leave merely another codebase — or a sharper corpus that makes the next project easier and better?

The territories, in one paragraph

Client evidence, decisions and outputs stay inside the client boundary. Transferable method — generic tests, exception classes, pricing conditions, failure shapes — may cross, after rights clearance, abstraction and a named human approval. Automatic upward promotion is a breach rather than a feature.

It belongs in this chapter specifically because "we learn from every engagement" is one clause away from "we pool everybody's data", and buyers know it. A variance model that cannot say which territory each row lives in is a confidentiality incident waiting for a procurement question.

The comforting fact, and it is worth saying because the anxiety is disproportionate: almost everything in the seven rows is method, not client data. "Bespoke-drafting share predicts disposition load" is a fact about contract estates in general. It does not describe anybody's contracts.

The culture problem, without moralising

Traditional professional services rewards knowledge concentration, and rewards it correctly given how the incentives are set. If you are the only person who knows the obscure trick, your utilisation rises, your indispensability rises, and your promotion prospects rise. "Please document everything you know for the benefit of the organisation" is an invitation to reduce your own leverage, and asking politely has been tried for about thirty years.

The status move is from I solved the hard case to I solved the hard case once and made sure nobody needs me for that class again.

Which only works if the firm actually measures the second thing. That is what the escalation-rate and disposition-density rows are for, and it is Chapter 21's subject.

One more thing about why this failed before, because the usual diagnosis is a character diagnosis and it is wrong. The joins were economically broken. Nine links stood between a lesson and its reuse — notice it was reusable, articulate it, abstract it out of client-specific language, document it, classify it, distribute it, persuade busy people to read it, have one of them remember it at exactly the right future moment, apply it correctly — and every one of those links was a person's afternoon. A nine-link chain of voluntary human effort has a throughput close to zero.

AI shortens the first six and eliminates links seven and eight, which is the structural discontinuity rather than a speed improvement: nobody has to be persuaded to read anything, and nobody has to remember at the right moment, because the next engagement's machinery loads the material as a starting condition. Link nine — applying it correctly — is still judgement, which is why the barbell survives and why none of this is an automation story.

One adjacency, named and left

The ritual that harvests higher-order learning after an engagement closes — who runs it, what it perturbs, how the sessions are structured — is a sibling's subject. This book specifies what the variance model must record. Not the review that fills it.

Key takeaways

  • Seven rows, and every one must change a decision. A row that changes nothing is a report.
  • No material departure from expected results is accepted without an explanation — before the engagement closes.
  • Engagement two with the same client is priced from that client's history. Make the lever visible, and make it move both ways.
  • Closing the book at acceptance systematically understates cost in one direction. Review at a fixed lag.
  • The joins were economically broken, not culturally. AI eliminates the two links that never worked; link nine is still judgement.

The model says what to record. The next chapter says how you tell whether recording it is working — and the answer is eight numbers, six of which are a column in a file you already keep.

21
Part VI: The Loss History

The Eight Records

Is your fixed price a product, or is it being subsidised by heroics? Both look identical from outside, and they diverge at exactly the moment you try to scale.

Two firms. Same offer, same fee, same client satisfaction scores, same margin, same partner confidence. One of them has typed exception classes, corrected drivers and a reusable harness. The other has four extremely good people.

Nothing on a profit-and-loss statement distinguishes them. Nothing in a client reference distinguishes them. They diverge on exactly two occasions: when the firm tries to run six engagements instead of two, and when one of the four people resigns.

Eight records tell them apart. Most of them are a column.

The programme

The eight-record measurement programme, with cost to start recording
# Record Where the data is Improving looks like Cost to start
1Remove-AI classificationA judgement, once per offer, revisited annuallyNot a trend — a category you can defendOne conversation
2Variance ownershipThe event log you already keep, plus a class columnClass 1 share rising while total events fallA column
3Disposition densityDisposition log ÷ census outputFalls — eventually; see belowA column
4Verification scalingThe verification ratio from Chapter 13Ratio flat or falling as breadth risesA timesheet field
5Boundary integrityClass 3 count ÷ total surprisesFalls, and the remainder is cleanly evidencedFree — record 2, filtered
6Band economicsFinance, joined to band assignmentVariance within a band narrowingA join, plus stable bands
7TransferStaffing records against the escalation logScarce-expert dispositions per paid unit fallingDeliberate instrument
8Customer valueRenewal language; who is named in the contractPurchases that survive the originator's absenceDeliberate instrument

Say which of these are nearly free out loud, because a programme presented as uniformly expensive does not get started — and a programme that does not get started is worth precisely as much as no programme.

Six of the eight are a column in a file you already keep. The other two are a decision.

Populated across the two specimen engagements

Engagement one from Chapter 15, and engagement two — the re-priced one from Chapter 18. With honest asterisks, because a populated table with eight good results would be evidence of nothing except optimism.

1 — Remove-AI classification

Both engagements: AI-dependent, not AI-constituted. Say why plainly rather than flattering the offer. If the machine cognition vanished, the assessment could still be produced — narrower, slower, at a much worse margin, over a fraction of the estate — but the recognisable commercial promise would survive.

That matters commercially rather than taxonomically: an AI-dependent offer's defensibility rests on cost position rather than on existence, and cost positions are contestable by anyone renting the same models.

2 — Variance ownership

Engagement one: sixteen events — nine Class 1, five reserve draws, three Class 3 (two clean, one after argument), one Class 4.

Engagement two: fewer total events, and the composition shifted. Surprises that had been reserve draws the first time arrived pre-typed and were absorbed as interior — because the scanned-fax family was now a census-time detection and the buyer-side dependency had a stated window.

That shift is the flywheel, visible in a table. Not a story about learning — a change in a distribution.

3 — Disposition density

Rose. Between engagement one and engagement two, the number of consequential human decisions per hundred machine-processed units went up.

Interpret it rather than explaining it away — see the section below, because this is the number that kills measurement programmes.

4 — Verification scaling

Improved, and this is the clearest evidence of compounding in the pair. The acceptance harness from engagement one was reused with local cases swapped in rather than rebuilt, so verification effort fell in absolute terms while the estate was larger.

It is also the least celebrated result, because nobody notices a harness they did not have to build. Which is exactly why it needs a number rather than a feeling.

5 — Boundary integrity

Engagement one: three of sixteen surprises reached the perimeter, one only after an argument that took a written rule to settle. Engagement two: cleaner.

And the reason is uncomfortable rather than flattering. Boundary integrity improved because the pricing improved — corrected drivers meant fewer things were surprising — not because change control got better. Attributing it to the change control would have been the easy read and the wrong one.

6 — Band economics

Cannot yet be computed. Two engagements in one band is not a population, and reporting a margin figure off n=2 would be exactly the fabrication this book spends a chapter warning against.

What it would take: enough comparable units in a stable band that the variance means something — and, per Chapter 6, bands that have not been redefined in the meantime.

7 — Transfer

Partially observable and honestly weak. Engagement two ran with less originator involvement in extraction, adapter work and acceptance — and with the same originator involvement in the disposition of bespoke clauses, which was the expensive part.

The machinery transferred. The judgement did not. That is a real result and it is not a good one.

8 — Customer value

Ambiguous, which is the honest reading at n=2. The buyer bought the offer. The buyer also asked for a named person in the contract.

Both facts are real, and the second one is the more predictive.

The rule for reading that table, stated once: the records that could not be computed are as informative as the ones that moved. A programme that only reports the rows with good news is a marketing artefact wearing an instrument's name.

The one that rises before it falls

Its own section, because misreading it kills programmes at exactly the wrong moment.

Disposition density typically goes up first.

The mechanism: typing exposes judgements that were previously being made silently and badly — by whoever was closest, at speed, without being recorded as a decision at all. Before the instrument they were not dispositions. They were assumptions with a person attached. After it, they are counted.

So the first movement is not the process getting worse. It is the measurement getting honest, and the number rising is evidence that it worked.

A firm that reads that as failure will abandon the programme in month three, which is the worst available outcome: they now have neither the old comfortable ignorance nor the new instrument, and they have learned that measuring things makes them look bad.

Honest exposure or genuine deterioration?

They look the same for one engagement. The tie-breaker is variance, not level.

Exposure: density rises, and the variance of outcomes falls. Escalations become predictable. The same class of call keeps appearing — which is a taxonomy waiting to be written.

Deterioration: density rises and variance rises. New classes keep appearing, and the calls do not resemble each other.

Expect two to three engagements before the trend turns. Say that to the partner group in advance rather than after the first bad number, because the conversation is very different in those two positions.

The interpretation rule

If those measures improve across engagements, the square is real. If they do not, the fixed price is being subsidised by heroics.

And the reason heroics are dangerous rather than merely unsustainable is worth being exact about: they do not appear on any P&L line. A heroically-delivered engagement and a well-machined one produce the same revenue, the same margin, the same satisfaction score and the same case study.

So the firm does not notice the difference until the heroes leave — at which point the offer stops working and nobody can say why, because the cause left the building with them and it was never written anywhere.

How this joins the firm-level proof

These eight are the delivery-level instrument. The firm-level standard is three migrations: production — scarce-expert dispositions per paid bounded unit trending down; budget — customers moving spend from the old unit into the new commitment; and capability renewal — the system creating more people capable of future judgement than it consumes.

Records 3 and 7 feed migration one directly. Record 8 is the leading indicator of migration two. If you are running both instruments, that is where they join, and running the eight without knowing that is how a delivery metric ends up in a board pack answering a question nobody asked.

And the anti-performance discipline, which is the real reason to prefer records to narrative: each must be observable in ledgers the firm already keeps, never in stories. A vendor who answers with logos is answering a different question — and that standard applies to my own claims in this book, which is what the next chapter is for.

If you can only start with one

Variance ownership. Record 2.

Three reasons, in order of increasing importance. It is the cheapest — a column. It is the input to four of the others: records 3, 5, and partly 6 and 7 all derive from a classified log. And it is the only one that changes a delivery lead's behaviour on the day it is introduced, because classifying an event requires knowing the rule, and knowing the rule changes what you do about the event.

The second one to add is record 4, the verification ratio, because it is one timesheet field and it answers the loudest objection in this whole book — that independent verification is unaffordable.

Everything else can wait a quarter. There is nothing virtuous about starting all eight, and a firm that tries usually finishes none.

Key takeaways

  • A product and subsidised heroics are indistinguishable on a P&L and diverge only when you scale or when somebody resigns.
  • Six of the eight records are a column or a field. Say so, or the programme never starts.
  • Report the rows you cannot compute. A table of eight good results is evidence of optimism, not compounding.
  • Disposition density rises first because measurement gets honest. Falling variance is the tell that it is exposure rather than decay.
  • Start with variance ownership: cheapest, feeds four others, and changes behaviour on day one.

Eight records, two engagements, four honest gaps. Which raises the question the next chapter exists to answer without flinching: what does this book actually prove, and what would show that it is wrong?

22
Part VI: The Loss History

Basis Risk, and the Book's Own Limits

Everything in this book replaces adjudication with observable triggers. That move has a price, it is well documented in the profession I have been borrowing from, and it has a name.

Recorded perimeter values. Published band drivers. Typed reserve draws. An oracle named at signature. Every instrument in this book takes a question that would otherwise be settled by argument at the moment of pain, and settles it by comparison against something written down in advance.

Parametric insurance makes exactly that move, at scale, in a regulated market — and, unlike most advocates of measurement, names its price honestly.

The structure being borrowed

A parametric contract "typically specifies (1) the payment amount; (2) the trigger (a pre-determined parameter based on observable data); and (3) an impartial third party to verify that the trigger was met." 36

Contrast that with traditional indemnity, where "the policyholder documents their losses and submits a claim after an event, the insurer reviews the claim, and an adjuster assesses and validates the claim before payment is made." That is the difference between a typed trigger and a change request, described by somebody who was not writing about consulting.

The benefits are real and both sides of a trade have to be stated for the trade to be honest. Clear triggers "reduce policy disputes", and they make premiums calculable for events that occur rarely. Payment lands in weeks rather than months or years.

And note the third element in that specification, because this book has been arguing for it since Chapter 13: an impartial third party to verify that the trigger was met. The independence of the oracle is not my invention either. It is standard construction in a real insurance product.

The cost has a name

Compensation from parametric policies is not linked to actual losses, so the claim payment may be higher or lower than the losses incurred. This is known as basis risk. — Congressional Research Service, IN1267036

The worked failures are unsparing, which is what makes them useful. A city insured on barometric pressure "might not be able to claim if the damage was due to storm surge rather than wind." And the New Orleans School District held parametric wind cover for 2024 — the winds from Hurricane Francine "did not meet the 100 mph trigger and the policy did not pay out despite damage to school facilities."

Your variance schedule inherits exactly that exposure, in both directions. Type a trigger badly and you will absorb something you should have re-contracted. Type it the other way and a buyer faces a re-contract for something that did not really hurt them.

I do not need to invent an example, because this book already contains one. Chapter 15's Event 4: volume and density were treated as independent variables, so the correct application of one rule silently invalidated the assumption behind another, and the reserve absorbed a cost that should have been a band conversation three weeks earlier. That is basis risk with a delivery accent — the trigger did not correlate closely enough with the loss it was standing in for.

The mitigations, transferred

From the same literature, and they translate without modification:

  1. Choose triggers highly correlated to the thing you actually care about. The CRS example is exact: barometric pressure is measurable and is not what damages a building. Contract count is measurable and is not what costs money.
  2. Name the impartial verifier up front. Chapter 13's four forms. The point is that the verifier is chosen at signature, not at dispute.
  3. Revise triggers between engagements, never during one. A trigger changed mid-engagement is not a trigger; it is a negotiation with a version number.
  4. Check your triggers for dependency. Ours, from Chapter 15's miss. A schedule whose classes are individually correct can still be wrong as a system.
Typing does not make the world tidy. It makes the disagreement about the world happen at a better time, in a cheaper form.

Which is the honest summary of the whole method, and the version I would defend in front of somebody hostile: a rule that claims to eliminate judgement is lying. A rule that concentrates judgement into the week before signature, where it can be exercised calmly by people who can still walk away, is the best available deal.

That is a considerably smaller claim than "AI makes fixed price safe". It is also the only one that survives contact with a real engagement.

What this book actually has

Now the part I would want to read first if somebody handed me this book, and the part most books like this leave out or dress up.

Evidence status

I do not have loss ratios by band across a multi-client cohort. Nobody in this field does, and no amount of framing changes that.

Nobody has published a correlation coefficient for delivery failures across engagements sharing a model vendor. The Bank of England and the Financial Stability Board establish the mechanism. Nobody has the magnitude. Any number in Chapter 17 would have been mine and would have had to be framed as mine — which is why Chapter 17 gives shapes and controls instead.

What I do have: the mechanism; my own pre-AI fixed-price history; one worked specimen; one simulation; one re-priced offer; one refusal.

What that is: architecture plus specimen. It is not portfolio evidence, and the difference is not rhetorical — it is the difference between this is how the instrument works and this is what the instrument produced across a market.

And the measurement programme in Chapter 21 exists precisely because the evidence does not. If I had the loss ratios I would have published the loss ratios. The programme is the honest substitute, and it is a better gift to a reader than a number they could not have checked.

The frame that makes that admission structural rather than apologetic: this is the doctrine applied to itself. The first external engagements are the experiment, and this chapter is the pre-registered protocol — the tests named before the results exist, so that nobody, including me, can move the goalposts afterwards.

What is deliberately not cited

Four things would have helped this argument and are not in it anywhere. Naming them makes the omission visible rather than silent, which is the same discipline Chapter 20 asks of a delivery lead closing an engagement.

  • Circulating outcome-based-pricing adoption percentages for professional services. Widely quoted, and they trace only to aggregator content with no primary study behind them.
  • The widely repeated fixed-price-versus-T&M success-rate split. The paper is real and its qualitative conclusion is cited in Chapter 2; the two percentages could not be confirmed against the paywalled text.
  • The construction contingency percentage range. It exists in vendor explainers and not in a standards body, so it appears here as convention rather than as a figure.
  • The "19% slower" developer productivity result as a standing fact. It appears in Chapter 13 only paired with its authors' own February 2026 statement that their follow-up data gives an unreliable signal.

Seven falsifiers

Each one with what you would observe and what you would do about it — because a falsifier without a response is a disclaimer.

Seven falsifiers for the book's central claim, with observations and responses
Falsifier What you would observe What you would do
Preflight metrics do not predict cost or disposition loadBand assignment uncorrelated with actual effort across a stable populationChange the driver set. If no driver works, the offer is not censusable — decline the fee shape
Exhaustion and re-banding stay frequent after two revisionsDraws hitting the ceiling on most engagements, twice correctedThe band or the trigger set is wrong. If neither fixes it, residual uncertainty dominates — hybrid
Margin variance stays high after all costsWide dispersion within a band that does not narrowThe population is not homogeneous. Split the band, or admit it is not a class
Principal escalations do not fall by the second engagementSame people, same density, engagement twoTransfer has not happened. The offer productised the expert, not the service
Dispositions scale with item countDisposition density flat as breadth risesThe machine absorbed breadth but not judgement. The economics do not compound
Buyers value flexibility over the certainty premiumLosing on fee shape rather than price; buyers asking for T&MThe market is not buying certainty for this work. Sell the unit, not the shape
Fixed boundaries remove the work that creates valueClean delivery, low renewal, "useful but not what we needed"The perimeter is in the wrong place. Re-cut the promise, or it is a tidier contract rather than a product

The reading rule: any one of those, sustained across a stable population, is a signal worth acting on. Two or more and the honest response is a hybrid commitment or a decline rather than a defence of the model — which is why Chapters 9 and 18 exist and why they are not optional.

The strongest case against everything above

Given its full weight, because a strawman here would undo the chapter.

Delivery cost may be driven less by measurable breadth and more by novel consequential judgement. Suppose that is true for a given offer. Then:

  • census variables will never predict cost;
  • bands will never stabilise margins;
  • reserves will be exhausted repeatedly;
  • "exceptions" will become disguised change requests;
  • scope will be narrowed until the offer loses its value;
  • or the supplier will preserve the relationship by silently burning margin.

Under those conditions fixed price should not prevail, and everything in this book is an elaborate way of arriving at a bad commercial position with better paperwork. The honest answer is a hybrid — fixed onboarding and core outcome, metered consumption, paid dispositions, explicit capacity reservation, capped exceptional work — or, in the extreme, to decline.

My answer is not a rebuttal, because none is available from an armchair. It is a test, and it runs per offer rather than per industry. The same firm can hold one offer squarely inside the territory and another one outside it, and treating that as an inconsistency is the mistake rather than the finding.

The strongest form of the claim

Compile a bounded commitment when complexity is measurable and residual risk is containable. Price the certainty premium. Retain every exception as underwriting knowledge. And where those conditions do not hold, use a hybrid unit — or decline the risk.

Not "fixed price will prevail". That was never the claim, and a book that made it would have been easier to write and worth less.

Where the method does not apply at all

Six cases sit outside the territory, published rather than re-derived here: physical scarcity; non-delegable human authority; open-ended intent; external dependency; unbounded liability; and true exploratory research. Each has a diagnostic, a shrink and a decline condition.

One is worth restating because it is the most expensive misreading of everything in this book: cheap analysis does not shrink a consequence tail. Machine breadth reduces the probability of missing something. It does nothing whatsoever to the magnitude of what happens when you do. Any commitment whose downside you cannot describe in a sentence and cap in a clause is outside the territory, no matter how good the analysis has become.

And the fence from the falsifier work, which bounds this book's own enthusiasm and which I would rather state than have quoted back at me: a promise may attach only to a state the supplier can keep and evidence. A promise too big to keep is also too big to falsify.

One adjacency, named in a clause and left where it is: whether market structure makes an AI-native consultancy a net absorber of other firms' hardest exposures — adverse selection at the industry level rather than the offer level — is a genuinely interesting question and a different book.

Key takeaways

  • Typing buys cheaper, earlier disagreement and costs basis risk. Name it, because the alternative is discovering it.
  • Choose triggers correlated to what you care about; name the verifier up front; revise between engagements, never during.
  • This book is architecture plus specimen, not portfolio evidence — and the measurement programme exists because of that.
  • Seven falsifiers, each with a response. Two or more sustained, and the answer is a hybrid or a decline.
  • Cheap analysis does not shrink a consequence tail. That is the misreading that costs the most.

Part VI has been written for a supplier. The same instrument, read from the other side of the table, produces six questions — and publishing them turns out to be in my interest rather than against it.

23
Part VII: What You Do About It

The Buyer's Six Questions

Two credible suppliers, the same fixed price, the same references. Nothing in either proposal tells you which one measured its exposure and which one is brave.

Put yourself in the buyer's chair for a chapter. Both proposals are professional. Both firms have done work like this. Both numbers are within a few per cent of each other, and both documents contain a section called Our Approach.

Six questions separate them, and none of them is "what's included?"

1. What did you measure before you quoted — and can I see the drivers?

Good answer. A census, with published band drivers, and your own numbers shown against the thresholds. Better still if the census ran on the same machinery that will deliver the work, because then the measurement is not a sales artefact.

Worrying answer. "We've done a lot of these." That is experience, it is genuinely valuable, and it is not measurement. Experience prices the average engagement. A census prices yours.

This is question one because everything downstream depends on it. A band assigned from a sales call is an estimate wearing a product's clothes, and every other mechanism in the proposal is decoration on top of it.

2. Which variance do you own, and which comes back to me?

Good answer. A written schedule with named classes and named triggers. The supplier can tell you, in advance, what kind of surprise produces silence, what kind produces a visible draw, and what kind produces a conversation.

Worrying answer. "We'll work with you on scope." That sentence means the classification will be made later, under pressure, by whoever is more determined on the day — which is a prediction about temperament rather than a commercial term.

The follow-up worth asking: what is your default? If the answer is anything other than "we absorb it unless a named field moved", then absorption is a decision somebody has to be brave enough to make — and bravery is not distributed evenly across a delivery team at ten o'clock at night.

3. What is in your reserve, what draws it, and will I see the counter?

Good answer. Named trigger classes, a published consumption rule, and a shared surface. A supplier who offers you the counter is a supplier who has one.

Worrying answer. "We build in contingency." Contingency you cannot see is money you are paying for and cannot audit — and, worse, it means the supplier's absorptions will be invisible to you too. You will never learn what they carried, which means you cannot value it at renewal.

The reciprocal. Ask what happens to the unconsumed portion. What is not drawn is not consumed is a defensible answer if the reserve is a priced absorption obligation, and the supplier should be able to explain why in one sentence. If they cannot, they are holding contingency and calling it a reserve.

4. Who signs acceptance — and is the oracle independent of the people who did the work?

Good answer. Held-out cases, a buyer-selected sample, reconciliation against a system neither party authored, or live receipts over a period. Named at signature, in the schedule, with a form you can picture.

Worrying answer. "Our QA process is very thorough." Thoroughness is not independence, and the two are routinely confused by people acting in complete good faith. A meticulous producer checking its own work is still a producer checking its own work.

The version that costs you nothing to ask for: let me pick the sample, after you have finished. A supplier who hesitates has told you something important. A supplier who agrees has just made your acceptance meaningful for the price of an afternoon.

5. What would count as failure? Say it out loud.

Good answer. A sentence describing a result that would mean the engagement did not succeed — one a competent third party could evaluate, that could come back negative, and that would change what you do next.

Worrying answer. Anything about process, effort or satisfaction. "If you weren't happy with the outcome" is not a failure condition; it is a mood.

The rule I would give any buyer: a supplier who cannot answer this has not sold you an outcome. They have sold you an activity with an optimistic adjective.

And the follow-up: what happens commercially if it fails? An acceptance test with no commercial consequence attached is a ceremony, however precisely it is worded.

6. What have you declined, and why?

The one nobody asks, and the most informative of the six.

Good answer. A specific opportunity, a reason expressed in the offer's own vocabulary — which perimeter fields could not carry a recorded value — and, the tell that it is real, what they offered instead.

Worrying answer. "We're pretty selective." Or, more revealingly, genuine surprise at the question. A supplier who has never been asked it has never had to think about where their offer ends.

Why it works: a supplier who has never declined anything has a band that predicts nothing, because every opportunity has been forced into it. The declination log is the shape of the boundary, and a firm with no boundary has no product — only a fee and a willingness.

The seventh, for AI-delivered work

What do all your engagements share, and what happens to mine when it changes?

Almost no buyer asks this yet. They will — and the firms that can answer it will have built the register in Chapter 11 rather than discovering the question in a procurement meeting.

A good answer contains: the shared components, the pinning posture, how behaviour change is detected and how fast, whether rollback exists, and what is excluded with a typed state rather than a disclaimer.

A worrying answer is reassurance about model quality. The question is not whether the model is good. It is what happens on the day it changes.

What to do with unsatisfying answers

Not walk away. Ask for the shrink.

A supplier who cannot answer questions one to four for the whole engagement can usually answer them for a smaller one. Ask what they can bound. Buy that. Let the filled perimeter make the rest quotable — by them, or by somebody else.

The symmetry is worth noticing, and it is a decent test of whether a mechanism is real rather than rhetorical: the same move works from both sides of the table. The supplier shrinks because the perimeter will not hold. The buyer shrinks because the answers will not hold. Same instrument, same conclusion, arrived at independently.

Why I publish these against my own interest

The obvious objection is that these questions make me easier to interrogate, and they do.

The answer is commercial rather than noble. A market that asks these questions rewards firms that have the machinery. I would rather compete there than in a market where the winner is whoever sounds most confident — because in that market, the machinery is a cost with no return, and the rational strategy is to skip it and get better at proposals.

Predictability is a premium attribute, not a discount. But only in a market that can tell it apart from confidence, and these six questions are how a buyer tells them apart.

And the reflexive version, which this book has to accept rather than dodge: these questions apply to me. My own offers are designed hypotheses about value, falsifiable only by paid engagements, and I have published them at that status rather than as validated prices.

Question six, answered about myself: the opportunity in Chapter 18. Five of eight perimeter fields could not carry a recorded value, three of those sat outside both parties' control, the promised state had been redefined twice in three meetings, and I sold the bounding instead — for a fraction of the revenue that was on the table. A chapter that tells buyers to demand a declination and does not supply one has failed its own test.

The buyer who is not a person

One paragraph, because a sibling owns the argument and because it is the reason this chapter will matter more next year than it does now.

Publish enough structure — eligibility, required inputs, price band, valid outcomes, exclusions, evidence returned, acceptance — that an authorised customer agent can determine fit without a discovery call.

Here is the connection nobody expects: an underwritten offer is legible by construction. The census fields, the band drivers, the variance schedule, the exclusions and the acceptance design are exactly the structured facts an agent needs to evaluate fit. A firm that has done the work in this book has already published its own machine-readable qualification criteria without intending to.

And the consequence for everybody else is not softened by time. If the future buyer arrives through an agent, the amorphous proposition is not merely weak. It is invisible.

Key takeaways

  • Six questions separate a measured promise from a brave one, and none of them is "what's included?"
  • Every worrying answer is something a decent, competent firm says in complete sincerity. That is what makes them worth listing.
  • "Let me pick the sample after you have finished" costs the buyer nothing and changes what acceptance means.
  • Ask what they declined. A supplier who has never declined anything has a band that predicts nothing.
  • Unsatisfying answers call for the shrink, not the walk-away — the same move that works from the supplier's side.

The buyer now has an instrument. The seller needs a different one — because the failure this whole book is written against is not incompetence. It is the vocabulary, adopted without the machinery.

24
Part VII: What You Do About It

Underwriting-Washing

Within a year, every firm you compete with will describe its pricing as measured, banded, reserved and independently verified. Most of those descriptions will be true of the proposal and false of the operating reality.

I have written this failure up once before under a different name. Title adoption without the machinery is FDE-washing — the certification cohort graduates, the LinkedIn titles update, and six months later no production deployment has changed shared firm infrastructure.

The same failure is arriving with a new costume, and this book has handed everybody the vocabulary for it.

The definition

Underwriting-washing: calling a price fixed without the census, the bands, the reserve, the exclusions, the independent oracle and the loss history.

The recognisable shape is not a bad firm. It is a good firm whose proposal says productised and whose operating reality is a senior person's Thursday intuition with a confident cover page and a schedule copied from the last deal.

The reason it spreads is structural rather than moral: every element is expensive to build and free to claim, and until Chapter 23's questions become normal there is no buyer behaviour that separates the two. In that market, the rational strategy is to skip the machinery and get better at proposals — and rational strategies get adopted.

Why it is worse than honest time-and-materials

Time-and-materials puts the variance where the client can see it. Both parties know what they are in. The mechanism is crude and it is not dishonest, and when reality turns out messier than anybody thought, the mess appears on an invoice somebody can argue about.

Underwriting-washing hides retained variance behind a number that implies it was measured. The buyer believes they bought certainty. The supplier believes they sold a product. Neither of them is holding an instrument that could tell them otherwise.

Which produces the real cost, and it is not the bad quarter. The firm is not primarily deceiving the client. It is deceiving itself — because a fixed price with no records produces no signal, so the firm's own belief that the model is working is unfalsifiable. Nobody can say how much risk capital was contributed last quarter. Nobody can say whether engagement twenty was safer than engagement two. The absence of bad news is read as the presence of good news, which is the one inference a firm should never make about its own product.

And then it scales. A thing that cannot be measured gets scaled on the strength of the fact that nothing has gone visibly wrong — and the mistake gets made at the largest available size.

The ten-point audit

A recurring review, not a launch-day poster. Each item is a yes or no, and each one has a named artefact behind it — because "we do that" is not evidence, and the entire subject of this chapter is the gap between claiming and doing.

The ten-point underwriting-washing audit, with the artefact required for each item
# The check The artefact
1A preflight census runs on the sensors that will deliver the work — not on a sales callThe census output for the last three quoted engagements
2Band drivers are published; a buyer can see why they are in their bandThe band assignment page a buyer actually received
3A written variance schedule exists, four classes, named triggersThe schedule, in the signed contract — not the playbook
4Interior variation is the default; absorbing requires nobody to be braveThe classification rule, and the absence of a hundred-row change log
5The reserve has trigger classes, a published rule, a visible counter, an owner, a window and a defaultThe counter surface a buyer can open, and the exhaustion clause
6Exclusions are operational, stated at signature, and they change the priceTwo quotes with different exclusions and different numbers
7The acceptance oracle is not controlled by the producer, and you can name its formThe held-out set, sampling protocol or reconciliation source, in the schedule
8An accumulation register exists and someone reviews itThe register, with a date on the last review and a filled detection column
9You can point to a declination or a re-band in the last two quartersThe declination log, with reasons and reopen conditions
10Band economics are reviewed on a cadence; a losing band is fixed in the product definitionThe review minutes, and a band that changed as a result

The scoring rule, stated so it cannot be gamed by generosity: an item counts only if the artefact exists and somebody outside the delivery team could find it without asking. That second clause does most of the work. Plenty of firms have all ten artefacts somewhere in somebody's folder.

The tells

The audit will be gamed by anybody who reads it as a checklist, which is why the tells matter more. These are things you overhear rather than things you measure.

  • A change log with a hundred rows. Length is a symptom. It means the default is wrong and every surprise is being negotiated.
  • A reserve that is never drawn. Either it is mispriced, or draws are being absorbed silently to keep the counter clean — and the second is worse, because it looks exactly like the first.
  • An exclusions list that has not changed in a year. Exclusions are learned. A static list is a template.
  • A band that always wins. Either the band is mis-set, or work is being steered into it, and both are pricing failures wearing a good result.
  • An acceptance criterion nobody has ever failed. Chapter 12's entire argument, arriving as a metric.
  • Nobody can name what engagement two started with. If the answer is "the team knew what they were doing", the flywheel is familiarity.
  • "We absorbed it, they're a good client" — said with pride. The single most reliable indicator on this list, and the easiest to hear in an ordinary delivery meeting.

The monthly operating questions

A quarterly audit is too slow to catch drift, and a weekly one does not get run. Seven questions, monthly, in an operating review:

  • Which trigger fired last month, and what did it teach?
  • Which band lost money, and what changed in the product definition as a result?
  • Which exception became a band driver?
  • What did we decline, and is the reason in the log?
  • Which shared component changed under us — and how did we find out?
  • Which absorption did we make that nobody recorded?
  • What is engagement-two lead share on active work?

That last one is imported deliberately from the practice-OS material, because it is the question that connects this audit to the transfer proof — and it is the only one on the list that cannot be answered by producing an artefact. Somebody has to know who actually led the work.

Who holds the audit

This is the chapter's own honesty clause, and it is this book's argument applied to its own instrument.

A rubric handed to the party being audited becomes a target. That is Chapter 12, arriving here, and it applies to the ten points above exactly as much as it applies to an acceptance test.

So the audit belongs with somebody who does not carry the delivery number — a partner outside the practice, a commercial lead, a non-executive, or practice leads on rotation auditing each other. Not because delivery people are dishonest. Because they are optimising for a different thing, and that is the whole finding of Part IV.

The artefact requirement is what makes it survivable even when it has to be self-administered. A process somebody can describe is easy to imagine into existence under time pressure. An artefact somebody outside the team could find is considerably harder.

And the reflexive application, briefly, because a paragraph of humility would be less credible than a sentence of specificity: this audit applies to my own firm, and Chapter 22 has already published what my honest score looks like — architecture and specimen, with band economics uncomputable at n=2.

What a passing score does and does not mean

It does mean: the machinery exists, and your claims about your own pricing are checkable by somebody who did not write them.

It does not mean: that the offer is profitable, that the bands are right, or that the square is a product. Those are Chapter 21's eight records, and they take engagements rather than an afternoon.

The distinction matters because there is a specific failure available here that looks like success. A firm that passes the audit and never reads the records has built the instruments and never looked at them. That is a better failure than the alternative — the instruments are there when somebody finally asks — and it is still a failure, because instruments that nobody reads change no decisions, and a row that changes no decision is a report.

Key takeaways

  • Underwriting-washing spreads because every element is expensive to build and free to claim.
  • It is worse than honest T&M because the firm's own belief that the model works becomes unfalsifiable.
  • Ten checks, each with a named artefact somebody outside the delivery team could find without asking.
  • The tells beat the checklist: a hundred-row change log, an undrawn reserve, an unchanged exclusions list, a criterion nobody has failed.
  • The audit belongs with somebody who does not carry the delivery number — this book's own argument, applied to itself.

The audit says what is missing. The last chapter says what to do about it — three moves, one sentence in the next proposal, and one line that the whole book has been earning.

25
Part VII: What You Do About It

What Changes on Monday

Three moves, one paragraph in the next proposal, and a quarterly hour. All of it runs against the book of work you are already holding.

You are holding some number of live engagements at different weeks of their clocks. At least one fixed price signed in the last quarter. And underneath all of it, a shared kernel that nobody has ever drawn a picture of.

Everything below runs against that. This week. Without a project, a tool selection, or anybody's permission.

Move 1 — Classify one completed engagement

One afternoon

  1. Take the last engagement that closed. Not the interesting one — the last one, because selecting for interest is how this exercise gets ruined before it starts.
  2. List every surprise: everything that made somebody say hang on. Emails count. Corridor conversations count. The list will be longer than your change log, and that gap is the first finding.
  3. Put each item in exactly one class, using the card from Chapter 15. Default to interior. Interior is a residual — it is what remains when every named field failed to move.
  4. Total the Class 1 column at your own internal cost.

What you get: the risk capital your firm contributed on one engagement without recording it. Multiply by your engagement count for the year and you have a number nobody in your business has ever seen.

The second finding, which is usually worth more than the number: the events where two people in the room classify differently. Each one is an underspecified rule, and each one is a sentence you can write today that settles that class permanently. Four sentences is a good afternoon's work.

Move one is first because it changes the most conversations for the least effort, and because it requires agreement from nobody outside the room.

Move 2 — Build the accumulation register

Ninety minutes

  1. One row per shared component across your live book. Include the ones you did not build, and the ones you have forgotten you depend on — that second category is where the surprises live.
  2. For each: who depends on it; what changes it without asking you; what you would observe if it changed badly; how long detection takes; whether you can roll back and how fast.
  3. Then look only at the detection column.

What you get: the rows where the honest answer is a client would tell us. Those are your unpriced catastrophe exposures, and until this morning they were not on any list in your firm.

In most firms, the uncomfortable row is the dependency where nothing fails — the output stays well-formed and becomes quietly wrong across every engagement at once. That row is the point of the exercise. If every row is comfortable, you have not listed the vendor-hosted ones.

Then decide, per row: control it, exclude it, or hold capital against it. There is no fourth option, and reserve for it is the wrong answer for the reasons Chapter 17 works through.

Move 3 — Start the eight records, beginning with variance ownership

A column in the delivery record

Record 2 first. Three reasons, in increasing order of importance: it is the cheapest; it is the input to four of the other seven; and it changes a delivery lead's behaviour on the day it is introduced, because classifying an event requires knowing the rule, and knowing the rule changes what you do about the event.

Add record 4 — the verification ratio — next quarter. One timesheet field, and it is the answer to the loudest objection in this book.

Leave the rest until you have a population. Records without a population produce arguments rather than findings.

And warn the partner group in advance that record 3, disposition density, will move the wrong way first. That is measurement getting honest, not delivery getting worse — Chapter 21 has the tie-breaker, and the conversation is very different held before the first bad number than after it.

The one thing to change in the next proposal

Name the oracle.

Write down, in the schedule, who verifies completion and how: held-out cases, a buyer-selected sample, reconciliation against a system neither party authored, or live receipts over a period.

It costs a paragraph. It is the highest-leverage change available in this entire book, for a reason worth stating plainly: it is the one thing a competitor cannot copy without actually doing it. Every other claim in a proposal can be asserted. An oracle either exists in the schedule with a named form and a named signer, or it does not.

The cheapest version, if you do nothing else: the buyer selects thirty items after generation, from the declared population, and we score them in front of you. One afternoon of exposure, and your acceptance means something for the first time.

There is a commercial return on that paragraph that arrives faster than the doctrinal one. Buyers who have been burned by an over-confident supplier recognise a real acceptance instrument immediately — and the ones who have not been burned yet will remember which of the two proposals contained one.

The quarterly hour

Four lines. It takes an hour and it is the entire governance overhead of everything in this book.

  • Review band economics. Which band lost money, and what changed in the product definition — rather than in somebody's effort.
  • Review the register. What changed under you, and how you found out.
  • Review declinations. What you refused, why, and whether any reopen conditions have since become true.
  • Review which exceptions became band drivers. This is the flywheel, expressed as a number of rows.

If a quarter passes with nothing in the third or fourth line, that is the finding — and it is a finding you can act on before it becomes a year.

Why start imperfectly today

A loss history compounds, and the compounding starts on the first entry. There is no version of this where waiting is cheaper.

The specific trap is worth naming because it is nearly universal: firms defer the programme until they can do it properly — and "properly" means a schema, a tool selection and a project, none of which produce a single row. Meanwhile the engagements that would have populated it close, the team moves on, and the evidence disperses into people's memories where it is unrecoverable within about six weeks.

The honest floor: a column and a shared document beat a system that does not exist. Every instrument in this book was designed to be startable with a spreadsheet, deliberately, because an instrument nobody starts is worth exactly nothing.

I am inside the blast radius

My own offers are designed hypotheses about value, falsifiable only by paid engagements. My proof is a specimen, a simulation and a refusal. Band economics on my own offers are uncomputable at n=2, and I said so in Chapter 22 rather than in a footnote.

The seven falsifiers apply to me first, and they were written before the results exist so that nobody — including me — can move the goalposts afterwards.

That is not a weakness of the doctrine. It is the doctrine, applied to itself: measure your exposure, name what you cannot know, and publish the test before the answer.

What all of this is for

Not a tidier contract. Not a better estimate. Not a way of feeling more confident about a number.

The point is narrower and harder: the ability to say which variance you can safely own — and to have a record that proves it, in a form that a buyer can interrogate, a competitor cannot copy, and your own next engagement can load before anybody starts work.

Everything else in this book is machinery for that one sentence.

The square is not drawn around the project. It is earned by underwriting the variance.
REF
Sources & Evidence

References & Sources

The evidence base behind every claim — primary research, industry analysis, and technical specifications

Research Methodology

This ebook draws on primary research from standards bodies, independent research firms, enterprise technology vendors, and consulting firms. Statistics cited throughout have been cross-referenced against primary sources.

Frameworks and interpretive analysis developed by Scott Farrell / LeverageAI are listed separately below — these represent the practitioner lens through which external research is interpreted, and are not cited inline to avoid self-promotional appearance.

LeverageAI / Scott Farrell — Practitioner Frameworks

The interpretive frameworks, architectural patterns, and practitioner analysis in this ebook were developed through enterprise AI transformation consulting. The articles below are the underlying thinking behind those frameworks. They are listed here for transparency and further exploration — not cited inline, as this is the author's own analytical voice.

Scott Farrell — AI-Native Service Architecture

The Square: a stable commercial perimeter around an adaptive machine-scale production interior; the eight perimeter fields; rigid-fluid-rigid topology; the interior may be horrendously irregular and none of it reaches the buyer (ch3 #49cbfe)

https://leverageai.com.au/wp-content/media/articles/226-ai-native-service-architecture.html

Scott Farrell — Buy Certainty First

Typed uncertainty: unknowns become bounded terminal states with a precise assertion and next consumer rather than unbounded labour; not observed does not mean does not exist; fixed price without a theory of uncertainty is gambling with better stationery (ch8 #a09dc9)

https://leverageai.com.au/wp-content/media/articles/204-buy-certainty-first.html

Scott Farrell — The Terminal Value Doctrine: Professional Services

The offer ladder; the successor knowingly prices the responsibility and variance the supplier retains, with the insurer's vocabulary — eligibility bands, complexity pricing, paid dispositions, reservations — doing the work the blended rate used to fake; the underwriting machinery beneath that sentence is a later book's (ch15 #de2950)

https://leverageai.com.au/wp-content/media/articles/231-terminal-value-doctrine-professional-services.html

Scott Farrell — Fog Is a Race Between Two Clocks

Four ways the money moves: unsold capacity, fully sold capacity, outcome pricing, demand expansion; the verification ratio; pricing a bounded commitment when delivery variance is real is its own discipline and a sibling subject (ch11 #51f7d3)

https://leverageai.com.au/wp-content/media/articles/232-fog-is-a-race-between-two-clocks.html

Scott Farrell — Preparedness Is the Product

Response commitments and the four refusals — sell concrete service acts with cut-offs and refuse the promises whose physics you do not control; delay compensation and project-loss indemnity are risk-transfer instruments that may amount to financial products or contracts of insurance depending on structure; customers who need insurance products should buy them from parties authorised to sell them (ch7 #69811b)

https://leverageai.com.au/wp-content/media/articles/214-preparedness-is-the-product.html

Scott Farrell — AI-Constituted Services

The Fixed-Price Envelope: automated preflight census, volume bands, included finding counts, the Flex Reserve, unsupported-source rules and the boundary clause; AI makes the cost curve flatter, not flat, and changes which variable drives it; the machine's breadth is nearly free and the human's dispositions are the metered resource; machine-measured product configuration rather than time-and-materials estimation (ch9 #f8a5b5)

https://leverageai.com.au/wp-content/media/articles/202-ai-constituted-services.html

Scott Farrell — Boundary Mutation, Not Change Request

What change control already knows: five precedent families including NEC compensation events, World Bank PPP variation classes, ITIL standard changes, PMI contingency and management reserve, and target-cost pain/gain share; a rule that claims to eliminate judgement is lying, and one that concentrates judgement into the week before signature is the best available deal (ch7 #85a101)

https://leverageai.com.au/wp-content/media/articles/228-boundary-mutation-not-change-request.html

Scott Farrell — Two Falsifiers

What a falsifier may attach to: a promise too big to keep is also too big to falsify; the keepable-and-evidenceable fence, with the two interesting failure cases of keepable-not-evidenceable and evidenceable-not-keepable; outcome contracting has its own documented gaming taxonomy of cherry picking, creaming and parking; the test to carry into a drafting session is whether, if it fires, you can tell whether you caused it (ch7 #90289b)

https://leverageai.com.au/wp-content/media/articles/229-two-falsifiers.html

Scott Farrell — Custom Software Verification

Custom software did not lose on build cost alone — it lost on verification cost; a non-developer commissioning custom software faces an unsolvable principal-agent problem, unable to inspect the work, tell a good developer from a confident one, or tell six months of progress from six months of invoices; it is a trust failure and it is structural, because it does not matter how honest your developer is if you have no independent way to verify the honesty (ch2 #5a9ee8)

https://leverageai.com.au/wp-content/media/articles/105-custom-software-verification.html

Scott Farrell — Hidden Gates

The measure that stops measuring: Goodhart 1975, Strathern 1997 and Campbell's law correctly attributed; a quality gate is a measure, and handing it to the worker as its target converts a quality check into a specification to be satisfied by the cheapest available route; specification gaming as the reinforcement-learning name for the same failure, where capability makes it worse rather than better (ch2 #55d1ab)

https://leverageai.com.au/wp-content/media/articles/94-hidden-gates.html

Scott Farrell — The Learning Subsidy

First-engagement learning ring-fenced as a pre-approved, capped, expiring subsidy tracked on a separate capability ledger with named assets and kill conditions, proved only when engagement two runs on ordinary staff with measurable reuse gains

https://leverageai.com.au/wp-content/media/articles/230-the-learning-subsidy.html

Scott Farrell — Forward-Deployed Practice OS

Three kernels, one ownership map: three ownership territories with enforced promotion gates, where confidential material does not auto-climb and automatic upward promotion is a breach rather than a feature; title adoption without the machinery is FDE-washing, and transfer is the product test (ch9 #52ef90)

https://leverageai.com.au/wp-content/media/articles/167-forward-deployed-practice-os.html

Primary Research & Standards Bodies

acquisition.gov — Federal Acquisition Regulation 16.202-1 [1]

"A firm-fixed-price contract provides for a price that is not subject to any adjustment on the basis of the contractor's cost experience in performing the contract. This contract type places upon the contractor maximum risk and full responsibility for all costs and resulting profit or loss."

https://www.acquisition.gov/far/subpart-16.2

Federal Register, Executive Order, 5 May 2026 — Promoting Efficiency, Accountability, and Performance in Federal Contracting [3]

"Use of any non-fixed-price contract, including a cost-reimbursement contract, a time-and-material contract, a labor-hour contract, or any other non-fixed-price type of contract under Part 16 of the Federal Acquisition Regulation, must be justified in writing by the contracting officer to the agency head." A review of FY2024 spending identified approximately $120 billion obligated on cost-reimbursement consulting contracts alone

https://www.federalregister.gov/documents/2026/05/05/2026-08900/promoting-efficiency-accountability-and-performance-in-federal-contracting

Magne Jørgensen, Parastoo Mohagheghi and Stein Grimstad, International Journal of Project Management 35(8), 2017, pp. 1573-1586, DOI 10.1016/j.ijproman.2017.09.003 — Direct and indirect connections between type of contract and software project outcome [4]

The use of fixed price contracts is connected with a higher risk of project failure compared to time and materials types of contracts

https://www.sciencedirect.com/science/article/abs/pii/S0263786317301813

American Academy of Actuaries, Committee on Risk Classification — Risk Classification Statement of Principles [5]

"Though an individual exchanges the uncertainty of occurrence, timing and magnitude of a particular event for the certainty of a fixed price, that exchange in no way makes the uncertain known. Nor need it… But it should find a way of establishing a fair price for assuming it."

https://actuary.org/wp-content/uploads/2025/05/risk.pdf

Casualty Actuarial Society — Statement of Principles Regarding Property and Casualty Insurance Ratemaking [8]

"Principle 1: A rate is an estimate of the expected value of future costs"; "Principle 2: A rate provides for all costs associated with the transfer of risk"; costs "include claims, claim settlement expenses, operational and administrative expenses, and the cost of capital"; "Ratemaking is prospective because the property and casualty insurance rate must be developed prior to the transfer of risk"

https://www.casact.org/sites/default/files/2021-06/Statement%20of%20Principles%20Regarding%20P&C%20Casualty%20Insurance%20Ratemaking_2021.pdf

Government Outcomes Lab, Blavatnik School of Government, University of Oxford — Outcomes-based contracting [11]

Outcomes-based contracts risk "'cherry picking', where eligible individuals are not referred or accepted onto a service if they seem unlikely to achieve payable outcomes, 'creaming', in which providers focus their efforts on those individuals who are easiest to help, and 'parking', which is neglect of those who may be more difficult to achieve outcomes with"; and "it can be difficult to set simple, measurable outcomes that align effectively with complex social problems"

https://golab.bsg.ox.ac.uk/the-basics/outcomes-based-contracting

U.S. Government Accountability Office, September 2005 — Defense Management: DOD Needs to Demonstrate That Performance-Based Logistics Contracts Are Achieving Expected Benefits (GAO-05-966) [12]

Only 1 of 15 program offices had performed a business case update; in that case it determined that the performance-based logistics contract did not result in expected cost savings and the weapon system did not meet established performance requirements; program officials typically relied on cost and performance data generated by the contractors' information systems and had not determined whether contractor-provided data were sufficiently reliable

https://www.govinfo.gov/content/pkg/GAOREPORTS-GAO-05-966/html/GAOREPORTS-GAO-05-966.htm

Bank of England — Financial Stability in Focus: Artificial intelligence in the financial system, April 2025 [15]

"A reliance on a small number of providers for a given service could also generate systemic risks in the event of disruptions to them, especially if is not feasible to migrate rapidly to alternative providers"; "under a scenario in which customer-facing functions have become heavily reliant on vendor-provided AI models, a widespread outage of one or several key models could leave many firms unable to deliver vital services"; "From a systemic risk perspective, the potential for AI-based participants to take increasingly correlated positions is an important consideration"; the July 2024 CrowdStrike update named as the demonstration

https://www.bankofengland.co.uk/financial-stability-in-focus/2025/april-2025

Financial Stability Board — The Financial Stability Implications of Artificial Intelligence, 14 November 2024 [16]

"Most LLMs are trained using the same underlying architecture and many are trained, at least in part, on common sources of web crawl data. The homogenisation in training data and model architecture can lead to correlated outputs, which could amplify market stress"; and "the interaction of AI-related third-party dependencies and market concentration among technology and AI service providers could increase domestic and international interconnections… exposing FIs to losses arising from operational impairments and supply chain disruptions affecting key vendors"

https://www.fsb.org/uploads/P14112024.pdf

Reuters, via Malay Mail, 21 July 2024 — Microsoft says CrowdStrike mayhem took out 8.5 million Windows devices [17]

Microsoft estimated the faulty CrowdStrike update affected approximately 8.5 million Windows devices

https://www.malaymail.com/amp/news/money/2024/07/21/microsoft-says-crowdstrike-mayhem-took-out-85-million-windows-devices/144464

Charles Goodhart (1975), "Problems of Monetary Management: The UK Experience"; Marilyn Strathern (1997), "'Improving ratings': audit in the British University system", European Review 5(3) p.308 — Goodhart's law [22]

Goodhart's original: "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." Strathern's compression: "When a measure becomes a target, it ceases to be a good measure."

https://en.wikipedia.org/wiki/Goodhart%27s_law

Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, Shane Legg, Google DeepMind — Specification gaming: the flip side of AI ingenuity [23]

"Specification gaming is a behaviour that satisfies the literal specification of an objective without achieving the intended outcome"; "Even for a slight misspecification, a very good RL algorithm might be able to find an intricate solution that is quite different from the intended solution, even if a poorer algorithm would not be able to find this solution… This means that correctly specifying intent can become more important for achieving the desired outcome as RL algorithms improve"

https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/

David Manheim and Scott Garrabrant — Categorizing Variants of Goodhart's Law (arXiv:1803.04585) [24]

"the importance of Goodhart effects depends on the amount of power directed towards optimizing the proxy, and so the increased optimization power offered by artificial intelligence makes it especially critical for that field"

https://arxiv.org/abs/1803.04585

Yanuo Ma, Ben Kereopa-Yorke and Ben Schultz, June 2026 — Building to the Test: Coding Agents Deliver What You Check, Not What You Requested (arXiv:2606.28430) [25]

Two production Copilot CLI agents re-implement a React Fluent-UI data table in Angular as a reusable library under a hidden 222-test Playwright oracle across 18 runs and three oracle-availability conditions: "Without the oracle, the library is present but unfinished, revealed by scores. With the oracle in the loop, the score reaches near-perfect, but from a demo holding the tested behavior directly, the library left dead or absent. We call this building to the test; the broader disposition behind both we call validation self-awareness. The agent does not, on its own, validate what it ships as a user would."

https://arxiv.org/abs/2606.28430

Monte MacDiarmid, Benjamin Wright, Jonathan Uesato et al., Anthropic and Redwood Research — Natural Emergent Misalignment from Reward Hacking in Production RL [26]

"We show that when large language models learn to reward hack on production RL environments, this can result in egregious emergent misalignment… the model generalizes to alignment faking, cooperation with malicious actors, reasoning about malicious goals, and attempting sabotage when used with Claude Code, including in the codebase for this paper"

https://assets.anthropic.com/m/74342f2c96095771/original/Natural-emergent-misalignment-from-reward-hacking-paper.pdf

Linford & Company LLP — SOC Report Types: Type 1 vs Type 2 SOC Reports/Audits [27]

A Type 1 SOC report "is as of a point in time… It only covers the design effectiveness of the internal controls"; a Type 2 report "covers a period of time… the operating effectiveness of the internal controls over time"

https://linfordco.com/blog/soc-report-types-1-vs-2/

Elliot Kim, Avi Garg, Kenny Peng, Nikhil Garg — Correlated Errors in Large Language Models (arXiv:2506.07962, ICML 2025) [28]

On one leaderboard dataset "models agree 60% of the time when both models err", and "larger and more accurate models have highly correlated errors, even with distinct architectures and providers"; in the judge setting each judge systematically inflates the accuracy of models less accurate than itself due to correlated errors

https://arxiv.org/abs/2506.07962

Shashwat Goel, Joschka Struber, Ilze Amanda Auzina, Karuna K Chandra, Ponnurangam Kumaraguru, Douwe Kiela, Ameya Prabhu, Matthias Bethge, Jonas Geiping — Great Models Think Alike and this Undermines AI Oversight (arXiv:2502.04313) [29]

As model capabilities increase it becomes harder to find their mistakes, and model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures

https://arxiv.org/abs/2502.04313

Lianmin Zheng et al. — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv:2306.05685) [30]

Strong LLM judges can match both controlled and crowdsourced human preferences well, "achieving over 80% agreement, the same level of agreement between humans"

https://arxiv.org/abs/2306.05685

International Accreditation Service — Understanding ISO/IEC 17020 Handbook [31]

"Rock solid demonstrations of impartiality require the IB to ensure that its own staff are not involved in any aspect of ownership, design, manufacture, or other relationship as regards the object of inspection or its manufacturer / supplier"; the categorisation of inspection bodies as Type A, B or C is essentially a measure of their independence, and third-party inspection requires Type A

https://www.iasonline.org/wp-content/uploads/2021/01/Tab-1-01-Understanding-ISOIEC-17020-Handbook.pdf

DORA, Google, 2025 — Balancing AI tensions: Moving from AI adoption to effective SDLC use [32]

"The verification tax: Time saved writing is often re-spent auditing"; "While AI successfully accelerates initial code generation and reduces the friction of starting new tasks, the time saved in creation is frequently re-allocated to auditing and verification"; "Verification is a fundamentally different cognitive task than creation"; "Velocity gains for an individual author frequently translate into a significantly increased cognitive load for the reviewer"; "higher AI adoption is associated with an increase in both software delivery throughput and software delivery instability"

https://dora.dev/insights/balancing-ai-tensions/

METR, 10 July 2025 — Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity [33]

Randomised controlled trial with 16 experienced open-source developers across 246 tasks on their own repositories: developers took 19 per cent longer with early-2025 AI tools, having forecast a 24 per cent speed-up and believing afterwards they had been sped up by 20 per cent

https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/

METR, 24 February 2026 — We are Changing our Developer Productivity Experiment Design [34]

"we believe that the data from our new experiment gives us an unreliable signal"; METR attribute the unreliability to selection effects including developers declining to participate without AI tools and avoiding submitting tasks they believed AI would handle well, and state the 19 per cent figure likely does not reflect current conditions

https://metr.org/blog/2026-02-24-uplift-update/

Casualty Actuarial Society — Statement of Principles Regarding Property and Casualty Loss and Loss Adjustment Expense Reserves [35]

"An actuarially sound loss reserve for a defined group of claims as of a given valuation date is a provision, based on estimates derived from reasonable assumptions and appropriate actuarial methods for the unpaid amount required to settle all claims, whether reported or not"; "The true value of the liability for losses or loss adjustment expenses at any accounting date can be known only when all attendant claims have been settled"; a reserve estimate should take into account the degree of uncertainty inherent in its projections

https://www.casact.org/sites/default/files/2021-04/statement_of_principles_Loss_Loss_Adjustment%20_Expense%20_Reserves_2021.pdf

Congressional Research Service — Parametric Insurance for Natural Disasters: Frequently Asked Questions (IN12670) [36]

"A contract for parametric insurance typically specifies (1) the payment amount; (2) the trigger (a pre-determined parameter based on observable data); and (3) an impartial third party to verify that the trigger was met (for example, the National Hurricane Center)"; in traditional disaster indemnity insurance "the policyholder documents their losses and submits a claim after an event, the insurer reviews the claim, and an adjuster assesses and validates the claim before payment is made"; "the use of a clearly defined trigger may make it easier for the insured to understand the coverage provided and reduce policy disputes"

https://www.everycrsreport.com/reports/IN12670.html

Major Consulting Firms

Business Insider — AI is reshaping how McKinsey makes money [2]

Kate Smaje, global leader of technology and AI at McKinsey: "Outcomes-based pricing didn't start because of AI, but the type of work AI transformation demands suits it." About a quarter of McKinsey's global fees now come through performance-based arrangements

https://www.aol.com/articles/ai-reshaping-mckinsey-makes-money-115132273.html

Industry Analysis & Vendor Research

Munich Re — Insure AI / aiSure [6]

"aiSure is a suite of comprehensive coverage for AI systems designed to address a wide area of AI-related risks for AI providers and corporate adopters caused by AI performance errors, including: Contractual liabilities; Own damages/financial losses; and Legal Liabilities"; performance warranties enable providers to indemnify clients for losses directly related to AI errors

https://www.munichre.com/en/solutions/for-industry-clients/insure-ai.html

Armilla AI — Insurers launch cover for losses caused by AI chatbot errors [7]

"Armilla Insurance Services is a Coverholder at Lloyd's. Affirmative AI Liability Insurance is underwritten by certain underwriters at Lloyd's."

https://www.armilla.ai/resources/insurers-launch-cover-for-losses-caused-by-ai-chatbot-errors

MPUG, quoting PMI definitions — Contingency Reserve and Management Reserve [9]

Contingency reserve is time or money allocated in the schedule or cost baseline for known risks with active response strategies; management reserve is an amount held outside of the performance measurement baseline reserved for unforeseen work that is within scope of the project, and drawing on it follows the change control process

https://mpug.com/contingency-reserve-management-reserve

Procore — Lump Sum Contracts in Construction [10]

Construction contingency is sized by project complexity, design completeness and site conditions; for fixed-cost builders contingencies are built into the overall project cost as a percentage of the contract sum; changes in design or deviation from the original plans require a change order paid by the owner, and change orders are the primary driver of contingency depletion; contingency covers unplanned costs nobody anticipated while an allowance covers planned scope not yet fully specified

https://www.procore.com/library/lump-sum-contracts

Troy Edwards and Peter Tolson, A&O Shearman — Cost reimbursable vs. lump sum turnkey construction contracts: the many routes to bankability [13]

"In truth, there is no such thing as an absolute fixed price contract"; lump-sum contractors often include significant contingencies in their pricing and are naturally incentivized to seek opportunities to reopen the fixed price where those contingencies prove insufficient; lump sum may still be preferable for well-defined, low-risk projects where scope and owner requirements are clear from the outset

https://www.aoshearman.com/en/insights/cost-reimbursable-vs-lump-sum-turnkey-construction-contracts-the-many-routes-to-bankability

Aayush Sharma, SignalFire — Beyond the billable hour — How AI is reshaping margins and models at law firms [14]

Hourly billing persists in complex, high-stakes or open-ended matters because the scope keeps evolving and the risk is asymmetric and dynamic, so pricing these engagements upfront requires embedding significant risk premiums, which often makes fixed-fee structures impractical

https://www.signalfire.com/blog/ai-is-redefining-billing-hours-at-law-firms

David Jones, Cybersecurity Dive, 25 July 2024 — CrowdStrike disruption direct losses to reach $5.4B for Fortune 500, study finds [18]

"Parametrix said the global IT outage linked to Crowdstrike will likely cost the Fortune 500, excluding Microsoft, at least $5.4 billion in direct financial losses"; "Cyber insurance will only cover 10% to 20% of the losses, based on large risk retentions and policy limits at many companies"; "CyberCube estimates the cyber insurance market will face preliminary insured losses of between $400 million and $1.5 billion, potentially the single worst loss in the cyber insurance sector over 20 years"

https://www.cybersecuritydive.com/news/crowdstrike-cost-fortune-500-losses-cyber-insurance/722396/

Amazon Web Services — Summary of the Amazon DynamoDB Service Disruption in the Northern Virginia (US-EAST-1) Region [19]

"The incident was triggered by a latent defect within the service's automated DNS management system that caused endpoint resolution failures for DynamoDB"; the event ran from 11:48 PM PDT on 19 October to 2:20 PM PDT on 20 October with three distinct periods of customer impact

https://aws.amazon.com/message/101925/

The Stack, quoting Lloyd's of London — Lloyd's cyber insurance exclusions flag systemic risk fears [20]

"The ability of hostile actors to easily disseminate an attack, the ability for harmful code to spread, and the critical dependency that societies have on their IT infrastructure, including to operate physical assets, means that losses have the potential to greatly exceed what the insurance market is able to absorb"

https://www.thestack.technology/lloyds-cyber-insurance-exclusions-state-backed-systemic-risk/

Lloyd's of London, 14 May 2024 — Market Bulletin Y5433, State-backed cyber-attack wordings [21]

"we wish to reiterate that policy language should be clear so that the scope of cover is understood by all the parties and the exposure is properly assessed and monitored by syndicates"; Lloyd's issued guidance to ensure that where state-backed cyber-attack coverage is underwritten it is done in a controlled and measurable way

https://assets.lloyds.com/media/6335bcb0-e2a2-4378-8328-1ddf54828f2f/Y5433.pdf

About This Reference List

Compiled August 2026. All URLs verified at time of compilation. Regulatory documents and standards specifications are subject to revision — check primary sources for the most current versions.

Some links to academic papers and vendor research may require free registration. Government and standards body publications are freely accessible.

Run it on your own last engagement

Take the last one that closed. List every surprise — everything that made somebody say hang on. Put each in exactly one of the four classes. Total the Class 1 column at your own internal cost.

That number is the risk capital your firm contributed without recording it. Almost nobody has it, and it takes an afternoon.

Scott Farrell · LeverageAI · leverageai.com.au