AI-Native Service Architecture
The Square, the Barbell, the Flywheel and the Membrane
Freeze what the customer buys. Leave the method generative.
Prove the promise at hard edges. Compound between engagements — without pooling client truth.
By the end of this book you can
- ✓ Draw your own square — eight perimeter fields, filled
- ✓ Place the barbell inside it, and say what the interior may and may not do
- ✓ Name what the flywheel has to show by engagement two
- ✓ Write the rights line the membrane needs, before the engagement starts
- ✓ Say which work must shrink the promise — or be declined
TL;DR
- • Fix the perimeter, free the interior. A named promise, authoritative inputs, exclusions, a commercial band, a time boundary and valid acceptance states stay frozen. Inside that square, the production method is free to search, generate, test, discard and regenerate.
- • Four structures, one system. The square makes commitment possible; the barbell makes interior latitude safe; the flywheel makes the economics improve; the membrane makes that improvement legitimate. Remove one and you get a specific, nameable collapse — not a slightly worse service.
- • Standardise the compiler, not the answer. Repeatability moves up a layer into the machinery — so the same square can produce more specificity rather than more sameness.
The Amorphous Shape
Every service you have ever sold had a shape. Nobody drew it — and the customer has been paying for the wobble.
Ask a partner to draw the shape of a consulting engagement. Not the Gantt chart, not the org chart of the delivery team — the shape. What you get back, if they are honest, is a blob. It bulges where a source system turned out to be undocumented. It stretches sideways where a stakeholder changed their mind in week five. It grows a lump where somebody spent three days chasing an edge case that nobody had budgeted for.
______
__/ \____
___/ \__
/ \____
project
That shape is not the work's fault. It is a contract term.
What time and materials actually says
Strip the professionalism off a time-and-materials engagement and the sentence underneath is this: we don't know exactly what reality will require, so you buy access to our people while we discover it. That is not a billing convention. It is a decision about who owns the supplier's production variance, made once, decades ago, and inherited ever since.
Watch how the shape stretches. A source turns out to be messy — more days. The requirements change — more days. An integration behaves strangely — more days. The consultant misunderstood something — more days. Someone has to investigate an edge case — more days. Every one of those is a real event. Every one of them lands, in this arrangement, on the buyer.
And so the machinery grows to manage it: project managers, resource plans, estimates, timesheets, burn rates, sprint plans, change requests, and long arguments about whether some newly discovered piece of reality was ever "in scope". Left unfettered, as Scott puts it, a project reaches out across the organisation and you get change requests — which is why you have so many project-management hours asking where are we going and what is happening.
The customer effectively owns a significant portion of the supplier's production variance.
That sentence is worth sitting with, because it is a reallocation claim rather than a complaint. Nobody is behaving badly. The variance is genuinely there, it genuinely has to be financed, and somebody has to hold it. The only question is who — and until recently there was only one economically sane answer.
Who has been paying for the wobble?
Everyone, and not quietly. The Project Management Institute's 2018 Pulse of the Profession found that 52 per cent of projects completed in the previous twelve months experienced scope creep or uncontrolled changes to scope — up from 43 per cent five years earlier1. Note the direction of travel. That deterioration happened during exactly the period in which the profession was getting more certified, more tooled and more methodological about project management.
Scope creep, over five years
of completed projects experiencing scope creep or uncontrolled scope change, five years earlier
by PMI's 2018 Pulse of the Profession — more discipline, worse containment
The two bad answers
Firms have tried to escape this, and the escapes have names. Bravado: quote a fixed price and eat the variance. Padding: quote a fixed price at triple the expected cost, price yourself out of every competitive deal, and still carry the tail.
This is not an AI-era discovery. It is the oldest lesson in construction contracting, where the lump-sum turnkey model has been stress-tested for a century. The contractor takes the cost and performance risk — and yet owners "may still face additional costs through change orders", because in truth, as A&O Shearman put it, there is no such thing as an absolute fixed price contract2. Contractors respond by pricing in significant contingencies; and where those contingencies prove insufficient, they are "naturally incentivized to seek opportunities to reopen the fixed price"2.
Two escapes, one destination
Bravado
- • Quote fixed against an unmeasured estate
- • Absorb whatever arrives, out of pride or hunger
- • Margin erodes silently until it doesn't
- • The price reopens anyway, late and badly
Padding
- • Price the fear rather than the work
- • Lose the competitive deals to someone braver
- • Win the ones nobody else wanted
- • Still carry the tail you couldn't describe
Neither is an architecture. Both are the same shapeless project with a different number on the front.
The diagnosis underneath
The entry transaction is broken by design, not by bad luck. Estimation and negotiation happen simultaneously, over unpaid speculative labour, so neither party can trust the number. Reality is discovered during delivery and turns into change requests, margin erosion and arguments about assumptions. And the number itself is unreviewable, because the senior who produced it was not estimating in any engineering sense — they were running a private database join in their head and emitting a price. Two capable seniors can produce two wildly different numbers for the same brief, both locally rational, with no substrate that would let a third person adjudicate between them.
We are not going to re-derive either of those diagnoses here; they have their own treatment. What matters for this book is the consequence. If the commercial object is opaque, every surprise becomes a negotiation, and negotiation is the most expensive way ever devised to handle a normal Tuesday.
Why it was tolerable — and what changed
For most of the history of professional services, this arrangement was approximately correct. Production really was the scarce input. Absorbing an extra thirty hours of thinking work meant absorbing thirty billable hours of an expensive person's life. A supplier who offered to hold that variance was offering to hold something genuinely expensive, and pricing it honestly would have killed the deal.
That premise moved. Querying a model that scores GPT-3.5-equivalent performance on MMLU fell from $20.00 per million tokens in November 2022 to $0.07 by October 2024 — a more than 280-fold reduction in roughly eighteen months3. Why cognition became abundant, what it does to strategy, and how long the curve holds are all arguments worth having; they are not this book's argument, and they get one paragraph here on purpose.
Key Insight
AI does not make the cost curve flat. It makes it flatter — and it changes which variable drives it.
That distinction is the hinge. The variance that historically killed fixed pricing — how many documents, how messy, how many candidate approaches, how many iterations before something passes — now lands mostly on cheap parallel machine work. What is left is the number of consequential judgments a human has to own. One of those is absorbable. The other has to be metered.
Which means the question who owns the production variance? has stopped being an inherited default and become a design decision. That is the whole book.
What is coming, so you are not waiting for it
The next chapter puts four structures on the table and shows what breaks when each one is removed. Part II carries a single production-shaped service all the way through them, twice, and then breaks it on purpose. And if you are already composing the objection some work simply cannot be bounded — you are right, and Chapter 15 is where that lives. It is not an afterthought; it is where refusal gets its own method.
One more thing before the architecture. Nothing here says the shape of the work has to be regular. The interior of a good engagement is irregular, wasteful and full of dead ends, and it should be. Irregular production and irregular commitment are two different things. They were only ever welded together for economic reasons — and those reasons have changed.
The amorphous shape was never the work. It was the contract.
Four Structures, One System
Square around the engagement. Barbell inside the square. Flywheel between the squares. Membrane around the learning.
Here is the whole architecture, before any of it is explained:
Square around the engagement.
Barbell inside the square.
Flywheel between the squares.
Membrane around the learning.
Four nouns you have no reason to trust yet. By the end of this chapter you will be able to draw them, and — more usefully — to break them.
Four views of one service
These are not four metaphors for the same thing, and they are not four best practices you can adopt in whatever order suits the quarter. They are four views of one service at different scales.
The square — the commercial boundary
A bounded promise, a price or unit, eligibility, exclusions, valid terminal states and acceptance conditions. It is what the buyer signs and what the supplier defends.
Its job: make commitment possible.
The barbell — the effort distribution
Heavy on intent and falsifiability at the beginning. Very light on prescribing the production middle. Heavy again on verification and consequential judgement at the end.
Its job: make the interior latitude safe.
The flywheel — what happens between squares
A completed engagement leaves reusable learning that changes the machinery the next engagement runs on: exceptions, evaluations, decision rules, delivery tooling.
Its job: make the economics improve rather than repeat.
The membrane — the boundary around learning
Client truth stays client truth. Reusable patterns cross outward only by abstraction, rights and a human-gated promotion decision.
Its job: make the improvement legitimate.
The canvas
Drawn together, with the write-back loop that turns one engagement into the starting position of the next:
Heavy
- • intent
- • boundary
- • falsifiers
- • kill conditions
- • wrong-if-broken
Light
- • generative middle
- • explore / build
- • retry / regenerate
- • research / branch
- • the machine does its thing
Heavy
- • proof
- • deterministic checks
- • human disposition
- • acceptance evidence
- • accepted outcome
↓ between the squares
extract reusable learning → de-identify / abstract / review → firm substrate, tests, tools → next square starts here
The deletion test
A composition claim is only worth making if it can be falsified. So here is the falsifier for this one: remove each structure on paper and name what specifically breaks. If a piece can be removed without breaking the other three, it was decoration and this book has oversold it.
| Remove… | Collapse mode | What the firm actually experiences |
|---|---|---|
| the square | Time-and-materials ambiguity | Every irregularity becomes a negotiation. The buyer finances the wobble again, and the firm's most senior people spend their week defending a scope line rather than making a call. |
| the barbell | Supervised automation | A human reviews every unit of machine output. You have paid for the machine and kept the labour cost — and the review is worse, because attention spread evenly is attention absent where it matters. |
| the flywheel | Repeated bespoke heroics | Every engagement starts cold. The same three people carry it. The second one costs what the first one cost, and the case study is the only artefact that survives. |
| the membrane | Client-data leakage | "We learn from every engagement" quietly becomes "we pool everybody's data". One clause in one contract, or one procurement questionnaire, ends the practice — and deservedly. |
Four different failures. That is the evidence that these are four different structures rather than one idea wearing four hats. You will run this test yourself in Chapter 17. Chapter 13 runs two of these collapses all the way to the bottom, on the book's own specimen, because a framework that cannot fail its own candidate is a brochure.
Three rules, and a family resemblance
Tight promise. Loose production. Hard acceptance.
If that rhymes, it is supposed to. It is the commercial sibling of a delivery rule this corpus already runs on: tight intent, loose method, hard verification. The resemblance is the point — the same insight, applied one layer up. It is also, as the next chapter argues at length, the reason the two get confused, and the confusion is expensive enough to deserve its own chapter.
Key Insight
The perimeter is productised. The interior is generative.
What this book is not doing
Two pieces of scope honesty, so nobody waits for them.
This book does not decide which commercial object should exist. Working out whether a candidate service has real customer friction behind it, whether AI genuinely constitutes it rather than decorating it, and whether it can compress a labour-priced unit of sale — that is qualification work with its own gates. This book starts after that choice has been made, when a candidate service is being constituted and somebody has to decide what shape it is.
Nor does it build three things it will name as it goes: the detailed contract test for boundary changes; the semantics of what makes an engagement wrong to start versus wrong to accept; and the finance decision behind deliberately under-earning on engagement one. Each of those is a seam, each gets one sentence where it appears, and none of them is smuggled in here at half-strength.
The size of the claim
Scott's own bound on this is worth having early, before the confidence starts:
"I'm not saying this is the only pattern. I'm saying it's probably a pretty good reusable construct for AI-native successor products and businesses — sort of a pattern, or a template."
That is the right altitude. This is a construct, tested against one generalised specimen and transplanted into a second domain in Chapter 14. It is not a law of services, and Chapter 15 is entirely devoted to the territory where it must shrink the promise or decline the work.
What the construct does do — and this is why it earns a book rather than a diagram — is tell someone designing almost any AI-native service where to put rigidity and where to deliberately allow chaos: rigid at the promise, fluid in the machine, rigid at proof.
The next four chapters take one structure each. We start with the one the buyer actually signs.
The Square: A Stable Unit of Commitment
The square is not “fixed price”. It is a perimeter explicit enough that surprise has to speak the product’s language before it can reach the invoice.
Eight questions. Read them flatly, before any of them is explained.
The perimeter fields
A service that cannot answer those eight questions is not a square. It is a wish with an invoice schedule.
Rigid, fluid, rigid
The outside is deliberately rigid. The inside is deliberately fluid. Then the exit boundary becomes rigid again.
Rigid at entry
promise · price or unit · time boundary · authoritative inputs · terminal states · exclusions · acceptance tests · authority
Fluid inside
search · decomposition · generation · retries · parallelisation · tool choice · code generation · research · testing · regeneration · internal replanning
Rigid at exit
evidence · acceptance · human disposition · delivered state
Definition
A stable commercial perimeter around an adaptive, machine-scale production interior.
Why the perimeter has to be somewhere
There is a reason a contract cannot simply enumerate everything and be done with it, and it is older than any of this. The incomplete-contracts literature — Grossman, Hart and Moore, whose work earned Hart and Bengt Holmström the 2016 Nobel — begins from the observation that contracts "cannot specify what is to be done in every possible contingency", and that at the time of contracting "future contingencies may not even be describable"4. Hart frames the resulting trade-off precisely: "The benefit of a rigid contract is that it fixes expectations, avoiding arguments. But it may not perform well when there is uncertainty. A flexible contract can adjust to the state of nature, but there is also room for arguments."5
Read that as a dilemma and you are stuck with the industry's two bad answers. Read it as a placement problem and it dissolves — because expectations and uncertainty do not live in the same place. Expectations live at the boundary: what is promised, what arrives, what counts as done. Uncertainty lives in the middle: how many sources, how messy, how many attempts, which approach. Put rigidity where expectations live and flexibility where uncertainty lives, and you have a topology rather than a compromise.
That placement was not economically available while the middle was made of expensive human hours. It is available now, which is the only thing that has actually changed.
The critical misreading to head off immediately: the square does not mean the work inside it is square. The interior may be horrendously irregular. It can explore dead ends, discover that an assumption was wrong, regenerate a whole approach, run ten parallel attempts, throw away eight of them, read another hundred documents, and rewrite its own architecture twice. None of that reaches the buyer. That is the change.
Absorber, not hedge
It is tempting to describe AI here as a hedge, and the instinct is right but the word is wrong. A financial hedge offsets risk. AI does not guarantee your risk decreases. What it does is radically reduce the marginal cost of responding to many forms of variance — which makes it elastic capacity, or more precisely a variance absorber.
Be exact about which variance. Historically dangerous questions — how many documents do we have to read? how messy are they? how many candidate solutions should we explore? how much code needs writing? how many mappings need checking? how many alternatives should we test? how many iterations until this passes? — increasingly fall on cheap parallel machine work rather than on expensive linear expert labour. An extra thirty hours of thinking need not mean thirty more billable consultant hours. The machine can simply do more work inside the same commercial envelope.
What remains scarce is the number of consequential judgments a named human must own. That is the metered resource, and it is the one the perimeter has to price. The machinery for doing that — measuring the input surface before quoting, assigning a band, including a defined quantity of human disposition work, and holding a typed reserve for defined classes of surprise — already exists as an organ and sits inside this perimeter. This book places it; it does not re-derive it.
Isn't this just fixed-price consulting?
No — and the correction matters enough to be a rule rather than a caveat.
Correction
Fixed price is a consequence, not the essence. The deeper invariant is a stable unit of commitment.
A stable unit of commitment can be a fixed-price decision product. It can also be a per-project activation, a verified decision, an annual continuity subscription, an assessed estate, an enrolled fleet, a per-machine service, a protected period, a guaranteed response commitment, or a capacity tier. What it increasingly must not be is "however many hours our internal process happens to consume."
So why does fixed price keep showing up? Because it is unusually strong evidence that a supplier has made its own complexity legible enough to take responsibility for it. A firm that can quote a fixed number against a measured input surface has demonstrated something a brochure cannot. That makes it a signal, not a religion — and a candidate that can only be sold as open-ended hours has not yet crossed into being a commercial object at all.
The three moves that answer the objection properly: the invariant is the unit, and this chapter has just listed ten of them that are not fixed fees; the fee shape is not what changed, the affordable variance is; and the test is whether your perimeter can absorb an interior surprise without a conversation. If it cannot, you have a price, not a square.
So when is it a change request?
This is the most useful thing the perimeter does. It reclassifies surprise.
"If I promise to analyse your estate and one part turns out to require twenty agent-hours rather than two agent-hours, that's my problem. The square absorbs it."
Two categories, different in kind
Interior variation — the supplier's problem
- • An approach fails its test and the system generates another
- • An integration needs three regeneration attempts rather than one
- • The evidence base needs rebuilding because the first assembly was wrong
- • Fifteen thousand pages instead of three thousand
- • Two authoritative sources disagree and reconciliation takes four passes
None of these is a change request. The square absorbs them, and the band was priced knowing they happen.
Boundary mutation — a commercial event
- • The buyer changes the promised outcome
- • The input estate crosses the agreed volume band
- • A new source class is introduced
- • The required authority changes
- • A physical constraint changes
- • A regulatory obligation appears
- • The buyer moves the deadline
- • A typed reserve is exhausted
Each of these moves a named perimeter field, and each reopens the commercial conversation with evidence attached.
"So I'd replace a lot of change-request logic with boundary-mutation logic."
That is a considerably cleaner distinction than the one most delivery organisations run on, and it is available to a reader immediately: for every surprise in your current engagement, ask which named perimeter field moved. If none did, it was interior. The detailed contract test for boundary changes — how reserves are typed, how consumption is measured, what exhaustion triggers — is a seam that gets its own treatment; the point here is only that the two categories are different in kind and must never be handled by the same process.
How unknowns leave the square honestly
A perimeter that can only emit success is a liar. The exit vocabulary has to carry the states reality actually produces: directly mapped, partially mapped, not observed, insufficient evidence, inaccessible within the audit boundary, unsupported source type, ambiguous and requiring a human decision, explicitly excluded from this phase. Unknowns become typed deliverables rather than unbounded labour, and one clause carries more weight than the rest combined: not observed does not mean does not exist. If your product allows that slide, you have created both a liability and a false map.
Again: referenced, not re-taught. Typed uncertainty is an organ and it lives inside this perimeter. What the square adds is the reason it must be there — a rigid boundary is only honest if the states crossing it are honest.
The perimeter now exists. The obvious next question is why anyone should trust what happens inside it. Before that can be answered, one confusion has to be dragged out and shot.
Topology Is Not Cadence
“Waterfall per increment” describes a production cadence. The square describes a commercial topology. They coexist — and conflating them wrecks both.
Two sentences, both of which sound reasonable, both of which are the same mistake in opposite directions:
Misreading A
"We've gone fixed outside, so we've frozen requirements for the year."
Misreading B
"We work in two-week increments, so we can't commit to a bounded outcome."
Each answers a question from one layer with a fact from the other. This chapter exists to make the layers impossible to confuse, because both of these misreadings are expensive and they are expensive in different currencies.
Two layers, two questions
Cadence answers: how does the interior learn? It is a delivery rhythm — the size of a slice, when feedback arrives, what gets verified before the next slice commits. Topology answers: what did the buyer buy? It is a commercial shape — where rigidity sits, what absorbs variance, what reopens a price.
A firm can change its cadence every quarter without touching its topology. A firm can change its topology without altering a single ceremony. They interact, but they are not the same object, and they are owned by different people for different reasons.
The cadence framework this corpus already carries runs each increment through understand → specify → design → generate → verify → deploy → learn, under tight intent, loose method, hard verification — with learning written back so that engagement two starts higher than engagement one. That is good practice at the delivery layer and this book does not re-derive it. It is worth saying why the temptation to merge the two is so strong: the frameworks rhyme. Tight intent, loose method, hard verification; tight promise, loose production, hard acceptance. Same shape, different altitude. Rhyme is exactly what makes two things easy to mistake for each other.
Misreading A: the perimeter as a requirements freeze
Here is how it goes wrong. The perimeter gets drawn properly — promise, inputs, exclusions, acceptance. Then somebody reads "the promise is stable" as "our understanding is complete", and discovery gets pushed into a design phase with a sign-off gate at the end of it.
What follows is predictable. The interior stops learning, because learning would require admitting the design was incomplete. Evidence that arrives in week six is treated as a threat rather than as information. And by the end, the team is defending the promise instead of keeping it — which is a subtly different job and a much worse one.
The correction is one sentence: the promise is frozen; the world model is not. The interior is supposed to change its mind. That is what the latitude is for. Nothing in a fixed perimeter says the supplier must know, on day one, how the answer will be produced.
Misreading B: increments as an excuse not to commit
The mirror image, and more common. A delivery lead who is honest about their own process points out — correctly — that nobody can predict what week six will contain. From this they conclude that the firm cannot promise anything, and the commercial team goes back to selling attention.
The correction is also one sentence: unpredictable method is not unpredictable promise. The reason the buyer cannot get a commitment is not that the work is uncertain. It is that nobody separated the two kinds of uncertainty — the kind that lives in how, which the supplier should absorb, and the kind that lives in what, which is exactly what the perimeter fixes.
The same engagement, two layers
| Commercial topology | Production cadence | |
|---|---|---|
| What is fixed | The promise, the inputs, the exclusions, the acceptance states | The slice discipline: specify, generate, verify before the next slice commits |
| What is free | Everything about how the state is produced | Which increment comes next, chosen on evidence |
| Who owns it | Commercial and practice leadership | Delivery leadership |
| What "a change" means | A named perimeter field moved — a commercial event | The next increment goes somewhere else — a normal Tuesday |
| What "done" means | Acceptance evidence exists and could have failed | This slice is verified and deployed |
| What failure looks like | The price reopens without evidence; the promise is renegotiated | A large wrong system, verified late |
What actually died, and what didn't
Scott's version of the history is blunt:
"We put all this stuff on top to manage the delicate genius. That model's gone away."
That is too narrow as a history of Agile, and the correction arrived in the same conversation that produced this book — which is worth showing rather than tidying away, because the correction is the useful part:
The correction
The part of Agile that survives AI is its deepest epistemic insight: users still don't completely know what they want until they see something. That uncertainty does not disappear because coding became cheap.
What depreciates is different: the labour-management layer.
Sprint capacity, story points, developer availability, backlog sizing, velocity, estimation, hand-offs, ceremonies, resource scheduling — a great deal of that machinery exists to allocate expensive, scarce human production capacity. When generation becomes cheap, a large wrong system can appear as quickly as a large right one, and the binding work moves upstream to framing and downstream to verification. Ceremony whose economic foundation was scarce implementation should have to re-earn its keep. The world-learning loop — build something, expose it to reality, learn, change direction — keeps its keep easily.
Notice that both of those are cadence claims. Neither of them tells you what to sell.
Why the conflation is expensive
Two costs, and they are real money.
It lets ceremony masquerade as a commercial answer. A buyer asks what they will have and when, and receives a description of your internal rhythm: two-week sprints, a demo every fortnight, a steering committee monthly. None of that is a commitment. The buyer usually cannot articulate why it feels unsatisfying, so they ask for a discount instead, and the supplier concludes the market is price-sensitive.
It makes commercial discipline look like an attack on iteration. Delivery teams who hear "fixed perimeter" reach for their Agile principles and resist, because the last three times someone fixed something it was a Gantt chart with a year on it. They are defending the right thing at the wrong layer, and the argument that follows is unwinnable by anyone.
The test
If your answer to "what is the shape of the commitment" mentions a ceremony, a sprint length or a stand-up, you have answered the wrong question.
Is this semantics? Only if you ignore the consequences. Misreading A costs you the interior's search advantage — the exact thing you bought the machine for. Misreading B costs you the ability to sell anything except time. Two opposite, expensive failures produced by one blurred distinction is the definition of a load-bearing one.
The layers are separated. Now the question that Chapter 3 left hanging can finally be asked: with the perimeter fixed and the interior free, what stops the interior from being ungoverned?
The Barbell: Heavy Ends, Loose Middle
A fixed perimeter with a free interior sounds like a governance failure with better branding. It is the opposite — and the reason is where the weight sits.
Take the objection seriously first, because it is the objection every risk committee will make and most of the people making it have earned the right. A supplier that fixes what it owes and then says how we get there is largely our problem has just described, to a compliance-minded reader, an unsupervised black box with a signature underneath it.
The answer is structural rather than reassuring. The barbell is not less governance than supervised automation. It is differently placed governance — and in practice it is more.
Heavy, light, heavy
The pattern that emerged from AI-assisted software work is not the model writes, then a senior carefully reads every line. It is closer to a barbell: heavy thought before implementation, cheap implementation in the middle, heavy verification after.
Left plate — design intensity
- • Intent, promises, non-goals
- • Alternatives and layers
- • Assumptions and failure cases
- • Acceptance tests
- • Recorded rejections
Middle — execution
- • The interior does its thing
- • Exhaust accumulates
- • Exceptions route to humans
- • No theatrical completeness
- • Nobody inspects every unit
Right plate — verification
- • Independent lenses
- • Deterministic checks
- • Adversarial review
- • Evidence coverage
- • Intent regression
The recognition that this is the same object as the square, seen from inside, is the moment four ideas became one architecture:
"The more we talk about the square shape of the fixed-price product, the more it looks like the barbell. Heavy on intent, heavy on testing, light on what happens in the middle — to let AI do its thing."
The square says what is fixed. The barbell says why that is affordable: you spend expensive precision at the boundaries, and then you deliberately stop micromanaging.
What latitude actually permits
"Latitude" is a word that means nothing until it is enumerated. Here it is, in the form you could hand to a delivery lead:
The interior is permitted to
Choose its own method and sequence. Search widely rather than narrowly. Generate multiple candidates and run them. Discard what fails. Parallelise. Select and build its own tools. Regenerate after failure rather than patching. Replan internally without asking. Branch when an unexpected constraint appears.
Latitude is not
Permission to change the promise. Permission to accept a consequential state. Permission to decide what counts as evidence. Those three are the perimeter, and the perimeter is not inside the machine's authority.
The authority split, written down
If this is not written down, it will be improvised under pressure, and it will be improvised differently by each person. So write it down.
| The machine is authorised to | The machine is not authorised to |
|---|---|
| Inspect the situation and identify assumptions | Select the binding design or commit the firm to it |
| Produce counter-proposals and layer shifts | Change formal status or approve a milestone |
| Compare options against prior evidence | Commit expenditure or vary the commercial band |
| Surface failure shapes and stakeholder effects | Communicate commitments to the client |
| Supply falsifiers and evidence for a decision | Alter live operational systems |
| Nominate a disposition for a human to make | Bypass accountable professional judgement |
Key Insight
The customer buys precision of intent and precision of outcome. They do not buy precision of internal procedure.
What this deletes from your contract
Traditional statements of work do almost the reverse of this. They try to protect the supplier by describing activities: conduct five workshops, interview twelve stakeholders, develop the design, run the analysis, hold weekly status meetings, prepare the report. Every one of those is a description of the middle — the exact part the buyer should not be buying, and the exact part the barbell says to leave alone.
Cut them. Here is what goes in their place:
The replacement
Here is the state we will establish. Here are the evidence boundaries. Here are the valid uncertainty states. Here is what would invalidate the engagement. Here is what constitutes completion. How we exhaustively get there is largely our problem.
Six sentences. A contract built on them is thinner in the middle and much heavier at the ends than the document it replaces — which is the whole design.
Where to stay prescriptive
The rule underneath all of this is to be precise about purpose and to stop over-specifying procedure. But latitude is a judgement about a class of work, not a blanket setting, and pretending otherwise would make this whole argument naive. Stay prescriptive when the output feeds a machine and must match an exact schema; when there is a single correct answer, such as a calculation or a lookup; and when a constraint is safety- or correctness-critical, in which case it is a hard constraint stated flatly rather than latitude.
Most over-specification is the accident of dictating the parts that need judgement because dictating feels safer — when it is precisely what is holding the quality down.
The regulator's version of the same shape
For a class of work this stops being a design preference. The EU AI Act requires high-risk AI systems to be designed so that natural persons can effectively oversee them — understanding the system's capabilities and limitations, remaining aware of automation bias, correctly interpreting output, deciding not to use the system, and being able to intervene or halt operation — with separate verification by at least two natural persons for certain uses6.
Read what that statute actually describes. It is not a requirement to review every unit of output. It is a requirement for comprehension at the front and intervention capability at the back: plates, not a uniform supervision layer. Where your service touches a high-risk use, the barbell is not your architecture — it is the law's.
Pitfall: supervised automation
Remove the barbell and this is what you get. A human reviewing every unit, because nobody froze intent at the front and nobody owns a consequential acceptance at the back. It costs the labour the machine was meant to remove, and it buys worse assurance: attention spread evenly across everything is attention absent from the things that matter.
The tell is a review process whose workload scales linearly with machine output. If your quality plan gets more expensive every time the model gets faster, you have built the wrong shape.
So the answer to the risk committee is not trust us. It is: here is frozen intent with the rejected alternatives preserved; here are the deterministic checks; here are the independent lenses; here is the named human who signs the consequential dispositions. A committee that insists on uniform review is buying the appearance of control at the price of the control that would actually work.
The left plate is loaded and the middle is free. Everything now rests on the right plate — which is where most firms have never actually built anything.
Hard Acceptance: The Outcome Is Not the Test
An engagement everyone was happy with, delivered on time, praised in the steering committee — and no artefact anywhere that could have said no.
Picture the closing meeting. The work landed. The client is pleased. The partner is pleased. The invoice goes out and nobody argues. Now ask one awkward question: what, exactly, was accepted?
In most engagements the honest answer is a feeling, held by several senior people at roughly the same time. That is not a defect of those people. It is a missing instrument.
The distinction that fixes it
Definition
The outcome is what was sold. The test is the instrument by which we can legitimately say the outcome occurred.
If I sell you a verified current-state architecture, the architecture is the outcome. The evidence coverage, the reconciliation checks and the human dispositions are the tests by which I can prove it.
Collapse those two and "outcome-based" reverts to vague consulting language, because there is nothing left that could fail. Keep them apart and three things become possible that were not possible before.
The buyer can tell completion from exhaustion. Right now, most engagements end because the calendar ended. A test that can pass or fail replaces "we've run out of weeks" with "the state exists, here is the evidence".
The supplier can defend the perimeter. A square with no closing event is a square with a soft edge; anything can be argued back into it after the fact. The acceptance test is what makes the boundary a boundary at the moment it matters most.
Disputes stop resolving to seniority. Without an instrument, disagreement about whether the promise was kept is settled by whoever is more senior or more determined. With one, it is settled by looking.
The sequence
Promise → Falsifier → Latitude → Evidence → Acceptance
Walk the arrows. The promise names the bounded state. The falsifier names what would make this the wrong engagement, or the wrong result, before anyone builds anything. Latitude is everything the interior is then free to do. Evidence is what the interior must produce as it goes, not reconstruct afterwards. Acceptance is the event at which a named human says the state exists.
Without the second arrow, "outcome based" can still turn into vague consulting language. The semantics of falsification — what makes an engagement wrong to start versus wrong to accept, and how the two differ — is a distinct piece of doctrine with its own treatment. This book places the falsifier in the sequence and stops there.
The test has to be able to fail
A work package should close against an observable receipt: a reconciliation within a stated tolerance, a report matching an agreed definition, a data product that exists under named criteria. Not "stakeholders are happy". If completion resolves to satisfaction, the discipline of the entire square evaporates at the last mile — and it evaporates at the exact moment the money changes hands.
Three diagnostics you can run on your current acceptance criteria this afternoon:
- Could a competent third party run it? If it requires the author to interpret it, it is a preference, not a test.
- Could it come back negative? Write the failing result as a sentence. If you cannot, there is no test.
- Would a negative result change what the buyer does? If not, you are measuring something nobody cares about, which is its own kind of theatre.
Pitfall: criteria that cannot fail
"The final report is delivered and presented to the steering committee." → Delivery is an event, not a state. Repair: name the state the report must establish, and the coverage threshold it must reach.
"Stakeholders confirm the recommendations are actionable." → Unfalsifiable by construction. Repair: name the decision the recommendation supports, and the evidence class required for that decision to be defensible.
"All identified issues are addressed." → Circular: identified by whom, against what boundary? Repair: coverage against a declared boundary, with typed unknowns for whatever was not observed.
Two instruments, not one
The right-hand plate carries two different tools, and naming them separately is what stops "human in the loop" from being a mood.
Deterministic checks
- • Coverage — was every declared source actually read?
- • Reconciliation — do the numbers tie within tolerance?
- • Completeness — is every required field populated or explicitly typed?
- • Provenance — does every claim carry an openable pointer?
- • Conformance — does the output match the agreed schema?
Machinery. It passes or fails. Nobody's seniority is involved.
Human dispositions
- • The consequential calls, made under authority
- • Each one bounded, named and recorded
- • The machine may investigate, synthesise and nominate
- • The machine may not become a disposition
- • The machine may not mutate an authoritative state
Judgement. It is the metered resource the band was priced against.
Honest unknowns are part of the product
A perimeter that can only emit success is a liar, and Chapter 3 gave the exit vocabulary that fixes it: unknowns become typed deliverables rather than unbounded labour, with not observed meaning exactly that and never "does not exist". What is worth adding here is the buyer psychology. Clients often hear typed unknowns as incompleteness — until they have been burned once by an over-confident report, at which point they become the strongest advocates for typing you will ever meet. Incomplete false confidence is the risk. Typed incomplete knowledge is the asset.
But AI is fast — can't acceptance be light?
No. And there is now a measurement rather than an intuition behind that.
METR ran a randomised controlled trial with experienced open-source developers working on their own repositories. Using early-2025 AI tooling, they took 19 per cent longer than without it. Beforehand they forecast a 24 per cent speed-up. Afterwards they believed they had been sped up by 20 per cent7.
Forecast, felt, measured
speed-up developers forecast before the trial
speed-up they believed they had experienced afterwards
actual measured change — they were slower
METR, randomised controlled trial, 246 tasks. The authors describe it as a snapshot of one relevant setting with early-2025 tooling, not a universal law — and this book uses it that way.
Be careful with what that does and does not show. It is one population, one class of work, one moment in a fast-moving capability curve. It does not show that AI does not work. What it shows is the gap between felt and measured — and that gap is the entire reason the right-hand plate exists. If your acceptance evidence is a demo and a good feeling, you have built a machine for generating confident wrongness at scale.
The structural version of the argument is simpler than the empirical one, and does not depend on any study: the faster the interior, the more output arrives per unit of human attention, so the cheaper it becomes to be wrong at volume. Speed is a property of production. Acceptance is a property of consequence. They do not trade against each other.
None of this is a house invention. Treating evaluation as a first-class, continuous part of the lifecycle — extending test-driven and behaviour-driven practice to cope with non-deterministic behaviour and post-deployment drift — is a recognised engineering discipline with its own literature8. The commercial version of that discipline is the right-hand plate.
If you are on the buying side
Four questions, before you sign a bounded promise:
- What is the acceptance evidence, specifically, and what form will it arrive in?
- Who signs it — by name and role, on both sides?
- What result would count as a failure? Ask them to say it out loud.
- What happens commercially if it fails?
A supplier who cannot answer the third question has not sold you an outcome. They have sold you an activity with an optimistic adjective.
The square is now closed and provable. One engagement can be sold, delivered and accepted — which is exactly where most firms stop, and exactly where the economics have not yet happened.
The Flywheel: What Engagement Two Has to Show
Compounding is not “we captured some learnings”. It is a measurable change in the starting position of the next engagement — and one engagement cannot evidence it.
"You might have had centres of excellence. You might have had community chats across the company. People were encouraged to share — but it didn't really work. Everyone wanted to be the hero."
Every partner over forty has lived that sentence. The temptation is to conclude that organisational learning is a fantasy and that firms are simply collections of individuals who happen to share a logo. That conclusion is wrong, and the correct version is more useful.
The joins were economically broken
Organisations genuinely tried. Communities of practice. Post-implementation reviews. Reusable methods. Knowledge bases. Centres of excellence. Methodology teams. Document libraries. Lunch-and-learns. The attempts were real, sustained and often well funded. So the honest diagnosis is not that organisational learning did not exist. It is that the joins were economically broken.
Look at what actually had to happen for one lesson to reach one future colleague:
Nine links, every one of them a person's afternoon
- Notice it was reusable
- Articulate it
- Abstract it out of client-specific language
- Document it
- Classify it
- Distribute it
- Persuade a large group of busy people to read it
- Have one of them remember it at exactly the right future moment
- Apply it correctly
Every link was human labour. A nine-link chain of voluntary human effort has a throughput close to zero, which is why so much knowledge management became archaeology.
Key Insight
Nobody has to remember to go and read it. The next run loads the substrate.
Be precise about which links change. AI shortens links one to six — noticing, articulating, abstracting, documenting, classifying, distributing all become cheap. It eliminates links seven and eight, and that is the structural discontinuity: nobody has to be persuaded to read anything, and nobody has to remember at the right moment, because the next engagement's machinery loads the material as a starting condition. Link nine — applying it correctly — is still judgement. Which is why the barbell survives.
The substrate is in the cost curve
That changes what the firm's knowledge base is. It stops being the company library and becomes part of the cost curve of the product.
Concretely: an exception observed in engagement one becomes a typed class, a decision rule, a regression test and a routing condition. Engagement two loads all four before anyone starts. The senior who resolved it the first time is not consulted the second time, and the third engagement does not know the exception ever existed as a problem. A library is a cost centre. A substrate is a cost driver.
Count the right thing
The industry's default denomination for this is hours saved. Professionals surveyed for Thomson Reuters' Future of Professionals Report 2025 projected time savings of five hours a week within the next year — about 240 hours annually, an average annual value of roughly $19,000 per user9. Those are real forecasts from real practitioners, and they are the wrong instrument for this argument.
The reason is mechanical rather than sceptical. Saved time arrives as confetti — eleven minutes here, twenty minutes there, scattered across a lot of people and a hundred small tasks — and confetti does not consolidate. Nobody hands their reclaimed eleven minutes back. It gets absorbed into a slightly longer lunch and one more meeting that expands to fill the gap.
Count what the engagement adds in kind instead: decisions that got better because the context to make them well was finally in the room; institutional knowledge recovered from heads and dead drives; questions answered that nobody previously knew to ask; the density of the substrate itself; and reuse across futures that did not exist when the asset was built. None of those rows can be inflated by a self-reported baseline, which is precisely why they survive a sceptical finance partner.
Buying the learning
"You might run the first few iterations of your AI successor product inside the square at a loss, or near break-even, to gain the substrate for future runs. To some extent, you're just buying the learning."
That is a real strategy and it has an obvious failure mode, which is why it comes with the hardest condition in the book attached:
Key Insight
A first engagement at break-even or a loss is rational if the loss is purchasing reusable capability. It is irrational if it is purchasing heroics.
And it has a duration. Not five years, not two years, not even one: the first two engagements, perhaps per industry, and then the ledger has to show something. How that subsidy is funded, accounted for and killed if it fails is a distinct piece of work with its own treatment. What this book needs is narrower — that the flywheel has an economics, and that the economics has a kill condition.
Two seniors, eight weeks, barely break-even
✓ An excellent investment, if it leaves
- • A functioning delivery vessel
- • A usable set of evaluation cases
- • Newly typed exception classes
- • Reusable deterministic tools
- • Corrected pricing-band logic
- • Reusable acceptance tests
- • A better qualification gate
- • Escalation classes encoded so they do not recur
✗ A low-margin consulting job, if instead
- • Two very good people worked very hard
- • Everything was solved by hand, brilliantly
- • The client was delighted
- • The lessons were shared over drinks
- • The case study is the only surviving artefact
Same P&L line. Completely different asset position.
What engagement two has to show
Eight quantities. This is the measurement set; Chapter 12 fills it in for the book's specimen as a worked model.
| Measure | What "improving" looks like |
|---|---|
| Estimate error | The gap between the band assigned and the effort consumed narrows, because the first engagement corrected the band drivers |
| Production variance absorbed | More irregularity handled inside the square without reaching the buyer at all |
| Human disposition load | Fewer consequential calls per completed unit — eventually. It usually rises first; see below |
| Exception classes | Fewer novel classes; more arriving pre-typed with a known route and owner |
| Acceptance evidence | The harness is reused with local cases rather than rebuilt from scratch |
| Reusable artefacts | Named things left behind, in locations the next run actually reads |
| Recurring escalation rate | Falls for the classes that were fossilised; flat for genuinely new ones |
| Scarce-expert density | Consequential dispositions per completed unit trends down without quality collapsing |
The qualitative form of this is that engagement two is the only proof: transfer means ordinary capable staff lead materially more of it because shared infrastructure changed, not because the same heroes worked late again. The quantitative honesty check is a ratio — paid bounded units divided by scarce expert dispositions — with strict discipline on both numerator and denominator. Both are organs. This book requires them; it does not re-derive them.
One engagement proves delivery, not compounding
Say it plainly, because it is the claim most likely to be over-sold. A successful first engagement demonstrates that the promise can be kept. It demonstrates nothing at all about whether the next one will be cheaper, and treating it as evidence of a compounding curve is how firms end up scaling heroics into the installed base.
What would falsify the compounding claim, concretely: escalation returning at the same density with the same people; exception classes that are still tribal knowledge rather than typed routes; pricing rules that bend whenever a principal is in the room; and a "playbook" that turns out to be a slide deck. Any two of those and the flywheel is a story.
The gate to apply at the end of every engagement is a single question: did we leave merely another codebase — or a sharper corpus that makes the next project easier and better? A loop that piles findings into chat history is linear activity. A loop whose outputs become structured substrate that the next loop reads as priors is compounding. That is the difference between an activity and an asset.
The part that is a culture problem
Traditional professional services rewards knowledge concentration. If you are the only person who knows the obscure trick, your utilisation rises, your indispensability rises, your promotion prospects rise. "Please document everything you know for the benefit of the organisation" is then asking people to reduce their own leverage, and asking politely has been tried for thirty years.
The highest-status act becomes not I solved the hard case, but I solved the hard case once and made sure nobody needs me for that class again.
That only works if the firm actually measures the second thing, which is what the escalation-rate and disposition-density rows are for. Scale that across a large bench and the argument changes shape entirely — but bench-scale transfer machinery has its own treatment and is not re-taught here.
A normal consultancy monetises accumulated expertise repeatedly as labour. An AI-native consultancy converts each paid engagement into infrastructure that reduces the amount of expertise that must be re-spent on the next one.
Which brings us to the sentence that makes all of this dangerous. We learn from every engagement is one clause away from we pool everybody's data — and buyers know it.
The Membrane: Three Territories, One Gate
“We learn from every engagement” is one clause away from “we pool everybody’s data”. The membrane is what makes the first true without making the second true.
Somewhere in a procurement review, a general counsel will ask the question in its sharpest form: when you say you learn from every engagement — learn what, exactly, and where does it go?
The answer needs to be architectural, because a reassuring answer is worth nothing:
"We deliver the project for the client, but we recapture the learnings — not on the client's data, on how the project went and the shape of the project. And we obviously have to de-identify."
Why "we anonymise it" is not an architecture
Anonymisation is a transformation applied to something that has already crossed. It says nothing about what may cross, who decides, in what form, or what happens to the candidates that are refused. It is a promise about handling, offered in place of a boundary — and a promise about handling is exactly what a buyer cannot audit.
A membrane is different from a policy. It has territories, a gate, and a set of decisions that each candidate must terminate in.
Three territories
1. Client truth
Their data, systems, people, decisions, evidence, context. Stays client-bound according to the engagement and the contract. There is no clever architecture that makes this cross; if your compounding depends on it crossing, stop reading and go and speak to your lawyers.
2. Engagement learning
What happened in this particular delivery: exceptions, failures, decisions, traces, project evidence. Much of this should also remain inside the engagement boundary. It is the territory people forget, because it feels like ours rather than theirs — and a trace of how a specific client's decision was reached is frequently more revealing than the data underneath it.
3. Promotable capability
An abstracted pattern that no longer depends on the client's confidential state:
- • "When this condition occurs, use this test."
- • "This source class needs this adapter."
- • "This pricing band underestimates disposition load when ambiguity exceeds these conditions."
- • "Do not make this architectural promise without this evidence."
- • "This delivery sequence failed because these constraints interact."
- • "This is a new exception class."
Only this territory is a candidate to cross. And even then: not automatically.
The gate
Five conditions, in order. A candidate that fails any one of them does not cross.
| Condition | What failure looks like |
|---|---|
| De-identification — client, people, systems and identifying detail removed | A "generalised" pattern that three people in the industry could immediately attribute |
| Abstraction — does the pattern survive without the specifics? | Stripping the fingerprints destroys the meaning, which means it was never a pattern |
| Transferability — would this help a different engagement? | A one-client peculiarity promoted because someone found it interesting |
| Clearance — contractual and confidentiality position checked | Nobody read the clause, and the answer is discovered during a renewal |
| Human approval — a named owner decides | Promotion by accretion: it ended up in the substrate because nobody stopped it |
Five dispositions, one of them mandatory
Every candidate ends in exactly one of these.
Local-only
The value is real only inside one client's context, or the fingerprints cannot be removed without destroying meaning. What ships: nothing outward — but keep the process prior ("how we found this") and the reason promotion was refused, so the next person does not re-litigate it blindly.
Configurable
The capability is shared; legitimate variation belongs in data or policy rather than in forked code paths. Test: can two deployments share one implementation and differ only by configuration?
Internal primitive
It helps delivery repeatedly — a harness, an evaluation suite, an exception taxonomy, a checklist — but must not be sold or supported as a customer-facing promise. Many of the most valuable field lessons live here.
Supported platform
Only when recurrence, strategic fit, clean ownership, de-identification and contract clearance, operability, documentation and an accountable promotion owner all clear. Missing any of them means you have hope with a release tag.
Reject
Wrong layer, unsafe, uneconomic, off-strategy, or a one-client vanity feature dressed as a platform need. Preserve the rejection. Rejections teach the boundary; a deleted backlog ticket teaches nothing.
Two rules make this a decision system rather than a backlog. First: "maybe later" is not a disposition — it is a deferred decision, and it must carry a revisit trigger. Second: the reject is an output, not an absence. A firm that cannot show you its rejections is a firm whose substrate is accumulating by accident.
The write-back path
From exception to next engagement
- machine Engagement one → exception observed
- human Significance verified — is this a pattern or an accident?
- machine Client details stripped
- machine Pattern abstracted into a condition and a route
- human Disposition decided by a named owner
- machine Firm substrate promoted — test, tool or rule updated
- machine Proposal, delivery, pricing and qualification systems see it
- machine Engagement two starts from there
Six machine steps, two human ones. Those two are the gate — and they are the only two that cannot be automated without dissolving the boundary.
This is not training on client data
The analogy that gets reached for is the model providers' bargain, and it is worth putting side by side with what a service firm should actually be doing.
The model provider's bargain
"Give us your interaction data and we'll subsidise inference, because the data improves our future model." The learning goes into weights. It is not inspectable, not reversible, and not attributable.
The service firm's version
"We may accept lower early product margin because the engagement improves our external organisational learning substrate." The learning goes into a wiki, an evaluation set, an acceptance harness, a pricing configuration, delivery code, adapters, tests, a pattern ledger, offer definitions, qualification rules, prompts and a toolchain.
You very often do not need to train a model at all. And that is not a limitation — it is the better mechanism, for five reasons that all matter commercially. External learning is immediately inspectable: you can read it. It is immediately reversible: a bad promotion can be removed this afternoon. It is attributable: you know which engagement produced it and who approved it. It is client-boundary aware: it lives at a level of abstraction the contract permits. And it is portable across model upgrades: when a better model arrives next quarter, all of it still applies, and applies harder. It is soft-weight learning at organisational scale.
Reciprocity, not extraction
The ethical test is not complicated. The customer receives useful delivery now, with explicit learning rights, clear IP and de-identification boundaries. Anything else is undisclosed vendor research funded by someone who thought they were buying a service.
This is also, usefully, where the professions already are. Practitioner guidance on ABA Formal Opinion 512 and equivalent state rules treats generative AI as non-lawyer assistance subject to supervision obligations, and advises that client confidential information only be entered into platforms that contractually commit to zero data retention and zero training on customer inputs10. Standard vendor terms now routinely cover deletion of customer data, prompts, outputs and derived training data, certification of deletion, and audit rights11. The membrane is what it looks like when a firm builds that line into the service rather than bolting it onto the master agreement afterwards.
The clause, in shape
What a firm should be able to put in front of a client before engagement one
What stays: your data, systems, documents, decisions and the record of this engagement. Named, and bounded to this engagement.
What may leave, and in what form: abstracted patterns that carry no identifying detail — a condition, a test, a rule, an adapter, an exception class — and nothing else.
Who approves: a named owner on our side, with a decision recorded per candidate.
What you get in return: the delivery you paid for, unaffected; and the benefit of everything previously promoted through this same gate.
What happens to refusals: recorded and retained as our own learning about the boundary, never as your content.
The wording is your lawyers' job. The structure is the architect's.
"If nothing crosses, nothing compounds"
Plenty crosses. Just not the client's truth. Look at what actually made engagement two cheaper in every example in this book: an exception class, a test, an adapter, a routing rule, a corrected band driver, a preserved rejection. Those are shape facts, not content facts. They describe how work behaves, not what a particular client's business contains.
And if your compounding genuinely does require the client's confidential content — say so out loud, to yourself first. You have a data business, not a service architecture, and it should be priced, contracted and governed as one.
The four structures now exist. What is missing is the thing that closes them.
Done Twice: The End-to-End Definition of Done
Two projects, same delivery quality, same happy client, same invoice. One of them closed. The other one just stopped.
Project A ended in week eight. The state existed, the evidence held, the client signed, the invoice cleared. Three months later a different team started something adjacent and began by reconstructing what Project A had already worked out.
Project B ended in week eight, on the same terms. Three months later, the adjacent team started with a typed exception class, a reconciliation rule and a corrected band driver already loaded, and finished a fortnight earlier than anyone had budgeted.
Both projects were delivered. Only one of them was closed.
The definition
Definition — the end-to-end Definition of Done
The customer received the promised, falsifiably accepted outcome and the provider disposed the resulting learning: local-only, reusable configuration, internal primitive, promoted doctrine, or explicit rejection.
That is a tougher definition of done than the industry currently runs on, and it is deliberately conjunctive. Half of it is not a partial pass.
The two questions
Did we keep the customer's promise?
What, if anything, should the organisation never have to learn again?
Notice the asymmetry, because it explains why the second question is almost never answered. The first question has an owner in every firm on earth — the engagement lead, the account partner, the delivery manager, someone whose bonus depends on it. The second question typically has no owner at all. It is nobody's deliverable, it appears in nobody's utilisation, and it is therefore the first thing to disappear when a project runs late.
The first question is what the client bought. The second is what the firm bought. A firm that only tracks the first will keep being genuinely surprised that its second engagement costs what its first one did.
What "disposed" means, precisely
Not "discussed". Not "captured". Not "noted in the retro". Disposed means each candidate has terminated in exactly one of the five dispositions from Chapter 8, with a named owner and a date — and where the answer is "not yet", the deferral carries a revisit trigger rather than a shrug.
That is a low bar in effort and a high bar in discipline. A candidate that has been thought about for twenty minutes and assigned to a category is disposed. A candidate that has been discussed warmly by five people for an hour is not.
The checklist a delivery lead can actually run
This is the point at which the canvas stops being a diagram and becomes something you fill in. Every item traces back to one of the four structures.
| Structure | Closure conditions |
|---|---|
| Square | Promise met as written, not as remembered · every boundary crossing recorded and dispositioned · exclusions held rather than quietly eroded · terminal state declared, including any typed unknowns that remain |
| Barbell | Frozen intent still traceable to the delivered state · rejected alternatives preserved with the reasons they lost · every consequential disposition signed by the named authority |
| Right plate | Acceptance evidence exists, is observable, and could have failed · deterministic checks passed and recorded · any failure written down rather than smoothed away |
| Flywheel | Candidates nominated · exception classes typed with routes and owners · artefacts the next engagement will load, named and located |
| Membrane | Every candidate disposed · de-identification and clearance evidenced · rejections preserved with reasons · the client-facing statement of what left the engagement is true as written |
The learning half of this has a lineage worth naming: a definition of done that adds design-doc update, conformance check, learning extraction, canon update and promotion consideration to the traditional list. What is new here is not the learning gate. It is that the learning gate and the commercial gate close the same object, at the same moment, with both signatures required.
Why both halves have to close together
Close only the promise and you have a delivery method. The engagement was profitable or it was not; next time you start from the same place; the flywheel is a diagram in a deck. This is the overwhelmingly common failure, and it does not feel like a failure at the time — it feels like a well-run project.
Close only the learning and you have something worse: a research programme that a client accidentally funded. That is precisely what the membrane chapter exists to prevent, and it is why the two halves are one checklist rather than two.
Pitfall: the retrospective nobody loads
A retrospective produces a document. A disposition produces a change to the substrate that the next run reads without being asked. The test is not whether the meeting happened, whether people were candid, or whether the notes were good. The test is whether anything downstream is different tomorrow — a rule that now exists, a test that now runs, a route that now resolves without a human.
If the honest answer is "the team learned a lot", you held a meeting.
Who runs it
The delivery lead runs the commercial half. A named substrate owner runs the learning half. Both signatures are required before the engagement is called closed, and both are visible — not to the client necessarily, but to whoever is accountable for the offer.
In smaller firms those two roles are the same person, and that is fine as long as it is said out loud. A person signing both halves is a person with a conflict, and a knowing conflict handled with a checklist is a great deal safer than an unrecognised one handled by good intentions on a Friday afternoon.
Key takeaways — Part I
- • The square makes commitment possible; the barbell makes interior latitude safe; the flywheel makes the economics improve; the membrane makes the improvement legitimate.
- • Interior variation is absorbed. Boundary mutation is a commercial event. Nothing else is a change request.
- • Commercial topology and production cadence are different layers, and answering one with the other is expensive in both directions.
- • An acceptance test that cannot fail is not a test.
- • One engagement proves delivery. Only the second can prove compounding.
- • Only abstracted, rights-cleared, human-approved patterns cross the membrane — and every candidate terminates in exactly one of five dispositions.
- • An engagement is closed when both questions are answered, and not before.
That is the claim, complete. A topology with four structures, a closure condition, and four predicted collapse modes. What has not been shown is that any of it runs.
The rest of this book is one service, carried all the way through that checklist — and then broken on purpose.
The Specimen: Promise, Perimeter and Authoritative Inputs
The buyer's actual sentence is “we need to decide whether to do this”. Everything that follows is an answer to: what state must exist for that decision to be defensible?
Specimen discipline — read this before Part II
The running specimen is generalised: a fixed-price board-decision product for a knowledge business. No firm, no offer name, no price, no engagement-specific detail. It is production-shaped — a service of a kind that exists and can be built — not a case study of one that was.
Where a number is not available, this book writes the shape instead of a plausible figure. Chapter 12's two-engagement comparison is explicitly a worked illustrative model with stated assumptions, not measured client data. That constraint is not modesty; a commercial architecture that cannot survive its own evidence standard is just another proposal.
What has actually been asked for
A board is circling a decision. Not a strategy, not an assessment, not a workshop series — a decision, with money and reputation behind it, which somebody will have to defend in a room where the mood has changed. What they say out loud is: we need to decide whether to do this.
The design question that turns that into a square is a different sentence, and it is the one this chapter answers eight times: what state must exist at the end for that decision to be defensible?
1. The promise
An evidence-backed decision state, whose valid outcomes are proceed, reshape or stop.
Read what that excludes, because the exclusions are the promise. It is not a recommendation: the supplier does not promise to prefer an option. It is not a report of an agreed page count. It is not a series of workshops with a synthesis at the end. It is not "clarity", "alignment" or any other noun that cannot fail. It is a state: the decision options are enumerated, each carries evidence at an agreed coverage, the material uncertainties are typed rather than smoothed, and every consequential interpretation has a named human's signature on it.
The hardest and most important part of that promise is the third valid outcome.
Key Insight
If "do not proceed" cannot be a successful paid outcome, you have not sold a decision. You have sold a wedge into implementation.
This is a commercial point before it is an ethical one. A buyer who suspects that the answer was decided at the proposal stage discounts everything that follows, and the discount is applied to the fee. A supplier whose incentives permit "stop" can charge for independence. A supplier whose incentives forbid it is selling a sales process with a research budget attached, and the good buyers work that out by the second engagement.
2. Authoritative inputs
"Authoritative" is a designation, not a compliment. Somebody has to be named, and the naming has to happen before the work starts, because the alternative is a week-six argument about whose spreadsheet was real.
| Input class | Who warrants it | If it is absent |
|---|---|---|
| Declared business intent — what the decision is for | The decision sponsor, in writing | Blocking. Without it there is no promise to bound; see the open-ended-intent case in Ch 15. |
| Current operating position — what is true today | A named operational owner per domain | Typed as "not observed within boundary" — never as "does not exist" |
| Constraint set — regulatory, contractual, capability | Legal or compliance owner; capability owner | Typed as an explicit assumption carried into the decision, visible to the board |
| Dependency surface — what this decision waits on | The sponsor, plus each dependency's owner | Typed as a dependency risk with a named holder; may become a boundary mutation |
Notice the third column. Every input class has a defined behaviour when the input does not arrive, and none of those behaviours is "we'll work around it". That column is what stops a missing input from silently becoming absorbed interior effort.
3. Exclusions that are operational, not decorative
Decorative exclusions are the dark twin of the authored statement of work: paragraphs listing what is out of scope with no operational teeth. A working exclusion names systems, work types, time periods and decision rights, connects to a typed state, and survives contact with a change request.
✗ Decorative
- • "Implementation is out of scope."
- • "Stakeholder management remains the client's responsibility."
- • "Third-party data is excluded."
- • "Detailed financial modelling is not included."
Each of these loses its first argument.
✓ Operational
- • Building, configuring or migrating anything — the decision pack ends at the decision, and any construction is a separately priced object.
- • Obtaining consent or access the sponsor cannot grant — if access is refused, the affected area returns "inaccessible within boundary".
- • Any source class outside the declared set — these are excluded from this phase, priced separately, or the product is redesigned. They are not "had a quick look at".
- • Persuading a dissenting executive — the pack carries evidence and typed disagreement; it does not carry a political outcome.
4. The commercial band, and what drives it
The band is product configuration, not partner mood. The mechanism — measure the input surface before quoting, using the same machinery that will deliver, so measurement is a by-product rather than a paid discovery phase; assign a band; include a defined quantity of human disposition work; hold a typed reserve for defined classes of surprise — is an organ that this book places rather than re-derives.
What is this book's job is naming the band drivers for a decision product, because they are not the same as for an estate assessment:
- Decision options in play. Two options is a different product from six. Each live option carries its own evidence obligation.
- Authoritative sources to be reconciled. Not the volume of data — the number of independent sources that must be made to agree, or to disagree in a typed way.
- Ambiguity rate in the declared intent. How much of the stated purpose can be parsed into testable claims on first reading. This predicts disposition load better than volume does, and it is measurable in the first week.
- Sign-off parties. How many named authorities must accept the state. Each one adds evidence requirements, not just calendar time.
This book publishes drivers rather than prices, for the simple reason that it does not have prices to publish. Inventing a rate card here would violate the only rule that keeps the rest of it usable. But publish your drivers to your buyers even when you keep your cost model private: opacity about what drives a band recreates exactly the budget-dance distrust the square exists to end.
5. The time boundary
A decision product needs a hard one, and the reason is not project hygiene. The boundary is what converts "we'll look into it" into an option with an expiry. A board that knows the state will exist on a date can plan around that date; a board that has commissioned an exploration will keep the question open until something else forces it shut, and the cost of that openness is almost never attributed to the exploration that caused it.
Practically: the time boundary is also what makes the reserve meaningful. A commitment with no end has nothing for a reserve to be measured against.
6. Valid acceptance states
Not one state. Several, and all of them are successes.
- Proceed — the decision pack is accepted and supports going ahead, with the residual risks typed.
- Reshape — the pack is accepted and supports a materially different option than the one that was assumed at the start. This is frequently the most valuable outcome and the hardest to sell in advance.
- Stop — the pack is accepted and supports not proceeding. The client avoided a bad investment; the supplier demonstrated independence; expensive delivery capacity was not consumed by a poor project.
Plus the typed unknowns that may legitimately remain inside any of those states — not observed, insufficient evidence, inaccessible within boundary, ambiguous and requiring a decision by a named party. A buyer's counsel will look specifically for the clause that says not observed does not mean does not exist, and its absence is a fair reason to distrust everything else in the document.
7. Authority
Two named authorities, on opposite sides of a line that must not blur. The supplier's authority warrants the state of the evidence: coverage, reconciliation, provenance, and the correctness of each typed interpretation. The client's authority owns the decision itself.
The supplier can warrant that the evidence supports proceeding. It cannot warrant that proceeding is right, because that depends on risk appetite, timing, politics and a dozen other things it does not control. Any promise that blurs the two is a promise about something outside the supplier's control — which is the definition of an unbounded outcome, and Chapter 15's territory.
The filled perimeter
| Field | The specimen's entry |
|---|---|
| Promise | An evidence-backed decision state; valid outcomes proceed, reshape or stop |
| Unit and band | One decision, banded on options in play, sources to reconcile, intent ambiguity rate and sign-off parties |
| Time boundary | A fixed date, with the reserve measured against it |
| Authoritative inputs | Declared intent; current operating position; constraint set; dependency surface — each with a named warrantor and a defined absent-state |
| Terminal states | Proceed / reshape / stop, each with typed residual unknowns |
| Exclusions | Construction; access the sponsor cannot grant; undeclared source classes; political outcomes |
| Acceptance tests | Coverage, reconciliation, signed dispositions, and a terminal state the client's own authority can test — written out in Ch 11 |
| Authority | Supplier warrants the evidence state; client's named authority owns the decision |
What the buyer sees on day zero
The machine's edges, disclosed by the seller rather than discovered in week three: which source classes the product cannot read; what the acceptance evidence will physically look like; exactly which events would reopen the commercial conversation; and what a "stop" outcome would be delivered as, so nobody is surprised by success arriving in an unpopular shape.
Disclosing the edges on day zero costs a small number of deals and saves all of the disputes.
"Our buyers can't specify a decision that precisely"
Many cannot — and that is an opportunity rather than an objection. If the buyer cannot bound the decision, then the bounding is the first paid unit: a smaller square whose promise is a stable, testable decision question rather than an answer to one. A supplier who can reliably do that bounding holds an advantage over one who can only respond to briefs that arrive already clear.
But hold the failure mode in view too. If, after that work, the intent still will not hold still — if every conversation produces a different definition of success — this is not a bounding problem. It is the open-ended-intent case, it does not belong in a square, and Chapter 15 says what to do about it.
The perimeter exists and is signed. Now the interesting half: what actually happens inside it.
Inside the Square: Latitude, Proof and Disposition
The interior is supposed to look wasteful from the outside. It can be free because two different proof instruments are waiting at the exit — and neither of them is the machine.
Day one inside the perimeter does not look like a workplan. Nobody is executing a sequence of agreed activities; there is no phase gate at the end of week two called "as-is complete". What happens instead is closer to a search.
What the interior actually does
The interior assembles the evidence base from the four declared authoritative input classes, reconciling as it goes and recording provenance for every claim it lands. It builds a current position and tests that position against the sources rather than against anyone's memory. Then it starts generating: candidate framings of the decision, each one a structure of options, evidence obligations and consequences. It runs those framings against the evidence — where does this framing require a fact we do not have? which options does it make indistinguishable? Framings that fail are discarded and new ones generated, not repaired. When an undeclared constraint surfaces, it branches, re-runs the affected tests, and carries on.
Seen from the outside this is wasteful. Eight framings were built and six thrown away. The evidence base was assembled twice because the first assembly used a source that turned out not to be authoritative. A whole line of analysis died on a constraint discovered in week three. If a client watched that on a timesheet, they would ask why they were paying for the six.
They are not. That is the entire commercial point of the square: the buyer bought a state, and the search that produced it is the supplier's business.
What the square absorbs, and what crosses
✓ Absorbed — interior variation
- • An authoritative source arrives in a form nobody anticipated and needs a new adapter written on the spot
- • Two sources disagree and reconciliation takes four passes instead of one
- • A candidate framing survives three tests and dies on the fourth
- • The evidence base is rebuilt because the first assembly was structurally wrong
- • The declared constraint set turns out to have an internal dependency nobody mentioned, requiring the whole option space to be re-tested
All real. All expensive in machine terms. None of them is a change request, and the band was set knowing they happen.
→ Crosses — boundary mutation
- • The buyer adds a second decision to the same engagement
- • A new party's sign-off becomes required
- • The declared intent changes — a different question is now being asked
- • An authoritative source is withdrawn, or access to it is refused
- • The decision date moves
Each moves a named perimeter field. Each triggers the same routine.
The routine for a crossing is worth stating, because most firms do not have one and improvise under relationship pressure: stop; name the field that moved; quantify the delta against the original measurement; present the options. Uplift the band, draw the reserve, narrow the boundary, or stop. Four options, presented with evidence, in the same week the crossing happened — not accumulated into a difficult conversation at month end.
Two instruments at the exit
Chapter 6 named them. Here is what they actually are on this specimen.
| Deterministic check | What it asserts on this specimen |
|---|---|
| Coverage | Every declared authoritative source was read, or is explicitly typed as inaccessible within boundary |
| Reconciliation | Where two sources bear on the same quantity, they tie within a stated tolerance or the disagreement is typed and surfaced |
| Completeness | Every decision option carries an evidence position; no option is left implicitly unevaluated |
| Provenance | Every claim in the pack carries an openable pointer back to the source that supports it |
| Conformance | The pack matches the agreed structure, so a reader can find the same thing in the same place every time |
None of those requires seniority. They pass or they fail, and they fail loudly. What they cannot do is decide anything — which is why the second instrument exists.
Human dispositions are the consequential calls, made under named authority. On this specimen they are things like: this source disagreement is material to option B and resolves in favour of the operational record; this constraint is a genuine blocker rather than a preference; this residual uncertainty is acceptable at board level and this one is not. The machine may investigate, synthesise and nominate every one of those. It may not become one, and it may not mutate an authoritative state.
The disposition surface is the product's face
The headline on the review interface is not "the machine found forty things." It is "forty calls still need your name on them."
That inversion does three commercial jobs at once. It makes the scarce resource visible to both sides, so nobody is surprised by how much senior attention the engagement requires. It prevents the machine's positions from laundering themselves into signed scope, because an undisposed position is visibly undisposed. And it is the thing the band was priced against — included disposition units are the meter, so showing the meter is honest rather than awkward.
A dashboard that says "zero decisions outstanding" on day one, before anyone has looked at anything, is the product working correctly. The discipline is borrowed from a different specimen in this corpus; the offer here is not that one, so take the discipline and leave the counts.
Exceptions in flight
Something genuinely novel will appear. On this engagement it is a source-disagreement pattern: two authoritative records that should agree, differ systematically in one direction, and the difference turns out to be an artefact of how one of them is compiled rather than a real discrepancy.
The routine is: type it, route it, resolve it, and require it to leave something behind. The rule underneath is not optional politeness — every escalation should reduce the probability of the next equivalent escalation.
One exception, and what it leaves
What happened
Two authoritative sources bearing on the same quantity disagree systematically. A senior spends most of a day establishing that the difference is a compilation artefact, not a discrepancy, and disposes the finding in favour of one source with the reasoning recorded.
What it must leave behind
- • A typed exception class with a name and a detection condition
- • A deterministic test that flags the signature automatically next time
- • A decision rule stating which source wins under which conditions
- • A routing trigger — when this fires, who looks at it and at what seniority
Without those four, a day of senior time bought one answer. With them, it bought a class.
If escalations only produce private answers in a chat thread, you have built a help desk. The difference between scale and exhaustion is whether the fossil ships.
The acceptance test, written out
Here is the specimen's failable test, in the form a competent third party could run:
The decision pack is accepted when
- Every declared decision option carries an evidence position at the agreed coverage threshold, or is explicitly typed as not evaluable and why.
- Every reconciliation between authoritative sources passes within the stated tolerance, or the disagreement is typed, surfaced and dispositioned.
- Every consequential disposition carries a named signature and recorded reasoning.
- The recommended terminal state — proceed, reshape or stop — is supported by an argument that the client's own named authority can test against the evidence without the supplier in the room.
It fails when any option lacks an evidence position without being typed; any material reconciliation is unresolved and unsurfaced; any consequential interpretation is unsigned; or the client's authority reads the argument and cannot trace it to the evidence. A failed acceptance is not a dispute — it is a defined state with a defined remedy, and the remedy is inside the band.
That last sentence is what most acceptance criteria are missing. If the test cannot fail, the square has a soft edge at the exact moment the money moves. And a test whose failure has no defined consequence is a test in name only.
What is nominated — and what does not cross
Engagement one ends. Four candidates go to the membrane, and every one of them is a nomination, not a promotion:
- The source-disagreement exception class, with its detection test and decision rule.
- A correction to a band driver: intent ambiguity predicted disposition load better than option count did, and the band arithmetic should reflect that.
- A reconciliation rule for a class of records that behaves the same way across organisations of this kind.
- A rejected approach: one framing that looked strong and failed for a structural reason worth keeping, so nobody spends a week on it again.
Nothing crosses in this chapter. The gate in Chapter 8 decides — de-identification, abstraction, transferability, clearance, a named human's approval — and at least one of these four might well come back as local-only. That is a legitimate result, not a failure of the engagement.
"If the interior is free, how do you estimate it?"
You do not estimate the interior. That is the point, and it is worth being blunt about because it is the question that most often causes a firm to give up and go back to hourly.
You measure the input surface before promising, you assign a band from published drivers, and you meter the scarce resource inside the band — consequential dispositions, which are the thing that actually costs. The interior's cost is a portfolio property, not a per-engagement prediction. Which is precisely why a single engagement is a bad place to learn your band, and why the next chapter is about the second one.
Engagement one is delivered and accepted. Everyone is pleased. That is exactly the moment at which nothing has been proved about the economics.
Engagement One, Engagement Two
The case study is already in draft. The honest post-mortem is quieter — and it is the only document that can tell you whether you have a product.
What this chapter is, before any number appears
This is a worked illustrative model with stated assumptions. It is not measured client data, it is not a portfolio average, and nothing in it should be quoted as evidence about any real engagement.
What it demonstrates is the shape of a compounding curve and — far more usefully — what you would have to measure to know whether you were on one. Where a real figure is unavailable, and most are, the shape is written instead of a plausible number.
The version that gets written up, and the version that is true
The case study writes itself. Fixed price held. Board decision delivered on the date. Client quotable. Somebody wants a logo on a slide by Friday.
The honest retrospective is quieter. The two most experienced people on the offer owned every hard call. When the source-disagreement exception surfaced, it went to one of them within an hour, and it went to them because the routing rule at that point was "ask her". The band was defended, but it was defended by absorbing more interior effort than anyone had modelled. Field staff participated. They did not lead.
You proved an augmented exceptional team. You did not prove a product.
The assumptions
State them before the table, because a comparison without stated assumptions is a story with columns.
- Same offer, same band. A decision product of comparable size for a comparable knowledge business. Comparing a small engagement to a large one and calling the difference "learning" is the oldest way to fake this.
- Different client. A second engagement with the same client compounds relationship knowledge, not machinery, and the two are easy to confuse.
- No model or tooling generation change between the two. This is the important one. If a materially better model arrives between engagement one and engagement two, everything improves and you will attribute it to your substrate. That mistake is close to universal.
- Engagement two is run by ordinary capable staff, not by the originating pair. If the same two people run both, the comparison measures their learning curve, which leaves when they do.
The comparison
| Dimension | Engagement one | Engagement two | What the movement means |
|---|---|---|---|
| Estimate error | Band assigned from first-generation drivers; consumed effort ran materially above it | Narrower, because the drivers were corrected — intent ambiguity now weighted ahead of option count | The band is becoming configuration rather than estimation. This row moves first and fastest. |
| Production variance absorbed | High, and unmeasured — nobody logged what was absorbed | Higher, and logged — more irregularity handled without reaching the buyer | Absorption rising is good. Absorption rising unmeasured is how margin disappears silently. |
| Human disposition load | Apparently low — because most dispositions were made informally and never counted | Apparently higher, because it is instrumented for the first time | The counter-intuitive row. See below — this is where firms quit. |
| Exception classes | All novel; each resolved by a principal; none typed at the time | Three arrive pre-typed with detection tests and routes; the genuinely new ones are fewer and different | The clearest single indicator that the substrate is real rather than aspirational. |
| Acceptance evidence | Harness built during the engagement, partly after the fact | Harness reused; only the local cases are new | Reuse here also improves quality: the same checks run on both, so the two are comparable at all. |
| Reusable artefacts | Four nominated (Ch 11); three promoted, one held local-only | Two nominated — fewer, and narrower | Falling nominations are expected and healthy. Rising nominations by engagement four would suggest the offer is not converging. |
| Recurring escalation rate | Every class escalated at least once | Falls for the fossilised classes; flat for genuinely new ones | Measure by class, never in aggregate — an aggregate hides exactly the signal you need. |
| Scarce-expert density | Unknown — not instrumented | Measured for the first time; a baseline, not yet an improvement | Engagement two often produces the first honest number, which is a result in itself. |
The row that makes firms quit
Key Insight
Transfer work usually raises measured disposition cost before it lowers it, because instrumentation reveals the true denominator.
Engagement one's disposition load looked low. It was not low; it was invisible. Consequential calls were made in corridors, in review meetings, in the margins of a draft, by people senior enough that nobody thought of it as a disposition. Engagement two puts a disposition surface in front of the team, and suddenly there are forty of them with names attached.
At that moment a firm faces a choice that has nothing to do with technology. It can accept that the number went up because the measurement got honest, or it can conclude that the new way of working "created more overhead" and quietly stop counting. Firms that panic here choose narrative over control — and having chosen it once, they never get a real number again.
Three curves
Illustrative units — shapes, not measurements
The normal services curve
Engagement 1: 100 · 2: 100 · 3: 100 · 20: 100
Units of expert effort. Revenue scales by adding people.
The weak "AI productivity" firm
Engagement 1: 80 · 2: 80 · 3: 80
A twenty per cent improvement, banked once. Still basically linear — and this is where most firms currently are.
The compounding shape
Engagement 1: high human novelty + product construction ↓ write-back
Engagement 2: less novelty, more substrate ↓ write-back
Engagement 3: exceptions narrower again ↓ …
Not because every engagement becomes identical — because the recurring parts stop being novel, and the scarce human work migrates towards the genuinely new edge.
What actually moved, and why
Numbers without a mechanism are numerology. Here is the mechanism, traced back to the four artefacts nominated at the end of Chapter 11.
The source-disagreement exception class, with its detection test and decision rule, was promoted. In engagement two the test fired in week one, the routing trigger sent it to a mid-level consultant with the rule attached, and it was disposed in twenty minutes rather than most of a day. That single artefact moves three rows: exception classes, recurring escalation rate and scarce-expert density.
The corrected band driver was promoted into the pricing configuration. Engagement two's band was assigned with intent ambiguity weighted ahead of option count, which is why estimate error narrowed. That artefact moves one row, but it is the row a finance partner looks at first.
The reconciliation rule was promoted as an internal primitive and folded into the harness, so it ran automatically rather than being remembered. That moves acceptance evidence, and quietly moves production variance absorbed, because a check that runs automatically absorbs variation that would otherwise have surfaced as a question.
The rejected framing was preserved. Nobody in engagement two spent a week on it. That is invisible in every row of the table, which is precisely why rejections have to be preserved deliberately rather than being expected to justify themselves.
What did not move
Three things stayed flat, and each says something.
Genuinely novel exception classes. The second client's constraint surface was different, and threw up two classes nobody had seen. That is not a failure of the flywheel; that is the flywheel doing its actual job, which is to make the recurring parts cheap so that scarce attention lands on the new edge.
Authority-dependent dispositions. The calls that require a named human under consequence did not fall and are not supposed to fall. If that row ever collapses towards zero, the correct response is alarm rather than celebration: something consequential is being decided without anybody owning it.
Anything that depended on the originating pair's relationship. The parts of engagement one that went smoothly because a principal had credibility with a particular executive did not transfer at all. That is the cleanest possible demonstration of the difference between a substrate and a hero.
A flat row is information. Every row flat is a verdict.
The only proof that counts
The qualitative test is transfer: ordinary capable staff lead materially more of engagement two because shared infrastructure changed, not because the same heroes worked late again. The quantitative test is the ratio — paid bounded units divided by scarce expert dispositions — and it lies immediately if either side is undisciplined. Both are organs; this book requires them rather than re-deriving them.
Five ways to fake this
- • Counting free pilots and demos in the numerator.
- • Moving scarce work off the books as "sales support" or "R&D".
- • Redefining what counts as a disposition halfway through, to protect a narrative.
- • Comparing unlike engagements — a small estate against a large one — without banding.
- • Declaring victory after one heroic run.
What this has and has not shown
It has specified a measurement set, and a mechanism by which each row would move. It has named the artefacts that do the moving, and traced each to the rows it touches. That is enough for a reader to instrument their own offer.
It has not demonstrated a market-wide effect, a margin curve, a rate of improvement, or anything about the third engagement. n = 2 is not a portfolio, and the honest sentence remains the one from Chapter 7: one successful engagement proves delivery, not compounding.
So name the falsifier. If by engagement three the scarce-expert density has not moved and the count of novel exception classes has not fallen, the compounding hypothesis is wrong for this offer. The correct response is repair or demotion with a date — not more sales. A firm that will not accept demotion when the ratio stagnates will eventually accept margin collapse instead.
And the obvious objection deserves a straight answer. You have just described a learning curve; every business has one. True — but a learning curve lives in people and leaves when they do. This lives in artefacts an ordinary consultant loads on day one. The test that separates them is exactly the assumption this chapter started with: run engagement two without the originating pair, and see which rows move.
The architecture has now been shown working. What is worth more than another success is watching it fail.
The Failure Specimen: Two Ways the System Breaks
A framework that only shows itself succeeding has demonstrated nothing except the author’s imagination. Here are two of the four collapse modes, walked all the way down.
The engagement went well. The machine found something genuinely valuable and genuinely generalisable. Eleven months later a different team rediscovered the same thing from scratch, at the same cost, using the same senior person.
Nothing went wrong in delivery. Something was missing in the architecture.
Failure A — the missing membrane
The setup. Engagement one of the decision product, run inside a properly drawn square with a loaded barbell. During reconciliation the interior discovers a repeatable pattern: a class of source disagreement with a reliable resolution route, plus the deterministic test that detects its signature. It is real, it recurs across organisations of this kind, and it does not depend on anything confidential about this particular client. It is exactly the sort of thing the flywheel exists to capture.
What was missing. No learning-rights clause in the engagement contract. No de-identification path defined before the work began. No named owner for promotion. No disposition step in the definition of done. Four absences, none of which cost anything to fix beforehand, all of which are impossible to fix afterwards.
The consequences, in order
- The pattern stays inside the engagement boundary. The only defensible reading of the contract is that everything arising from the engagement belongs to the engagement.
- The principal who found it becomes its only carrier — not by choice, but because there is nowhere legitimate to put it.
- Engagement two rediscovers it, consuming a scarce disposition that should have cost nothing.
- The recurring escalation rate for that class does not fall.
- Scarce-expert density does not move.
- The compounding hypothesis fails — and it fails invisibly, because delivery looked fine both times and both clients were happy.
That last point is what makes this failure dangerous rather than merely annoying. There is no bad quarter, no angry client, no incident review. The firm simply never gets the second-engagement economics it built the whole architecture for, and the reason is buried in a contract nobody reopens.
The two bad exits. Breach quietly and hope — which is how it usually goes, one Slack message and one "well, it's abstracted enough" at a time, until somebody senior discovers during a renewal that the firm has been reusing something it had no right to. Or forget deliberately, stay linear, and watch a competitor who wrote the clause pull ahead.
The repair
The clause before the engagement, not after. What stays, what may leave and in what form, who approves, what the client gets in return, what happens to refusals. A paragraph in a document that is being signed anyway.
The abstraction step, designed in. A pattern that cannot be stated without the client's specifics is not promotable and never was. Strip the client; keep the condition and the route. If the meaning does not survive, the correct disposition is local-only — and that decision is itself worth recording.
A disposition council with field, product and operations in the room. Thirty minutes. One nomination, forced to exactly one of the five dispositions, with a revisit trigger if it is deferred.
The promotion gate. De-identification, abstraction, transferability, clearance, a named human's approval.
The rule. A pattern that cannot legally cross is not an asset. Rights are part of the architecture, not part of the paperwork — and the asymmetry is brutal: establishing them costs a paragraph before the work and is impossible after it.
Failure B — the unmeasured boundary
The setup. Same offer, sold well, fixed commitment, competent team. The band was assigned from a sales conversation rather than from a measured input surface, because the buyer was in a hurry and running the preflight measurement "would have delayed the start by a week".
What was missing. No measurement, therefore no band drivers with values attached, therefore no delta available when reality arrived, therefore no language in which surprise could speak. The perimeter existed on paper and had nothing underneath it.
Two paths from an unmeasured start
✗ What actually happened
- • Interior variance is real and gets absorbed — correctly, and invisibly
- • Boundary movement is also real: two more sign-off parties appear, a source class nobody declared turns up in week three
- • Nothing distinguishes the two, because nothing was measured at the start
- • The supplier absorbs both, and congratulates itself on not raising a change request
- • Margin erodes silently, then quickly
- • The supplier eventually has to reopen the price
The buyer experiences that reopening as the old change request in new clothes — and the trust the square was supposed to buy is spent in a single meeting.
✓ What a measured start produces
- • The same two events occur — reality does not care about your process
- • Both are recognised as boundary mutations within a week, because a named field moved against a recorded value
- • A delta is presented with evidence: uplift the band, draw the reserve, narrow the boundary, or stop
- • The buyer chooses, in possession of the facts, early
Same reality, same money at stake, entirely different relationship — because surprise had a language.
This is not an AI-era novelty and it should not be presented as one. Lump-sum contractors have lived it for a century: they price significant contingencies precisely because reality moves, and where those contingencies prove insufficient they are "naturally incentivized to seek opportunities to reopen the fixed price" — because, in truth, there is no such thing as an absolute fixed price contract2.
The square does not repeal that physics. What it does is give the reopening a governed language and an evidence base — or, without measurement, it does not, and you are back to bravado with a schedule.
The repair
Measure before promising, with the machinery that will deliver. Then measurement is a by-product of the work rather than a paid discovery phase the buyer has to be talked into.
Publish the drivers even if you never publish a price. A buyer who understands why they are in this band will accept a delta against it. A buyer who does not will hear any delta as a renegotiation.
Type the crossings. Each named perimeter field, each with a defined consequence, agreed before anyone needs it.
Make exhaustion produce a conversation with a delta attached, not a surprise at month end. The full contract test for boundary changes — how reserves are typed, how consumption is measured, what exhaustion formally triggers — is a seam with its own treatment.
The rule. A fixed price without a measured input surface is not a square. It is bravado with a schedule, and the schedule makes it worse, because it delays the moment of honesty until the money is already spent.
Reading these against the collapse table
Failure A is the membrane row from Chapter 2. Failure B is the square row. The other two have already been walked in mechanism, so they get a paragraph each rather than another full case.
Without the barbell, this same specimen becomes supervised automation: a reviewer on every unit of machine output, an engagement whose review cost rises in proportion to how much the machine produces, and no economics at any volume. Chapter 5 walked why that shape costs more and assures less.
Without the flywheel, it becomes repeated bespoke heroics: both engagements start cold, the second costs what the first did, and the only artefact that survives is a case study. Chapter 12 walked exactly which rows stay flat when that happens, which is a more useful description than another narrative would be.
| Missing structure | Collapse mode | Walked in |
|---|---|---|
| Square | Time-and-materials ambiguity, arriving late as a reopened price | Failure B, above |
| Barbell | Supervised automation | Chapter 5 |
| Flywheel | Repeated bespoke heroics | Chapter 12 |
| Membrane | Learning stranded, or client-data leakage | Failure A, above |
Four structures, four distinct collapse modes, each recognisable from somebody's real history. That is what the composition claim from Chapter 2 was asserting, and this is the chapter that pays for it.
A framework that cannot fail its own candidate is a brochure
There is a lineage for running a favoured candidate through the gates until it fails on purpose. The habit is worth keeping for one reason: falsifiability is a feature of a framework, not an embarrassment. A book that only shows its architecture succeeding has demonstrated the author's imagination and nothing else.
The predictable objection to both failures above is that they are execution mistakes rather than architectural ones. An execution mistake is one a better team would not make. Both of these were made by good teams doing everything their process asked of them — there was no step in the method that said "write the learning clause" or "measure before quoting", so nobody skipped a step. A failure that survives competence is the definition of a missing structure rather than a missing effort.
What honest governance does next
The same response fits both failures: repair or shrink, with a date, protected from sales enthusiasm.
Repair means funded design work with a named owner — write the clause, build the census, define the gate — on a schedule, not on somebody's evenings. Shrink means a narrower promise you can actually keep while the machinery catches up: fewer decision options, a tighter input boundary, a smaller commitment at a smaller price.
And demotion is a success when it prevents an offer from consuming senior time indefinitely while producing neither margin nor substrate. The market will not do this for you. Buyers will buy heroics while they last.
Two structures removed, two specific collapses, two repairs. The architecture has now been shown working and shown failing. What remains is whether any of it belongs to one trade.
The Domain Strip: The Same Canvas, a Different Trade
Remove the profession, the buyer, the offer and the present moment. If what is left is a fixed-price consulting argument, this book has failed — and should say so.
The strip test is not a rhetorical device. It is the check that separates a portable layer from a playbook with ambitions. Take the architecture, delete the trade it was discovered in, delete the buyer, delete the specimen, delete the market moment that made it interesting. What survives?
The transplant chosen here is deliberately unkind: a regulated engineering condition assessment. An asset owner under a compliance obligation needs a defensible statement of the condition of a physical estate, carrying a licensed sign-off, on a statutory reporting cycle. It breaks three things the decision product never had to face — physical inspection, licensed authority, and a statutory consequence for being wrong.
The same eight fields, filled twice
| Field | Board-decision product | Regulated condition assessment |
|---|---|---|
| Promise | An evidence-backed decision state; proceed, reshape or stop | A defensible condition position with typed unknowns, carrying a licensed sign-off |
| Authoritative inputs | Declared intent; operating position; constraint set; dependency surface | Asset register; physical observation; prior inspection record; manufacturer and regulatory specification |
| Exclusions | Construction; access the sponsor cannot grant; undeclared sources; political outcomes | Remediation works; destructive testing; assets outside the declared boundary; anything requiring an outage that was not granted |
| Band drivers | Decision options; sources to reconcile; intent ambiguity rate; sign-off parties | Asset count and class mix; access difficulty; record quality; statutory scope |
| Time boundary | A fixed decision window, chosen by the parties | The statutory reporting date — externally fixed, and therefore harder |
| Terminal states | Proceed / reshape / stop, each with typed residual unknowns | Compliant / non-compliant with typed defects / insufficient access to assert |
| Acceptance | Coverage, reconciliation, signed dispositions, a testable argument | Coverage against the declared asset boundary; reconciliation to the register; a licensed signature |
| Authority | Supplier warrants the evidence; the client's authority owns the decision | Supplier warrants the evidence; the licensed engineer owns the assertion — and is not substitutable |
What changes
The interior stops being purely cognitive. Somebody has to physically stand in front of an asset. That cost curve does not move with token prices, which changes the band drivers and shrinks the class of variance the square can absorb. The interior is still generative — route planning, prioritisation, defect interpretation, reconciliation against messy records — but a portion of the middle is now physics, and physics is priced separately.
The right-hand plate acquires a person who cannot be replaced by machinery. A licensed signature is a legal object, not a quality step. In the decision product the supplier's authority warrants the state of the evidence; here the licensed engineer's assertion is a regulated act with personal consequence attached. The barbell's right plate gets heavier, and its weight is set by statute rather than by design preference.
The exception classes change character. In the decision product the exceptions are about analysis — source disagreement, ambiguous intent, undeclared dependency. Here they are about access and asset condition: the door was locked, the outage was cancelled, the asset is not where the register says it is, the condition is worse than any category in the taxonomy. Access exceptions are also the ones most likely to become boundary mutations rather than interior variations, which means the crossing routine from Chapter 11 runs more often and has to be faster.
What does not change
Key Insight
Every field's content changed. The schema did not. That is what it means for a layer to be portable.
The perimeter/interior split holds. The barbell's shape holds — heavy at frozen intent, heavy at the licensed assertion, light in the middle where route planning and interpretation happen. The write-back mechanism holds. The membrane's three territories hold.
And one thing gets easier, which is worth noticing because it is counter-intuitive. In the assessment domain the flywheel's payload is almost entirely non-confidential: the defect taxonomy, the access playbook, the reconciliation rules for messy registers, the evidence templates that survive a regulator's review. Not one of those is client truth. The membrane is a lighter obligation here, not a heavier one — which suggests that the domains where compounding is hardest are not the regulated ones but the ones where the valuable pattern is entangled with the client's commercial position.
A second, shorter strip: legal due-diligence review
Square
A bounded review of a defined document population against a declared question set, within a declared data-room boundary, producing typed findings by issue class. Terminal states include "insufficient disclosure to assert" — which is a finding, not a failure.
Barbell
Heavy at the question set and the materiality thresholds, which is where the value of the review is actually determined. Light through the review itself. Heavy at the privileged opinion, which a named lawyer owns and cannot delegate to a machine.
Flywheel
The issue taxonomy, the clause-pattern library, the review harness, and the materiality rules corrected by each deal. All of it shape rather than content — and all of it directly reduces the next review's senior hours.
Membrane
The sharpest version of the problem, because privilege makes the boundary a professional obligation rather than a commercial preference — and because practitioner guidance already requires platforms committed to zero retention and zero training on client inputs10. A firm that gets the membrane right here has a defensible answer everywhere else.
Where the strip gets uncomfortable
It would be easy to keep going and produce a clean sweep, so here is a domain where it is genuinely hard: open-ended creative or design work, where the intent will not hold still by definition. The client does not yet know what they want, and coming to know it is the value of the engagement. A perimeter drawn around a moving intent is either false or so loose that it is not a perimeter.
That is not resolved here, and it should not be resolved with a clever move. It belongs to the next chapter, which is entirely about the territory where the honest answer is to shrink the promise or decline it. Better to hand a reader a real edge than a suspiciously clean sweep.
Pitfall: transplanting the content instead of the schema
The commonest way this goes wrong is copying the entries from one domain into another: importing "proceed, reshape or stop" into an assessment product where the real terminal states are compliance categories, or importing decision-option band drivers where the real drivers are asset counts and access difficulty. The schema transplants. The entries never do, and a filled canvas borrowed from someone else's trade is worse than a blank one, because it looks finished.
The honest limit
This is a design transplant, not a delivery record. It demonstrates that the schema survives the strip — that the eight fields can be filled meaningfully in a domain with a different buyer, a different risk surface and a different legal shape. It does not demonstrate that anyone has run this offer at scale in either trade, and it would be trivially dishonest to imply otherwise.
The claim being made is architectural. It is falsifiable in the ordinary way: take a bounded knowledge service you know well, try to fill the eight fields, and see whether the exercise produces clarity or nonsense.
One objection deserves an answer before we leave. Regulated work cannot be fixed-priced. Parts of it demonstrably are: condition assessments, statutory inspections and certification work are routinely bounded and routinely quoted. What cannot be bounded is the remediation that follows — which is exactly why it sits in the exclusions column, drawn before the engagement rather than argued afterwards. The architecture's contribution here is the placement of that line.
The schema survives two strips and gets uncomfortable in a third. That discomfort is the next chapter's subject.
Where the Architecture Must Shrink or Decline
The offers that quietly destroy a practice are almost never the ones that were badly delivered. They are the ones that should never have been quoted.
"We won't quote that."
Most practice leads can remember the last time they said it. Rather fewer can remember the last time they said it to a large, enthusiastic, well-funded buyer with a relationship attached. And that is the population that matters, because the engagements that damage a firm are rarely the badly delivered ones. They are the ones that were accepted while somebody in the room already knew the promise could not be bounded.
Key Insight
The rule is not AI → fixed price. It is: AI expands the territory in which complexity can be bounded, configured and priced as a product — and the edge of that territory is a design output, not an act of courage.
The reasoning is the one from Chapter 3, taken to its conclusion. AI cheaply absorbs cognitive variance. It does not absorb all variance. The square breaks wherever the residual uncertainty lives somewhere tokens cannot reach — and there are six such places.
The six cases
| Case | Diagnostic question | Shrink to… | Decline when… |
|---|---|---|---|
| Physical scarcity trucks, stock, technicians, laboratories, buildings, working capital |
Does any commitment in the perimeter depend on a resource that does not scale with tokens? | Bound the physical component separately, priced against real capacity; keep the cognitive component in the square | The physical component dominates and cannot be reserved or contracted |
| Human authority regulated principal, clinician, board member, licensed engineer |
Is the promised state achievable without one specific person's judgement, on their timetable? | Promise the state up to the authority boundary; make the sign-off a client responsibility with a named party and a date | The authority sits outside both parties' control entirely |
| Open-ended intent success keeps being redefined |
Has the declared intent survived two contacts with evidence without changing shape? | Sell the bounding itself — a smaller square whose deliverable is a stable, testable intent | The buyer will not commit to a decision they would actually act on |
| External dependency a third party on their own clock |
Does the critical path run through an organisation that is not a party to this contract? | Exclude the dependency explicitly; type the state it produces ("blocked pending third party") as valid and deliverable | The dependency is the work — you would be selling someone else's timetable |
| Unbounded liability a consequence tail you cannot describe |
Can the consequence tail be bounded — and is it bounded in writing? | Cap it, carve out the uncapped classes, and narrow the promise to states you can evidence | The cap is refused and the tail is real |
| True exploratory R&D no testable terminal state exists yet |
Can you write an acceptance test that could come back negative? | Sell a bounded exploration with typed outputs, including "we may establish that this is not possible" as a successful terminal state | Someone wants a fixed promise on a discovery |
Three of them, in detail
Physical scarcity is the easiest to spot and the easiest to underestimate. Machine cognition improves the cognitive and coordination curve; it does not repeal physics, and a beautifully reasoned commitment that no operational network can honour is a cognitively elegant work of fiction. The shrink move here is usually clean: two commercial objects instead of one, with the physical one priced against real capacity and the cognitive one inside a square. The failure is refusing to split them because splitting looks less impressive in a proposal.
Unbounded liability is where the cheap-cognition argument does the most damage, because it feels like it should help. It does not. Without a contractual limitation, professional liability is unlimited and can exceed the cover maintained under the professional indemnity policy behind it12. Read that against the architecture: nothing in the square, the barbell, the flywheel or the membrane changes the size of a consequence tail. Better analysis lowers the probability of an error. It does not lower the cost of the one you make. A supplier who accepts an uncapped obligation because the analysis is now cheap has confused two different quantities, and will discover the difference exactly once.
Open-ended intent is the most common and the most seductive, because the buyer is usually delightful and genuinely wants the work. The diagnostic is unsentimental: has the declared intent survived two contacts with evidence without changing shape? If every time you show them something the definition of success moves, there is no perimeter to draw — and drawing one anyway produces an engagement that is simultaneously over-committed and under-scoped. The shrink is to sell the bounding: a smaller square whose entire deliverable is a stable, testable question. That is a real product and it is frequently the most valuable thing you could sell them.
The physics that does not change
It is worth grounding all of this outside our own corpus, because a firm's own doctrine is a poor witness in its own defence. Lump-sum contractors, who have been doing bounded commitments for far longer than any of us, price significant contingencies into their bids and are "naturally incentivized to seek opportunities to reopen the fixed price" where those contingencies prove insufficient — because there is no such thing as an absolute fixed price contract2.
Do not read that as a counsel of despair, and do not read it as permission either. The square does not abolish the pressure to reopen. It governs it: a measured start, published drivers, typed crossings and a reserve mean that the reopening happens early, with evidence, as a commercial conversation rather than as a betrayal. The firms that get hurt are the ones that believed the perimeter was a force field.
A product that cannot say no is not a product
The moral version of this argument is easy and unpersuasive. The commercial version is better.
Forcing out-of-band work into a fixed price does three things, all of them expensive. It destroys the band's meaning for every future buyer, because the band no longer predicts anything. It teaches your own sales system that drivers are negotiable, which reintroduces the private-judgement pricing the square was built to replace, only now with a product's name on it. And it converts a product back into a bespoke project with a product's price and a project's cost — which is the worst combination available.
There is a fourth cost, quieter and worse. A lane full of forced exceptions cannot teach you anything. The whole measurement apparatus from Chapter 12 depends on comparable units; a half-comparable population produces numbers you cannot act on, and a firm that cannot measure its own product cannot improve it.
The shrink move, generalised
Four steps
- Identify which perimeter field is unstable. Not "this feels risky" — which of the eight fields cannot be answered, or cannot be held?
- Remove the commitment that depends on it. Precisely that one, not a defensive haircut across everything.
- Re-home the removed part as one of three things: a client responsibility with a named owner, a separate commercial object, or a typed terminal state.
- Re-check that what remains is still worth buying. This step is the one people skip.
That last step deserves its own sentence, because it is where the method earns its keep. A shrunken promise nobody wants is a decline that has not admitted itself yet — and the cost of discovering that after the contract is signed is much higher than the cost of saying so in the room.
"Our competitors will just say yes"
Some will. Some of them will win the deal and lose the money, and that is a market you can wait out — particularly since the buyer who was over-promised to is a buyer who will be looking for someone credible in eighteen months.
But answer the objection honestly too, because a purely principled response is not usable. Refusing costs revenue now, and a firm without a pipeline cannot afford principles. Which is precisely why the realistic move is almost always the shrink rather than the walk-away, and why this chapter's method exists: shrinking is the harder skill and the one nobody teaches. A supplier who can reliably convert an unboundable request into a smaller boundable one keeps the relationship, keeps the revenue, and keeps the band's integrity at the same time.
What refusal buys
Three returns on saying no
Band integrity
Your bands keep predicting things, which means your pricing keeps being configuration rather than negotiation.
A competence signal
Buyers who have been burned by an over-confident supplier recognise a boundary immediately, and they pay for it. Predictability is a premium attribute, not a discount.
A measurable product
A clean population of comparable units is the only thing that makes the flywheel's measurement set mean anything.
Pitfall: cheap analysis does not shrink a consequence tail
This is the single most expensive misreading of the whole architecture. Machine breadth reduces the probability of missing something. It does nothing whatsoever to the magnitude of what happens when you do. Any commitment whose downside you cannot describe in a sentence and cap in a clause is outside the territory, no matter how good the analysis has become.
Key takeaways
- • The claim is territorial, not universal: AI expands where complexity can be bounded, configured and priced. It does not make everything boundable.
- • Six cases sit outside: physical scarcity, human authority, open-ended intent, external dependency, unbounded liability, true exploratory R&D.
- • Each has a diagnostic, a shrink, and a decline condition. Run the diagnostic before the proposal, not during delivery.
- • The shrink is the skill. The walk-away is the fallback.
- • A product that cannot say no stops being able to measure itself — and a product that cannot measure itself cannot compound.
The edges are drawn. Inside them, one payoff has been implied for fifteen chapters and never collected.
Standardise the Compiler, Not the Answer
Say “productised” to a good practice lead and watch them flinch. They are right about the history and wrong about the mechanism.
Bronze, silver, gold. Three tiers of the same work with fewer knobs on the cheaper ones. A "framework" that is a slide with four boxes and a template underneath. Every experienced practitioner has watched productisation arrive as a project to make the work more generic, and has correctly concluded that what gets standardised is quality.
They are right about what happened. They are wrong about why — and the why is what has changed.
What productisation used to require
To predict the labour, you had to reduce variance in the customer's answer. That was not laziness; it was arithmetic. Composing a genuinely fitted answer required expensive human cognition per instance, and per-instance cognition is exactly the thing a fixed price cannot afford. So you standardised the deliverable, templated the analysis, constrained the questions, and priced the result.
The predictability was real. It was purchased with fit — and the buyer who most needed the service, the one whose situation was genuinely unusual, was the one the package served worst.
The honest ancestor
Mass customisation got partway out of this trap decades ago. The term was coined by Stan Davis in 1987, developed by B. Joseph Pine II in 1993, and defined as producing goods and services "to meet individual customers' needs with near mass production efficiency"13. Standardise modules; let combinations vary; ship something that fits better than a catalogue item without costing what bespoke costs.
Say clearly what that achieved and what it did not. It achieved genuine variety at scale in domains where the answer could be assembled from parts. It did not achieve novel answers, because the combinations still came from a catalogue — and the reason the catalogue existed was that composing something outside it required a person to think, at a cost per instance that broke the economics.
That constraint is the one that moved.
The inversion
| Old productisation | AI-native productisation | |
|---|---|---|
| The core move | Reduce variation in the customer's answer | Allow variation in the customer's answer; standardise the machinery that absorbs it |
| What is standardised | The deliverable | The compiler, the proof harness, the exception classes, the boundary logic |
| Where repeatability lives | In the output | In the machinery |
| What specificity costs | An expensive exception, quoted separately or refused | Nothing extra — it is the standard path |
| Where margin comes from | Doing the same thing again more efficiently | The recurring parts ceasing to be novel |
| What the buyer receives | A known artefact, adapted at the margins | A fitted answer, produced through a known and provable path |
Why this follows from the topology
This is not an extra claim bolted onto the architecture. It is a consequence, and the derivation is three lines:
- The perimeter is fixed, so commitment is stable and the firm can be paid for a bounded state.
- The interior is generative, so the answer is free to fit whatever the evidence actually is.
- Repeatability therefore cannot live in the output — it has nowhere left to go except the machinery.
The inversion is what the topology means, once you notice where repeatability ended up. Which also explains why firms that adopt the square without the flywheel end up disappointed: they fixed the perimeter, freed the interior, and then had nowhere to bank the repeatability, so every engagement was a fresh act of intelligence at a fixed price. That is a business model with one very good year in it.
The four things you actually standardise
1. The compiler
The governed path from raw client evidence to an intermediate representation the rest of the machinery can operate on: what gets read, how it is atomised, what provenance travels with each unit, how claims are formed and reconciled. You standardise the path to understanding — never the conclusion.
Buildable version: the ingestion adapters, the atomisation rules, the provenance schema, and the reconciliation logic. Written once, configured per engagement.
2. The proof harness
The deterministic checks and acceptance instruments from Chapters 6 and 11: coverage, reconciliation, completeness, provenance, conformance. Reused on every engagement; the thresholds are configuration, not code.
Buildable version: a test suite that runs against the intermediate representation and fails loudly, plus the evidence pack it emits.
3. The exception classes
The typed taxonomy from Chapter 11, each class carrying a detection condition, a decision rule, a route and an owner. New exceptions are nominated, disposed and typed — never improvised twice.
Buildable version: a ledger with one row per class, and a detection test per row that runs inside the harness.
4. The boundary logic
The band drivers, the crossing definitions, the reserve triggers and the exclusions. This is what makes the perimeter reproducible rather than negotiated per deal — the difference between a product and a series of similar-looking contracts.
Buildable version: the census metrics with published thresholds, the list of named crossings, and the exclusion set written operationally.
Standardise those four and the answer is free. That is the whole move, and it has been named before in this corpus in its sharpest form: you productised the path, not the answer.
Same square. Different thing inside every time.
Economies of specificity
The economic name for the consequence is economies of specificity: when recomputation becomes cheap, value shifts from reproducing for the average to fitting each actual customer, context or mission. And it connects directly back to Chapter 12 — the same machinery becomes more profitable and more reliable across successive engagements, not merely repeatable, because every engagement leaves the compiler, harness, taxonomy and boundary logic slightly better than it found them.
That's not merely better contracting. It's a compiler architecture applied to commerce.
The demand side is already moving
Two signals, offered as evidence about the market rather than about the framework.
McKinsey's global managing partner, describing the firm's own model: "we're migrating pretty quickly away from, let's call it pure advisory work, which was a lot of the origins of our firm and a fee-for-service model… It's moving to much more of an outcomes-based model where we say, 'Look, let's identify a joint business case together and we will underwrite the outcomes of that business case'"14. Whatever one thinks of that firm, it is not a marginal player experimenting at the edges of professional-services pricing.
And from a different profession entirely, the gap stated as a number: 71 per cent of clients would prefer to pay a flat fee for their entire case, while hourly billing remains the most common model, offered by 71 per cent of law firms15.
The same number, on both sides of the table
of clients would prefer to pay a flat fee for their entire case
of law firms still offer hourly billing as their most common model
Clio, Legal Trends Report. Buyers want a stable unit of commitment. Most suppliers still sell time. The gap is the opportunity — and the architecture is how you cross it without gambling.
"Productised means generic"
Myth vs reality
The myth
Productising a service means making the output the same for everyone. Specificity and predictability trade against each other, and you have to choose.
The reality
Genericness was a consequence of expensive per-instance cognition, not a property of productisation. Remove the cost that produced the constraint and the constraint disappears. Under machine breadth, specificity can be the standard path rather than the expensive exception.
The test
If two of your engagements produced substantially the same deliverable, you standardised the answer. If they produced different deliverables through the same machinery, you standardised the compiler.
Run that on your last two engagements before you run anything else. It takes ten minutes and it tells you which business you are actually in.
The payoff is collected. What remains is the part you can do on Monday.
Designing Your Own Square: A Field Method
Take one offer you are currently selling — ideally the one you are least sure about — and put it next to the canvas.
No recap. Pick the offer. Not a hypothetical, not your best one; the one that has been quietly bothering you.
Step 1 — Run the deletion test
Half a page each. The value is not in the answers you write; it is in which of them you cannot write.
1. No square
If we removed the bounded promise and sold access to people instead, what changes?
Good answer: several specific things the buyer would lose — a date, a defined end state, a risk they no longer carry. Bad answer: "not much." What the bad answer tells you: you were already selling time and calling it an outcome, and your competitors can see that too.
2. No barbell
If nobody froze intent up front and nobody owned a consequential acceptance at the end, who would be inspecting what?
Good answer: a named review load you can count in hours. Bad answer: "the team would just be careful." What it tells you: your governance is a culture rather than a structure, and cultures do not survive volume, turnover or a bad quarter.
3. No flywheel
Name three things engagement two will start with that engagement one did not have.
Good answer: three artefacts, each with a location a person could open right now. Bad answer: "the team will know more." What it tells you: your second engagement will cost what your first one cost, and your pricing has no downward pressure available to it.
4. No membrane
Write the sentence you would say to the client's general counsel describing what leaves the engagement, and in what form.
Good answer: a sentence you could say out loud in the room without adjusting it. Bad answer: hesitation, or a paragraph. What it tells you: do not promote anything yet — and write the clause before the next engagement starts, because you cannot write it afterwards.
Step 2 — Fill the perimeter
The eight fields from Chapter 3. Answer what you can. Mark what you cannot. Do not soften a blank into a sentence — a blank is information and a plausible sentence is not.
The fill
The blanks are the design work. They are also, in almost every firm that runs this exercise, exactly where last year's margin went.
Two blanks are so common they are worth predicting. Authoritative inputs: nobody has ever been named, so week six contains an argument about whose spreadsheet was real, and the cost of that argument is absorbed as interior effort. Valid acceptance states: there is only one, and it is "the client is happy" — which means the engagement has no closing event and can therefore be reopened by anyone, at any time, at no cost to them.
One rule while you are here: publish your band drivers even if you never publish a price. Keeping your cost model private is fine. Keeping the drivers private recreates precisely the budget-dance distrust the square was built to end — a buyer who cannot see why they are in this band will hear any adjustment as a renegotiation.
Step 3 — Build three instruments, in this order
First: the failable acceptance test
Everything downstream depends on there being an observable event that closes the engagement. And writing it does something else valuable immediately: it exposes whether the promise was ever real. If you cannot write a test that could come back negative, stop here. The rest of this is decoration until that sentence exists.
Cost: a day. Effect: changes what you can sell, this quarter.
Second: the exception taxonomy
It converts surprise into product language. It starts empty and it is populated by escalations — which means it costs almost nothing to run, provided every escalation is required to leave something behind: a detection condition, a decision rule, a route, an owner. Without that requirement you are running a help desk with a spreadsheet.
Cost: accretes for free. Effect: the first row that fires automatically pays for the whole thing.
Third: the promotion gate and its clause
You cannot promote what you have no right to promote, and the clause has to exist before an engagement rather than after it. Which means writing it now costs a paragraph in a document you are already signing, and writing it later costs the pattern.
Cost: a paragraph and a conversation. Effect: the difference between Failure A in Chapter 13 and a flywheel.
Step 4 — Log one metric before the next engagement
The metric
Scarce-expert dispositions per completed unit.
Numerator discipline: paid, bounded units the
customer actually bought. Not pilots. Not demos. Not "influenced" work. Not every AI-assisted task
inside a legacy day-rate project.
Denominator discipline: genuinely scarce expert dispositions — material
judgements made under authority and consequence. Not every human touch. Not junior assembly the
machine should have absorbed. Not project administration. And not excluding the founder's
evenings, which is the most common way this number is quietly falsified.
Then the warning from Chapter 12, repeated deliberately because this is the moment firms give up: instrumenting this properly usually makes the number look worse first, because it reveals dispositions that were previously invisible. That is the measurement getting honest, not the work getting harder. A firm that flinches here never gets a real number again.
Step 5 — Four refusals for the first ninety days
✗ Don't scale on engagement one
It proved delivery. It did not prove compounding, and scaling now multiplies heroics and brand risk simultaneously.
✗ Don't promote without a gate
A pattern that crossed because it seemed useful is a liability with a release tag on it.
✗ Don't publish a band you cannot defend
If you cannot say why this buyer is in this band, you have reintroduced private judgement at the front of the product you built to replace it.
✗ Don't answer a commercial question with a cadence fact
If your answer to "what will I have, and when" mentions a ceremony, you have answered the wrong question.
The four seams this book did not build
An honest map of the edges, one line each. None of these is smuggled in at half strength anywhere in the preceding sixteen chapters.
- The change-control contract test. How boundary mutations are typed, how reserves are consumed and measured, and what exhaustion formally triggers. Chapter 3 names the distinction; the contract mechanics are a separate treatment.
- Falsifier semantics. What makes an engagement wrong to start versus wrong to accept, and how the two differ. Chapter 6 places the falsifier in the sequence and derives nothing.
- The finance decision behind engagement one. How a deliberate early under-earning is funded, accounted for and killed. Chapter 7 establishes that the flywheel has an economics and stops there.
- Why cognition became abundant. Chapter 1 assumes it in a paragraph with one cited price curve, because arguing it properly is a different book.
"We can't do all of this at once"
You should not. The order in Step 3 exists because it is the cheapest possible sequence: the acceptance test costs a day and changes what you can sell; the taxonomy accretes for free if you make escalations pay for it; the clause is a paragraph in a document that is already being signed. Nothing here requires a programme, a transformation office or a budget line. It requires somebody to write four documents and one metric definition.
The architecture, once more
Square around the engagement. Barbell inside the square. Flywheel between the squares. Membrane around the learning.
And the closure condition — an engagement is not operationally closed until it has answered two questions:
Did we keep the customer's promise?
What, if anything, should the organisation never have to learn again?
The first question is what every firm already tries to answer. The second is the one with no owner — and it is the one that decides whether the whole architecture is a better way to run a project or a different kind of business.
That second question is what turns fixed-price productisation from a one-time commercial trick into a genuinely compounding business model.
Which returns us to the drawing this book started with. The blob with many edges and many curves, the one every service has had and nobody drew, the one the buyer has been financing since the invention of the timesheet. It was never the work that was shapeless.
It was the contract.
References & Sources
The evidence base behind every claim — primary research, industry analysis, and technical specifications
Research Methodology
This ebook draws on primary research from standards bodies, independent research firms, enterprise technology vendors, and consulting firms. Statistics cited throughout have been cross-referenced against primary sources.
Frameworks and interpretive analysis developed by Scott Farrell / LeverageAI are listed separately below — these represent the practitioner lens through which external research is interpreted, and are not cited inline to avoid self-promotional appearance.
Primary Research & Standards Bodies
Project Management Institute / Catherine Elton — Scope Patrol: Scope Creep Is On The Rise As Stakeholder Expectations Increase (PM Network 32(7), 38-45) [1]
PMI 2018 Pulse of the Profession: 52 per cent of projects completed in the last 12 months experienced scope creep, up from 43 per cent five years earlier
https://www.pmi.org/learning/library/scope-creep-rising-11308
Stanford HAI — Artificial Intelligence Index Report 2025, Chapter 1: Research and Development [3]
GPT-3.5-equivalent querying fell from $20.00 to $0.07 per million tokens between November 2022 and October 2024 — a more than 280-fold reduction in about 18 months
https://hai.stanford.edu/ai-index/2025-ai-index-report/research-and-development
Grossman, Hart and Moore — Incomplete contracts (Grossman & Hart 1986; Hart & Moore 1990; Hart 1995) [4]
Contracts cannot specify what is to be done in every possible contingency; at the time of contracting, future contingencies may not even be describable
https://en.wikipedia.org/wiki/Incomplete_contracts
Lindau Nobel Laureate Meetings — Oliver Hart: Incomplete contracts and the theory of the firm [5]
"The benefit of a rigid contract is that it fixes expectations, avoiding arguments. But it may not perform well when there is uncertainty."
https://www.lindau-nobel.org/oliver-hart-incomplete-contracts-and-the-theory-of-the-firm
METR — Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity [7]
Randomised controlled trial: developers took 19 per cent longer with early-2025 AI tools, having forecast a 24 per cent speed-up and believing afterwards they were 20 per cent faster; N = 246 tasks
https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
Ke, Xie, Lu, Zhu, Xing and Ruiz — Evaluation-Driven Development of LLM Agents: A Process Model and Reference Architecture (arXiv:2411.13768) [8]
Embeds continuous evaluation across the lifecycle, extending TDD and BDD to the non-deterministic behaviour and post-deployment evolution of LLM systems
https://arxiv.org/html/2411.13768v2
B. Joseph Pine II — Mass customization (Stan Davis, Future Perfect, 1987; B. Joseph Pine II, Mass Customization, 1993; definition per Tseng & Jiao, 2001) [13]
"producing goods and services to meet individual customers' needs with near mass production efficiency"
https://en.wikipedia.org/wiki/Mass_customization
Industry Analysis & Vendor Research
A&O Shearman — Cost reimbursable vs. lump sum turnkey construction contracts: the many routes to bankability [2]
"In truth, there is no such thing as an absolute fixed price contract"; contractors price significant contingencies and are incentivised to reopen the fixed price
https://www.aoshearman.com/en/insights/cost-reimbursable-vs-lump-sum-turnkey-construction-contracts-the-many-routes-to-bankability
Thomson Reuters — Future of Professionals Report 2025 — the data speaks: what changed in AI adoption [9]
Professionals projected time savings of five hours a week within the next year — about 240 hours a year, an average annual value of roughly $19,000 per user
https://www.thomsonreuters.com/en/insights/articles/the-data-speaks-what-has-changed-in-ai-adoption-trends-this-year
GC AI — AI Legal Ethics in 2026: 6 Cases, 4 Rules, 1 Policy Template (practitioner analysis of ABA Formal Opinion 512 and state guidance) [10]
Client confidential information should only be entered into AI platforms that contractually commit to zero data retention and zero training on customer inputs
https://gc.ai/blog/ai-legal-ethics
Andrew S. Bosin LLC — Technology Lawyer Explains AI Vendor Contracts for Startups (2026 Guide) [11]
Standard AI vendor clauses now cover deletion of customer data, prompts, outputs and derived training data, certification of deletion, and audit rights
https://www.njbusiness-attorney.com/technology-lawyer-ai-vendor-contracts-startups
JMD Ross Insurance Brokers — Professional services contract clauses – Some key points [12]
"without a contractual limitation, liability is unlimited and could exceed the level of cover maintained under your PI policy"
https://www.jmdross.com.au/wp-content/uploads/2018/02/Professional-services-contract-clauses.pdf
Clio — Legal Trends Report — Is Flat Fee Billing Becoming the Norm in Law? [15]
"71% of clients prefer to pay a flat fee for their entire case, and 51% want to pay flat fees for individual activities within their case. Still, hourly billing is the most common, offered by 71% of law firms."
https://www.clio.com/guides/flat-fees-legal-trends
LeverageAI / Scott Farrell — Practitioner Frameworks
The interpretive frameworks, architectural patterns, and practitioner analysis in this ebook were developed through enterprise AI transformation consulting. The articles below are the underlying thinking behind those frameworks. They are listed here for transparency and further exploration — not cited inline, as this is the author's own analytical voice.
Scott Farrell — Buy Certainty First
The broken entry transaction #bb6111; the SOW author as join algorithm #4b124b; the compiled SOW and failable acceptance tests #1cecbb; typed uncertainty #a09dc9; the pricing envelope #7a3521; specimen discipline, n=1 stated as n=1 #b781ee
https://leverageai.com.au/wp-content/media/articles/204-buy-certainty-first.html
Scott Farrell — AI-Constituted Services
The Fixed-Price Envelope: census, bands, included finding counts, Flex Reserve; AI makes the cost curve flatter, not flat — ch9 #f8a5b5; the consulting compiler — ch7 #b9af10
https://leverageai.com.au/wp-content/media/articles/202-ai-constituted-services.html
Scott Farrell — FDE Delivery Looks Like Waterfall Per Increment
The production cadence — understand, specify, design, generate, verify, deploy, learn under tight intent, loose method, hard verification #ef061c; learning write-back so engagement two starts higher than engagement one #3a3429
https://leverageai.com.au/wp-content/media/articles/171-fde-delivery-looks-like-waterfall-per-increment.html
Scott Farrell — AI-Native Successor Offer
Fixed price is a strong signal, not a requirement #27ee47; the successor-unit vocabulary #252640; delivery physics, economics and the transfer gate #75d7af; scarce-expert elasticity #75106a; the worked negative case #9a7f8f
https://leverageai.com.au/wp-content/media/articles/213-ai-native-successor-offer.html
Scott Farrell — Governance Barbell
Heavy plan, cheap middle, heavy verify #ae0640; tight intent, loose method and the authority split #7b0e9e
https://leverageai.com.au/wp-content/media/articles/140-governance-barbell.html
Scott Farrell — The North Star Prompt
Tight intent, loose method — be precise about purpose, stop over-specifying procedure, and stay prescriptive where output feeds a machine or a constraint is safety-critical — ch7 #0aa4c6
https://leverageai.com.au/wp-content/media/articles/70-north-star-prompt.html
Scott Farrell — Wiki Is CapEx
Saved hours arrive as confetti and never consolidate #847bd2; denominate the return in capability, not hours saved #b81c22; the capability ledger #c24980
https://leverageai.com.au/wp-content/media/articles/112-wiki-is-capex.html
Scott Farrell — Forward-Deployed Practice OS
Compile once, escalate with fossils — every escalation should reduce the probability of the next equivalent escalation #7aaf5a; engagement two is the only proof, and the de-identified, transferable, human-gated write-back #6822aa
https://leverageai.com.au/wp-content/media/articles/167-forward-deployed-practice-os.html
Scott Farrell — Experience Is Compressed Priors
The compounding test — did we leave merely another codebase, or a sharper corpus that makes the next project easier and better? #721eb4
https://leverageai.com.au/wp-content/media/articles/150-experience-is-compressed-priors.html
Scott Farrell — A Blueprint for Future Software Teams
Definition of Done v2.0 and its five learning gates #e9e403; knowledge promotion from personal to team to organisational canon #e9c8b1
https://leverageai.com.au/wp-content/media/articles/29-blueprint-future-teams.html
Scott Farrell — The FDE as Paid Product Discovery
The five dispositions — local-only, configurable, internal primitive, supported platform, reject #d885bf; run one nomination end to end through a disposition council #8bd6db
https://leverageai.com.au/wp-content/media/articles/172-the-fde-as-paid-product-discovery.html
Scott Farrell — Five Postures of an AI-Native Consultancy
The offer foundry's six functions; recurrence nominates, humans promote; "build nothing" as a successful paid outcome — ch6 #def85e
https://leverageai.com.au/wp-content/media/articles/210-five-postures-ai-native-consultancy.html
Scott Farrell — One Problem One Offer
Bundled package versus compiler; standardise the compiler, not the answer; productisation as economies of specificity — ch3 #ddb5f1
https://leverageai.com.au/wp-content/media/articles/212-one-problem-one-offer.html
Regulatory Frameworks & Compliance
European Union — Article 14: Human oversight, Regulation (EU) 2024/1689 (EU AI Act) [6]
High-risk AI systems must be designed so natural persons can effectively oversee them, interpret output, remain aware of automation bias and intervene or halt operation; certain uses require separate verification by at least two natural persons
https://artificialintelligenceact.eu/article/14
Major Consulting Firms
Bob Sternfels, Global Managing Partner, McKinsey & Company — Where McKinsey—and Consulting—Go From Here (HBR IdeaCast, January 2026) [14]
"we're migrating pretty quickly away from, let's call it pure advisory work... It's moving to much more of an outcomes-based model where we say, 'let's identify a joint business case together and we will underwrite the outcomes of that business case.'"
https://hbr.org/podcast/2026/01/where-mckinsey-and-consulting-go-from-here
About This Reference List
Compiled August 2026. All URLs verified at time of compilation. Regulatory documents and standards specifications are subject to revision — check primary sources for the most current versions.
Some links to academic papers and vendor research may require free registration. Government and standards body publications are freely accessible.