AI-Native Service Architecture · The change-control seam
Boundary Mutation, Not Change Request
Typing the perimeter so surprise has to declare itself. AI broke the link between “the work got harder” and “the work got more expensive” — and your change-control clause is still keyed to the half that broke.
Somewhere this week, a delivery lead is drafting a change request for work their own machine did.
The engagement was sold as a fixed commitment. Week three, the team discovers that one part of the client’s estate is far messier than the sales conversation implied. The agents that were supposed to read three thousand pages read fifteen thousand. An integration that was supposed to take one pass took three. A whole approach was generated, tested, failed, and thrown away, and a second one was generated in its place. None of it took a human much longer than it would have anyway — most of it happened while the delivery lead was in a different meeting. But it is visibly more work than the plan described, and the statement of work has a clause about that.
So the change request gets written. And the client — who bought certainty — learns that the fixed price was soft, that the supplier’s surprises are the buyer’s problem, and that “fixed” is a marketing word.
Nobody in that story behaved badly. The delivery lead followed the process. The process is what is wrong. It was designed for a world where a different method meant more expensive human hours, and it has not noticed that the world changed.
The claim
Implementation movement is not contract movement. A commercial change exists only when a named perimeter field moves against a recorded value, or a typed reserve exhausts. Everything else — the search, the retries, the regeneration, the ordinary discovery that reality is messier than the deck — belongs to the supplier, because that is what the price already bought.
This piece is about the seam between those two things: the change-control test for an AI-native service. The architecture it sits inside — a stable commercial perimeter around an adaptive, machine-scale production interior — is set out elsewhere and is not rebuilt here1. That book drew the perimeter and then said, twice and in writing, that the detailed contract test for boundary changes — how reserves are typed, how consumption is measured, what exhaustion triggers — is a seam that gets its own treatment.
This is the treatment. What follows is one artefact and one rule: the Boundary Mutation Matrix, which types the perimeter at signature; and three dispositions that replace the generic change request. Then an engagement event log classified against them, including the cases that were genuinely disputed and how they were settled. Then the case where the rule refuses to fix the price at all.
Your change clause is keyed to a broken proxy
Start with what conventional change control actually tests, because it is more specific than “scope” and the specificity is the problem.
In the PMI tradition, the entry test is the baseline. Perform Integrated Change Control is “the process of reviewing all change requests, approving or rejecting changes, and managing changes to deliverables, project documents, and the project management plan”2, and the rule for what enters that funnel is blunt: every change to a baseline needs approval, which is what “keeps scope, schedule, and cost under control and preserves traceability”2. A baseline is a plan of work. So under the world’s most widely taught change-control standard, method movement and perimeter movement enter the same pipeline.
The commercial contract does the same thing with better manners. A change-control clause “aims to regulate change and exclude the possibility of informal, and perhaps inadvertent, variations being made to an agreement orally, or by conduct”3, and a good one is procedurally complete: proforma request and response templates, “time limits for the steps in the change process and the consequences of non-compliance with those deadlines”3, a signed Change Control Note at the end. Read a dozen of them and you will notice what is missing every time. They specify who may request, in what form, by when, and documented how. They almost never specify what makes something a change in the first place.
That absence was survivable for thirty years because there was a reliable proxy sitting underneath it. When production was made of expensive human hours, “the method deviated from the plan” and “the cost went up” were, for practical purposes, the same sentence. Deviation from the described method was cheap to observe and tightly correlated with the thing anyone actually cared about. So contracts keyed on deviation, and the arrangement worked.
The mechanism
Change control has a billing proxy: an observable it uses to stand in for “cost has moved”. The proxy was deviation from the estimated method. Machine-scale production decoupled that proxy from the thing it was proxying for. The gauge still moves; it no longer moves with the quantity it was calibrated against.
How decoupled? Concretely enough to price. Inference cost for a fixed level of capability fell from “$20.00 per million tokens in November 2022 to just $0.07 per million tokens by October 2024” — “a more than 280-fold reduction in approximately 18 months”4. At the task level, one current account of agentic coding puts a full session — “a single agentic task where the AI reads through a codebase, implements a feature across multiple files, runs tests, and iterates through failures”5 — at roughly “$6.00 per session on Opus versus $0.60 on Composer 2 Standard”5. Those are the author’s own illustrative figures rather than measured client data, and should be read as an order of magnitude rather than a rate card. But the order of magnitude is the entire argument. Set six dollars beside one senior consultant’s morning. That gap is the thing your change-control clause has not been told about.
Be precise about which variance moved, because over-claiming here is how a good rule gets discredited. Cheap cognition absorbs breadth: how many documents, how messy, how many candidate approaches, how many retries. It does not absorb everything, and the evidence on net delivery outcomes is genuinely mixed. Google’s DORA programme, surveying nearly five thousand technology professionals, found that “AI adoption does continue to have a negative relationship with software delivery stability”6 even as throughput improved — “AI accelerates software development, but that acceleration can expose weaknesses downstream”6. The marginal cost of another attempt collapsed. The marginal cost of verifying the attempt did not.
Which is exactly why the rule that follows has a hard boundary at both ends and a free interior in the middle, rather than a general presumption that the supplier absorbs things.
The instrument that measured the old world has stopped working
There is a strange piece of corroboration for this from the most rigorous attempt anyone has made to measure AI’s effect on developer productivity. In July 2025, METR ran a randomised controlled trial on experienced open-source developers and found that “when developers are allowed to use AI tools, they take 19% longer to complete issues” while believing they had been sped up by 20%7. That result was widely cited — including, at the time, by us.
In February 2026 METR retired it. The original page now carries a banner: “These results are out of date… We believe these historical results no longer reflect the current impact of AI models on open-source developer productivity”7. Anyone still quoting the 19% figure in 2026 is quoting a superseded result.
But the reason METR changed the experiment is more interesting than the number ever was. They could no longer make elapsed human time mean anything:
“Some developers reported it was challenging to report time-spent in completing tasks when they used agentic tools, because they would often work an unrelated task while waiting for the agent to complete its work.”
— METR, We are Changing our Developer Productivity Experiment Design, February 20268
And they could no longer get people to accept the counterfactual: “30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI”8.
Read that as a commercial finding rather than a methodological one. A research organisation with randomisation, screen recordings and hourly incentives could not keep human-hour accounting meaningful for AI-assisted work. A professional-services firm, with none of those instruments, is billing against it.
Three dispositions, not one change request
The replacement is not a better change-request process. It is a classification made before any of this happens, with exactly three outcomes and no fourth.
| Disposition | What it means | Commercial effect | Who decides |
|---|---|---|---|
| Interior variation | Method, prompt, model, code, sequencing, regeneration, analysis route, ordinary implementation surprise — while every perimeter field holds its recorded value. | None. The supplier absorbs it. The band was priced knowing it happens. | Nobody. It is the default, and the default is silence. |
| Typed surprise | A pre-declared exception class occurs inside the perimeter: access narrowed after contract, an unsupported source the client still needs read, exception density above the band’s assumption. | Reserve draws down by a published rule, visibly, until its band is spent. Unconsumed reserve is not consumed. | Delivery lead, against the published consumption rule. Logged, not negotiated. |
| Boundary mutation | A named perimeter field has moved against its recorded value — or the reserve has exhausted. | Re-contract. Stop, name the field, quantify the delta, present four options: uplift the band, extend the reserve, narrow the boundary, or stop. | A named commercial owner on each side, within a stated window. |
The middle row is the one that is new, and it is worth being explicit about why it earns its own disposition rather than sitting as a bullet inside the third. In the parent architecture, reserve exhaustion appears in the list of things that cross the perimeter, which is correct as far as it goes1. But a reserve draw is structurally unlike both of its neighbours: money moves and nothing reopens. It is the only state in the system where the commercial agreement stays closed while the supplier’s absorption is visibly consumed. Collapse it into interior variation and the buyer never sees the meter until it has run out. Collapse it into boundary mutation and you have re-contracted for a delayed access approval, which is exactly the behaviour the rule exists to prevent.
None of the three is an invention from nothing. The reserve is inherited: the Fixed-Price Envelope already defines “a named, priced absorption layer for typed surprise” that is “drawn against typed events, visible to both sides” and where “what is not drawn is not consumed”9. This piece does not re-teach that instrument. It types the trigger, which is the part nobody has written down.
The default has to be “no commercial event”
There is a design detail here that decides whether any of this survives contact with a delivery organisation, and it is not the taxonomy. It is which disposition is the default.
Build a classifier with no default and you get a change request for everything, for the same reason an assistant with no threshold narrates every event and trains you to ignore it. The fix is structural: a null option that is always available and always compared against. In our own work on when an agent should stay silent, the pattern is called stand-pat, and the framing transfers exactly — almost every event is a capture on the board: available, tempting, and safe to leave alone; a very few are checks you are genuinely not allowed to ignore; and the entire skill is telling those two apart10.
So: interior variation is the default, and it requires no decision by anybody. A surprise becomes commercial only by earning its way out of the default, and it earns its way out by moving a named field. That is the whole rule. Everything below is the work of making “moving a named field” observable.
The Boundary Mutation Matrix
Eight fields. Each one gets three things at signature: a recorded starting value, an observable mutation trigger, and a pre-agreed commercial response.
The design rule that governs every row is a destructive test, and it should be applied to your own matrix before a client ever sees it:
The trigger test
A trigger that requires judgement at dispute time has failed. If settling whether the field moved needs an argument, a partner’s memory of the negotiation, or a reading of intent, you have not typed the field — you have renamed the dispute. Rewrite it as a comparison against a recorded value, or admit in the contract that this field is untyped and price accordingly.
Six of the eight rows below are the perimeter fields of the parent architecture, re-cut for change control1. Two are additions this rule needs and the parent does not have as fields: buyer-controlled dependencies and consequence/liability class. The parent treats both as reasons to decline an engagement. Under a change-control rule they have to be live rows, because both of them move during delivery, and when they move the supplier’s position changes more than any amount of extra analysis can compensate for.
| Perimeter field | Recorded at signature | Observable mutation trigger | Commercial response |
|---|---|---|---|
| 1. Promised state | The bounded state that will exist at the end, written as a testable sentence, plus the valid terminal states. | The written promise sentence would have to change to describe what is now being asked for. Test: read the signed sentence aloud; does the current request fit inside it without adding a clause? | Re-contract. This is the one field where latitude has no room at all — method can move, the promise cannot. |
| 2. Authoritative input estate | The declared source classes, the named systems, who warrants each one, and the access status of each at signature. | A source class not on the declared list is required; or a declared source is withdrawn, replaced, or its warrant moves to a different party. | Re-contract, or type as excluded. Adding a source class is a mutation; more volume within a declared class is row 3. |
| 3. Volume / band | The census counts that assigned the band, each with its measured number, and the band’s upper threshold published. | A census metric crosses its published band threshold. Purely arithmetic: re-run the census, compare to the recorded value. | Band uplift at the published rate. Not a negotiation — a re-configuration. |
| 4. Buyer-controlled dependencies | Every input the buyer owes: access grants, environments, nominated people, decisions, third-party cooperation — each with an owner’s name and a date. | A dated buyer obligation passes its date unmet, by the number of days stated in the schedule. | First occurrence: typed surprise, reserve draws. Beyond the reserve band, or where the critical path moves: re-contract with a revised time boundary. |
| 5. Authority and access | Who signs what, on which side; which classes of action may be autonomous; the security and privilege model the work runs under. | A new signatory or approving party is required; or an approval class moves from one side of the table to the other; or the privilege model narrows after contract. | Re-contract. An added approver changes the disposition path and therefore the metered resource. |
| 6. Consequence / liability class | The named liability cap, the carve-outs as a closed list, the standard of care, and the intended use of the deliverable. | The deliverable’s intended use changes; a regulator or a fitness-for-purpose obligation enters; a carve-out is requested outside the closed list; a cap is asked to move. | Re-contract, always — and re-underwrite. This field is never absorbed and never reserved. See below. |
| 7. Acceptance rule | The observable event that closes the engagement, written so that it can fail, plus who runs it. | A new acceptance condition is proposed; or the agreed test is declared insufficient by the party that agreed it; or the named runner of the test changes. | Re-contract. A failed acceptance under the agreed test is not a mutation — it is a defined state whose remedy sits inside the band. |
| 8. Fixed time boundary | The end date, what the date means commercially, and which party’s clock each dependency runs on. | Either party asks to move the date; or a dated buyer obligation (row 4) has pushed the critical path past it. | Re-contract. Compression is a mutation as much as extension — a shortened boundary changes the disposition schedule. |
Row 6 has different physics from the other seven
Every other row can, in principle, be traded: a bigger band for a bigger fee, a later date for a narrower promise. Row 6 cannot, and the reason is insurance rather than commerce.
Without a contractual cap, professional liability is not large — it is unbounded. As one Australian professional-indemnity broker puts it: “without a contractual limitation, liability is unlimited and could exceed the level of cover maintained under your PI policy”11. And some obligations are not insurable at any price: “fitness-for-purpose obligations should be avoided as they are generally uninsurable under a PI policy provided to a professional services provider”11.
Caps hold, and they hold at levels that make the point vividly. American engineering practice records a $50,000 limitation enforced even though it “accounted for only 8% of the designer’s fee”, and another that limited recovery “to only $550,000 out of a $9.5 million jury verdict”12. The rationale is stated without embarrassment: design professionals’ “fees do not cover the potential that they can be liable for virtually unlimited financial exposure if there is a claim”12.
Now the part that belongs in this article rather than a liability textbook. In the first of those cases, the cap had been set at a percentage of the original fee; scope was then added through addenda, nobody reopened the cap, and by the end the cap was 8% of the fee. The court enforced it anyway, and said why:
“The failure of (the contractor) to address or renegotiate the limitation of liability clause during the execution of each addendum has made the term of the contract more burdensome than previously anticipated… This court is unwilling to allow (the contractor) to avoid a term of the contract simply because it has become more burdensome due to its own failure to renegotiate.”
— Zirkelbach Construction Inc. v. DOWL LLC, as reported in ASCE’s Civil Engineering12
That is a litigated example of a perimeter moving while the contract stood still. It is also the reason the matrix is two-sided rather than a supplier’s convenience: the party who failed to notice the movement is the party who wore it, and in that case it was the buyer of the services who ended up holding the exposure they thought they had transferred.
Typed at signature, not argued at dispute
Everything above collapses if the fields have no recorded values. This is not a theoretical failure mode — it is the ordinary one.
The failure looks like this. A firm sells a bounded engagement well, staffs it competently, and assigns the band from a sales conversation rather than a measured input surface, because the buyer was in a hurry. Delivery starts. Interior variance happens and is absorbed, correctly and invisibly. Boundary movement also happens — two more sign-off parties appear, a source class nobody declared turns up in week three. Nothing distinguishes the two, because nothing was measured at the start. The supplier absorbs both and congratulates itself on not raising a change request. Margin erodes silently, then quickly. And then the supplier has to reopen the price after all, at which point the buyer experiences the reopening as the old change request in new clothes13.
The diagnosis is one sentence: no measurement, therefore no recorded values, therefore no delta when reality arrived, therefore no language in which surprise could speak13.
This is also the honest answer to the obvious objection — that a matrix just moves the argument rather than removing it. It does move the argument. That is the point. It moves it from week three of delivery, when both parties have spent money and one of them is embarrassed, to the week before signature, when neither has. And it changes the form of the argument from interpretation to comparison.
This is not novel, and the book is stronger for saying so
Three established contract families already do parts of this, and the differences are more instructive than the similarities.
NEC contracts enumerate their triggers. Under NEC, “a compensation event is a term used… to mean an event which can affect the cost to the Client of the work being carried out, the time when the works will be completed, or both”, and — crucially — “a compensation event is the only way in which these can be changed. There are no other ways in which a Contractor can claim additional payment… or be allowed additional time”14. Clause 60.1 lists the events. The list is exhaustive by construction, and it includes client-side failures — “a failure by the Client, the Project Manager or Supervisor to take an action which the contract requires them to do”14 — which is precisely why row 4 of the matrix exists. What NEC does not do is the move made here: it still assesses each event’s cost and time effect individually. It prices the work. The matrix prices the boundary.
Infrastructure procurement already has a three-class model. The World Bank / APMG PPP guidance splits variations into categories, and the first has no procedure at all: “in circumstances where a proposed variation involves no additional costs for either party, no formal variation procedure is required”15. The second is pre-rate-carded — the contract “can require the private partner to provide a schedule of rates for a range of likely small works at the beginning of each year”15. The third is a full re-contract. That is interior / reserve / mutation in an infrastructure setting. The guidance even recommends typing at signature: where variations “can be foreseen to a reasonable degree before the signing… the government should explore the feasibility of requiring the private partner to commit to pricing pre-specified variations as part of the” agreement15. And it states the classification question outright — contract managers must “verify that a variation request is actually a change and not covered under the existing agreement and pricing structures”15.
IT service management already pre-authorises a class of change. ITIL 4’s standard change is “a low-risk, repeatable, and pre-authorized change that follows a documented procedure” and “often require[s] little or no additional approval”, frequently automated16. ITIL got there to unblock a change advisory board rather than to fix a commercial rule, and it classifies by risk of the change rather than by which commercial variable moved. But it establishes the principle that pre-classifying a change type out of the approval path is ordinary engineering practice, not a supplier land-grab.
So the honest positioning is narrow: professional services is the outlier. The world’s most scrutinised long-term contracts already refuse to treat every deviation as a commercial event. We never adopted the discipline, because while production was human, we did not need to.
The event log: one engagement, every surprise classified
A matrix that has only ever been applied to clean cases proves nothing. What follows is a complete event log for a single bounded engagement, with every variation classified — including the four that were genuinely contested and the rule used to settle each one.
What this is, and what it is not
This is a worked model with stated assumptions, not measured client data. The engagement is a generalised bounded board-decision product: a fixed commitment to produce an evidence-backed decision pack whose valid terminal states are proceed, reshape or stop. Assumptions: a measured input surface at signature; four declared authoritative source classes; a band that includes twenty consequential human dispositions; a named reserve of five access exceptions. No prices appear, because none of the figures here are market data and inventing them would make the rest untrustworthy. Where a number would be evidence, the shape is written instead.
| # | Event | Disposition | Why — and the field, if one moved |
|---|---|---|---|
| 1 | Week 1. The evidence base is assembled, then rebuilt from scratch: the first assembly leaned on a source that turned out not to be authoritative. | Interior | No field moved. The declared source list was unchanged; the supplier misread its own list. Rework of the supplier’s own error is the clearest interior case there is. |
| 2 | Week 1. Access to the second system is granted four days after the date in the schedule. | Reserve | Row 4 trigger fired: a dated buyer obligation passed its date. Under the band, one access exception drawn. Reserve at 1 of 5. Critical path unaffected, so row 8 did not move. |
| 3 | Week 2. An authoritative source arrives in a structure nobody anticipated; an adapter is written on the spot. | Interior | The source class was declared; only its shape surprised us. Shape is method. Class is perimeter. |
| 4 | Week 2. Two declared sources disagree systematically on a material quantity. Reconciliation takes four passes; a senior spends most of a day establishing that the difference is a compilation artefact rather than a real discrepancy. | Interior | No field moved — but this consumed one of the twenty included dispositions, so it is metered even though it is not commercial. It also had to leave a fossil behind: a typed exception class, a detection test, a decision rule and a routing trigger17. |
| 5 | Week 3. The team concludes that regenerating the analysis pipeline from an improved specification beats patching the existing one, and throws away nine days of machine output. | Interior | Textbook. Regeneration over patching is a production decision inside an unmoved perimeter18. The buyer bought a state; the search that produced it is the supplier’s business. |
| 6 | Week 3. A fifth source class appears that nobody declared: an operational system holding data material to two of the decision options. | Mutation | Row 2. A source class not on the declared list. Delta quantified against the recorded list within the week; four options presented. Buyer chose to narrow the boundary and type the system as excluded from this phase. |
| 7 | Week 4. Agent read volume runs to roughly five times the planning estimate because the declared documents were far denser than the census sample suggested. | Interior | Row 3 was checked and did not move: the census counted documents in declared systems, and that count was accurate. Density is not a band driver. See dispute A. |
| 8 | Week 4. Security review adds a two-week privilege-approval cycle before sensors can run in the second environment. | Reserve | Pre-declared exception class. Reserve at 2 of 5. Flagged as a row 8 watch item because a repeat would move the critical path. |
| 9 | Week 5. The buyer asks for a second business unit to be included. The requested interface, deliverable and decision look identical to the first. | Mutation | Rows 2, 5 and 6 all moved: a different data authority warrants the estate, a different approver signs, and a different regulator applies. Identical UI, different perimeter. See dispute B. |
| 10 | Week 5. A candidate framing survives three tests and dies on the fourth; a replacement is generated. | Interior | The interior is supposed to look wasteful from outside. Eight framings built, six discarded, and none of it reaches the invoice. |
| 11 | Week 6. A named buyer decision-maker is unavailable for eleven days; two dispositions queue behind them. | Reserve | Row 4 again: a nominated person is a dated buyer obligation. Reserve at 3 of 5. The eleven days are visible in the log rather than absorbed silently. |
| 12 | Week 6. An undeclared internal dependency in the constraint set forces the whole option space to be re-tested. | Interior | The constraints were declared; their interaction was not understood. Understanding constraints is the work. |
| 13 | Week 7. The buyer’s legal team asks that the decision pack be able to be relied on by a lender in a financing process. | Mutation | Row 6, and the only row where the answer is always the same. Intended use changed, therefore the consequence tail changed, therefore the cap and carve-outs are being asked to move. Re-contract and re-underwrite, or decline the reliance. See dispute D. |
| 14 | Week 7. Access to a third environment is narrowed after contract; a workaround costs four machine-days. | Reserve | Reserve at 4 of 5. The workaround cost is irrelevant to the classification — the trigger is the access change, not the effort. |
| 15 | Week 8. A sixth access exception occurs: a fourth environment requires a privilege model the supplier does not hold. | Mutation | Reserve exhaustion. Exception six against a five-exception band. This does not become a margin dispute; it becomes a scheduled decision. Walked through below. |
| 16 | Week 9. Acceptance is run and one option fails its evidence-coverage threshold; the remedy takes three days. | Interior | A failed acceptance under the agreed test is a defined state with a defined remedy, and the remedy sits inside the band17. If a failed test were a change request, the test would not be a test. |
| 17 | Week 9. The buyer asks whether the engagement can also “design and launch three new offers” once the classification work lands. | Mutation | Row 1. Read the signed promise sentence aloud — it describes classifying a current estate. The requested work does not fit inside it without adding a clause. The promise moved; latitude cannot absorb it. See dispute C. |
Seventeen events. Eight interior, four reserve draws, five boundary mutations. In a conventional change-control regime, a defensible reading of the same log would have produced somewhere between eight and thirteen change requests, because events 1, 3, 4, 5, 7, 10, 12 and 16 all involve visibly more work than the plan described.
The four disputes, and the rule that settled each
The clean rows are not the proof. These are.
Dispute A — five times the reading volume (event 7)
The argument. Delivery said this was obviously a volume event: the band exists to price volume, and volume was five times the estimate. Commercial said the band was assigned on document count and document count was correct.
The rule that settled it. A band driver is whatever the published census actually counts, and nothing else. Ours counted documents in declared systems; it did not count pages, tokens or density. So the recorded value did not move, and the event was interior. The supplier absorbed it, and it was the right answer for the wrong-feeling reason.
What it changed. The classification stood, and the census was wrong. Both of those things are true at once, and separating them is the discipline. Density went into the band-driver review as a candidate metric for the next contract — because when the same exception class keeps appearing, the fix is to promote it into a band driver rather than leave it as reserve folklore19. What you may not do is retro-fit a driver mid-engagement to recover margin. That is the behaviour that destroys the band’s meaning for every future buyer20.
Dispute B — the second business unit that looked identical (event 9)
The argument. The buyer’s position was reasonable and sincerely held: same screens, same deliverable, same decision, a modest increment of work. Their delivery counterpart agreed. Our commercial lead did not.
The rule that settled it. Classification runs on fields, not on resemblance. Three fields moved: row 2 (a different party warrants the authoritative estate), row 5 (a different executive signs the disposition), row 6 (a different regulator, therefore a different consequence class). The user interface is not a perimeter field and never was.
Why the buyer accepted it. Not because the supplier had a better argument, but because the fields had recorded values with names against them, and one of the names was theirs. This is the single strongest practical case for typing at signature: at dispute time the conversation was two people reading a table, not two people recalling a meeting.
Dispute C — “you have all the analysis anyway” (event 17)
The argument. The buyer’s point had real force: the machine had already read everything needed to design the new offers, so the marginal cost of doing it was small. Under this article’s own logic — the supplier absorbs what the machine can do cheaply — why not absorb this?
The rule that settled it. Because absorption is bounded by the promise, not by the cost. Row 1 is the field where cheapness is irrelevant: classifying a current estate and designing three new offers are different promised states, with different acceptance tests and a different consequence class if they are wrong. The fact that the marginal machine cost was low is exactly the trap. Cheap to produce is not the same as cheap to be wrong about.
What made it easy to say. The answer was not “no”. It was “that is a second commercial object and here is its shape” — which is a better commercial outcome than absorbing it would have been, and a far better one than an argument at month end.
Dispute D — lender reliance (event 13)
The argument. Delivery saw a document-handling question: the pack already exists, so letting one more party read it costs nothing. That is true of the reading and false of everything else.
The rule that settled it. Reliance changes who can sue, for what, and under which standard. It is row 6, and row 6 is never absorbed and never reserved. Better analysis lowers the probability of an error; it does nothing at all to the cost of the one you make20. Cheap cognition is silent on the consequence tail.
Outcome. Re-contract with a named reliance party, a revised cap, and the carve-outs restated as a closed list — because a generous cap means little if the waiver strips out the losses anyone would actually claim21. Had the buyer declined to re-contract, the correct answer was to decline the reliance and keep delivering the original engagement.
Exception six: reserve exhaustion is a decision, not a dispute
Event 15 is where most fixed-price engagements quietly go wrong, so it is worth walking rather than naming.
The band included a named reserve of five access exceptions. Four had been drawn, visibly, each one logged against its trigger. In week eight a sixth arrived. In a conventional engagement this is the moment the supplier starts absorbing quietly and the relationship starts corroding — or the moment a change request lands without warning and the buyer feels ambushed. Neither is necessary, because the exhaustion of a reserve is the most predictable event in the entire engagement. It has a counter on it.
The exhaustion protocol
- The counter is published from day one. Reserve consumption appears on the same surface as disposition consumption. “Access exceptions: 4 of 5” is visible to the buyer in week seven, not disclosed in week eight.
- A named owner on each side. Written into the schedule at signature, with roles not people where possible. Not “the parties will discuss”.
- A stated window. The decision happens within five business days of the draw that exhausts the band — in the same week, not accumulated into a month-end conversation.
- A delta, with evidence. What was recorded, what is now true, how much of the band it consumed, and what the remaining schedule looks like under each option.
- Four options, always the same four. Uplift the band. Extend the reserve at a published rate. Narrow the boundary — type the fourth environment as excluded and deliver the rest. Or stop, with the work to date delivered in its typed state.
- The default if nobody decides. Stated in the contract, because unowned decisions default to whoever is least able to refuse. Ours: work continues on everything not blocked by the exhausted class, and the blocked portion is typed inaccessible within boundary until the decision lands.
In the worked engagement the buyer chose to narrow the boundary. The fourth environment was typed as excluded from this phase, and the decision pack shipped saying so — which is the point of typed uncertainty: an unknown became a deliverable rather than an argument19. The parent doctrine already says exhaustion should produce “a commercial conversation with a census delta and a recommendation: uplift band, extend reserve, narrow boundary, or stop”22. What is added here is the owner, the window, and the stated default — because a conversation with no owner and no clock is how the silent absorption starts.
What a reserve draw has to leave behind
A draw that produces only an answer is an unpriced leak. Every draw should deposit four things: a typed exception class with a detection condition, a deterministic test that flags the same signature automatically next time, a decision rule for which way it resolves, and a routing trigger naming who looks at it and at what seniority17. Without those, a day of senior time bought one answer. With them, it bought a class — and the next contract can price it as a band driver instead of a reserve event.
Which variable actually costs you
The matrix implies a claim about unit economics, and the claim should be checked rather than asserted: that implementation churn no longer deserves to be the billing proxy, because it no longer predicts cost.
The honest way to test that is the contribution shape our own qualification work already uses23 — willingness to pay, minus machine and infrastructure cost, minus human disposition cost, minus sales and onboarding, minus physical delivery capacity, minus liability and commitment risk, minus exceptions and failure remediation. What follows is a sensitivity read across that shape, stated as directions and magnitudes rather than invented percentages. Nobody has published measured figures for this and I am not going to manufacture them.
| Variable | Who owns it | Effect on engagement cost when it moves | Should it be a billing trigger? |
|---|---|---|---|
| Implementation churn (retries, regenerations, discarded approaches) | Supplier | Real but small in absolute terms, and superlinear in session length rather than linear — retries at high context cost more per attempt than early ones5. Averages out across a portfolio. | No. It is a portfolio property, not a per-engagement prediction, and it is the supplier’s own search. |
| Consequential dispositions (calls a named human must own) | Shared | The dominant controllable cost. Each one occupies scarce senior attention that does not parallelise and cannot be regenerated. | Yes — this is the meter. Band it, publish the included count, show the counter. |
| Consequence / liability class | Buyer’s use, supplier’s exposure | Step function, not a slope. Adding a reliance party or a fitness-for-purpose obligation can exceed the entire fee — and may be uninsurable11. | Yes, absolutely — and it is a re-underwrite, not a repricing. |
| Buyer-controlled dependencies | Buyer | Costs calendar rather than machine time, and calendar is where the supplier’s capacity is actually consumed. Delay also idles the scarce humans, which is the expensive resource. | Yes — but as a reserve draw first, because a few slipped dates are normal and re-contracting for them is absurd. |
| Authority and access | Buyer | Adding an approver adds dispositions and lengthens every loop that touches them. A narrowed privilege model can invalidate the whole production route. | Yes. It changes the metered resource directly. |
| Volume within a declared class | Buyer’s estate | Mostly absorbed by machine breadth — unless it pushes disposition count up, which is the real transmission channel. | Only at a published threshold, and only because volume correlates with dispositions, not because volume itself is expensive. |
| Exception density | Reality | The most under-costed line in every early product. Early economics look attractive precisely because exception tails have not appeared yet23. | Yes, via the reserve — and if it recurs, promote it into a band driver. |
Read down the “who owns it” column and the design falls out of it. Every variable the supplier controls is absorbed. Every variable the buyer controls is typed. The one variable nobody controls — exception density — is the one that gets a reserve, because a reserve is precisely the instrument for shared exposure to reality.
Which also answers the sharpest objection to the whole scheme: isn’t this just a supplier deciding what it feels like absorbing? No — because the supplier absorbs exactly the variables it controls, and gets no relief at all on the ones it does not. A rule that only ever ran in the supplier’s favour would have absorbed row 4 as a gesture of goodwill and quietly repriced row 3. This one does the opposite.
Where the rule refuses to fix the price
The matrix has a self-test, and it fails honestly. Work through the eight rows for a candidate engagement and ask, for each: can I write an observable trigger for this? Where the answer is no, you have not found a hard row — you have found a field you cannot bound, and a fixed price over an unbounded field is not a commitment, it is a wager.
Counterexample: the engagement this rule refuses
A buyer wants a fixed-price commitment to “get us to a decision on our regulatory posture” across a group whose subsidiaries are still being restructured.
Row 1 — promised state. The decision has been redefined twice in three meetings. There is no sentence to record, so there is nothing to compare against later.
Row 2 — authoritative estate. Which entity’s records govern depends on a restructure that has not completed. Nobody can warrant the estate today.
Row 5 — authority. The approving body will exist after the restructure. Its composition is unknown.
Row 6 — consequence class. The regulator that will apply is one of the things the restructure decides.
Verdict. Four of eight rows cannot carry a recorded value, and three of the four are outside both parties’ control. This is not a hard engagement. It is an unclassifiable one: no recorded value means no delta, no delta means no trigger, and no trigger means every surprise resolves into an argument. Do not fix this price.
The shrink, which is the actual answer. Sell the bounding. A smaller fixed commitment whose entire deliverable is a stable, testable question and a recorded perimeter: which entity governs, which regulator applies, who approves, and what decision is actually being asked. That is a real product, it is boundable today, and it makes the larger engagement quotable in a way that no amount of confidence would have.
Refusing is not a moral posture, and a purely principled version of this argument is not usable by anyone with a pipeline. The commercial case is stronger: forcing out-of-band work into a fixed price destroys the band’s meaning for every future buyer, teaches your own sales system that drivers are negotiable, and converts a product back into a bespoke project with a product’s price and a project’s cost20. And the realistic move is nearly always the shrink rather than the walk-away.
None of which is AI-era novelty. Construction lawyers have been saying the quiet part for decades: lump-sum contractors “often include significant contingencies in their pricing”, are “naturally incentivized to seek opportunities to reopen the fixed price” where those contingencies prove insufficient, and — the sentence to keep — “in truth, there is no such thing as an absolute fixed price contract”24. The same source notes that lump-sum “may still be preferable for well-defined, low-risk projects where scope and owner requirements are clear from the outset”24, which is the refusal test in a lawyer’s mouth.
So the matrix does not abolish reopening. Nothing does. It makes reopening early, typed and evidenced instead of late, adversarial and improvised. That is a smaller claim than “AI makes fixed price safe” and it is the only one that survives contact with a real engagement.
The honest cost of typing: basis risk
There is a mature commercial instrument built entirely on observable triggers, and it is worth borrowing from because it has already paid for the lesson. Parametric insurance pays “a pre-agreed amount based on the magnitude of the event, as opposed to the size of losses”, and a contract typically specifies “(1) the payment amount; (2) the trigger (a pre-determined parameter based on observable data); and (3) an impartial third party to verify that the trigger was met”25. The payoff is that eliminating claims assessment lets money move “in a matter of weeks… versus months or years”25, and that “the use of a clearly defined trigger may make it easier for the insured to understand the coverage provided and reduce policy disputes”25.
The price is a named thing:
“Compensation from parametric policies is not linked to actual losses, so the claim payment may be higher or lower than the losses incurred. This is known as basis risk.”
— Congressional Research Service, Parametric Insurance for Natural Disasters, May 202625
Their worked failure is instructive: a school district’s parametric wind policy did not pay after a hurricane because the winds “did not meet the 100 mph trigger… despite damage to school facilities”25. A city insured on barometric pressure could find itself uncovered when the damage came from storm surge.
The Boundary Mutation Matrix inherits exactly this exposure, and dispute A above is a small instance of it: the trigger said the band had not moved, and the supplier absorbed something a better-chosen trigger would have caught. That is the trade. The mitigation is the same one the insurance literature gives — choose triggers “highly correlated” to what you actually care about, name the verifier up front, and revise the triggers between engagements rather than during them.
Say the basis risk out loud in the contract. A rule that claims to eliminate judgement is lying; a rule that concentrates judgement into the week before signature, where it can be exercised calmly, is the best available deal.
What to do with this on Monday
None of this requires a new contract template to start.
- Run the classification backwards on your last completed engagement. List every surprise. For each, ask which named perimeter field moved. Count the ones you cannot answer without an argument. That count — not the number of change requests you raised — is the honest measure of whether your perimeter was typed or merely written.
- Write the eight recorded values for your live engagement, today. Even retrospectively. A recorded value written in week four is worth more than none, and the act of writing them will surface at least one field you cannot answer.
- Write an observable trigger for each field and delete the ones that need judgement. The deletions are the finding. They are the fields that will be litigated, and they are where your next contract needs work.
- Publish one counter. Reserve consumption, on the same surface as everything else the buyer sees. Visible absorption is worth several times silent absorption, and it costs nothing to show.
- Name the exhaustion owner and the window. Two names and five business days, in the schedule. This is the cheapest clause in the whole scheme and it prevents the most expensive failure.
- Take the next unclassifiable opportunity and shrink it rather than pricing it bravely. Sell the bounding as its own commitment. Shrinking is the harder skill and the one nobody teaches20.
What all six have in common is that they move judgement earlier. That is the entire mechanism, and it is why the rule is worth more than its taxonomy: not because three categories are better than one, but because the categories force a set of decisions to be made while both parties are calm, informed and not yet committed.
Fixed price was never the invariant — it is a strong signal that a supplier has made its complexity legible enough to take responsibility for it, and a signal is not a religion26. The invariant underneath is a stable unit of commitment. What this rule protects is the stability: a promise that does not move when the method does, and does move — promptly, with evidence, in front of the person who owns the consequence — when the boundary really has.
Stop writing change requests for work your machine should absorb. Start writing down which eight things, if they move, mean the deal has changed.
A note on scope: acceptance semantics — what counts as evidence, and how a test earns the right to fail an engagement — are treated only here as one perimeter field among eight. They deserve, and will get, their own treatment. Any contract language suggested above is illustrative drafting to show the shape of a clause, not legal advice; commercial counsel review is required and jurisdictions differ materially.
References
- Scott Farrell, LeverageAI. “AI-Native Service Architecture — The Square, the Barbell, the Flywheel and the Membrane.” — The eight perimeter fields, the interior/boundary split, and the deferred seam: “The detailed contract test for boundary changes — how reserves are typed, how consumption is measured, what exhaustion triggers — is a seam that gets its own treatment.” https://leverageai.com.au/wp-content/media/articles/226-ai-native-service-architecture.html
- Project Management Knowledge. “Perform Integrated Change Control.” — “Perform Integrated Change Control is the process of reviewing all change requests, approving or rejecting changes, and managing changes to deliverables, project documents, and the project management plan”; “Every change to a baseline does. This keeps scope, schedule, and cost under control and preserves traceability.” (Practitioner reference describing the PMBOK process, not PMI’s own text.) https://project-management-knowledge.com/definitions/p/perform-integrated-change-control
- Oracle Law Global. “The benefits and pitfalls of a contract’s ‘change control’ clause.” — “it aims to regulate change and exclude the possibility of informal, and perhaps inadvertent, variations being made to an agreement orally, or by conduct”; “It should also specify time limits for the steps in the change process and the consequences of non-compliance with those deadlines.” https://oraclelawglobal.com/news/general/the-benefits-and-pitfalls-of-a-contracts-change-control-clause
- Stanford HAI. “Artificial Intelligence Index Report 2025,” Chapter 1: Research and Development. — “dropped from $20.00 per million tokens in November 2022 to just $0.07 per million tokens by October 2024 (Gemini-1.5-Flash-8B)—a more than 280-fold reduction in approximately 18 months.” https://hai.stanford.edu/assets/files/hai_ai-index-report-2025_chapter1_final.pdf
- Casey Harding, Vantage. “The Hidden Cost Driver in Agentic Coding Sessions in 2026.” — “A single agentic task where the AI reads through a codebase, implements a feature across multiple files, runs tests, and iterates through failures can consume more tokens than a week of casual usage”; “$6.00 per session on Opus versus $0.60 on Composer 2 Standard”; “Three failed attempts at turn 40 don’t just cost 3x a single turn.” (Author’s illustrative model, not measured client data.) https://www.vantage.sh/blog/agentic-coding-costs
- Nathen Harvey and Derek DeBellis, Google Cloud / DORA. “Announcing the 2025 DORA Report” (survey of nearly 5,000 technology professionals). — “AI adoption does continue to have a negative relationship with software delivery stability”; “AI accelerates software development, but that acceleration can expose weaknesses downstream.” https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report
- Joel Becker, Nate Rush, Beth Barnes and David Rein, METR. “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity” (10 July 2025). — “When developers are allowed to use AI tools, they take 19% longer to complete issues”; and the 2026 banner on the same page: “These results are out of date… We believe these historical results no longer reflect the current impact of AI models on open-source developer productivity.” https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- Joel Becker, Nate Rush, Tom Cunningham, David Rein and Khalid Mahamud, METR. “We are Changing our Developer Productivity Experiment Design” (24 February 2026). — “Some developers reported it was challenging to report time-spent in completing tasks when they used agentic tools, because they would often work an unrelated task while waiting for the agent to complete its work”; “30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI.” https://metr.org/blog/2026-02-24-uplift-update/
- Scott Farrell, LeverageAI. “AI-Constituted Services.” — The Fixed-Price Envelope: “A named, priced absorption layer for typed surprise… the reserve is drawn against typed events, visible to both sides, and what is not drawn is not consumed”; “AI does not make the cost curve flat. It makes it flatter.” https://leverageai.com.au/wp-content/media/articles/202-ai-constituted-services.html
- Scott Farrell, LeverageAI. “Stand Pat.” — The structural null option: “the default output is nothing, and only a genuine, must-answer event earns an interruption”; “The entire skill is telling those two apart.” https://leverageai.com.au/wp-content/media/articles/101-stand-pat.html
- JMD Ross Insurance Brokers. “Professional services contract clauses — Some key points.” — “without a contractual limitation, liability is unlimited and could exceed the level of cover maintained under your PI policy”; “Fitness-for-purpose obligations should be avoided as they are generally uninsurable under a PI policy provided to a professional services provider.” https://www.jmdross.com.au/wp-content/uploads/2018/02/Professional-services-contract-clauses.pdf
- Michael C. Loulakis and Lauren P. McLaughlin. “Limitation of liability clauses are like kryptonite,” Civil Engineering, American Society of Civil Engineers, December 2021. — “their fees do not cover the potential that they can be liable for virtually unlimited financial exposure if there is a claim”; a $50,000 cap enforced at “only 8% of the designer’s fee” (Zirkelbach Construction Inc. v. DOWL LLC, 2017); “$550,000 out of a $9.5 million jury verdict” (Taylor Morrison of Colorado Inc. v. Terracon Consultants Inc., 2017). https://www.asce.org/publications-and-news/civil-engineering-source/civil-engineering-magazine/article/2021/12/limitation-of-liability-clauses-are-like-kryptonite
- Scott Farrell, LeverageAI. “AI-Native Service Architecture” (the unmeasured-boundary failure specimen). — “No measurement, therefore no band drivers with values attached, therefore no delta available when reality arrived, therefore no language in which surprise could speak”; “The buyer experiences that reopening as the old change request in new clothes.” https://leverageai.com.au/wp-content/media/articles/226-ai-native-service-architecture.html
- Peter Higgins, NEC Contracts. “Clause 60 — compensation events.” — “A compensation event is the only way in which these can be changed. There are no other ways in which a Contractor can claim additional payment for carrying out the works or be allowed additional time”; “A failure by the Client, the Project Manager or Supervisor to take an action which the contract requires them to do.” https://www.neccontract.com/news/clause-60-%E2%80%93-compensation-events
- APMG / World Bank Group. “PPP Certification Guide — Variation Management.” — “In circumstances where a proposed variation involves no additional costs for either party, no formal variation procedure is required”; “the PPP agreement can require the private partner to provide a schedule of rates for a range of likely small works at the beginning of each year”; “verify that a variation request is actually a change and not covered under the existing agreement and pricing structures.” https://ppp-certification.com/ppp-certification-guide/7-variation-management
- Sophie Danby, ITSM.tools. “Change Enablement in ITIL 4: Definition, Practice & Best Approaches.” — “A standard change is a low-risk, repeatable, and pre-authorized change that follows a documented procedure. Standard changes often require little or no additional approval and are frequently automated.” https://itsm.tools/change-enablement
- Scott Farrell, LeverageAI. “AI-Native Service Architecture” (inside the square: latitude, proof and disposition). — “type it, route it, resolve it, and require it to leave something behind”; the four deposits of an escalation; “A failed acceptance is not a dispute — it is a defined state with a defined remedy, and the remedy is inside the band.” https://leverageai.com.au/wp-content/media/articles/226-ai-native-service-architecture.html
- Scott Farrell, LeverageAI. “Waterfall Per Increment.” — “Regeneration over patching — because fresh generation beats accumulated patches”; “Invest in spec quality… Bottleneck: specification clarity.” https://leverageai.com.au/wp-content/media/articles/44-waterfall-per-increment.html
- Scott Farrell, LeverageAI. “Buy Certainty First” (typed uncertainty). — “Unknowns become typed deliverables instead of unbounded consulting labour”; “When the same exception class appears repeatedly, promote it into a first-class state or a band driver rather than leaving it as endless Flex Reserve folklore.” https://leverageai.com.au/wp-content/media/articles/204-buy-certainty-first.html
- Scott Farrell, LeverageAI. “AI-Native Service Architecture” (where the architecture must shrink or decline). — “Better analysis lowers the probability of an error. It does not lower the cost of the one you make”; forcing out-of-band work into a fixed price “destroys the band’s meaning… teaches your own sales system that drivers are negotiable… converts a product back into a bespoke project”; “shrinking is the harder skill and the one nobody teaches.” https://leverageai.com.au/wp-content/media/articles/226-ai-native-service-architecture.html
- GC AI. “Limitation of Liability Clause: Caps, Carve-Outs, and Examples.” — “A generous-looking cap means little if that waiver strips out the losses you would actually claim”; “Tie carve-outs to a closed list rather than open-ended categories.” https://gc.ai/clauses/limitation-of-liability
- Scott Farrell, LeverageAI. “Buy Certainty First” (the pricing envelope). — “Exhaustion does not produce silent unpaid work. It produces a commercial conversation with a census delta and a recommendation: uplift band, extend reserve, narrow boundary, or stop”; “The envelope made surprise speak product language.” https://leverageai.com.au/wp-content/media/articles/204-buy-certainty-first.html
- Scott Farrell, LeverageAI. “AI-Native Successor Offer” (Gate 6, unit economics). — The contribution shape including “liability and commitment risk” and “exceptions and failure remediation”; “True contribution collapses when disposition is costed and exception tails appear.” https://leverageai.com.au/wp-content/media/articles/213-ai-native-successor-offer.html
- Troy Edwards and Peter Tolson, A&O Shearman. “Cost reimbursable vs. lump sum turnkey construction contracts: the many routes to bankability” (11 August 2025). — “In truth, there is no such thing as an absolute fixed price contract”; “Contractors will naturally be incentivized to seek opportunities to reopen the fixed price, particularly where the contractor’s cost contingencies prove to be insufficient”; “LSTK contracts may still be preferable for well-defined, low-risk projects where scope and owner requirements are clear from the outset.” https://www.aoshearman.com/en/insights/cost-reimbursable-vs-lump-sum-turnkey-construction-contracts-the-many-routes-to-bankability
- Congressional Research Service. “Parametric Insurance for Natural Disasters: Frequently Asked Questions,” IN12670, updated 19 May 2026. — “A contract for parametric insurance typically specifies (1) the payment amount; (2) the trigger (a pre-determined parameter based on observable data); and (3) an impartial third party to verify that the trigger was met”; “Compensation from parametric policies is not linked to actual losses… This is known as basis risk”; the New Orleans School District policy that “did not meet the 100 mph trigger.” https://www.everycrsreport.com/reports/IN12670.html
- Scott Farrell, LeverageAI. “AI-Native Successor Offer.” — “Fixed price is a strong signal, not a requirement. It is valuable because it demonstrates that complexity has become sufficiently legible to configure and bound.” https://leverageai.com.au/wp-content/media/articles/213-ai-native-successor-offer.html
