Leverage AI

AI-Native Services

AI-Native Service Architecture

📖 This article has an expanded ebook edition — read the full ebook.

The square, the barbell, the flywheel and the membrane — how to sell a fixed commitment around a generative machine, and make each engagement pay for the next.

Freeze what the customer buys. Leave the method generative. Prove the promise at hard edges. Compound between engagements — without pooling client truth.

TL;DR

Every service you have ever sold had a shape. Nobody drew it.

Ask someone to sketch the shape of a consulting engagement and you will get something with many edges and many curves. It bulges where a source system turned out to be undocumented. It stretches sideways where a stakeholder changed their mind. It grows a lump where somebody had to investigate an edge case for three days. It is an amorphous shape — many edges, many curves — and it is difficult to explain, which is precisely why so much of a professional-services firm's payroll is spent explaining it.

        ______
     __/      \____
 ___/              \__
/                     \____
       project

Time-and-materials contracting is not a billing convention. It is a decision about who owns that shape. It says, roughly: we don't know exactly what reality will require, so you buy access to our people while we discover it. A messy source means more days. A change in requirements means more days. A misunderstanding means more days. Behind those days sit project managers, resource plans, burn rates, timesheets, change requests, and arguments about whether a newly discovered piece of reality was ever in scope.

The customer effectively owns a significant portion of the supplier's production variance.

Everyone has been paying for that arrangement, and there is data on how well it works. PMI's 2018 Pulse of the Profession found that 52 per cent of projects completed in the previous twelve months experienced scope creep or uncontrolled changes to scope — up from 43 per cent five years earlier1. The trend went the wrong way during exactly the period the industry was getting better at project management.

The obvious response — quote a fixed price and stop arguing — has its own long, documented casualty list. Mature contracting practice is blunt about it. In lump-sum turnkey construction, the contractor takes the cost and performance risk, "although in practice, owners may still face additional costs through change orders and significant disputes can arise… In truth, there is no such thing as an absolute fixed price contract."2 Contractors respond by pricing in "significant contingencies", and where those contingencies prove insufficient they are "naturally incentivized to seek opportunities to reopen the fixed price."2

So the industry has two answers and both are bad. Bravado: quote fixed and eat the variance. Padding: quote fixed at triple and price yourself out of the deal while still carrying tail risk. Neither is an architecture. Both are the same shapeless project with a different number on the front.

Why no contract can enumerate the middle

There is a reason this failure is stubborn, and it predates AI by forty years. The incomplete-contracts literature — Grossman, Hart and Moore, whose work earned Hart and Bengt Holmström the 2016 Nobel — starts from the observation that "contracts cannot specify what is to be done in every possible contingency. At the time of contracting, future contingencies may not even be describable."3

Hart frames the resulting trade-off precisely: "The benefit of a rigid contract is that it fixes expectations, avoiding arguments. But it may not perform well when there is uncertainty. A flexible contract can adjust to the state of nature, but there is also room for arguments."4

Read that again, because the whole architecture is a refusal of the choice it offers. Rigid or flexible is only a dilemma if the contract is one undifferentiated thing. It stops being a dilemma the moment you notice that expectations and uncertainty do not live in the same place. Expectations live at the boundary — what is promised, what is delivered, what counts as done. Uncertainty lives in the middle — how many documents, how messy, how many attempts, which approach. Put rigidity where expectations live and flexibility where uncertainty lives, and you have a topology instead of a compromise.

That was never economically available before, because the middle was made of expensive human hours. Absorbing an extra thirty hours of thinking work meant absorbing thirty billable hours. It is available now because the price of cognition collapsed: querying a model at GPT-3.5-level performance on MMLU fell from $20.00 per million tokens in November 2022 to $0.07 by October 2024 — a more than 280-fold reduction in roughly eighteen months5. That is the whole premise, and it earns exactly one paragraph here.

Key insight

AI does not make the cost curve flat. It makes it flatter, and it changes which variable drives it. The variance that historically killed fixed pricing — how many documents, how messy, how many alternatives to test — now lands mostly on cheap parallel machine work. What remains is the number of consequential judgments a human must own. Meter that, and the rest can be absorbed.

The architecture: four structures, one system

Square around the engagement. Barbell inside the square. Flywheel between the squares. Membrane around the learning.

Those four are not four metaphors for the same thing, and they are not four best practices you can adopt in whatever order suits your quarter. They are four views of one service at different scales, and each one exists because a specific thing collapses without it.

StructureWhat it doesCollapse mode if removed
Square
the commercial perimeter
Makes commitment possible: a bounded promise, authoritative inputs, exclusions, a commercial band, a time boundary, valid acceptance statesTime-and-materials ambiguity. Every irregularity becomes a negotiation; the buyer finances the wobble again
Barbell
the effort distribution
Makes interior latitude safe: heavy at frozen intent, heavy at consequential acceptance, deliberately light betweenSupervised automation. Humans inspect every unit, and you have paid for a machine while keeping the labour cost
Flywheel
between engagements
Makes the economics improve: exceptions, evals, rules and delivery machinery from engagement N raise the starting substrate of N+1Repeated bespoke heroics. Every engagement starts cold; the same experts carry it; the curve is flat
Membrane
around the learning
Makes the improvement legitimate: only de-identified, abstracted, rights-cleared, human-approved patterns cross outClient-data leakage. "We learn from every engagement" becomes "we pool everybody's data" — and one clause ends the practice

Four different failures means four different structures. This is the deletion test, and it is the honest way to check whether a composition claim is real: take each piece out on paper and say what specifically breaks. If a piece can be removed without breaking the other three, it was decoration.

The perimeter is productised. The interior is generative.

What the square actually contains

The outside is deliberately rigid: promise · price or unit · time boundary · authoritative inputs · terminal states · exclusions · acceptance tests · authority. The inside is deliberately fluid: search · decomposition · generation · retries · parallelisation · tool choice · code generation · research · testing · regeneration · internal replanning. Then the exit boundary becomes rigid again: evidence · acceptance · human disposition · delivered state.

Tight promise. Loose production. Hard acceptance. It is the commercial sibling of the delivery rule we already run on — tight intent, loose method, hard verification — and the family resemblance is not an accident. It is the same insight applied one layer up.

Fixed price is a consequence, not the essence

This is the point at which most readers hear "fixed-price consulting" and either get excited or get frightened. Both reactions miss it. The invariant is not the fee shape; it is a stable unit of commitment. That could be a fixed-price decision product, a per-project activation, a verified decision, an annual continuity subscription, an assessed estate, an enrolled fleet, a protected period, a guaranteed response commitment or a capacity tier. What it increasingly must not be is "however many hours our internal process happens to consume."

Fixed price matters because it is unusually strong evidence that a supplier has made its own complexity legible enough to take responsibility for it. That makes it a signal, not a religion.

Change requests become boundary mutations

Once the perimeter is real, the most useful thing it does is reclassify surprise. In conventional delivery, almost any newly discovered complexity can become a change request. Under this topology, ordinary production complexity must not.

If I promise to analyse your estate and one part turns out to need twenty agent-hours rather than two, that is my problem; the square absorbs it. If an integration needs three regeneration attempts rather than one, my problem. If the first architecture fails its test and the system generates another, my problem. Those are interior variations.

A commercial event occurs only when something crosses the perimeter: the buyer changes the promised outcome; the input estate crosses the agreed volume band; a new source class appears; the required authority changes; a physical constraint changes; a regulatory obligation lands; the deadline moves. Those are boundary mutations, and they reopen the commercial conversation with evidence rather than with a grievance. (The detailed contract test for boundary changes — typed reserves, exhaustion rules, the mutation matrix — is a separate piece of work; the point here is only that the two categories are different in kind.)

Topology is not cadence

One distinction has to be made explicitly or the whole thing gets misfiled. "Waterfall per increment" — specify, generate, independently verify inside each slice; deploy, observe, learn between slices — describes a production cadence. The square describes a commercial topology. They coexist happily. They are not the same layer, and conflating them wrecks both.

A firm that hears "fixed outside" and freezes its production cadence into a twelve-month requirements phase has misread the architecture completely. A firm that hears "short increments" and concludes that its commercial commitment must therefore be open-ended has misread it in the other direction. The interior can iterate as fast as it likes. What it cannot do is renegotiate the promise every time it learns something.

The part of Agile that survives is its deepest epistemic insight: users still don't completely know what they want until they see something. The part that depreciates is the labour-management layer — sprint capacity, story points, velocity, estimation, hand-offs — machinery whose economic premise was that human implementation capacity is the scarce resource you must ration. When generation is cheap, a large wrong system appears as fast as a large right one, and the binding work moves upstream to framing and downstream to verification. That is a cadence consequence. The commercial consequence is different and sits above it.

The barbell: put the expensive attention at the ends

The square explains what is fixed. The barbell explains why that is safe.

Heavy thought before implementation. Cheap implementation in the middle. Heavy verification after. On the left plate: intent, promises, non-goals, alternatives, assumptions and failure cases, acceptance tests, recorded rejections. On the right plate: independent lenses, deterministic checks, adversarial review, evidence coverage, intent regression. In between: latitude.

The customer buys precision of intent and precision of outcome. They do not buy precision of internal procedure.

Traditional statements of work do almost the reverse. They protect the supplier by describing activities — conduct five workshops, interview twelve stakeholders, run analysis, hold weekly status meetings, prepare report. Those are descriptions of the middle: the exact part the buyer should not be buying. AI-native contracting can be comparatively thin there, and thick at the ends:

What replaces the activity list

Here is the state we will establish. Here are the evidence boundaries. Here are the valid uncertainty states. Here is what would invalidate the engagement. Here is what constitutes completion. How we exhaustively get there is largely our problem.

Two objections arrive immediately, and both deserve a straight answer.

"Our risk committee will never accept latitude in the middle." The barbell is more governed than supervised automation, not less. Supervised automation puts a thin, tired human layer across every unit of output, which is expensive and — as the evidence keeps showing — not especially reliable. The barbell concentrates authority where consequences concentrate. Where a use is genuinely high-risk, this stops being a design preference: EU AI Act Article 14 requires high-risk systems to be designed so that natural persons can effectively oversee them, interpret their output, remain aware of automation bias, and intervene or halt operation — with separate verification by at least two natural persons for certain uses6. That is a barbell written into statute.

"But AI is fast, so acceptance can be light." No. In METR's randomised controlled trial, experienced open-source developers working on their own repositories took 19 per cent longer with early-2025 AI tools than without — while forecasting a 24 per cent speed-up beforehand and believing they had been sped up 20 per cent afterwards7. It is one setting, 246 tasks, and a snapshot of a fast-moving capability; METR say so themselves. But the gap between felt and measured is the entire reason the right-hand plate exists. If your acceptance evidence is a demo and a good feeling, you have built a machine for generating confident wrongness at scale.

The outcome is not the test

One small distinction does a lot of work here. The outcome is what was sold. The test is the instrument by which you can legitimately say the outcome occurred. If you sell a verified current-state architecture, the architecture is the outcome; the evidence coverage, the reconciliation checks and the human dispositions are the tests that prove it. Collapse the two and "outcome-based" quietly reverts to vague consulting language, because nothing can fail.

The sequence, then, is: promise → falsifier → latitude → evidence → acceptance. And the acceptance test has to be able to fail. If completion resolves to stakeholder satisfaction, the discipline of the whole square evaporates at the last mile.

The flywheel: engagement two is the only proof

Everything so far describes one engagement. The commercial argument only becomes interesting across two.

Organisations have always tried to learn from delivery. Communities of practice, post-implementation reviews, centres of excellence, knowledge bases, lunch-and-learns. It is not true that organisational learning didn't exist. It is true that the joins were economically broken. Somebody had to notice a lesson was reusable, articulate it, abstract it out of client-specific language, document it, classify it, distribute it, persuade a large group to read it, have one of them remember it at exactly the right future moment, and apply it correctly. Every link in that chain was human labour, which is why so much knowledge management became archaeology.

AI changes essentially every link — and the last one structurally. Nobody has to remember to go and read it. The next run loads the substrate.

So the wiki, the eval set, the acceptance harness, the exception taxonomy and the pricing rules stop being the company library and become part of the cost curve of the product. That is a different claim from "we captured some learnings", and it comes with a brutal condition attached:

The condition

A first engagement at break-even or a loss is rational if the loss is purchasing reusable capability. It is irrational if it is purchasing heroics.

Suppose engagement one takes two senior people eight weeks and barely breaks even. That is an excellent investment if it leaves behind a working delivery vessel, a set of eval cases, newly typed exception classes, reusable deterministic tools, sharper pricing-band logic, reusable acceptance tests and a better qualification gate. It is a low-margin consulting job if the two seniors simply worked incredibly hard, solved everything by hand, made the client happy and told everyone what they learned over drinks.

Which means engagement two is the only proof. Transfer means ordinary capable staff lead materially more of the second engagement because the first one improved shared infrastructure — not because the same heroes worked late again. If escalation returns at the same density, if exception classes are still tribal, if pricing rules still bend whenever a principal is in the room, you ran a successful project and did not install a practice. And one successful engagement never evidences compounding. It evidences delivery. Say so plainly, and name what the second engagement would have to show: lower recurring novelty, falling scarce-expert density, typed rather than tribal exceptions, and reusable artefacts an ordinary consultant actually loads.

A normal consultancy monetises accumulated expertise repeatedly as labour. An AI-native consultancy converts each paid engagement into infrastructure that reduces the amount of expertise that must be re-spent on the next one.

There is a cultural consequence worth naming. Traditional professional services rewards knowledge concentration: if you are the only person who knows the trick, your utilisation, indispensability and internal status all rise. Under this architecture the highest-status act changes from I solved the hard case to I solved the hard case once and made sure nobody needs me for that class again. That is a different profession, and pretending it will happen by exhortation is how the last twenty years of knowledge management went.

The membrane: three territories and one gate

"We learn from every engagement" is one clause away from "we pool everybody's data", and buyers know it. So the boundary has to be architectural, not reassuring.

There are three territories, and they are not the same:

  1. Client truth. Their data, systems, people, decisions, evidence, context. Stays client-bound according to the engagement and the contract.
  2. Engagement learning. What happened in this particular delivery — exceptions, failures, decisions, traces, project evidence. Much of this should also remain inside the engagement boundary.
  3. Promotable capability. An abstracted pattern that no longer depends on the client's confidential state: when this condition occurs, use this test; this source class needs this adapter; this pricing band underestimates disposition load under these conditions; do not make this architectural promise without this evidence.

Only the third category is even a candidate for wider use — and not automatically. The gate is de-identification, abstraction, transferability, contractual clearance and human approval, in that order, with a named owner. Every candidate ends in exactly one disposition: local-only, configurable, internal primitive, supported platform, or reject. "Maybe later" is not a disposition; it is a deferred decision that needs a revisit trigger. And rejections are preserved, not deleted — they teach the boundary.

This is emphatically not model training on client exhaust. The model-provider bargain is: give us your interaction data and we'll subsidise inference because it improves our future model. The service firm's version should be: we may accept lower early margin because the engagement improves our external organisational learning substrate — wiki, evals, acceptance harness, pricing configuration, delivery code, adapters, tests, pattern ledger, prompts, toolchain. You very often do not need to train anything at all. In fact you probably shouldn't: external learning is inspectable, reversible, attributable, boundary-aware, usable by tomorrow's engagement, and portable across model upgrades. It is soft-weight learning at organisational scale.

This is also not a novel ethical position; it is where professional conduct already sits. Practitioner guidance on ABA Formal Opinion 512 and equivalent state rules treats generative AI as non-lawyer assistance under supervision obligations, and advises that client confidential information only be entered into platforms that contractually commit to zero data retention and zero training on customer inputs8. Standard vendor terms now include deletion of customer data, prompts, outputs and derived training data, plus certification of deletion and audit rights9. The membrane is what that looks like when you build it into the service instead of bolting it onto the MSA.

Done twice

Put the four structures together and the definition of "done" gets tougher than the industry's current one.

The end-to-end Definition of Done

The customer received the promised, falsifiably accepted outcome and the provider disposed the resulting learning: local-only, reusable configuration, internal primitive, promoted doctrine, or explicit rejection.

An engagement is not operationally closed until it has answered two questions. Did we keep the customer's promise? And: what, if anything, should the organisation never have to learn again? The first question is what the client bought. The second is what the firm bought. A service architecture that only closes the first one is a delivery method. Closing both is what makes it a business model.

Where the architecture must shrink the promise — or decline

If this reads as "AI makes fixed-price projects easy", it has been read wrong, and the correction is important enough to state as a rule.

AI cheaply absorbs cognitive variance. It does not absorb all variance. The square breaks when the residual uncertainty lives in:

The rule

Not: AI → fixed price. Instead: AI expands the territory in which complexity can be bounded, configured and priced as a product — and the edge of that territory is a design output, not an act of courage. Out-of-band work gets re-bounded or declined. A product that cannot say no is not a product.

The inversion: standardise the compiler, not the answer

Here is the part that should change what you do on Monday.

Historically, productising a service meant reducing variance in the customer's answer. Standardise the deliverable so you can predict the labour: templates, packages, bronze/silver/gold tiers that mostly amount to the same work with fewer knobs. Mass customisation — Stan Davis's term, developed by Joseph Pine — got partway out of this by standardising modules and letting combinations vary, producing goods and services "to meet individual customers' needs with near mass production efficiency"11. But the answer still came from a catalogue.

Old productisationAI-native productisation
Reduce variation in the customer's answerAllow variation in the customer's answer
Standardise the deliverableStandardise the machinery that absorbs variation
Repeatability lives in the outputRepeatability lives in the compiler, proof harness, exception classes and boundary logic
Specificity is an expensive exceptionSpecificity is the standard path
Margin comes from doing the same thing againMargin comes from the recurring parts ceasing to be novel

You standardised the path to understanding, not the conclusion. You productised the path, not the answer. And because the machinery improves every time it runs, the same square becomes more profitable and more reliable across successive engagements rather than merely repeatable.

Same square. Different thing inside every time.

The demand side is already moving towards this shape of promise, which is why the design question is urgent rather than theoretical. McKinsey's global managing partner describes the firm "migrating pretty quickly away from… a fee-for-service model" towards "much more of an outcomes-based model where we say, 'Look, let's identify a joint business case together and we will underwrite the outcomes of that business case'"12. In legal services the same gap is visible from the buyer's side: 71 per cent of clients would prefer to pay a flat fee for their entire case, while hourly billing remains the most common model, offered by 71 per cent of firms13. Buyers want a stable unit of commitment. Most suppliers still sell time. The architecture is how you close that gap without gambling.

What to do this week

Take one live offer and run the deletion test on paper. Four questions, half a page each:

  1. No square. If we removed the bounded promise and sold access to people instead, what changes? If the answer is "nothing much", you were already selling time and calling it an outcome.
  2. No barbell. If nobody froze intent up front and nobody owned a consequential acceptance at the end, who would be inspecting what? Count the review hours you would need. That number is what the barbell buys.
  3. No flywheel. Name three things engagement two will start with that engagement one did not have. If you cannot name three, the second engagement will cost what the first one cost.
  4. No membrane. Write the sentence you would say to the client's general counsel describing what leaves the engagement and in what form. If you cannot write it calmly, do not promote anything yet.

Then fill the perimeter fields for that offer — promise, authoritative inputs, exclusions, commercial band, time boundary, valid acceptance states — and mark the ones you cannot answer. The blanks are the design work. They are also, usually, exactly where last year's margin went.

Rigid at the promise. Fluid in the machine. Rigid at proof. And closed only when both questions are answered.

References

  1. Project Management Institute / Catherine Elton. "Scope Patrol: Scope Creep Is On The Rise As Stakeholder Expectations Increase." PM Network 32(7), 38–45. — "PMI's 2018 Pulse of the Profession® found that 52 percent of projects completed in the last 12 months experienced scope creep or uncontrolled changes to the project's scope—up from 43 percent five years ago." www.pmi.org/learning/library/scope-creep-rising-11308
  2. A&O Shearman. "Cost reimbursable vs. lump sum turnkey construction contracts: the many routes to bankability." — "In truth, there is no such thing as an absolute fixed price contract." / "Contractors will naturally be incentivized to seek opportunities to reopen the fixed price, particularly where the contractor's cost contingencies prove to be insufficient." www.aoshearman.com/en/insights/cost-reimbursable-vs-lump-sum-turnkey-construction-contracts-the-many-routes-to-bankability
  3. Grossman & Hart (1986); Hart & Moore (1990); Hart (1995), summarised in "Incomplete contracts." — "contracts cannot specify what is to be done in every possible contingency. At the time of contracting, future contingencies may not even be describable." en.wikipedia.org/wiki/Incomplete_contracts
  4. Lindau Nobel Laureate Meetings. "Oliver Hart: Incomplete contracts and the theory of the firm." — "The benefit of a rigid contract is that it fixes expectations, avoiding arguments. But it may not perform well when there is uncertainty. A flexible contract can adjust to the state of nature, but there is also room for arguments." www.lindau-nobel.org/oliver-hart-incomplete-contracts-and-the-theory-of-the-firm
  5. Stanford HAI. "Artificial Intelligence Index Report 2025", Chapter 1: Research and Development. — "The cost of querying an AI model that scores the equivalent of GPT-3.5 (64.8) on MMLU… dropped from $20.00 per million tokens in November 2022 to just $0.07 per million tokens by October 2024 (Gemini-1.5-Flash-8B)—a more than 280-fold reduction in approximately 18 months." hai.stanford.edu/ai-index/2025-ai-index-report/research-and-development
  6. European Union. Regulation (EU) 2024/1689 (EU AI Act), Article 14 — Human oversight. — High-risk AI systems must be designed so natural persons can oversee them, interpret output, remain aware of automation bias, decide not to use the system, and intervene or halt operation; certain uses require separate verification by at least two natural persons. artificialintelligenceact.eu/article/14
  7. METR. "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (July 2025). — "when developers use AI tools, they take 19% longer than without—AI makes them slower." Developers forecast a 24% speed-up and estimated a 20% speed-up afterwards; N = 246 tasks. metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
  8. GC AI. "AI Legal Ethics in 2026: 6 Cases, 4 Rules, 1 Policy Template" (practitioner analysis of ABA Formal Opinion 512 and state guidance). — "client confidential information should only be entered into AI platforms that contractually commit to zero data retention and zero training on customer inputs." gc.ai/blog/ai-legal-ethics
  9. Andrew S. Bosin LLC. "Technology Lawyer Explains AI Vendor Contracts for Startups (2026 Guide)." — Key provisions include "deletion of Customer Data within a defined period… deletion of derived training data (if possible), certification of deletion, return of customer datasets" plus audit rights. www.njbusiness-attorney.com/technology-lawyer-ai-vendor-contracts-startups
  10. JMD Ross Insurance Brokers. "Professional services contract clauses – Some key points." — "without a contractual limitation, liability is unlimited and could exceed the level of cover maintained under your PI policy." www.jmdross.com.au/wp-content/uploads/2018/02/Professional-services-contract-clauses.pdf
  11. Stan Davis, Future Perfect (1987); B. Joseph Pine II, Mass Customization: The New Frontier in Business Competition (Harvard Business School Press, 1993); definition per Tseng & Jiao (2001). — mass customisation as "producing goods and services to meet individual customers' needs with near mass production efficiency." en.wikipedia.org/wiki/Mass_customization
  12. Bob Sternfels (Global Managing Partner, McKinsey & Company), interviewed on HBR IdeaCast, "Where McKinsey—and Consulting—Go From Here", January 2026. — "we're migrating pretty quickly away from, let's call it pure advisory work… It's moving to much more of an outcomes-based model where we say, 'Look, let's identify a joint business case together and we will underwrite the outcomes of that business case.'" hbr.org/podcast/2026/01/where-mckinsey-and-consulting-go-from-here
  13. Clio. Legal Trends Report, presented in "Is Flat Fee Billing Becoming the Norm in Law?" — "71% of clients prefer to pay a flat fee for their entire case, and 51% want to pay flat fees for individual activities within their case. Still, hourly billing is the most common, offered by 71% of law firms." www.clio.com/guides/flat-fees-legal-trends