AI-Constituted Services
The Business That Can't Exist Without the Machine
"Find an AI use case" keeps producing trinkets — because it's a question about your existing workflow.
There is a third category of service: remove the AI, and the offer itself stops making sense.
What this book gives you
- ✓ A falsifiable test that sorts any AI offer into enabled, dependent, or constituted
- ✓ The hunting ground: economically suppressed services your market has never sold
- ✓ The standard architecture — Cognitive Workflow Recomposition — walked end to end on a working specimen, transplanted twice, and run honestly against the cases where it fails
Scott Farrell · LeverageAI · leverageai.com.au · August 2026
The Trinket Problem
The machines work. Organisations keep pointing working machines at the wrong target.
You have been in this workshop. The current processes go up on the wall — inspection cycles, monthly reporting, the service desk, the automation backlog — and someone asks the question that feels responsible and ends up small: "Where can AI help?" The room inventories what it already does. It nominates steps to speed up. A copilot for the analysts, a summariser for the meetings, a chatbot for the customers. Small, plausible, fundable.
Trinkets.
Eighteen months later comes the portfolio review. Traffic lights, mostly amber. Adoption dashboards standing in for earnings impact. Everyone privately concludes AI was overhyped; publicly, the firm is "on the journey". And the next strategy day opens with the same question that caused the problem, asked with more urgency.
The record says this is the normal outcome
The failure is not anecdotal — it is one of the best-measured phenomena in enterprise technology. MIT's NANDA initiative examined 300 public AI deployments alongside 150 leader interviews and found that about 95% of generative AI pilots deliver no measurable P&L impact. Its diagnosis was pointed: not model quality, but "the 'learning gap' for both tools and organizations" — flawed enterprise integration1. S&P Global's survey work found the share of companies abandoning most of their AI initiatives jumped from 17% to 42% in a single year, with the average organisation scrapping 46% of proofs-of-concept before they reached production2. RAND puts overall AI project failure above 80% — twice the rate of IT projects that don't involve AI — and names the leading root cause: stakeholders "misunderstand—or miscommunicate—what problem needs to be solved using AI"3.
The failure record, measured three ways
of generative AI pilots deliver no measurable P&L impact (MIT NANDA, via Fortune)
of companies now abandon most of their AI initiatives — up from 17% a year earlier (S&P Global)
AI project failure rate — double that of non-AI IT projects (RAND)
Read those three findings together and notice what they are not saying. They are not saying the models are weak. Every one of them locates the failure at the boundary between the machine and the organisation — the integration, the problem selection, the workflow. The machines work. Organisations keep pointing working machines at the wrong target.
Horse optimisation
I have a blunter name for the mechanism, from watching this play out across boardrooms and consultancies: horse optimisation. Ask someone how they want to use AI and they can only see horse optimisation — strapping intelligence to the inherited process and asking it to trot faster. Everyone's trying to automate the existing workflow. In every case I've seen where AI does useful work, it's doing something you're not doing. Changing the pattern.
The question guarantees the failure, because "where can AI help?" contains the old workflow inside it. The person answering inventories their current applications, their current teams, their current processes, their current backlog — and then nominates steps to accelerate. I've had this conversation with heads of IT who genuinely know a great deal about vendors, automation, data and security, and who have still filed AI neatly into the automation drawer. Their map is technically sophisticated and strategically incomplete at the same time. Nothing on it is wrong. What matters is what cannot appear on it: work that nobody does today.
Key Insight
"Where can AI help?" contains the old workflow inside it. The answers can only ever be faster versions of what you already do.
There is a prior question, and I've made the case for it at length elsewhere: before you grease a single cog, ask whether the machine should exist at all — whether this process would exist in this form if it were designed today. This book runs that question at a different altitude: not a process inside your business, but the services on your price list — and, more importantly, the ones that aren't.
The supplier side of the same failure
The firms selling AI to those workshops are running the same error at higher speed. I've watched consultancies respond to soft advisory demand by hunting for — their instinct, my words — tiny little trinkets of AI to sell to their clients. And they're bad at it, because there's no framework, no structure behind the hunt. Can we bolt AI onto last year's data project? Can we sell an AI governance deck? In no way, shape or form are they asking the question that matters: what's a new product we can build, with AI embedded, that we can't make now without AI?
That sentence is the hinge of this book. Firms keep asking how to sell AI through their existing services. The opportunity — the one the rest of these chapters build machinery for — is to rebuild the service through AI. Those are completely different ambitions. The first produces trinkets with better margins. The second produces offers that have no pre-AI ancestor at all.
"AI does not improve the service. AI permits the service to exist."
There is a class of offer for which that sentence is the literal, testable description — not marketing elevation, but an existence condition. Remove the AI from a trinket and you lose a convenience. Remove the AI from one of these services and the offer itself stops making sense: the price can't be promised, the coverage can't be delivered, the evidence can't be produced. Nothing is left to sell.
Where this book goes
Part I gives that class a name — the AI-constituted service — and a falsifiable test that sorts any offer, yours or a competitor's, into one of three categories. It then maps the hunting ground: the services missing from every catalogue because they were never rational to sell. Part II is the build: a standard architecture — Cognitive Workflow Recomposition — that manufactures such services, from line-item decomposition through typed uncertainty to a fixed price that is engineered rather than brave. Part III walks one working proof end to end. Part IV transplants the architecture into two other industries, argues why the resulting moat survives copying, and — because a test that can't fail is worthless — runs the whole framework against the cases where it breaks. Part V is the field guide: three moves you can start on Monday.
The workshop wall was never going to surface any of this. The interesting services aren't on it.
Three Kinds of AI Service, One Test
"AI-native" has stopped meaning anything. Here is a classification you can falsify.
Any vendor can bolt a model onto an existing product and claim the label "AI-native". Plenty do. The label costs nothing, asserts nothing checkable, and has accordingly inflated to worthlessness — which is a real problem for anyone trying to decide what to build, buy, or defend, because the difference between a trinket and a new business is precisely the thing the label was supposed to mark.
What's needed is a classification with teeth: one that a sceptic can run against your offer and prove you wrong.
I hit the need for it mid-sentence, talking through my own product. It's AI-enabled — but it's more than AI-enabled. It's AI allowed to exist. What's the word there? Without AI it wouldn't even be possible. The word that didn't exist yet is the one this book is named for: AI-constituted. Not "native", which is branding. Constituted — because the claim is about existence conditions. The AI doesn't decorate the offer or accelerate it. The AI is what the offer is made of.
The taxonomy
| Type | Remove the AI and… |
|---|---|
| AI-enabled | The service remains; it becomes slower or more expensive |
| AI-dependent | The promised price, speed or scale becomes uneconomic |
| AI-constituted | The entire offer and operating model cease to make sense |
Each row is a different commercial animal. An AI-enabled service is a productivity story: the consultants draft faster, the analysts summarise quicker, the service survives any model outage with grumbling. Most of what currently wears the "AI-native" label lives here. An AI-dependent service is a margin story: the offer's price point, turnaround or scale was set assuming the machine — remove it and the economics collapse, though the service itself could limp on at the old price. An AI-constituted service is an existence story: there is no version of the offer without the machine, because the promise itself — the fixed price against unmeasured complexity, the full coverage instead of sampling, the evidence chain behind every finding — was never deliverable by humans at any price a client would pay.
The strategic implications differ just as sharply: enabled buys parity (everyone gets the same copilots the same quarter), dependent buys cost position (real but fragile), constituted buys a market that didn't previously exist.
The counterfactual test
The classification is operational because the test is a thought experiment anyone can run today, against any offer, in two questions:
Remove the AI. Does the same offer still exist, only more slowly? Then it is AI-enabled.
Remove the AI. Do the fixed price, coverage, specificity and evidence promise collapse? Then it is AI-constituted.
The middle case falls out between them: if the offer survives but the price breaks, you're looking at dependence. And note what makes the test worth having — it can embarrass you. Run it honestly and it will classify most of your "AI transformation" portfolio as enabled, some flagship as dependent, and possibly nothing as constituted. A classification that can't produce an unwelcome answer can't produce a useful one either. Chapter 14 runs the test where it fails entirely, on purpose.
Running the test on a famous case
Klarna's AI assistant is the best-documented service-AI deployment in public. The company's own release reports the assistant handling two-thirds of customer-service chats in its first month — "doing the equivalent work of 700 full-time agents", resolution times down from 11 minutes to under two, and "estimated to drive a $40 million USD in profit improvement to Klarna in 2024"4. A genuinely impressive deployment, and real money.
Now apply the test. Remove the AI: does customer service still exist? Obviously — it costs more and answers slower, but the offer to the customer is unchanged. That is substitution: AI-dependent at best, arguably just enabled at scale. Which is a classification, not a criticism.
And here is why the classification matters: it predicts behaviour. By May 2025 Klarna was recruiting humans back into customer service, with its CEO conceding that "cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality"5. A cost-substitution offer can always snap back, because the pre-AI version of the service is sitting right there, waiting to be re-staffed.
Contrast the offer this book walks end to end in Chapter 10: a fixed-price, full-coverage, evidence-backed data readiness review. No consultancy has ever had that on its price list — not for lack of demand, but because promising it under human-labour economics was commercially insane. There is nothing to snap back to. Remove the AI and you don't get a slower version; you get no version.
Surface embedding versus constitutive embedding
The taxonomy also dissolves a confusion about what "embedding AI" means. There are two entirely different acts wearing the same phrase.
Two kinds of embedded AI
Surface embedding
- • A copilot, chatbot, summariser or widget inside an existing product or engagement
- • Workflow, commercial model and customer obligation substantially unchanged
- • AI is visible, marketable — and easy to copy
Constitutive embedding
- • AI inside the economics, cognition and production machinery of the offer
- • A new commercial product becomes possible at all
- • The customer may barely interact with the AI — and doesn't need to
What does the client of a constituted service actually experience? Not "AI". They experience a fixed price where pricing was previously dangerous. Comprehensive analysis where consultants previously sampled. Specificity where the firm previously standardised. Evidence rather than a confident recommendation. A shorter, clearer path to a defensible decision. The machine can be invisible to the buyer and indispensable to the producer at the same time.
"The strongest AI product is often one whose customer does not need to care that it uses AI."
That inverts the instinct most firms bring to AI positioning — the badge on the brochure, the "powered by AI" slide. If your customer has to care that it's AI, you're probably selling surface. When the AI has sunk all the way into the offer's existence conditions, the pitch stops mentioning it: the boardroom-safe version is simply, "This is not a traditional consulting service accelerated by AI. It is a service that could not previously be delivered with reliable scope, evidence and fixed-price economics."
Bottom Line
AI is not a feature of a constituted offer. It is the economic substrate that allows the offer to exist in productised form.
So the category exists, the test is runnable, and one public case has already demonstrated both halves of the dependent class — the win and the snap-back. The obvious next question: if constituted services are the prize, where do you find one? Not in your process map. The next chapter goes hunting where the offers aren't.
The Hunting Ground: Economically Suppressed Services
Every catalogue has a shadow — the services nobody sells because nobody could.
Every industry's service catalogue has a shadow: the offers that aren't in it. Not hidden inefficiencies inside current processes — whole services that clients would buy tomorrow and that no firm can rationally promise. You cannot see them by studying what you do today, because they are defined by what nobody does. And they are where AI-constituted services come from.
Call them economically suppressed services: valuable work absent from every catalogue because the cognition, coordination or uncertainty cost made it commercially irrational to offer.
The common signature across all of them: the promise scales with the client's mess. More spreadsheets, more systems, more exceptions — more senior hours. Human cognition priced by the hour cannot absorb that scaling, so the promise was never made. The services aren't failing in the market. They were never born.
Why cheap cognition changes the inventory
I've written a full treatment of the value ladder this belongs to, so one paragraph here and no more. The first version of AI value automates existing tasks — highest failure rate, as Chapter 1's numbers showed. The second applies ten to a hundred times more analysis to problems you already have: checking everything instead of sampling. The third — the frontier — "isn't about accelerating current work. It's about making entirely new categories of work rational to attempt for the first time." An economically suppressed service is a Version 3 opportunity wearing a price tag — the service-product face of the frontier.
Which means the discovery question changes. Not "where can AI help?" — Chapter 1 buried that one — but:
"What important thinking do you currently not perform because it would require too many experts, too much coordination or too much time?"
I use that question inside organisations as the opening move of a cognition scarcity audit — hunting the analysis that never happens because human labour made it uneconomic. This book runs the same question outward: not "what thinking does our organisation skip?" but "what services does our market not sell?" The audit finds internal absences; the hunting ground is the same absence, priced.
The statement of work: a suppression case study from inside consulting
Professional services carries a particularly expensive specimen of suppressed economics, and I have the scar tissue to describe it from the inside. When I ran my own consulting company, writing statements of work was so bad I had to write them all myself. I couldn't hand them to sales staff or senior technical staff — it would take too long and they'd be wrong. It limits how much work you can write up, and how much you can be bothered writing up.
The structural reading of that pain: the senior person writing an SOW is manually joining the client request, the account history, the firm's actual capability, prior architectures, available people, commercial risk and their own scar tissue — context the firm's systems do not reliably hold anywhere. So a senior person reconstructs it, from emails, meetings, colleagues, old proposals and instinct, every single time.
"The senior SOW author is the join algorithm."
That is why one partner quotes $200,000 and another quotes $600,000 for the same request. Both are locally rational — each assembled a different world, and there is no explicit substrate against which either answer can be tested. Economically, a large share of what firms book as "business development" is unpriced cognition spent reconstructing the same commercial object from scratch, over and over, for deals that mostly don't close.
And the unpriced join doesn't stay in pre-sales; it detonates downstream. Project Management Institute data puts only 62% of projects within their original budget, with roughly a third experiencing scope creep6. The risk in fixed-price work is usually not building the solution. It is discovering too late that the quoted scope was based on an incomplete understanding. Sane firms respond by retreating — time-and-materials, vague discovery phases, caveated estimates — and the suppressed service stays suppressed. Nobody sells "a measured, evidence-backed, fixed-price scope" because nobody can afford to produce one.
Why now: the market is repricing the old equilibrium
If the suppression were stable, this book could wait. It isn't. The consulting market is bifurcating, in public, in the accounts. KPMG Australia reported consulting revenues down 18% for FY25, citing "a significant reduction in the government use of consultants" and a rebalancing "with increased focus on technology transformation and AI"7. In the same market cycle, BCG grew to US$14.4 billion, with "AI- and tech-focused services now represent[ing] over 40% of BCG's total revenue… driven by 25% year-over-year growth in AI services"8.
Bifurcation, not collapse
KPMG Australia consulting revenue, FY25 — traditional advisory demand soft
BCG AI services growth, year over year — applied AI demand firm; AI/tech now >40% of revenue
The demand side is moving the same direction. McKinsey now takes about a quarter of its global fees through performance-based arrangements, and its global technology leader Kate Smaje puts it flatly: "This is a moment where many of the fundamentals of the professional services model are coming under challenge"9. Industry commentary goes further — one widely-read piece argues the Big Four's hourly billing model is "fundamentally incompatible with how AI transforms productivity"10 — attribute that as commentary, not research, but notice it's the buyers asking why the invoices haven't changed.
The one-line diagnosis I keep returning to: advice got cheap; verification did not. Polished recommendations that once took a pyramid of analysts now take an afternoon; what buyers will still fund is the path that survives architecture review, cyber, finance and production. A constituted service manufactures that path by construction — the evidence chain isn't an add-on, it's the delivery mechanism.
Venture capital has spotted the same repricing from the other side. Foundation Capital calls it "service-as-software" and sizes the shift at US$4.6 trillion — "responsibility for achieving the desired outcome sits with the company selling the service"11; a16z frames it as capital turning into labour — "out comes code that takes the role of labor"12. Both theses see software firms absorbing service markets. What neither supplies is a falsifiable classification of the offers, or a build recipe for service firms constituting new services from suppressed demand. Those are this book's Chapter 2, and everything from Chapter 4 on.
Key Insight
The hunting ground is external and measurable: services suppressed by breadth, coverage, frequency or uncertainty costs — in a market that is actively repricing unverifiable advice.
One more warning for the broad incumbent reading this comfortably. The historical advantage of nebulousness — enter through reporting, expand into platforms, then operating models, then transformation — depended on production requiring armies and buyers having no alternative. When a smaller competitor can pre-work the estate, assemble specialised cognition without a bench, and ship a fitted offer, breadth reverses polarity: hard to explain, hard to buy, hard to scope, hard to price, hard to know when it's finished. The nebulousness stops being flexibility and starts becoming friction rent.
So: the category (Chapter 2), the hunting ground (this chapter). Now the build. How do you actually manufacture a service that cannot exist without the machine? That is Part II, and it starts with the move I had to invent a name for.
The Build Move: Cognitive Workflow Recomposition
You cannot automate your way to a suppressed service. There is no workflow to automate.
Here is the awkward fact about every service in Chapter 3's inventory: nobody delivers it, so there is no process to improve. You cannot automate your way to a suppressed service. The build move has to manufacture a workflow that never existed.
And where a related workflow does exist — the manual discovery phase, the sampled audit, the handwritten scope — it deserves no deference at all.
"The old workflow was often not a purposeful design. It was the residue of limitations."
Run the list of limitations that actually shaped the workflows you've inherited. Software couldn't understand messy input, so humans re-keyed reality into forms. Systems couldn't join language-based evidence, so meaning travelled by meeting. Semantic matching was too expensive, so nobody reconciled anything exhaustively. Specialists couldn't inspect everything, so everything got sampled. Institutional memory lived in people, so context was rebuilt from scratch per engagement. Humans carried ambiguity between applications, because nothing else could carry it. Every one of those constraints has now lifted. A workflow shaped by them is a fossil of the constraints, and reproducing it — even efficiently — reproduces the fossil.
The definition
The move I use instead, and the only framework this book mints from scratch:
Cognitive Workflow Recomposition
Rebuild a business outcome as evidence-bearing units of cognition and authority; assign each unit to deterministic software, AI or a human; then recompose them through a governed application and a learning substrate.
Every phrase is doing work. Business outcome — the input is what should exist, not the process you have; you decompose the destination, not the journey. Evidence-bearing units — every unit carries its own provenance; a judgment that can't show its evidence isn't a unit, it's a vibe. Cognition and authority — two different things get placed: who thinks, and who decides; conflating them is how AI systems end up either impotent or ungovernable. Governed application — the recomposition is real only when software enforces it; a methodology deck enforces nothing. Learning substrate — each run improves the next, which is where the moat will come from (Chapter 13). The workflow that results may never have existed before. As I said when I first caught myself doing it: it's a workflow recomposition. The workflow wasn't there before.
The seven-stage pipeline
In practice, every recomposition I've built or designed lands on the same shape:
NATIVE BUSINESS EVIDENCE
documents, spreadsheets, emails, photos, systems, prior cases
↓
DETERMINISTIC ATOMISATION
line items, records, candidates, IDs, required checks
↓
COMPILED CONTEXT
wiki, prior decisions, relationships, exceptions, source evidence
↓
BOUNDED AI JUDGMENT
matching, interpretation, classification, proposed relationship
↓
HUMAN DISPOSITION
accept, modify, reject, escalate, request evidence
↓
DETERMINISTIC COMPILATION
scope, order, decision, workflow state, action or report
↓
OUTCOME + WRITE-BACK
receipt, correction, new pattern, improved future judgment
Native business evidence. The service starts from artefacts the client already possesses — workbooks, drawings, purchase orders, emails, systems. No re-keying reality into the vendor's forms; the mess is the input. This single property removes the adoption tax that kills most enterprise tooling.
Deterministic atomisation. Before any model sees anything, code establishes identity and splits the problem into governable units: line items, records, candidate matches, required checks. If the atoms don't exist, nothing downstream can be inspected, queued, disposed or audited — which is why this stage is code, not model. The units must be the same on every run.
Compiled context. The knowledge substrate the judgments will run against: prior decisions, known relationships, exceptions, source evidence, navigable and routed back to origin. Judgment without compiled context is a cold-start guess by a very confident stranger.
Bounded AI judgment. The machine does the semantic work — matching, interpretation, classification, significance — under bounds: a defined question, a scoped evidence slice, a structured output, and citation obligations. Bounded is the operative word. The model proposes; it holds no pen.
Human disposition. Accept, modify, reject, escalate, request evidence — a typed action on a typed unit, recorded, attributed. Not "human in the loop" as a vague safety blanket; a decision surface with named deciders. Chapter 5 gives this stage its own chapter, because it is the one everybody builds wrong.
Deterministic compilation. Only approved decisions compile into the deliverable — the scope, the order, the report. The model never writes the output; the compiler assembles it from dispositions. This is what makes the deliverable defensible line by line.
Outcome and write-back. Receipts out; learning in. Every run leaves the machinery better — a new source shape, a corrected pattern, a failure case. The compounding starts on run one.
Key Insight
AI handles messy perception. Code evaluates the resulting structure. People own what happens next.
That compressed rule is the pipeline in one breath, and Chapter 6 unpacks the placement logic behind it. For canon-watchers: this is the "should this process exist at all" question executed at service-design altitude — the reimagined level, not the assistive one — and distinct from redesigning an existing workflow, because the usual input here is an outcome with no incumbent workflow at all.
The design order — where most builders go backwards
The most important discipline in the whole move is sequence. The correct order:
- Buyer and business outcome — who pays, for what decision or result
- Ideal service method — what delivery would look like in a perfect world
- Evidence and decision contract — what gets proven, who decides what
- Placement — human / AI / software, unit by unit
- Governed application — the software that enforces all of the above
- Commercial product — name, price, envelope
Never the reverse — new AI capability in hand, hunting for somewhere to deploy it. That reverse order is precisely the trinket factory from Chapter 1, rebuilt with better parts. My own version of the discipline, from the project that taught me it: if you had the perfect world, what is the consulting you would do? What does that shape look like? Now — do you need AI to support that model? The fixed-price readiness product looked commercially irrational under labour economics; the design chose the better client outcome first, then built the cognitive machinery that made it viable. AI enters as the economic substrate of a chosen shape — never as a capability seeking a use case.
Specimen check: the Data Readiness Review that anchors this book is this pipeline, instantiated end to end — a client's estate and spreadsheets in, typed findings and a gated scope out. It runs in full in Chapter 10; the next five chapters first take the stages that carry the commercial weight — the line items, the placement, the compiler, the typed unknowns, and the price.
The Line-Item Decision Surface
The problem with the giant model answer isn't its quality. It's its shape.
Ask a frontier model for "the correct project scope" and you will get something fluent, structured, and frequently right. You will also get something ungovernable. You cannot inspect it part by part. You cannot accept forty of its claims and reject five. You cannot correct one mapping without regenerating the whole. You cannot test it, because it isn't decomposed into testable assertions. And you cannot attach accountability to it, because nobody can sign one paragraph of a monolith.
The problem is not the answer's quality. It is the answer's shape. And most enterprise AI disappointment is a shape mismatch: organisations can only adopt outputs they can dispute in parts.
What a governable unit looks like
The recomposition's answer is the line item — the atom Chapter 4's second stage manufactures. Here is the full anatomy, and why governance requires every property:
Stable identity
The unit can be referenced, tracked and audited across runs. Without identity there is no queue, no history, no accountability — just prose.
Bounded question
One narrow judgment — "do these two concepts correspond?" — not "assess everything". Narrow questions can be evaluated; broad ones can only be admired.
Minimum evidence set
The slice of evidence the judgment may see. Scoping evidence is scoping risk — and it makes "what did it look at?" a question with an answer.
Proposed result, with sources
The model's answer arrives carrying citations to the evidence it used. A proposal without provenance is an opinion wearing a lab coat.
Confidence or unresolved status
The unit can honestly say "insufficient evidence". Chapter 8 promotes these states into deliverables — here, note only that the schema permits honesty.
Human disposition
Accept / modify / reject / escalate — a typed, recorded action, not a nod in a meeting.
Accountable owner
A named person whose call it was. Signatures attach to units, not to monoliths.
Downstream effect
What the disposition unblocks or changes. Consequence is wired into the graph, not implied by a paragraph.
My working shorthand for the whole arrangement: turn the complex problem into line items, with a brain that helps assess the line items. The decomposition converts a vague intelligent answer into a decision surface — and the decision surface, not the model, is the adoptable unit of enterprise AI.
What decomposition buys, concretely
Five properties appear the moment the atoms exist, and none of them is available to the monolith. Partial acceptance: approve forty mappings, reject five, escalate two — the deliverable advances without an all-or-nothing argument. Precise correction: fix one judgment without touching the rest; the correction is itself recorded, attributed, and fed back. Testability: each judgment type gets its own evaluation — you can measure the matcher's accuracy separately from the classifier's, and improve them independently. Parallelism: breadth becomes machine-parallel; ten thousand bounded questions fan out where one giant question could only queue. Audit: every line of the final deliverable traces to a unit, its evidence, and the human who disposed it. The deterministic core never needs to understand language at all — it operates over structured facts and completed decisions.
The UI is an authority surface
Now the part that gets dismissed as cosmetics and is actually load-bearing. The interface over the decision surface is where the organisation can finally see: what the machine observed; what it inferred; what evidence it used; what it doesn't know; which findings matter; who may decide; what has been approved; and what each approval will cause.
Without that interface, you have an agent producing clever output into a chat window — impressive on demo day, organisationally weightless by Friday. With it, you have queues, decisions, exceptions, evidence, responsibilities, state, receipts, and a controlled path into the next action. The interface is what makes a newly manufactured hybrid workflow organisationally real: something a firm can staff, schedule, govern and bill against.
"45 findings still need your call."
That is the headline message on the specimen's dashboard, and it repays a moment's analysis. It does not say "AI found 45 things" — machine achievement, human as spectator. It says the machine has done the breadth, the evidence is attached, and forty-five bounded judgments are now waiting on someone with authority. The human decision is the organising object of the entire screen; the reports are downstream of the decision surface. Where a system puts its centre of gravity tells you who it thinks is in charge — and a system that centres the disposition queue is one an enterprise can actually trust with breadth.
Myth vs reality
✗ Myth
- • More AI autonomy = more value
- • Human review is friction to engineer away
- • The goal is the machine's answer, delivered faster
✓ Reality
- • Bounded proposals + typed dispositions are what enterprises can adopt
- • The disposition record is itself product — receipts a buyer pays for
- • Constraining the judgment is what lets the machine run at full breadth everywhere else
Takeaway
The decision surface — not the model — is the adoptable unit of enterprise AI. Build the atoms first; the intelligence has somewhere safe to land.
Specimen check: the Data Readiness Review's forty-five scope-bearing findings — 24 found, 6 partially supported, 15 not found — are exactly these atoms: each separately disposable, each with evidence attached, each disposition an immutable receipt. Chapter 10 walks the queue. First, though: who gets which judgment — code, model, or human? That allocation is a design act with reasons — placement judgment, the scarce skill I've argued deserves its own owner — and it's the next chapter.
Placement: Where Code, Model and Human Each Belong
The split among deterministic software, AI and human authority is itself the product.
The recomposition stands or falls on a decision most AI builds never consciously make: where each unit of work lives. The demo pattern defaults everything to the model. The legacy pattern defaults everything to code. Both defaults fail, in opposite directions — and the deliberate allocation between them is not plumbing. The split among deterministic software, AI and human authority is itself the product.
My rule of thumb, from building these systems: use AI in the spots where you don't have to enumerate all the rules and edge cases. Then you want deterministic code to support the governance.
What deterministic code owns: the binding work
Code takes everything where the system must behave identically twice. Identities and hashes — repeatability. Immutable evidence and resolvable pointers — provenance. Decomposition and mandatory checks — completeness. Queues and state transitions — order. Citation validity — honesty. Who may mutate what — authority. Atomic promotion of results — consistency. Decision receipts — accountability. Whether outputs may advance at all — consequence.
The principle underneath the list: wherever repeatability, authority, provenance or consequence is at stake, probabilistic components are the wrong material. Not because models are unreliable at judgment — because governance is a property of mechanisms, not of intentions, and only deterministic mechanisms can be tested into trustworthiness.
What AI owns: the bounded semantic work
The model takes the work that defeated every rules engine ever written: interpreting messy language; judging whether two differently-worded things denote the same concept; assessing whether evidence supports a claimed relationship; spotting ambiguity and exceptions; proposing matches and dispositions; explaining significance against compiled context. This is precisely the work whose edge cases cannot be enumerated — which is why it stayed human, expensive, and sampled for fifty years of enterprise software.
"Bounded" is doing heavy lifting, so make it concrete: a defined question, a scoped evidence slice, a structured output schema, citation obligations, and no write access to anything authoritative. The physics matters too — this placement keeps the machine where deployment conditions favour it: batch rather than live, artefact-producing rather than irreversible, reviewable before anything acts on it, blast radius limited by construction. I've made the full argument for that lane discipline elsewhere; here it's one sentence because the pipeline enforces it structurally.
The two physical states of AI
Here is the observation that reorganised my own thinking about these systems, mid-build: there's AI in a live form, interpreting and proposing — and there's a solidified version of AI in the deterministic code.
The live form is the runtime model: walking the compiled context, judging line items, proposing mappings under bounds. The solidified form is stranger and more valuable. At design time, a model — drawing on broad domain knowledge no single engineer holds — writes the deterministic extractor, its tests, its schema. Humans review the artefact. It gets versioned, hashed, deployed, and can be rolled back. From that moment it no longer improvises: it repeats a reviewed observation policy, identically, every run. The intelligence underwent a phase change:
probabilistic synthesis
↓ inspected code
↓ tested behaviour
↓ versioned sensor
↓ repeatable evidence
Why does this matter commercially? Because it is how these systems pass review in regulated environments in weeks rather than quarters. Design-time AI outputs route through governance that already exists — code review, testing, version control, rollback — while live AI decisions require governance most organisations haven't built. The formula (I've named it Synthetic SME in the longer treatment): organisational context × domain priors × code synthesis → reviewable, testable, versionable artefacts. The full sensor architecture — safe representations, privilege boundaries, what the model may never see — is its own piece of work and deliberately out of this book's scope; here it is one stage of the pipeline with a clear owner: reviewed code.
What humans own: consequence
The human takes what no machine should hold in this architecture: accepting or rejecting material mappings; resolving genuine ambiguity; applying commercial and ethical judgment; authorising anything that binds the organisation; owning the outcome.
Notice the asymmetry the placement creates. The human's work is disposition, not re-derivation. The machine prepared the evidence, framed the bounded question, attached the citations; the human spends judgment, not reading time. That asymmetry — judgment concentrated, reading distributed to the machine — is the economic event the next chapter measures.
The placement table
| Form | What it does | Privilege | Character |
|---|---|---|---|
| Design-time AI | Designs extractors, tests, schemas | Development context only | Broad, inventive |
| Deterministic sensor | Runs approved logic against privileged systems | Broad read, narrow operations | Repeatable, inspectable |
| Runtime AI | Interprets safe evidence, proposes mappings | Narrow model-visible world | Flexible, probabilistic |
| Human | Accepts, changes, rejects, escalates | Decision authority | Accountable |
| Deterministic compiler | Applies approved transitions, emits outputs | Approved structured state only | Binding, non-discretionary |
Key Insight
Live AI where rules can't be enumerated. Solidified AI where privacy and repeatability demand mechanism. Humans where consequence lands. Write the reason for every row.
Specimen check: in the Data Readiness Review, the model proposes mappings under citation validation; code owns identity, queues and promotion; the consultant owns dispositions; and the workbook extractor was AI-authored at design time, human-reviewed, and deterministic at runtime. Every row of the table above is running in Chapter 10. Next: what this allocation does to an engagement's economics — the consulting compiler.
The Consulting Compiler
Standardise the pipeline, not the client. Bespoke answers, standard machinery.
Professional services has spent a century impaled on one contradiction. A productised service means standardised output — the same methodology deck with the logo swapped. Consulting value means client-specific answers — which, under labour economics, means bespoke effort, unrepeatable delivery, and margins that depend on heroes. The industry's two resolutions are both fakes: templates pretending to be tailored, or heroics pretending to be repeatable.
The recomposition architecture dissolves the contradiction instead, and the mechanism deserves precision because it is the economic heart of the category.
Lower every client into the same object language
Every client begins as a different mess: different semantic models and reports, workbooks with different structures, requirements written in different dialects, wildly different maturity. The system does not simplify the mess. It lowers it — compiler's term, deliberately — into a small standard object language:
source evidence → observed estate → declared requirement → proposed mapping → finding → human disposition → scoped work item
Once a client's world is translated into those objects, the rest of the engagement follows a stable process — same queues, same gates, same report structures — while every answer inside the objects remains completely client-specific. The heterogeneity went into the translation, not into the delivery.
"You standardised the compilation pipeline, not the client's reality."
Or in its sharper twin form: you productised the path to understanding, not the conclusion. The compiler metaphor is exact, not decorative. A compiler accepts unbounded variety in source programs because it owns a fixed intermediate representation; it does not require all programs to be the same program. Under traditional consulting, each client is a new expedition — new maps, new guides, new casualties. Under the compiler, each client is a new source program passed through the same compiler.
The economic law underneath
This is economies of specificity — the inversion I've argued at book length: when recomputation is cheap, value shifts from reproducing a standard answer for the average customer to computing a fitted answer for each actual one. "Computed-made", rather than mass-produced or handmade: bespoke at the answer layer, standardised at the machinery layer.
The broader economy is already reorganising around the same inversion. McKinsey finds "companies that grow faster drive 40 percent more of their revenue from personalization than their slower-growing counterparts"13; MIT Sloan Management Review named the structural version "economies of unscale" — AI learning about individuals and tailoring at scale, dissolving the old advantage of standardised size14. What those treatments describe for products, the consulting compiler does for engagements.
The honest complexity claim
Now the correction I have to make to my own first instinct, because the honest version of this claim is also the stronger one. When I first felt the compiler working, I reached for the biggest available words: complexity goes from O(N) to O(1) — a thousand spreadsheets, no worries. That is not what happens, and saying it invites a debunking the architecture doesn't deserve.
A thousand spreadsheets still cost more machine work than ten. A thousand requirements still produce more candidate mappings and more findings. The complexity has not disappeared. It has been compiled — moved off the expensive substrate onto the cheap one:
Where the effort goes
The old shape — senior cognition carries everything
- senior effort ≈ sources
- × requirements
- × undocumented dependencies
- × possible relationships
- × interpretation cycles
The consultant reads broadly, holds the estate in their head, notices candidate relationships, resolves ambiguity, then explains it all.
The compiled shape — machine carries breadth
- machine effort ≈ evidence extraction + bounded searches + item-level judgments (scales with volume; cheap; parallel)
- senior effort ≈ material findings + genuine exceptions + consequential decisions
The system holds the many-to-many graph. The consultant sees one evidence-bearing case at a time.
"The unit of human work changes: from understanding the whole estate to disposing bounded findings."
That sentence is the entire economic event. Breadth becomes cheap parallel machine work. Depth becomes selective drill-down where a question earns it. Senior human attention — the scarce, expensive, unscalable input that suppressed the service in the first place — is reserved for the small set of places where interpretation, authority or consequence is genuinely material. The senior person stops being the graph-holder. The system holds the graph.
Why this is what makes a product possible
Repeatability without genericness is precisely what "productised service" was always supposed to mean and never could. Because the pipeline is invariant, the offer can have a name, a stable delivery process, trained operators, and — Chapter 9 — a priced envelope. Because the answers are computed per client, the deliverable never degrades into the template deck. The contradiction this chapter opened with doesn't get balanced or traded off. It gets dissolved: both sides, fully, at once.
Key Takeaways
- • Standardise the compilation pipeline, not the client's reality.
- • Complexity is compiled, not removed — machine effort still scales; senior effort collapses to dispositions.
- • Bespoke at the answer layer, standardised at the machinery layer: that is what a service product is.
- • Refuse the O(1) claim. The honest version survives scrutiny; the hyped one invites it.
Specimen check: in the retained runs, three utterly different workbooks — one matching the estate, one partially matching, one deliberately unrelated — went through one identical pipeline and produced three fully client-specific results: 14 of 14 direct mappings; ten direct plus three not-found; eight not-found. Same compiler, three source programs. Chapter 10 shows the runs.
The closing line of this chapter is the one I'd put on the compiler's nameplate: the system does not make every client simple. It makes every client's complexity enter through the same governed language. And once complexity enters through a governed language, something remarkable becomes possible: you can promise things about the unknown. That's next.
Typed Uncertainty Is the Product
The fixed-price promise doesn't require resolving every unknown. It requires typing them.
Everyone assumes a bounded promise needs complete knowledge — that "fixed price" must mean "we will understand everything about your estate, guaranteed." That assumption is exactly why the promise was never made. Complete knowledge of a client's mess is not purchasable at any price, and every consultant knows it.
The constituted service's contract says something different, and at first hearing it sounds like a trick: the engagement completes when every line item reaches a typed terminal state — and several of those states are flavours of "we don't know." It is not a trick. It is the most honest commercial promise in this book.
The terminal states
Here is the full type system, as the specimen implements it. Every unit ends in exactly one of these states; each state asserts something different, costs something different to produce, and feeds something different downstream:
| Terminal state | What it asserts | What it feeds |
|---|---|---|
| Directly mapped | The declaration corresponds to observed reality — coordinates and pages cited | Scope, as low-risk work |
| Partially mapped | Correspondence exists, with named gaps | Scope, with the delta explicit |
| Not observed | Searched; absent within the audit boundary | Build-new or investigate decisions |
| Insufficient evidence | The question is well-formed; the evidence isn't there yet | A formal information request |
| Inaccessible within the audit boundary | The boundary, not the estate, ended the inquiry | A boundary-extension decision |
| Unsupported source type | An honest tooling limit | The product backlog — not the client's risk register |
| Ambiguous — consultant decision required | The machine's proper stopping point | The disposition queue |
| Excluded from this phase | A decision — recorded as one | The next engagement's backlog |
"Unknowns become typed deliverables instead of unbounded consulting labour."
Look at what that sentence does to the engagement's economics. In the traditional model, every unknown is an open-ended liability: someone has to chase it, and the chasing is unbounded, which is why unknowns blow budgets and why PMI's base rates look the way they do. In the typed model, an unknown is a product: "not observed" is a finding the client pays for; "insufficient evidence" is a formal request that lands in the client's court; "consultant decision required" is a queue item with a named owner. The engagement doesn't promise omniscience. It promises that nothing ends in a shrug.
The promise, restated
So the fixed-price review is not promising "we will completely understand and solve every item in your data estate." It is promising: a defensible inventory of what is present, what maps, what does not, what remains unknown, what decisions are required, and what should enter the next scope. That is a deliverable a buyer can hold, audit, and act on — and, crucially, it is a promise whose cost the seller can actually bound.
Most consulting scopes do the opposite: they erase the uncertainty that produced them. The polished recommendation arrives confident, the assumptions buried in an appendix, the unknowns dissolved into professional tone. The specimen's reports do the reverse — they preserve the uncertainty as structure: every unreviewed model position labelled as exactly that, weak-signal assumptions flagged as weak-signal, and the count of consultant decisions taken so far displayed on the front page (in the retained run: zero — and the report says so). A buyer can see precisely what remains conditional. That is why the document survives hostile review.
This is a discipline I've argued before under a different name: an evidence package is "claim plus exhibit plus resolvable pointer plus a confession of what could not be verified" — the confession exists so that hallucination cannot compound silently, and honest gaps are the load-bearing part of credibility. Typed uncertainty is that authoring discipline promoted into a product feature: every finding a witness you can check, never an oracle you must trust.
The boundary clause
"Not observed" within the audit boundary is not proof of absence outside it.
That sentence belongs in every report the machinery produces, and the specimen builds it into its map responses and knowledge pages as an operating rule. It prevents the two worst epistemic failures at once. Without it, in one direction, a model invents presence — hallucinating that the thing exists because the question implied it should. In the other, a retrieval miss quietly becomes a fact about the world — "the system didn't find it, so it isn't there" — which is how audits produce confident wrongness. One sentence, structurally enforced, closes both doors.
Myth vs reality
✗ Myth
- • Unknowns in a deliverable are admissions of failure
- • A confident report sells better than an honest one
- • The consultant's job is to make uncertainty disappear
✓ Reality
- • Typed unknowns are what make the confident parts believable
- • Each unknown arrives pre-converted into the client's next decision
- • Preserved uncertainty is why the document survives hostile review
Bottom Line
The engagement doesn't promise omniscience. It promises that nothing ends in a shrug — every unit lands in a typed state with a named next step.
Specimen check: fifteen of the specimen's forty-five scope-bearing findings are "not found" — each one a deliverable, not an apology. And the single most persuasive number in the retained runs is the complete-miss workbook coming back with eight not-founds: the machine declining to hallucinate a fit. In a market drowning in confident AI output, the refusal is the sales demo.
Typed uncertainty tells the buyer what they're getting. It doesn't yet tell them what they'll pay. Turning the types into a price — an engineered one — is the envelope, and it closes the architecture.
The Fixed-Price Envelope
Not bravado, not padding. A measured, typed, priced set of boundaries.
"How can you possibly fix-price this? Our clients differ by an order of magnitude." It is the first question every consultancy asks about the readiness review, and it deserves a real answer, because the industry's two existing answers are both bad. Bravado: quote fixed and eat the variance — a strategy with a documented casualty rate. Padding: quote fixed at triple the expected cost — which prices you out of every competitive deal and still leaves tail risk.
The constituted service's answer is neither. It is an envelope: an engineered set of measured, typed, priced boundaries that turn the promise into an instrument.
First, why the naked promise fails — the base rate, briefly, since Chapter 3 laid it out: only 62% of projects finish within their original budget and a third experience scope creep, and the killer in fixed-price work is always the same — the quote was made against an unmeasured estate. Time-and-materials, vague discovery phases and caveat walls are all just ways of refusing the promise. The envelope is a way of making it.
The six parts of the envelope
1. The automated preflight census. Before anything is quoted, the machine measures the estate: workbooks, sheets, tables, formulas, connections, queries, semantic models, measures, requirement counts, early ambiguity rates. The census runs on the same deterministic sensors that will deliver the engagement — so the measurement is a free byproduct of the machinery, not a paid discovery phase. The single most important property: complexity is measured before the promise is made, by the thing that will have to keep the promise.
2. Volume bands. The census assigns the engagement to a priced band. This changes the client conversation from "how long is a piece of string?" to "you are a Band B estate — here is what Band B includes." Bands are product configuration, not estimation: the same move a cloud provider makes when it prices instances instead of quoting each customer a bespoke server.
3. Included finding counts. Each band includes a stated number of material findings or review cases. This is the subtle one, and the most important: it aligns the price with the engagement's true cost driver. After Chapter 7, we know what that is — not calendar time, not data volume, but disposition load: the number of bounded judgments a senior human must make. The machine's breadth is nearly free; the human's dispositions are the metered resource. So meter them, visibly, in the contract.
4. The Flex Reserve. A named, priced absorption layer for typed surprise: unusual access problems, exception density beyond the band's assumption. Not a contingency slush fund — the reserve is drawn against typed events, visible to both sides, and what is not drawn is not consumed. It is the contractual home for the honest fact that measurement flattens variance without eliminating it.
5. Unsupported-source rules. The tooling's limits are contract terms, not mid-engagement discoveries. Chapter 8's "unsupported source type" state gets a commercial path decided in advance: a manual quote, or an explicit exclusion. The client learns the machine's edges on day zero, from the seller.
6. The boundary clause. "Not observed within the audit boundary is not proof of absence outside it" — carried from the epistemics (Chapter 8) into the commercial terms, so the contract and the evidence model say the same thing.
"That is not old-fashioned time-and-materials estimation. It is machine-measured product configuration."
What AI actually does to the cost curve
Be precise here, because the envelope only works if the underlying economics are stated honestly. AI does not make the cost curve flat. It makes it flatter — and it changes which variable drives it. The variance that historically killed fixed pricing — how many spreadsheets? how messy? how many undocumented systems? — now lands mostly on cheap, parallel machine work. The residual variance — how many findings will need a senior human's call — is exactly what the included-finding count meters. So the envelope isn't padding for total ignorance, the way old fixed prices were. It prices a measured residual. The padding shrinks to a reserve; the reserve is typed; the types are visible. Each layer of the envelope exists because some specific uncertainty used to be unpriceable and now isn't.
The demand side is already moving toward this shape of promise. About a quarter of McKinsey's global fees now come through performance-based arrangements — clients arriving with the outcome they want and asking the firm to price against delivering it9. Outcome pricing is the pull; machine-measured scoping is what makes it survivable for the seller.
Two fixed prices — only one is an instrument
✗ The old fixed price
- • Quoted against an unmeasured estate
- • Padding priced in for unknown unknowns
- • Variance discovered during delivery, litigated at the end
- • Every surprise is a dispute
Outcome: the documented base rates — a third of projects over budget.
✓ The enveloped fixed price
- • Census before quote; band as configuration
- • The true cost driver (dispositions) metered visibly
- • Reserve drawn against typed events only
- • Tooling limits and epistemic boundaries in the contract
Outcome: a promise the machinery can keep — and show it kept.
One honesty note before the proof, stated here and argued fully in Chapter 14: the envelope logic is implemented intent for the specimen and design discipline elsewhere — census bands and Flex Reserves have not been market-tested across a client base. Where this book has no number (band prices, reserve percentages), it gives the shape and refuses to invent the figure. The claim being made is architectural: this is what makes a fixed price honest, and nothing less does.
Key Takeaways
- • Measure before you promise — the census is the same machinery that delivers.
- • Price the band, not the hours; meter the true cost driver — dispositions.
- • Reserve for typed surprise, not general fear.
- • Put the tooling limits and the epistemic boundary in the contract.
- • Fixed price done this way is the honest form — the reckless one is the meter running against client ignorance.
Specimen check: the specimen's version of the envelope's core promise — nothing enters scope undecided — is enforced by code, not policy: the scope compiler stays blocked until every material finding has a human disposition. A commercial gate implemented as a software gate. Which completes the architecture — category, hunting ground, recomposition, line items, placement, compiler, types, price. Time to watch the whole thing run.
The Specimen: A Data Readiness Review, End to End
One production-shaped instance, honestly labelled. The numbers come from retained records, not slides.
Before any of it impresses you, hear the label. What follows is one production-shaped specimen with synthetic retained runs — deployed, inspectable, and not yet client-proven. It is a specimen, not a prescription: it proves the method exists; discovery determines whether, where and how it applies to anyone else. Every number in this chapter comes from the system's retained records. None of them comes from a slide.
The offer, in one sentence: a fixed-price review that reads a Power BI estate, treats the client's business spreadsheets as declarations of a target architecture, and compiles the difference into an evidence-backed readiness decision and a scoped next engagement. The workbench behind it is called FDE BI. How it came to exist — the build, the corrections, the story — is a different book. Here, it is the running proof of Part II.
Why spreadsheets are the declaration
The design choice that makes the offer land with finance buyers deserves its own beat, because I learned it the hard way, inside a large insurance programme years ago. Call it an application and it needs rigour, a project, governance. Call it a spreadsheet and somehow it doesn't. They had massive databases living in spreadsheets — and two out of three meetings were about reconciling the spreadsheets to themselves, or to the real world. A farce, but an instructive one: the organisation's real requirements had migrated into workbooks precisely because workbooks evade governance.
So enterprise Excel is a governance escape hatch — and therefore a treasure. The workbook is part application, part requirements document, part historical receipt: databases disguised as cell ranges, business rules disguised as formulas, integration disguised as copy-and-paste. The recomposition treats each workbook as a declaration of the target — what the business believes it needs — and never as an oracle. Declared is a distinct authority state from observed, interpreted, decided, or scoped, and nothing in the machinery is allowed to collapse those states into one confident answer.
Stage 1 — Deterministic sensing
The engagement opens with reviewed code, not a model, reading the estate. The estate inventory catalogues the Power BI environment; the workbook extractor performs what I can only call spreadsheet archaeology — rendering formulas, sheet structure, tables and names, comments, validations and connection metadata as evidence, while ordinary business values stay out of the model's sight entirely. Secrets are masked. Opaque binaries are fingerprinted rather than dumped. And every category of extraction reports its own status: found, absent, truncated, or detected-without-a-full-decoder — the typed honesty of Chapter 8, applied to the sensor itself.
This is not an exotic posture; it is where the platform vendors themselves are pointing. Microsoft's scanner APIs catalogue "table and column names, measures, DAX expressions, mashup queries" as governed sub-artifact metadata — the sanctioned shape of reading an estate's structure without touching row-level data15. And Microsoft's own Document Inspector documentation enumerates what hides inside workbooks — hidden rows, columns and whole worksheets, invisible objects, custom XML16 — which is exactly why raw workbooks are never dumped into a model. One jurisdictional note completes the picture: under Australian guidance, identifiability is contextual — the same field can be harmless structure in one setting and personal information in another17 — so the safe representation is an engineered judgment, not a filter. The full architecture of that engineering is a sibling work's subject; for this walk, what matters is the output: a bounded, text-native, evidence-linked world the model is allowed to inhabit.
Stage 2 — Compiled context
The sensed estate becomes a navigable current-state knowledge base — pages, claims and typed relationships, every claim routed back to its source evidence. The engagement's brain, built before any judgment is asked for. When the model later says "this measure exists and serves this report", that statement carries a pointer a human can click and check.
Stage 3 — Bounded AI judgment: the blind comparison
Now the machinery Part II promised. Each workbook's evidence packet goes to a model that must reconcile it against the current-state knowledge — under deliberate blinding. No answer key. No fixture names, no expected counts. No visibility of prior targets, mappings, decisions or scope pages. The model must cite exact workbook coordinates and permitted current-state pages for every mapping it proposes; deterministic code validates every citation — and refuses to decide any mapping. Judgment stays with the model; honesty enforcement stays with code.
Three fixture workbooks went through the blind path in the retained runs, and the three results carry different meanings:
The three blind runs — retained results
Matching workbook: every requirement directly mapped — the machine finds what is there
Partial workbook: ten direct mappings, three not-founds — it discriminates
Complete-miss workbook: eight not-founds — it declines to hallucinate a fit
The third run is the one to stare at. A machine that returns "not found, not found, not found" when handed an unrelated workbook is demonstrating the property every AI buyer actually needs and almost none can verify: the refusal to manufacture agreement. In a market drowning in confident output, the refusal is the demo.
Stage 4 — Human disposition
"45 findings still need your call."
That is the organising message of the workbench's dashboard — not "AI found 45 things". The counts, labelled precisely because unexplained count differences are what corrode trust in an evidence product: 46 findings in total; 45 scope-bearing; one context-only. Of the forty-five: 24 found, 6 partially supported, 15 not found — and every single one an unreviewed model position until a consultant disposes it. Dispositions are immutable receipts. Consultant steering enters the system separately from workbook evidence, as attributed guidance, so nobody can later confuse what the machine concluded with what a human instructed.
Stage 5 — Deterministic compilation
Downstream of the decision surface, the compiler produces the commercial object: a scope draft of 21 proposed work items, each carrying explicit assumptions — two flagged as weak-signal — client responsibilities, dependencies, exclusions, and acceptance criteria that can actually fail. And the gate that makes the whole thing a product rather than a demo: scope compilation is blocked until every material finding has a disposition. The scope's states are kept unmistakable — model-generated draft, consultant-reviewed, client-agreed, ready for commercial issue — because the difference between those states is the difference between a working position and a promise. From there, the compiled scope enters an engagement chain built to carry verification rather than externalise it — the proof-carrying pattern this book's Chapter 3 already cited.
The counterfactual, applied
Remove the AI — stage by stage
- • No machine reading → full coverage of hundreds of workbooks reverts to sampling.
- • No machine matching → the many-to-many reconciliation reverts to hundreds of senior hours nobody will fix-price.
- • No compiled evidence chain → "evidence-backed" reverts to "trust our judgment".
- • No census, no typed states → the fixed price reverts to suicide or padding.
What collapses is not the margin. It is every load-bearing promise. By the Chapter 2 test: AI-constituted.
A fixed-price, full-coverage, evidence-backed readiness review has never been on a consultancy's price list. Not because clients wouldn't buy it — because promising it under human-labour economics was commercially insane. The machine doesn't improve this service. The machine permits it to exist.
What this specimen does not prove
Kept on the table deliberately: the retained runs are synthetic — real client estates are messier, more adversarial, and politically inhabited. Multi-user authority, client data policy and hardened packaging are named future increments, not present facts. And the hardest proof is entirely ahead: whether a second consultant can deliver this offer without the builder in the room. Chapter 14 owns that confession in full. What the specimen proves is narrower and still decisive: the architecture runs — every stage of Part II, live, gated, and inspectable, end to end.
Key Insight
The specimen's real output is not 45 findings. It is the allocation of work: the machine carried the breadth, code carried the honesty, and every consequential call waited for a human.
One working instance in one vertical, though, proves a product. The category claim needs the architecture to survive leaving home. Next: two transplants.
Recomposition Worked: Parts E-Commerce
Strip the domain. If the architecture only works where the author lives, it's a product, not a category.
State the objection at full strength, because it's a good one: "Very nice — for data consulting. Where the inputs are already digital, the client is technical, and you happen to own the tooling. Strip all of that away and what's left?" This chapter and the next answer by transplant. Two businesses I've actually sat across the table from, both stuck — at the time — in horse-optimisation conversations. Neither in data consulting. Same seven stages.
An honest label first, matching Chapter 10's discipline: these are design recompositions — architectures worked from real sales conversations, not shipped systems. The proof class is lower than the specimen's, and I'll say so again at the end.
The scene
A parts retailer. Heavily e-commerce, with a large incumbent platform, a substantial catalogue, and a heavy support load. The house doctrine, delivered to me with total conviction: customers need help using the catalogue; support exists to train them; and if you want to order from us, you use our e-commerce. The support queue was read as evidence that customers needed more training. They could not see that the friction was theirs.
The inherited frame — "how do we help customers use our website?" — assumes the customer should learn the vendor's ontology: part numbers, category trees, search syntax, substitution rules. Decades of e-commerce practice have normalised that assumption so completely that its cost is invisible. Every confused customer, every mis-ordered part, every support call is the cost — booked as a customer deficiency.
The recomposed outcome
Run the Chapter 4 discipline: ignore the current workflow, name the outcome. The outcome a parts buyer wants is not "successfully operate the vendor's website". It is "the right parts arrive". So recompose for that:
Let the customer express purchasing intent in whatever artefact they already possess — a purchase order, a bill of materials, a photo, an old invoice, free text — and make the supplier perform the translation.
The pipeline, instantiated stage by stage — deliberately in the same order as Chapters 4 and 10, because the repetition is the argument:
- Native evidence. A PO, a BOM, a photo of a failed component, an old invoice, an email thread. Whatever the customer already has is the order form.
- Deterministic atomisation. Code extracts lines, quantities, known part numbers; assigns each line an identity. The order becomes governable units before any model touches it.
- Compiled context. The catalogue, substitution rules, the customer's order history, prior resolutions — the knowledge the judgment runs against.
- Bounded AI judgment. Per line: semantic equivalence, variant identification, candidate substitutions — with evidence and unresolved questions attached to each line, never one grand guess at "the order".
- Human disposition. Ambiguous or high-consequence lines route to a person — the customer's confirmation, or the supplier's parts desk. Wrong-part risk is real money and freight; consequence stays human.
- Deterministic compilation. Approved lines compile into an order on the existing platform.
- Write-back. Every corrected match improves the matcher; every new artefact format becomes a supported input. The translation layer gets better with every order it survives.
"The customer no longer learns how to order from you. Your system learns what the customer is trying to order."
The counterfactual, run
Remove the AI. "Order by sending us whatever you've got" collapses immediately — deterministic parsing alone cannot survive the artefact variety; that's the pre-AI world, and it's exactly why the offer has never existed. Staff the translation with humans and you've rebuilt the support queue under a different cost centre, at a cost that kills the promise. There is no slower version of this offer. There is no offer. Constituted.
And notice what the recomposition does to the support queue's meaning. The workload was never proof that customers needed training. It was evidence that the customer was being forced to act as the translation layer — unpaid, error-prone, and resentful. The suppressed service was hiding inside the support budget all along, misfiled as a cost of doing business — the kind of accidental friction I've argued firms should treat as an attack surface rather than an overhead.
What survives of the incumbent system
Everything that was ever good at its job. The e-commerce platform keeps the catalogue, pricing, stock and transactions — the recomposition doesn't raze it; it stops using it as a cognitive burden to place on customers. This matters as a general principle: recomposition changes where the thinking sits, not necessarily which systems exist. The capital spent on the platform is preserved; the ontology-learning tax on customers is abolished.
The doctrine, echoed once each
- Line items (Ch 5): each order line is a bounded, disposable judgment with evidence attached
- Placement (Ch 6): platform keeps identity and transactions; AI judges only equivalence; humans confirm consequence
- Typed uncertainty (Ch 8): "no confident match — review" is a first-class line state, not a failure
- Envelope (Ch 9): lines per order and ambiguity rate are the census that meters the service
Takeaway
Every "customers struggle with our system" queue hides a translation burden the supplier could absorb — and the absorption is usually a constituted offer waiting to be built.
One transplant down. But parts matching is, frankly, the friendly case: digital artefacts, bounded consequence, no professional liability. The next chapter picks the hostile one — physical-world engineering, where a wrong scope line carries structural risk — and runs the same play.
Recomposition Worked: Concrete-Remediation RFQs
The hostile transplant: physical evidence, heavy standards, liability-adjacent judgment.
If parts matching is the friendly transplant, this is the hostile one, chosen deliberately. Concrete remediation: site evidence is physical, the standards are heavy, the judgments carry professional liability, and a wrong scope line can mean structural risk. If the architecture holds here, the category claim has teeth.
The scene
A firm of remediation specialists — serious engineers with deep expertise — drowning in complex RFQ and scoping work. They came to the conversation wanting their existing quoting process automated, and they led hard: they knew exactly what they wanted, and what they wanted was the same thing, faster. The classic faster-horse request, delivered with the full weight of domain authority. Trying to counter it live, in the room, did not work — more on that below, told against myself.
Why the faster-horse request fails here, specifically: the RFQ's cost was never the typing. A remediation scope is expensive because of the joins — site evidence against standards, defects against analogous past jobs, observations against risk memory — currently performed inside a senior engineer's head, one bid at a time. Chapter 3 named this pattern in consulting: the estimator is the join algorithm. Automating the document production while leaving the joins in the engineer's head automates the cheap part.
The recomposed outcome
Decompose the outcome — a defensible, priced remediation scope — not the current process. The stages, third time in this book, same order, different world:
- Native evidence. Site notes, drawings, photographs, emails, applicable standards, past jobs. The mess the firm already generates on every inspection is the input.
- Deterministic atomisation. Code atomises the evidence into defects, areas, constraints and evidence gaps — each with an identity, each traceable to its source photo or note.
- Compiled context. Remediation classes, the standards library, analogous completed work, failure history — the firm's institutional memory made navigable instead of anecdotal.
- Bounded AI judgment. Per item: likely remediation class, associated risks, candidate exclusions, information requests — every proposal citing the photo, the drawing or the standard it leans on.
- Human disposition. The engineer disposes each material item. The liability-bearing judgment stays with the licensed human — by design, not by afterthought. Hold this beat; Chapter 14 builds on it.
- Deterministic compilation. Approved items compile into the RFQ/SOW structure — every scope line traceable to a defect, a photo, a standard, or a recorded decision.
- Write-back. Every priced job — won or lost — enriches the analogous-work memory. The estimator's scar tissue becomes a firm asset instead of a retirement risk.
What changes commercially is not faster document writing. It is an evidence-backed scope-forming system that did not previously exist — and with it, the suppressed service surfaces: a bounded-price scoping product for remediation work, previously irrational because every site was an expedition and every quote a hostage to concealed conditions.
The doctrine ports without modification
- Typed uncertainty (Ch 8): "insufficient evidence" becomes a formal information request — a deliverable that used to be an awkward phone call. "Inaccessible within boundary" covers what the site visit couldn't reach.
- Envelope (Ch 9): the census is sites, assets, defect density, drawing coverage; the Flex Reserve absorbs concealed conditions — the industry's classic margin-killer, now a named, priced object.
- Line items (Ch 5): partial acceptance means the client can fund urgent structural items now and defer cosmetic ones — impossible while the scope was one prose lump.
- Placement (Ch 6): AI proposes classes and risks; code owns traceability; the engineer owns every consequential call.
The paragraph I owe you about the room
One honest beat about how that conversation actually went, because it teaches something the architecture alone doesn't. I tried to make this case in the room, live, against their framing — and it didn't work. Their language kept pulling the conversation back into their process; any immediate alternative sounded like an outsider's opinion set against decades of domain authority. My selling technique's got to improve a bit. The transferable lesson, in one clause: the first meeting's job is to earn the right to inspect the evidence — the recomposition is what you earn it for, not what you pitch cold. The full protocol for that entry transaction is another piece of work; this book's contribution is making sure that when you've earned the inspection, you know what to build.
The category tell
Three domains have now been walked: data consulting at full resolution (Chapter 10), retail intent translation (Chapter 11), engineering scope formation (this chapter). Look at what varied and what didn't.
Three domains, one architecture
What varied
- • The evidence: workbooks · purchase orders · site photos
- • The judgments: mappings · part equivalence · remediation classes
- • The humans: consultant · parts desk · licensed engineer
- • The output: readiness scope · compiled order · priced RFQ
What didn't
- • Seven stages, in the same order
- • Line items with evidence and identity
- • Bounded proposals, human dispositions, gated compilation
- • Typed unknowns; census-metered economics; write-back
The invariant is the architecture; the variation is everything the client actually sees. That is what "category, not product" means — and it satisfies the audit this book set itself at the start: remove Power BI, remove data consulting, remove the specific vendor moment, and the thesis survives.
Key Insight
The harder the domain looks, the deeper the suppression ran — and the more valuable the recomposition that ends it.
A category that transplants invites the obvious commercial question: if anyone can see the architecture — it's printed in this book — what stops a competitor from simply building it? The next chapter answers with the least romantic word in strategy: economics.
The Moat: They Can Copy the Offer by Tuesday
The brochure copies overnight. The delivery economics don't.
The day after you launch "Fixed-Price Data Readiness Assessment", assume it appears on a competitor's website. Same name by Tuesday. Same report headings by Friday. Screenshots that look convincingly similar within a quarter. If the offer were the asset, the category this book describes would be worthless the week it worked.
When I first saw this clearly, mid-conversation about my own product, it came out as a taunt: for competitors to copy the product, they need the same delivery machinery. And guess who's got the machinery? Thirteen chapters in, that swagger can now be cashed out as an argument.
Start from symmetry
The premise that makes moat-thinking urgent: frontier capability is symmetric. Every competitor rents the same models, from the same handful of labs, on the same day they ship — whatever advantage a release confers, it confers on everyone at once, which means it confers durable advantage on no one. Symmetry is the opposite of a moat. "We use AI" is a claim about the rented part. The moat, if one exists, must be made entirely of things that cannot be rented.
The hidden stack
Here is what the imitator cannot see from your brochure, enumerated. Each layer looks minor. Together they are the offer's existence conditions:
- Deterministic sensors and their adversarial tests — the archaeology code, hardened against the workbooks that break naive parsers.
- The safe-representation design — what the model may see, what it may never see, and the reasoning behind every line of that boundary.
- Content-addressed evidence with immutable receipts — the substrate that makes "evidence-backed" a mechanism instead of a adjective.
- The compiled context — the knowledge architecture judgments run against, and the machinery that builds it fresh per client.
- Blind reconciliation with citation validation — the anti-hallucination harness, including the discipline of hiding the answer key from your own system.
- The authority-state separation — observed / declared / interpreted / decided / scoped, enforced so no state silently becomes another.
- Disposition and authority controls — who may decide what, recorded how.
- The scope compiler and its gates — nothing advances undecided.
- The typed-uncertainty contract — the terminal states and their commercial paths.
- The operating doctrine — the accumulated judgment that decided all nine layers above, including the failures that taught them.
The moat, defined
A governed delivery system that makes an otherwise dangerous commercial promise economically repeatable.
Notice what that definition refuses to say. Not "proprietary AI". Not "unique data". The dangerous promise — fixed price, full coverage, evidence-backed — is public; anyone can make it. The moat is being able to make it and survive keeping it, engagement after engagement. I've argued the general form of this elsewhere: the product is the governed join — client knowledge, firm capability, operating doctrine, and an executable delivery path — and the moat is the composition, never any single component.
The gap widens on its own
The composition isn't just hard to copy. It compounds — Chapter 4's write-back stage, operating at business scale:
more engagements
→ another tested source shape, another extractor,
another failure pattern, another scope structure
→ lower delivery cost
→ safer fixed-price promise
→ more sales
→ more engagements
One boundary keeps the loop honest: what compounds is the machinery — extraction patterns, tests, failure cases, report structures — never the client's private evidence, which stays inside its engagement boundary. The reusable asset is method, not data.
Now put the imitator inside this picture. They start with your brochure; you start with the accumulated production system behind it. They can promise the fixed price — and then eat the variance the census would have measured. They can ship the confident report — and carry the risk the typed states would have surfaced. They can quote the offer — and absorb the first public failure the citation gates would have caught. Copying the offer without the composition is buying the risk without the instrument. Until they reproduce enough of the stack to match your delivery cost, scope confidence and evidence quality, they can imitate the claim but not safely honour it.
The two-product structure
The cleanest way to hold the whole argument, for a consultancy weighing this up:
"The readiness review is the product the consultancy sells. The delivery operating system is the product that makes it capable of selling it."
Two products, two owners of value. The first earns revenue; the second earns the right to the revenue. Selling the first without owning the second is how imitators get hurt. Owning the second without selling the first is doctrine without a business. The firms that win this category will hold both — and price them separately in their own heads even when the client only ever sees one.
What happens to advisory
The anxious question underneath every consulting conversation about AI: does this kill advisory? No — it relocates it. In a constituted service, the frameworks shape what the system notices; senior judgment defines the decision criteria and the disposition classes; consultants resolve the escalations the machine correctly refuses. The judgment is all still there. It has moved from the deliverable into the delivery system.
"Advisory becomes the control plane inside the product rather than the product delivered at the end."
What dies is not judgment. It's judgment sold as retyped narrative — the deck whose production cost was the price justification. Chapter 3's market data already showed which side of that split is growing.
The altitude this belongs at
One closing frame, because project selection is where this argument has to land. Building an AI-constituted service is not an efficiency project and shouldn't compete for budget as one. In the terms I've argued at board level: it is Car Construction — strategic capital into a compounding asset, selected because it defends what the firm will be worth when cognition, software and first-pass advice are cheap. When execution gets cheaper, the constraints on execution become more valuable — which is exactly why the governance, evidence and placement layers of the stack are the moat rather than the overhead.
Key Takeaways
- • Under capability symmetry, only unrentable assets defend — and the composition is unrentable.
- • The imitator buys the risk without the instrument: the promise without the machinery that survives it.
- • The loop widens the gap: every engagement upgrades the shared machinery, never the client's data.
- • Two products: the one you sell, and the one that lets you sell it.
- • Advisory relocates into the control plane. The deck dies; the judgment doesn't.
Every argument so far has run in the category's favour, which should make a careful reader suspicious. A test that always says yes is broken. The next chapter breaks it on purpose.
When the Machine Doesn't Constitute
A category that can't lose arguments can't win them either. This chapter runs the test where it fails.
Chapter 2 promised the counterfactual test was falsifiable. Time to spend the promise. There are three classes of service where the honest answer to "remove the AI" is not the one this book has been celebrating — and one confession about this book's own evidence that belongs in print rather than in a footnote.
Class 1: accountability is the product
Audit signatures. Legal opinions. Medical decisions. Engineering certifications. In each of these, the offer is constituted by a licensed human accepting liability — the signature is the product, and everything upstream of it is preparation. AI can assemble the evidence, draft the analysis, type the uncertainty, queue the judgments — and the offer still cannot exist without the signer, because what the client is buying is a person the law can reach.
The test has a mirror form for exactly this case: remove the human and watch what collapses. If the answer is "the offer", the service is human-constituted, and no amount of machinery changes its category. The nuance worth keeping: these services still benefit enormously from the architecture — evidence rails, typed states and decision surfaces make the signer faster, safer and more defensible. But that machinery is AI-enabled tooling for a human-constituted offer. Calling it "AI-constituted" would be precisely the label inflation Chapter 2 exists to stop — and it's why Chapter 12's engineer kept every consequential call. The taxonomy polices its author too.
Class 2: judgment density defeats the economics
The subtler failure. Some work decomposes beautifully into line items — and then nearly every line item turns out to be ambiguous, high-consequence and precedent-poor. Bespoke litigation strategy. Novel-situation crisis counsel. The first-of-kind deal. When the human must effectively re-perform each judgment anyway, the machine layer adds routing cost without absorbing breadth: you've built an expensive workflow wrapper around what remains artisanal work.
Here the recomposition discipline pays for itself in a way I didn't originally design: it is a build/no-build detector. Decompose the outcome honestly, write the placement column with reasons — and if that column reads "human" on nearly every unit, you have learned, before spending a dollar on software, that the offer is human-constituted. That is the method working, not failing. Build the evidence rails if they genuinely help the humans; do not pretend the machine constitutes the service.
The quantitative tell
Constituted economics need the disposition load to be a small fraction of the unit count. When dispositions approach the number of units, you've rebuilt the old cost curve with extra steps.
No fake threshold number — the sources don't give one and I won't invent one. But the shape is measurable in any honest decomposition, which is more than most build decisions ever get.
Class 3: the physics is against you
Live, irreversible, adversarial contexts — real-time customer-facing judgment with brand-scale blast radius — remain the machine's worst deployment lane, for reasons I've argued at book length elsewhere and won't rehearse. The external failure literature says the same thing from the field: RAND's fifth root cause of AI project failure is that "the technology is applied to problems that are too difficult for AI to solve… AI is not a magic wand"3. The failure statistics of Chapter 1 aren't an embarrassment to this book's argument. They're the boundary data — the map of where constitution was attempted against the physics.
Klarna, revisited
Chapter 2 left a thread hanging: Klarna's celebrated AI assistant, and then the 2025 re-recruitment of humans, the CEO conceding cost had dominated quality. Close the loop properly: snap-back is the signature of non-constitution. A substitution offer always has the old staffing model waiting in the wings — that is what makes reverting possible at all. A constituted offer cannot snap back, because there is no pre-AI version of the service to return to; the offer's history starts with the machine. The reversal doesn't weaken the taxonomy. It is the taxonomy, predicting behaviour in public.
The confession: what this book actually rests on
Now the part most business books bury. Stated flat:
This capstone rests on one production-shaped specimen with synthetic retained runs, plus two design recompositions. The architecture is demonstrated — deployed, gated, inspectable. It is not surveyed. n = 1, stated as n = 1.
Itemised, what remains unproven: real-client performance under messy, adversarial, politically inhabited estates. Multi-user authority and client data policy under production conditions. The envelope — census bands, included findings, Flex Reserve — under commercial fire. And the hardest one: the transfer test. Engagement one can always be heroics; the only proof that machinery rather than a person carried the work is engagement two — whether an ordinary consultant can deliver materially more because the shared infrastructure improved, not because the original builder worked late again. That test is ahead of the specimen, not behind it.
And because a confession without consequences is just mood, here are the falsification conditions — what would make this book wrong:
- If a firm delivers the same promise — fixed price, full coverage, evidence-backed — at sustainable economics without the composition, the moat claim falls. The stack would be decoration, not existence condition.
- If the specimen's economics fail on real engagements despite the composition, the flagship falls back from "proof" to "promising architecture", and this book's Part III overreached.
- If offers this taxonomy classifies as AI-constituted routinely snap back to human delivery, the test is mislabelling dependence as constitution, and Chapter 2 needs a sharper knife.
Why publish your own attack surface? Because it's the same discipline this book sold in Chapter 8 — the confession of what could not be verified is the load-bearing part of credibility — applied to the book itself. The claim ships with its own test suite. I'd rather be falsified precisely than believed vaguely; the first one teaches something.
Key takeaways
- • Run the mirror test: remove the human too — accountability-products are human-constituted, full stop
- • Judgment density is measurable: disposition load approaching unit count means the machine adds routing, not breadth
- • The recomposition doubles as a build/no-build detector — a "no" before the build is the method working
- • Snap-back reveals substitution; constitution has nothing to revert to
- • n = 1 stated as n = 1 — engagement two is the proof that matters
- • The category claim ships with its own falsification conditions
The boundaries are drawn, the confession is filed, and the test cuts in both directions. What's left is the part you do. One chapter: three moves, starting Monday.
Build One: The Category Ahead
Not a transformation programme. Three moves, in order, each producing an artefact.
No roadmap slide, no operating-model redesign, no platform purchase. Three moves you can start this week, in order — each one produces a concrete artefact, and each artefact is the input to the next move.
Move 1 — Classify the catalogue (output: a marked-up service list)
Run the Chapter 2 test across every offer you currently sell, in writing, one line each: remove the AI — what collapses? Expect the honest result and don't flinch at it: mostly AI-enabled, a little dependent, probably nothing constituted. That's information, not failure — it tells you where you're standing before you move.
Read the classifications strategically. Enabled buys parity: necessary, table stakes, and everyone gets the same copilots the same quarter — hold it, don't celebrate it. Dependent buys cost position: real but fragile, and remember the Klarna lesson — a substitution play can snap back the moment quality or politics demands it, because the old staffing model is still waiting. Constituted — if you genuinely have one — is a market position; Chapter 13's moat logic applies to it today, so start compounding.
Then do the disrespectful version: classify your competitors' offers. Every "AI-native" claim in your market that fails the test is a brochure, not a business — and knowing which is which changes how you sell against them.
Move 2 — Inventory the suppressed services (output: capability sentences)
Now hunt where Chapter 3 pointed: not your process map — the market's catalogue shadow. The question, run outward: what important thinking does this market not perform because it would require too many experts, too much coordination or too much time?
Four absence signatures make the hunt systematic. Where analysis is done only on samples — what would full coverage be worth, as a product? Where events are investigated only on escalation — what would always-investigated look like, priced? Where planning stress-tests a handful of scenarios — what would exhaustive stress-testing sell for? Where synthesis waits to be asked — what would unsolicited, evidence-backed synthesis be, as a subscription?
Force the output into one disciplined format — capability sentences: "We could promise [outcome] at [bounded price] if a machine carried [the breadth]." Half a dozen honest sentences beat forty brainstormed ideas. Then select one, using two filters: a named buyer with budget and a decision they can't currently purchase, and native evidence the client already possesses — because Chapter 4's pipeline starts from artefacts, and an offer that requires the client to produce new artefacts has already lost.
Move 3 — Recompose one outcome (output: a placement design, then the smallest shippable version)
The design checklist — nine steps, each one enforced by a chapter of this book:
- Name the buyer and the outcome. A person, a budget, a decision. The outcome, not a feature list — and in the design order of Chapter 4: ideal service first, machinery second.
- Decompose into line-item judgments. Units of cognition and authority the outcome needs — not steps your current process has. If your unit list reads like your SOP with IDs, start again.
- Place each unit, with written reasons. Code, model or human, per Chapter 6 — and treat the reason column as a deliverable. If "human" dominates the column, stop: you've found human-constituted work, and Chapter 14 just saved you the build.
- Type the uncertainty before building. Enumerate the terminal states — Chapter 8's list is the template. Every state is a deliverable with a named downstream consumer, or it isn't a state.
- Design the disposition surface. Who may decide each class of finding, what evidence they see, what their decision causes. The UI is the authority surface — budget for it like one.
- Gate the outputs in code. Nothing compiles without dispositions. The gate is software, not policy — Chapter 10's blocked scope compiler is the pattern.
- Census before pricing. Measure the complexity drivers with the same machinery that will deliver; band the price; meter the true cost driver — dispositions; reserve for typed surprise. Chapter 9 is the template.
- Ship the smallest version that keeps every property. Smallest scope — one workbook type, one judgment class, one buyer — but never fewer properties. Dropping the gates or the typed states to ship faster produces a trinket with better plumbing.
- Write back. Every run must leave the machinery better: a new source shape, a corrected pattern, a failure case. Chapter 13's compounding loop starts with run one, or it never starts.
Remember
Shrink the scope, never the properties. The properties are what make it constituted; the scope is just where you start.
Where this sits on the longer ladder
For orientation — because one constituted offer is a position, not an end state:
| Consultancy posture | What it sells |
|---|---|
| AI advisory | Advice about AI |
| AI trinkets | Small AI features attached to existing work |
| AI-enabled consulting | Existing engagements produced more efficiently |
| AI-constituted services | New commercial products that cannot exist without AI |
| AI-native consultancy | A repeatable system for discovering and launching those products |
This book takes a firm to the fourth rung: one constituted offer, designed, built, priced and defended. The fifth — the organisational system that repeatedly discovers, builds and launches such offers as a portfolio — is a different problem with different machinery, and it deserves its own treatment rather than a closing paragraph here.
The whiteboard, rewritten
Go back to where Chapter 1 started: the workshop, the processes on the wall, the question that could only ever produce trinkets — because it was a question about the existing workflow. The question that produces businesses was never on that whiteboard:
What would you promise if a machine could carry the breadth?
Somewhere in your market is a service nobody sells because it was never rational to sell. The machinery to constitute it now exists — this book just walked it end to end, priced it, transplanted it twice, and told you where it breaks. Stop selling AI as the subject of the engagement. Embed AI deeply enough that a better engagement can exist.
"AI does not improve the service. AI permits the service to exist."
One ask
Run the counterfactual test on your own catalogue this week — every offer, one line each: remove the AI, what collapses?
Then bring me the most interesting line. If you lead a consultancy or a services business and want to pressure-test an offer against the test — or design your first recomposition — I'm at leverageai.com.au.
References & Sources
The evidence base behind every claim — primary research, industry analysis, and technical specifications
Research Methodology
This ebook draws on primary research from standards bodies, independent research firms, enterprise technology vendors, and consulting firms. Statistics cited throughout have been cross-referenced against primary sources.
Frameworks and interpretive analysis developed by Scott Farrell / LeverageAI are listed separately below — these represent the practitioner lens through which external research is interpreted, and are not cited inline to avoid self-promotional appearance.
Primary Research & Standards Bodies
Fortune / MIT NANDA — MIT report: 95% of generative AI pilots at companies are failing [1]
About 5% of AI pilots achieve rapid revenue acceleration; failure attributed to the learning gap and flawed enterprise integration, not model quality
https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
CIO Dive / S&P Global Market Intelligence — AI project failure rates are on the rise: report [2]
Companies abandoning most AI initiatives rose from 17% to 42%; average organisation scrapped 46% of AI proof-of-concepts before production
https://www.ciodive.com/news/AI-project-fail-data-SPGlobal/742590/
RAND Corporation — The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed [3]
More than 80 percent of AI projects fail, twice the rate of non-AI IT projects; the leading root cause is misunderstanding or miscommunicating what problem needs to be solved
https://www.rand.org/pubs/research_reports/RRA2680-1.html
Project Management Institute — Pulse of the Profession 2021: Beyond Agility [6]
2021 global figures: 62% of projects completed within original budget; 34% experienced scope creep
https://www.pmi.org/-/media/pmi/documents/public/pdf/learning/thought-leadership/pulse/pmi_pulse_2021.pdf
MIT Sloan Management Review (Hemant Taneja with Kevin Maney) — The End of Scale [14]
Business in the century ahead will be driven by economies of unscale, in which traditional competitive advantages of size are turned on their head; AI can tailor products at scale
https://sloanreview.mit.edu/article/the-end-of-scale/
LeverageAI / Scott Farrell — Practitioner Frameworks
The interpretive frameworks, architectural patterns, and practitioner analysis in this ebook were developed through enterprise AI transformation consulting. The articles below are the underlying thinking behind those frameworks. They are listed here for transparency and further exploration — not cited inline, as this is the author's own analytical voice.
Scott Farrell — Stop Automating. Start Replacing
The Spock question: whether the process should exist at all precedes any automation decision; the old workflow is inherited constraint, not design
https://leverageai.com.au/wp-content/media/articles/24-stop-automating-start-replacing.html
Scott Farrell — Maximising AI Cognition and AI Value Creation
Version 3 of AI value: making entirely new categories of work rational to attempt for the first time, versus automating or scaling existing work
https://leverageai.com.au/wp-content/media/articles/27-maximising-ai-cognition.html
Scott Farrell — The Cognition Scarcity Audit
The audit method that hunts valuable analysis, vigilance and synthesis an organisation never performs because human cognition made breadth, depth or frequency uneconomic
https://leverageai.com.au/wp-content/media/articles/142-cognition-scarcity-audit.html
Scott Farrell — Proof-Carrying Transformation
Advice got cheap, verification did not: what buyers still fund is a path that survives architecture review, cyber, finance and production — the engagement model that retains verification rather than externalising it
https://leverageai.com.au/wp-content/media/articles/164-proof-carrying-transformation.html
Scott Farrell — Someone Has to Decide Where Intelligence Belongs
Placement judgment: the scarce skill of deciding, step by step, whether work belongs with deterministic software, a model, a human, or no intervention
https://leverageai.com.au/wp-content/media/articles/163-someone-has-to-decide-where-intelligence-belongs.html
Scott Farrell — The Lane Doctrine
Deploy AI where the physics supports it: latency-tolerant, artefact-producing, reviewable work with bounded consequences — batch the brain, ship the artefacts, govern like software
https://leverageai.com.au/wp-content/media/articles/47-the-lane-doctrine.html
Scott Farrell — The Simplicity Inversion
Governance arbitrage: design-time AI produces reviewable artefacts that route through existing SDLC governance, while runtime AI decisions require purpose-built enforcement; the Synthetic SME formula
https://leverageai.com.au/wp-content/media/articles/41-simplicity-inversion.html
Scott Farrell — The Team of One
Economies of specificity: when thinking becomes cheap, customisation becomes cheaper than standardisation — computed-made outputs recomputed per case through standard machinery
https://leverageai.com.au/wp-content/media/articles/20-team-of-one.html
Scott Farrell — Product of One
The evidence package: claim plus exhibit plus resolvable pointer plus a confession of what could not be verified — witness discipline that makes honest gaps load-bearing
https://leverageai.com.au/wp-content/media/articles/129-product-of-one.html
Scott Farrell — The Friction Attack Surface
Necessary vs accidental vs wrong friction: accidental friction should be compiled away, and a customer forced to act as the translation layer is accidental friction misread as customer failure
https://leverageai.com.au/wp-content/media/articles/113-friction-attack-surface.html
Scott Farrell — Forward-Deployed Engineering
Capability symmetry: every competitor rents the same models from the same labs on the same day, so the model cannot be the moat — durable advantage must live in what cannot be rented
https://leverageai.com.au/wp-content/media/articles/175-forward-deployed-engineering.html
Scott Farrell — The Forward-Deployed Practice OS
The product is the governed join of client knowledge, firm capability, capability kernel and executable delivery path; the moat is composition, not any individual component
https://leverageai.com.au/wp-content/media/articles/167-forward-deployed-practice-os.html
Scott Farrell — The Terminal Value Doctrine
Select AI projects by whether they defend or increase terminal value under cheap cognition and rising governance pressure; when execution gets cheaper, the constraints on execution become more valuable
https://leverageai.com.au/wp-content/media/articles/61-terminal-value-doctrine.html
Case Studies
Klarna (press release) — Klarna AI assistant handles two-thirds of customer service chats in its first month [4]
The assistant handled two-thirds of service chats, equivalent work of 700 full-time agents, resolution down from 11 minutes to under 2, estimated US$40M profit improvement in 2024
https://www.klarna.com/international/press/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month/
Industry Analysis & Vendor Research
CX Dive — Klarna changes its AI tune and again recruits humans for customer service [5]
CEO Sebastian Siemiatkowski: cost was too predominant an evaluation factor, resulting in lower quality; Klarna recruiting humans back into customer service
https://www.customerexperiencedive.com/news/klarna-reinvests-human-talent-customer-service-AI-chatbot/747586/
Consultancy.uk (commentary, Scott Lane) — Big Four business models face moment of reckoning with rise of AI [10]
Industry commentary: the hourly billing model is fundamentally incompatible with how AI transforms productivity; clients asking why lower costs are not reflected in invoices
https://www.consultancy.uk/news/amp/45185/big-four-business-models-face-moment-of-reckoning-with-rise-of-ai
Foundation Capital — AI leads a service-as-software paradigm shift [11]
AI companies leading a transition from software-as-a-service to service-as-software; a $4.6 trillion opportunity as the services market dwarfs software
https://foundationcapital.com/ai-service-as-software/
Andreessen Horowitz (Alex Rampell) — Input Coffee, Output Code: How AI Will Turn Capital into Labor [12]
Capital buys coffee, engineers and GPUs; out comes code that takes the role of labor — enterprise software spend is infinitesimal against the white-collar labour market
https://a16z.com/ai-turns-capital-to-labor/
Microsoft Learn — Metadata scanning overview — Microsoft Fabric [15]
Scanner APIs catalogue tenant metadata including table and column names, measures, DAX expressions and mashup queries as subartifact metadata, without row-level data
https://learn.microsoft.com/en-us/fabric/governance/metadata-scanning-overview
Microsoft Support — Remove hidden data and personal information by inspecting documents, presentations, or workbooks [16]
Workbooks can contain hidden rows, columns and worksheets, invisible objects and custom XML data not visible in the document itself
https://support.microsoft.com/en-us/office/remove-hidden-data-and-personal-information-by-inspecting-documents-presentations-or-workbooks-356b7b5d-77af-44fe-a07f-9aa4d085966f
Office of the Australian Information Commissioner — What is personal information? [17]
What is personal information varies depending on whether a person can be identified or is reasonably identifiable in the circumstances
https://www.oaic.gov.au/privacy/your-privacy-rights/your-personal-information/what-is-personal-information
Major Consulting Firms
KPMG Australia — KPMG releases annual impact report [7]
Consulting revenues down 18% for the year, driven by reduced government use of consultants; division rebalancing toward technology transformation and AI
https://kpmg.com/au/en/media/media-releases/2025/08/kpmg-releases-annual-impact-report.html
Boston Consulting Group (via PR Newswire) — BCG Reports $14.4 Billion in Revenue, Marking 22nd Consecutive Year of Growth [8]
Revenue rose to $14.4B in 2025; AI- and tech-focused services above 40% of total revenue, driven by 25% year-over-year growth in AI services
https://www.prnewswire.com/news-releases/bcg-reports-14-4-billion-in-revenue-marking-22nd-consecutive-year-of-growth-302751073.html
Business Insider (syndicated via Yahoo Finance) — AI is reshaping how McKinsey makes money [9]
About a quarter of McKinsey's global fees come from performance-based pricing; Kate Smaje: many fundamentals of the professional services model are coming under challenge
https://finance.yahoo.com/news/ai-reshaping-mckinsey-makes-money-195132745.html
McKinsey & Company — The value of getting personalization right—or wrong—is multiplying [13]
Companies that grow faster drive 40 percent more of their revenue from personalization than slower-growing counterparts
https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/the-value-of-getting-personalization-right-or-wrong-is-multiplying
About This Reference List
Compiled August 2026. All URLs verified at time of compilation. Regulatory documents and standards specifications are subject to revision — check primary sources for the most current versions.
Some links to academic papers and vendor research may require free registration. Government and standards body publications are freely accessible.