Build vs Buy · Verification · Second Edition

Don’t Buy Software, Build AI Instead

Cheap Code Was Never the Argument

Scott Farrell

LeverageAI — leverageai.com.au

Second edition · July 2026

After Reading This Ebook, You Will:

  • Know why custom software actually died — and why that determines what has and hasn’t changed
  • Itemise the four products inside a subscription and see which two AI repriced
  • Run the Leverage Test and the 50% diagnostic on one platform — and read your configuration as an exit specification
  • Build a verification harness a non-developer can read — and know the two things it does not cover
  • Leave with a one-page decision sheet, a counter-case at full strength, and five corrections to the first edition

TL;DR

01
Part One · Why the Old Rule Existed

The Slogan Went Mainstream. The Reasoning Didn’t.

The production cost of software is falling. The price of software keeps rising. The gap between those two facts is where this decision lives — and almost everybody is measuring the wrong side of it.

Two things are true at once, and both of them have dates attached.

The first: it has never been cheaper to make software. In December 2025, Andrej Karpathy — founding member of OpenAI, former director of AI at Tesla, a man with no vendor to sell you — posted that he had never felt this far behind as a programmer. “I have a sense that I could be 10X more powerful if I just properly string together what has become available over the last ~year,” he wrote, “and a failure to claim the boost feels decidedly like skill issue.”1 That is not a prediction. It is a confession, from the top of the field, that the ground has already moved.

Theo Browne, who runs production engineering teams, put operational numbers underneath it: “I am writing the majority of my code with AI now. I would say more than the majority, like 90%. And for the teams that I run, we’re at at least 70% AI generated code.” His view on timing was blunter still — “getting into it now isn’t getting into it early anymore. Getting into it now is getting into it late.”2

The second thing that is true: software has never cost more to buy.

The two ledgers, 2026

Making software — falling
  • • Karpathy, Dec 2025: “10X more powerful” sitting unclaimed on the table.
  • • Theo Browne: ~90% of his own code, 70%+ across the teams he runs.
  • • A third of surveyed builders have already replaced a bought tool with a built one.
Buying software — rising
  • • SaaS pricing up ~11.4% year on year against 2.7% G7 inflation.
  • • 79% of IT leaders met a price increase at renewal in the last twelve months.
  • • Software is the fastest-growing line in Gartner’s 2026 IT forecast, up 15.1%.

The prediction became a measurement

In January of this year I published an argument that said, roughly: the economics have inverted, stop renting software you could own. It was a forecast. Six months later it is a description.

Retool surveyed 817 builders in late 2025 and published the results in February 2026. Thirty-five per cent had already replaced at least one SaaS tool with something they built themselves. Seventy-eight per cent expected to build more of their own tools during 2026. Sixty per cent had shipped something outside IT oversight in the preceding year, and a quarter did it routinely.3

Now the caveat, in the same breath rather than in a footnote at the bottom of the page: Retool sells a tool for building things. The sample is Retool builders and customers — self-selected, already equipped, already inclined. Those numbers are evidence of direction among people able to act. They are not a population rate, and anyone who quotes them as one is doing what this book exists to argue against.

Directionally, though, there is no ambiguity. The behaviour is real, it is not marginal, and it is happening in organisations that have not read a single word about it.

Meanwhile the other ledger got worse in ways you can date. SaaS pricing rose roughly 11.4% year on year against an average G7 inflation rate of 2.7%.4 In a 2026 survey of IT leaders, 79% met a price increase at renewal, 78% met unexpected charges tied to consumption or AI features they had not asked for, and 61% cut planned projects because of unplanned software cost increases.5 Gartner’s 2026 forecast has software as the fastest-growing segment of IT spending, up 15.1% on the year.6

Software is supposedly becoming free. The software line in your budget keeps going up. Both of those sentences are true, and the space between them is the subject of this book.

The provocation, and the immediate bound

The version of this argument that travels is short and it belongs to a note I wrote to myself before any of the research existed:

Don’t buy software — recreate it. Don’t hire experts — write software for them.

There was a third line in that note: anything you’re paying for, replace with your own software.

As a provocation it is excellent. As a rule it is false, and I would rather say so on page one than let you find out in month seven. Some software should be bought forever, and the reason is not caution or sentiment or nobody-ever-got-fired conservatism. It is that what you are buying from those vendors is not code. Chapter 4 itemises exactly what it is instead.

Keep the provocation. It is doing the job a provocation should: it makes you look at a renewal you had stopped looking at. Just do not mistake it for the instrument.

What’s wrong with the popular reasoning

The version of this argument circulating everywhere is a cost argument. Code got cheap, so build. Every conference talk, every LinkedIn thread, every vendor-displacement pitch runs on it. My own first edition ran on it.

It is aimed at the wrong variable, and there are three specific reasons — each of which is a chapter later in this book.

Three reasons the cost argument fails

  1. It is aimed at a variable that never decided this question. Custom software did not die of build cost, and knowing what it actually died of tells you exactly what has and has not changed. Chapters 2 and 3.
  2. It treats a subscription as one product when it is at least four. AI has repriced two of them to near zero and left the other two entirely intact. Chapter 4.
  3. It gives you no way to know whether the thing you built is right. Which is the entire problem, historically, and the one place the evidence in 2026 is genuinely alarming. Chapter 9.

Here is the practical damage. A right answer held for a wrong reason does not announce itself. It fails silently, at the boundary, in both directions at once: it tells you to build things you have no way to verify, and it tells you to keep paying for verification you could now produce yourself. Both of those errors feel like prudence while you are making them.

Key Insight

Two organisations with identical cost spreadsheets should reach opposite conclusions, if one can specify and verify and the other cannot — and a spreadsheet has no column for that.

That sentence will recur. It is the reason the five-year total-cost-of-ownership model, which everybody has and everybody trusts, is not the instrument for this decision. It is a perfectly good model of the wrong thing.

The reader contract

This is a commercial argument about money, which means it is a genre in which fabricated statistics breed. Every buy-versus-build piece you have ever read has a table in it. Tables demand numbers. When the numbers do not exist, the temptation is overwhelming to supply plausible ones — and plausible ones are indistinguishable from real ones until somebody checks.

So, the contract for this edition:

How numbers work in this book

  • Every figure carries a named source and a date in the sentence that uses it, not in a bibliography you will never open.
  • Where a source does not exist, the book states the shape of the claim and says explicitly that it is doing so.
  • Where a number is our own framework rather than a measurement, it is presented as our analysis and never dressed up as research.
  • No LeverageAI page is ever cited as the source of a statistic. Citing yourself for a number turns an assertion into a footnote and changes nothing about whether it is true.

And the uncomfortable half of that contract: the first edition of this book broke every one of those rules. It cited a collapse in custom development cost to our own website. It ran a platform amortisation ladder — first project this much, second project that much, third project less again — in which not one figure had a source. It quoted win rates that came from nowhere at all.

Those are named individually and retracted in Chapter 6, where the business case lives, and consolidated into a ledger in Chapter 15. I am flagging them here so that nothing in this book reads as a claim I have not already audited.

What this book actually claims

The compressed version, which I will not repeat again until the end:

Stop pricing the build. Price the specification and the proof.

The route to it runs like this. Chapter 2 re-derives why custom software actually lost for thirty years, and shows that the standard explanation cannot account for the period it describes. Chapter 3 names what changed — two things, together, and only one of them is the one everybody talks about. Chapter 4 itemises what you are really buying from a vendor, which is where “keep buying” becomes a positive doctrine rather than an admission of defeat.

Part Two builds the instrument: where the customisation dollar actually compounds, what belongs on the asset line, and how to route the decision by how much time the work can be given to think rather than by which software category it falls into. Part Three runs it — extracting the specification you already own, building the harness that proves the result, choosing between forking and regenerating, and deciding what to encode rather than hire. Part Four asks whether the decision is even strategy, and what is shifting underneath it while you run the analysis. Part Five makes the strongest available case against the whole thing and grades this book’s own evidence rather than yours.

What you should get out of it is an instrument, not enthusiasm. If you finish this book more excited than you started and no better equipped to decide, I have wasted your afternoon.

Where we start

To know what has changed, you have to know what actually broke.

And the standard story about what broke — custom software was too slow and too expensive, so the packaged product won on price — cannot explain the thirty years it claims to describe. It cannot explain why the most widely used custom-software platform in history was thriving in every office on earth during the entire supposed death of custom software.

There is a better explanation. It is less comfortable, it is about trust rather than money, and it predicts precisely what would have to change before building your own became sensible again.

Key Takeaways

  • The behaviour is measured now, not predicted — but the best available numbers are seller-side, and that matters.
  • Production cost is falling while purchase price rises; the decision lives in that gap.
  • “Code got cheap, so build” is a cost argument aimed at a variable that never decided this question.
  • Every number in this book has a name and a date. Where one does not exist, the book says so.
02
Part One · Why the Old Rule Existed

It Died of Verification, Not Cost

For thirty years, building your own software was career suicide. Everybody knows why. Almost everybody is wrong — and getting the cause right is the whole game.

A good friend of mine — someone I had advised for years — once tried to build his own manufacturing system with custom software. He wasn’t a developer. He couldn’t lead the project. And he got totally ripped off.

That is not a rare story. It is roughly what happened to most organisations that tried the same thing. For about three decades, building your own software was career suicide, and everyone in the room knew it. You bought the package. You didn’t build.

I want to argue that the reason that era died is not the reason most people think — and that getting the cause right is the whole game, because the exact thing that killed custom software is the exact thing AI just reversed.

The standard story, and why it explains nothing

The standard story goes: custom software was too expensive and too slow to build, so packaged software won on price and speed. Spread the development cost across a thousand customers and each of them pays a fraction. Simple economics, sad but inevitable.

That story is not false. It is incomplete, and being incomplete is precisely why it cannot explain anything useful. It produces no prediction about when the situation would change. It offers no test for whether it has changed. And — as we will see at the end of this chapter — it flatly contradicts one of the largest facts of the period it claims to describe.

Key Insight

Custom software didn’t lose on build cost. It lost on verification cost.

The problem was never the price of the code

Go back to my friend and his manufacturing system. His problem was never that code was expensive. His problem was that he couldn’t tell whether he was getting what he paid for.

Look at the structure of what he was facing, because this is the load-bearing part of the whole book. A non-developer commissioning custom software confronts an unsolvable principal-agent problem:

  • They cannot inspect the work.
  • They cannot tell a good developer from a confident one.
  • They cannot tell six months of progress from six months of invoices.
  • Every status update is a claim they have no independent way to check.
  • By the time the gap between the claim and the reality becomes visible, the money is gone.

That is not a cost failure. It is a trust failure — and it is structural, not moral. It does not matter how honest your developer is if you have no independent way to verify the honesty. Good faith is not a control. The buyer in the middle is flying blind, and blind buyers get taken.

Notice what the word is. Not “misaligned incentives”, which is management-speak for the same thing with the teeth pulled out. The word is unverifiable. He could not check. Nobody in his position could check. That is the condition the entire era was organised around.

So why did the package win?

Once you see the death as a verification problem, the winner makes sense in a way the cost story never quite manages.

Packaged software won as a trust technology at least as much as an economic one. You could see the finished product before paying. You could run a demo on your own data. And — the piece that mattered most and got discussed least — ten thousand prior customers had already de-risked it. They had, collectively and at their own expense, done your verification for you.

“Nobody got fired for buying SAP” was never about SAP being good. It was about the buyer’s inability to verify anything else.

The old line — nobody ever got fired for buying IBM, then SAP, then whichever incumbent held the seat — is usually read as a joke about corporate cowardice. It isn’t. It is a precise description of how buyers manage verification risk when they cannot verify. If you cannot check the work yourself, you outsource the checking to the crowd. “Everyone else runs it” is the verification.

Myth vs Reality

The myth
  • • Custom software was too expensive to build.
  • • Packaged software won on amortised development cost.
  • • Prediction: bespoke returns when bespoke gets cheap.
  • • Test it against thirty years of end users building their own tools in spreadsheets, and it fails.
The reality
  • • Custom software was impossible for the buyer to check.
  • • Packaged software won as a trust technology — the crowd did your verification.
  • • Prediction: bespoke returns wherever the builder and the verifier are the same party.
  • • Test it against the same thirty years, and it holds.

The consequence closes the loop neatly. The package did not have to be the best tool. It only had to be the one whose correctness someone else had already paid to establish. Which is exactly why “build it yourself” stayed radioactive for three decades. It was never that the building was impossible. It was that the checking was impossible for the person who most needed to do it.

It is worth sitting with how much work that arrangement did. The vendor was not merely selling functionality; they were selling a position in an argument you could not otherwise win. When the audit committee asked whether the system was correct, “it’s the same system four hundred companies in this sector run” was a real answer. “Our developer says it’s finished” was not.

The fact the cost story cannot survive

Here is the test, and it is the reason I am confident about the diagnosis rather than merely attached to it.

Custom software supposedly died for thirty years. During those same thirty years, the most widely used custom-software platform in human history was flourishing in every office on earth.

VisiCalc arrived in 1979 — the first spreadsheet program for personal computers, and widely credited as the application that turned the microcomputer from a hobbyist’s toy into a serious business tool.7 Then Lotus 1-2-3. Then Excel. What all of them did was let end users materialise bespoke logic without developers. That is the whole trick, and it never stopped working.

Corporate IT then spent three decades trying to stamp out shadow spreadsheets and never could. Policies, platforms, governance programmes, data-quality initiatives, entire consulting practices. None of it worked, and the reason it did not work is the reason this book exists: the person with the problem kept building the tool anyway, because building it themselves was the only path that did not route through a developer they could not verify.

The two predictions, scored

The cost story predicts: bespoke logic dies when bespoke is expensive, and returns when it gets cheap.

The verification story predicts: bespoke logic thrives wherever the builder and the verifier are the same person, regardless of cost.

Only one of those matches thirty years of evidence. The spreadsheet was the one place ordinary people got to be build-first the entire time — not because it was cheap, but because the person who wanted the answer could see whether the answer was right.

Hold on to that, because it carries a warning as well as a precedent. What the spreadsheet lacked was never power — you can build astonishing things in Excel, and people have destroyed companies doing exactly that. What it lacked was a harness. No tests, no version control, no review, no way to prove that this month’s workbook still does what last month’s did. Shadow spreadsheets are simultaneously the precedent for everything in this book and its cautionary tale, and Chapter 9 is where that debt comes due.

Why I am confident about this

I am not theorising from the sidelines. On my very first accounting job I turned up to books of paper and managers with pencils, and I typed it all into one terminal of an eight-terminal Unix accounting box. It was hilarious. A year or so later I led the charge onto spreadsheets, because doing it on a proper computer seemed more fun. Then I watched spreadsheets explode, then Visual Basic, then PC apps — a genuine build-first era, the last one we had.

So when I say we are coming back around to custom software, I am a two-cycle witness. I was in the room for the first one. I watched what actually killed it, and it was not the price of the code.

The question this leaves

If the cause of death was verification, then the question for the next chapter is not has code got cheaper. Obviously it has; everybody can see that; it is the least interesting fact in the argument.

The question is: has verification changed?

Because if it has not, then what we are watching is a circle rather than a spiral, and every organisation about to build should expect my friend’s outcome delivered at higher speed and lower unit cost. Cheaper failure is still failure.

Key Takeaways

  • The buyer’s inability to check the work, not the price of building it, is what made custom software radioactive.
  • Packaged software won as a trust technology: ten thousand prior customers did your verification at their own expense.
  • “Everyone else runs it” was the verification — which is why the incumbent only had to be safe, not best.
  • The spreadsheet never lost, because the person who wanted the answer could see whether it was right. That is the whole mechanism, in miniature.
03
Part One · Why the Old Rule Existed

A Spiral, Not a Circle

Two things changed at the same time. Every article about this covers the first one. Almost none of them names the second — and only the second licenses a different outcome.

Let me put the objection at full strength before I answer it, because it is a good objection and the people making it are usually the people worth convincing.

If this is a circle — the same conditions coming round again, cheaper — then you should expect exactly the disaster from the last chapter, and you would be right to. A cheaper way to commission software you cannot check is a faster way to be ripped off. My friend’s manufacturing system would fail again. It would fail for less money and in less time, which is something, but nobody writes a book to help you lose faster.

So the burden here is specific. It is not enough to show that building got cheaper. Something has to have changed about the checking.

Key Insight

AI’s return of custom isn’t a circle. It’s a spiral — arriving at the same point from the opposite cause.

Change one: the loud one

Build cost collapsed. You already believe this, so I will be brief.

The shape of it is what matters, not the magnitude. The “developer” now costs almost nothing to fire, iterates in hours rather than quarters, and can be run in as many parallel instances as you have tabs open. That is a change in kind, not degree — the constraints that shaped every software decision of the last thirty years were built around a scarce, expensive, serial, human resource, and one of those four adjectives is now optional and three are negotiable.

I am deliberately not adding statistics here. Chapter 1 has Karpathy and Theo Browne, and Chapter 15 explains at length why I think the developer-productivity literature — in either direction — was never capable of settling a procurement question. If you feel the absence of a number in this section, that is the argument working.

One thing this change genuinely buys, and it is worth naming before we move on: a bad idea now surfaces in a night instead of a year. Hold that. It comes back at the end of the chapter, and it is doing more work than it looks like.

Change two: the one nobody names

The second change is the one that actually matters and almost nobody names: verification got restructured.

In the old world, verifying custom software meant trusting a person you could not audit. In the new world, you verify behaviour with tests.

You don’t trust the AI. You trust the tests.

The mechanism has a name — characterisation testing — and it is simple enough to state in one sentence: you treat the running system as its own oracle, capture what it actually does for known inputs, and turn those input/output pairs into executable, falsifiable assertions. Pass or fail is binary. It is not a matter of opinion, and it is not a matter of reading code.

The mechanics belong to Chapter 9, along with the two things this instrument provably does not cover. What this chapter owns is the structural claim, which is bigger than any test suite:

The substitution

What packaged software rented you was a crowd of ten thousand prior customers who had already done your checking.

What replaces it is a test suite you own and can read as a percentage climbing toward 100. Same job. Different technology. That is the spiral.

Consider what that does for the buyer in the middle — the person Chapter 2 left flying blind. It gives them an independent instrument. Not a better relationship with the developer, not a stricter contract, not a more detailed statement of work. An instrument, which they hold, which returns a number, and which does not care how confident anyone in the room sounds.

Put the two eras side by side and the shift is legible in one pass.

  Custom 1.0 (the era that died) Custom 2.0 (the era returning)
What was expensive Writing the code Knowing precisely what you want
What could be checked Nothing the buyer could inspect Behaviour, against a harness the buyer owns
Who carried the risk The buyer, blindly The buyer, with an instrument — and it sits in the specification now
Where failure surfaced In production, or in the invoices In the specification, before anything ships
Time for failure to become visible A year A night
What the buyer held afterwards A codebase nobody could read and one departed developer A specification and a test suite — and the code is regenerable
The gating skill Can you manage developers? Can you specify what you want?

That is the thing my friend never had: a way for a non-developer to verify a system’s behaviour without being able to inspect a single line of it.

Why the conjunction is the argument

Run the counterfactuals, because they are clarifying.

One change without the other

Cheap generation, no verification instrument

The honest advice would be: don’t. You would be producing more of exactly the thing that historically destroyed buyers, at higher volume, with the same inability to tell whether any of it is right. Chapter 9 has dated evidence for what that looks like in practice, and it is not encouraging.

Verification instrument, no cheap generation

You would have a better procurement process and the build economics of 2015. Genuinely useful — a harness improves any custom project — but not a category change. The economics still say buy for most of the portfolio.

It is the conjunction that moves the decision. Say it as a conjunction, and refuse to let either half carry the argument on its own. Half the commentary in this space takes the first column and calls it a revolution. The occasional sceptic takes the second and calls it a nothing-burger. Both are describing one leg of a two-legged thing.

What actually stayed the same

Now the honest part, and it belongs here in Part One rather than in a caveat somewhere near the end.

The ripoff risk did not disappear. It migrated up the stack, into the spec.

My friend would still fail today if he could not articulate what his manufacturing system was supposed to do. The difference is that now he would fail cheaper and faster — a bad specification surfaces in a night of generation instead of a year of invoices — and that speed is genuinely most of the improvement. But the failure mode itself relocated. It did not get solved.

The gating skill moved with it: from can you manage developers to can you specify what you want. You no longer need software engineers so much as you need people who can design what they want precisely enough that the software can materialise from it.

The bottleneck isn’t execution anymore. It’s goal formation.

I feel this personally every working day. I can do an almost unlimited amount of coding in a day — it writes, it tests itself, it debugs, it drives a browser to check its own work — and the only thing limiting the output is the ideas and the features I can think of clearly enough to ask for.

The consequence for the decision

Which brings us to the payoff of Part One, and the sentence the rest of this book is built on:

Bottom Line

Two organisations with identical cost spreadsheets should reach opposite conclusions if one can specify and verify and the other cannot.

A cost model has no column for that. It cannot represent it, cannot weight it, cannot be adjusted to include it. Which is why the five-year total-cost-of-ownership comparison — the artefact every procurement function in the world will produce for this decision — is not the instrument. Part Two builds a different one.

And because it is the single most useful sentence in this book, here it is long before the decision sheet at the end, so you can act on it this week:

The stopping rule

If you cannot write the tests, keep buying. That is not a failure of nerve, and it is not a temporary setback to be worked around with a bigger contingency. It is the instrument returning an answer, and it is the correct answer.

Where this leaves the buy side

Part One has established what killed custom software and what has changed since. What it has not done is itemise what you are actually paying a vendor for.

That matters more than it sounds, because the moment you accept that verification is now producible in-house, the obvious next thought is that the vendor is selling you something you no longer need. For one of the four things in that subscription, you would be right. For two of them, you would be catastrophically wrong — and the failure shows up at three in the morning in month seven, not in the business case.

Key Takeaways

  • Two changes happened together: build cost collapsed, and verification was restructured into tests the buyer owns.
  • Only the second one licenses a different outcome from 1995. Cheap generation alone would make the old failure faster, not rarer.
  • The risk did not vanish. It moved up the stack into the specification, where it now surfaces in a night rather than a year.
  • If you cannot write the tests, keep buying.
04
Part One · Why the Old Rule Existed

Four Products in One Subscription

You renew it every year. You have never itemised it. AI repriced two of the four things inside it to nearly nothing and left the other two exactly where they were.

Your invoice says one thing. Your budget line says one thing. Your renewal conversation is about one number.

What you are buying is at least four separate products, bundled so thoroughly that nobody in the organisation has ever priced them apart — and they have completely different exposure to what changed in the last chapter.

The decomposition

What’s in the subscription What it really is Repriced by AI?
Operational depth Global infrastructure, 24/7 incident response, fraud detection at scale, carrier and regulator relationships, disaster recovery No
Liability transfer Compliance certifications, audit posture, penetration testing, someone else’s name on the failure No
Verification by crowd Ten thousand prior customers who already established that it works Yes — replaceable
The software itself Fields, workflows, rules, reports, screens Yes — collapsed

Operational depth is the one people underestimate most reliably. The value isn’t in the code — it’s in the operational muscle. Global infrastructure presence. Compliance certification programmes that took years and a dedicated team. Real-time fraud and abuse detection that only works because the vendor sees traffic from everyone at once. Carrier and network relationships that cannot be bought at your volume. Building that would cost more than subscribing forever, and it would still be worse, because most of it is not a build problem at all.

Liability transfer is the one build advocates pretend is not a product. It is. When the vendor’s certification is in your audit pack, you have purchased a position in an argument you would otherwise have to win yourself, annually, in front of people whose job is to be unimpressed. That has a price. The price is not zero, and it does not appear anywhere in a total-cost-of-ownership model because it is not a cost — it is an avoided risk that only becomes visible on the day you no longer have it.

Verification by crowd is the row Chapter 2 was about. This is the one that just became replaceable, for the first time in thirty years, by an instrument you can own and read yourself.

The software itself is the part everybody talks about, and for the platforms that actually hurt it is the smallest line in the bundle.

Key Insight

Both failure shapes come from treating the bundle as one line item: over-buy and you pay for four when you need two; over-build and you quietly cancel the depth and the liability transfer without noticing until something breaks.

Three tiers, and only one of them is moving

The decomposition explains why. The tiers tell you where.

Tier 1 — Commodity plumbing. Keep buying.

Examples: authentication, payments, communications, CDN and infrastructure, basic storage.

Signature: real-time critical, deeply operational, cheap per unit relative to value delivered, effectively zero customisation — you are plugging in, not fitting in.

Why it doesn’t move: the vendor’s scale is a genuine advantage, not a story. Nobody should be generating their own payment fraud detection.

Tier 2 — Horizontal platforms. The flip zone.

Examples: CRM, IT service management, HR platforms, mid-market ERP, marketing automation.

Signature: expensive base cost, large customisation projects, internal administrators whose entire job is the platform, and “implementation partners” on retainer.

The uncomfortable naming: you are paying for translation, not software — converting “what you want” into “what the platform allows.”

Tier 3 — Industry verticals. Also the flip zone, with extra deception.

The promise: “Built for your industry.”

The reality: built for the average of your industry. The core platform is largely generic; the vertical is a skin over it.

The trap: the vendor optimises for their 80th-percentile customer, so everything that makes you distinctive is, by definition, customisation work — at premium pricing, justified by industry expertise.

There is a second-order effect on Tier 3 worth stating plainly, because it inverts the intuition. Fewer alternatives exist inside a vertical, which means less negotiating leverage. So the customer with the most specific needs — the one who most needs to move — has the least room to move. That is not an accident of the market. It is the market working exactly as designed.

Run those three questions on your own stack before you read any further. Most organisations discover the answers are not evenly distributed — there is usually one platform where all three answers are bad, and everyone already knows which one it is. They just have not been given permission to say it out loud in a meeting.

This is not anti-SaaS, and I want to be specific about why

The original SaaS promise was real and it was valuable: spread research and development costs across thousands of customers, get enterprise-grade security and compliance you could never build alone, receive continuous updates without a migration project, and rent operational expertise you would not otherwise have.

That promise still holds — for the categories it was designed for.

The failure was extension. Vendors took a model that works brilliantly for payments and authentication and applied it to customer relationship management and workflow. Those are categories where fit, not scale, determines the value, and where the vendor’s averaging is the problem rather than the product.

So let me bound the provocation from Chapter 1 explicitly rather than quietly walking away from it. “Anything you’re paying for, replace it with your own software” is a good provocation and a bad rule. This table is exactly where it stops. Two of the four rows are not going anywhere, and pretending otherwise is how build projects fail in ways that make the news.

What the vendors were actually selling

It is worth being precise about the moat, because it explains why it is pointed the wrong way for this particular attack.

Oracle did not win because their database was un-replicable. They won because they had a global sales force, an installed base and ecosystem lock-in. Salesforce did not dominate customer relationship management on technical superiority; they made it easy to buy and painful to leave. SAP won on the implementation partnership ecosystem rather than on product excellence.

Distribution was the moat. Not technology.

And distribution is a formidable moat against competing vendors. It is not a moat against a customer who has stopped needing a vendor for one specific thing. Nothing about a global sales force prevents an organisation with a specification and a test harness from building the workflow module they were going to configure anyway. The wall is enormous, expensive, and facing the wrong direction.

The one number I can defend

Here is where the first edition of this book had a table of dollar figures. Licence cost, implementation, internal administrators, consultant retainer, integration maintenance, total. It looked authoritative. Every number in it was invented.

So here is the single sourced magnitude I am willing to put in its place.

What enterprises actually spend on the software they buy

2–7×

licence cost, spent on implementation and integration

33%

average code reuse — teams rebuild about two-thirds of the functionality anyway

HFS Research with Unqork, 18 November 2025. Survey of 123 respondents at Global 2000–scale organisations, fielded September 2025.8

The study’s own framing is worth keeping, because it is the sentence that should govern every AI-assisted migration anyone attempts this year:

AI is not a silver bullet, it’s an amplifier of whatever already exists in your enterprise stack.
— Phil Fersht, CEO and Chief Analyst, HFS Research, in the same study

One sourced multiple is worth more than a page of confident arithmetic. And notice what it lets you claim honestly, as shape rather than magnitude: for configured platforms, the subscription is not the cost. The licence is the visible number and frequently the minority of the spend.

The precise ratio for your platform is not something a book can tell you. It is something you measure — and Chapter 5 is where you measure it.

Pitfall

The industry-vertical trap. “Built for your industry” produces a false sense of fit, which suppresses the customisation-burden question entirely. Nobody audits a platform that is supposed to already understand them. Meanwhile the customisation burden is often higher than a horizontal platform, because the vertical’s opinions are stronger and your deviation from its average is exactly where your margin lives.

What Part One leaves you with

You now know what killed custom software, what changed, and what you are actually buying. That is the doctrine. It is enough to stop you making the two obvious mistakes and not yet enough to make a decision.

The cost comparison will not decide it, and the tier table is a taxonomy rather than a trigger. Part Two supplies the instrument — starting with the question that replaces “is it cheaper to build?”, which turns out to be a question about where your money compounds rather than where it goes.

Key Takeaways

  • A subscription bundles four products. AI repriced two of them and left operational depth and liability transfer untouched.
  • Tier 1 commodity plumbing stays bought — the product is operational muscle, not code.
  • “Built for your industry” means built for the average of your industry, and your advantage is your deviation from it.
  • Distribution was the moat, and distribution does not defend against a customer who has stopped needing a vendor for one specific thing.
  • For configured platforms, the licence is not the cost. That is shape; the ratio is something you measure.
05
Part Two · The Instrument

Where Does the Customisation Dollar Compound?

The old question compared two prices at a moment. The new one compares two trajectories — and under monthly capability releases, the trajectory is the decision.

The old question was: should we build or buy this software? Which is cheaper upfront, which has the lower five-year total cost of ownership. That was exactly the right question when building was expensive, because when building is expensive the price is the risk.

The new question is different in kind, not in detail:

Where does your customisation dollar get AI leverage?

The swap matters because the old question compares two prices at a single moment, and the new one compares two trajectories. Under conditions where capability arrives on a monthly cadence, the trajectory is the whole decision. A dollar that compounds and a dollar that does not are different assets, even when they cost the same dollar.

Why the asymmetry is structural, not accidental

Money spent configuring a proprietary platform lives in closed configuration languages that agents can barely touch. The returns are linear and locked in. And here is the part that gets missed: model improvements deliver zero benefit to it. Next year’s model does not make your workflow rules better. It does not make your validation logic cleaner. The capability arrives and simply does not reach that part of your estate.

Money spent on spec-driven code does the opposite. The generation is largely automated, the returns compound, the artefact is portable, and every model release is a free upgrade the moment you regenerate.

Our own framework puts this at roughly 5% of the work receiving AI leverage on the platform side against roughly 80% on the code side — call it a sixteen-fold gap. That is our estimate, from our framework, not a measured external statistic, and I would rather say so in the sentence than have you discover it in a footnote. The precise ratio is arguable. The direction is not, and the direction is what the decision runs on.

The reason for the gap is worth dwelling on, because it is a genuinely dark joke. Platform vendors built proprietary configuration languages and closed ecosystems deliberately, as moats. Those moats were designed to make leaving expensive. They have turned out to also make improving expensive. The same wall that keeps you in keeps the leverage out, and the vendors cannot lower it without dismantling the thing that made them valuable.

What agents can and cannot reach

With code
  • • Parses structure, types and dependencies
  • • Refactors while preserving behaviour
  • • Generates and runs tests
  • • Integrates static analysis and reads the errors
  • • Converts a requirement into an implementation
With platform configuration
  • • Cannot navigate configuration user interfaces
  • • Cannot reliably reason about undocumented proprietary formats
  • • Cannot run platform tests, because there are none
  • • Cannot predict cascade effects across workflow rules
  • • Cannot regenerate anything, because there is no spec to regenerate from

Two kinds of debt, two different futures

Technical debt is not one thing. There are two kinds here and they have completely different trajectories, and the difference is not about code quality — it is about what tooling exists to manage the mess.

Dimension Platform configuration debt Codebase debt (with discipline)
Version controlSnapshot restores at bestMature: branching, merging, diffs
TestingManual and hopefulAutomated and systematic
RefactoringPoint-and-click archaeologyStructured and tool-assisted
AI assistanceMarginalSubstantial
DocumentationTribal knowledgeCode plus commit history
Change cost over timeEscalatingFlat or decreasing
Model upgrade benefitZeroFree improvement on regeneration
Error recoveryRestore a snapshot, if anyone remembersRevert a commit
Knowledge transferKey-person dependencyThe codebase speaks for itself

The historical point underneath that table is the one that makes it more than a comparison. Version control, automated testing, continuous integration — none of these appeared by magic. They emerged from code complexity, over decades, because the complexity became unbearable and people built tools to survive it.

Platform configuration inherited all of the complexity and none of the solutions. Fifty years of accumulated engineering discipline on one side; zero on the other. The configuration is every bit as intricate as the code it replaced — and it is intricate in an environment with no diffs, no tests, no rollback and no refactoring.

There is a pointed aside here that I will make once and not chase: Agile was never really a solution to complexity. It was a coping mechanism for “this is too hard to reason about, so let’s do a small piece and see.” Which raises an interesting question about what happens to the coping mechanism when the complexity becomes manageable again. That is a different book.

The two trajectories

Where each kind of debt takes you

Platform configuration

  • Year 1: clean, understood by the implementers, changes are fast and predictable.
  • Year 3: original implementers gone. Archaeology begins. “We can’t touch that” enters the vocabulary.
  • Year 5: major excavation for any change. One platform whisperer. Shadow documentation everywhere.
  • Year 7: ossification, and a re-implementation proposal — on the same platform, restarting the cycle.

Debt compounds forever, because nothing exists to pay it down.

Code, with a specification and a harness

  • Year 1: clean implementation, tests pass, changes are fast.
  • Year 3: normal debt accumulation, addressed by refactoring; the tests keep validating.
  • Year 5: cleaner than year 3, because regeneration keeps getting cheaper and model improvements arrive free.
  • Year 7: the specification is the asset; the implementation is a rendering of it.

Debt is manageable — conditional on the specification and the harness existing.

The second trajectory is not automatic. It is bought, with the artefacts in Chapters 6, 8 and 9.

State the condition loudly, because it is the honest half of the argument and it is where most of the failures live. The good trajectory is conditional on the specification and the harness existing. Without them, a self-built system tracks the first trajectory exactly — year three archaeology, year five key-person dependency, year seven ossification — with the vendor’s name replaced by yours and nobody to escalate to.

If you recognised four or more of those, you do not have a platform. You have an artefact whose behaviour is now determined by history rather than by design, and the person who understands the history is a resignation letter away from taking it with them.

The 50% diagnostic

Here is the trigger, and it is deliberately crude, because a crude instrument you actually run beats a sophisticated one you don’t.

If customisation burden exceeds 50% of the total cost of running a platform, run serious build analysis.

What counts as the “buy” component: the licence or subscription. That is it. What counts as customisation burden: implementation amortised over its useful life, consultants and contractors, internal administrator and developer time, training and certification, integration development and maintenance, and change-request overhead.

Customisation burden as a share of total platform cost

<30%

Clear buy

30–50%

Evaluate carefully

50–70%

Strong build candidate

>70%

Already building custom — badly

This is a diagnostic, not a verdict. It tells you where to spend the analysis; it does not tell you the answer. A platform sitting at 65% with nobody in the building who can own a specification is still a buy, and Chapter 14 explains why that is the method working rather than the method failing.

What “over 70%” actually means

Sit with the top band for a moment, because it is the sting in this chapter.

At over 70%, you are doing custom software development. Not considering it. Doing it. You have data models, business rules, integration contracts, access policies and a release process. What you do not have is version control, automated tests, code review, a specification, or the word “development” anywhere in the budget line that funds it.

You have taken on every cost of building bespoke software and forfeited every discipline that makes bespoke software survivable. And because it is filed under “platform administration” rather than “engineering,” nobody has ever asked the questions that would surface the problem.

Key Insight

If AI can’t help you with it, you’re building the wrong way — and if the burden is over 70%, you are already building, just without any of the equipment.

Chapter 9 opens on the same disease being performed by very senior people at partner rates, which is where it gets genuinely expensive.

One platform, not the portfolio

A practical instruction to close on. Run this on one platform. Your most expensive one, or the one where four of the six symptoms landed.

The portfolio version of this exercise is a strategy offsite: it produces a matrix, a heat map, a set of recommendations and no decisions. The single-platform version produces a number, an argument and a next step. One of those changes what happens next quarter.

What you will have when you finish is a trigger and a diagnosis. What you still will not have is anything to put in front of a chief financial officer who wants a business case with numbers in it — which is Chapter 6, and it opens with a retraction, because that is precisely where the first edition of this book put numbers it had made up.

Key Takeaways

  • Replace “is it cheaper to build?” with “where does the customisation dollar compound?”
  • Configuration receives almost no AI leverage and zero benefit from model improvements. That is structural, and vendors cannot fix it without dismantling their moat.
  • Both kinds of debt compound; only one has fifty years of tooling to pay it down — and only if you own a specification and a harness.
  • Over 50% customisation burden: run the analysis. Over 70%: you are already building custom software without the equipment.
  • Run it on one platform. The portfolio version is an offsite, not a decision.
06
Part Two · The Instrument

Price the Specification, Not the Build

The first edition of this book put a table of dollar figures in exactly this position. Every number in it was invented. Here is what belongs there instead.

Let me get the retraction out of the way first, in detail, because a vague admission of imprecision is humility theatre and a specific one is an author who went back and checked.

The January edition of this argument claimed that custom development which would have cost half a million dollars five years ago now costs fifty to a hundred and fifty thousand. It ran a platform amortisation ladder — first project two hundred thousand, second eighty, third forty. It described a vendor charging forty thousand a year and a bid win rate moving from twelve per cent to thirty-four.

Every one of those numbers was invented. Not modelled. Not sourced. Not disclosed as illustrative. Some of them were cited to our own website, which is worse than having no citation at all, because a footnote pointing at yourself dresses an assertion up as research and makes it harder rather than easier for a reader to check.

I want to be honest about how it happened, because the mechanism is more useful than the apology. A build-versus-buy argument creates enormous pressure to produce a comparison table. A comparison table demands numbers in every cell. When the numbers do not exist — and for this decision, at this moment, most of them genuinely do not — the temptation to supply plausible ones is close to irresistible. That is how a real argument acquires a fraudulent skeleton, and it is why this chapter is structured the way it is.

Stop pricing the build and start pricing the specification.

What actually goes on the asset line

Start from a claim that sounds like an engineering detail and is in fact a procurement position: there is source before the source code.

Intent and prompts compile, through an agent, into files we still call source. Relative to the package above them, those files are already compiled output. We have a new, earlier compiler in the chain, and almost nobody has moved their definition of “the thing we own” to match it.

The asset line

On it
  • Intent — what this is for, and what “good” means.
  • Design and constraints — the decisions that cannot be safely re-inferred, especially the ones that look arbitrary.
  • Worldview context — the organisational facts, exceptions and vocabulary the system assumes.
  • Acceptance and characterisation tests — the executable definition of correct.
  • Starting state — the repository and environment the generation began from.
Not on it
  • • The generated implementation, once it can be regenerated.
  • • Line counts, module counts, commit counts.
  • • The particular framework or model that produced it.
  • • Anything whose value disappears when a better model ships.

The failure this prevents has a name and it is depressingly common. A team treats the repository as sacred. Every module committed, pull requests orderly, branch hygiene immaculate. Six weeks later nobody can answer a simple question: why does the retry path wait exactly that long, and would a second agent run make the same choice? The chat is gone. The agent instructions were never filed. The acceptance criteria lived in someone’s head, and that head has moved on.

We say we have source control. We have intermediate-representation control.

And here is why that is a buy-versus-build point rather than an engineering aside:

Key Insight

Software you cannot regenerate is not ownership. It is a maintenance liability with your name on it instead of a vendor’s.

You have not escaped the trap in that case. You have renamed it, removed the support contract, and taken on the operational burden yourself. That is strictly worse than the platform you left.

The Delete Test: how you know the asset is real

“We own the specification” is a claim, and claims in this territory have a poor record. So make it falsifiable.

The Delete Test

  1. Pick one module with real judgement in it — not a data-access layer.
  2. Delete the implementation.
  3. Regenerate from the retained upstream package alone.
  4. Compare behaviour against the harness.

Pass: equivalent behaviour regenerates. The code was intermediate representation, and your asset claim is honest.

Partial or fail: judgement existed nowhere upstream. The test has just told you precisely which judgement to promote into the specification — which is more useful than a clean pass, and is the reason to run it early rather than at the end.

Run it once a quarter on a rotating module. Not once, at the close of the project, when the answer is academic and everyone is too tired to act on it. The discipline is worth more than the result: a team that knows the Delete Test is coming writes different specifications.

Two years ago this would have been an expensive ritual. It is now cheap, which is itself part of why the specification-as-asset position is available at all.

The accounting argument, which is the one a CFO can hear

There is a version of everything above that does not require anyone to care about software architecture, and it is the version that gets budget.

A workflow accelerator is operating expenditure. Its value is welded to the lifespan of the process it accelerates. And here is the trap: AI is the very thing shortening process lifespans.

You are strapping a rocket to a workflow whose remaining life the rocket itself is busy reducing.

A compiled asset is substrate, not process. When the workflow is redesigned — and in this era it will be, repeatedly — the accelerator dies with the old workflow and the specification, the tests and the encoded judgement carry over and feed the new one.

Apply that directly to the decision in front of you. A build justified purely as “this platform costs too much” is an operating-expenditure argument, and it dies the moment the process it supports is redesigned. A build justified as “we will own the specification of how this part of the business actually works” is a capital argument, and it survives the redesign because the redesign consumes it as an input.

Same project. Same code. Completely different question about whether the money should be spent, and completely different answer about who should sponsor it.

Do not denominate this in hours saved

Almost every business case for a build is written in labour hours. It is the wrong denominator, for three separate reasons that compound.

  1. The hours are usually confetti. Twenty minutes here, an afternoon there, spread across forty people. They never consolidate into a line anyone can recover, which means the saving is real to the individuals and invisible to the accounts.
  2. The case is self-defeating. If you do consolidate them properly, the return is funded by removing the jobs of the people whose cooperation the project requires. It is the only business case whose beneficiaries are its opponents.
  3. The baseline is gameable. It is self-reported before-and-after, produced by people who want the project approved. Everyone in the room knows this, which is why nobody defends the number when it is challenged.

Denominate it in capability retained instead. What can the organisation now decide, change, prove or answer that it could not before? What does it own afterwards? How much cheaper is the next one?

On a build-versus-buy paper that takes concrete forms: a portable specification of a business process that did not previously exist in any single place; a behavioural test suite that did not exist at all; a named person accountable for the behaviour of the system; and a change that can now be made in a day rather than a quarter. Those are auditable. Hours saved are an opinion with a spreadsheet attached.

The numbers this book will not give you

Four claims that matter, stated as shape, with an explanation of why no magnitude follows.

Shape, not magnitude

  • Build cost has fallen sharply. The direction is well evidenced — Karpathy’s own account, Theo Browne’s production numbers, Retool’s measured behaviour. No dollar range is asserted, because no dated primary source supports one and the range in the first edition came from us.
  • The second build costs less than the first. Structural: the specification, the harness, the platform components and the accumulated judgement carry over. No amortisation ladder is asserted, because the ratio depends entirely on how much of build one was genuinely reusable — which is the thing organisations most reliably overestimate.
  • Seat-based pricing scales with headcount; a built system’s running cost does not. That is a property of the two pricing models, not a measured comparison. It matters most for organisations expecting headcount to grow faster than usage, and least for the reverse.
  • For configured platforms, the licence is not the cost. The one sourced magnitude here is the 2–7× implementation multiple from Chapter 4, and it belongs there rather than being restated as though repetition made it stronger.

The rule that panel encodes is the one worth carrying out of the book: a sourced qualitative claim beats an invented quantitative one. Every time. In a board paper, in a vendor negotiation, in a conversation with an auditor.

And if you need a number for your own paper — you probably do — then measure your own platform in Chapter 5’s terms. It is a week or two of work, the result is specific to you, and it will be more accurate than anything any book could have told you.

The outcome you should actually expect

One more thing belongs in the business-case chapter, because it changes how the analysis should be funded.

The highest-frequency good outcome of this whole exercise is not an exit. It is a costed, credible build alternative — which is the only thing that has ever changed a renewal conversation. Vendors do not respond to dissatisfaction. They respond to a customer who can describe, in specifics, what they would do instead and what it would take.

That means the analysis pays for itself before the decision is made, in every branch, including the branch where you stay. Fund it accordingly: as procurement preparation, not as a project pre-approval.

You now know what to price and what to own. What you still do not know is where inside the organisation a build actually pays — which turns out not to be a software-category question at all.

Key Takeaways

  • The first edition’s numbers in this position were invented. They are retracted, individually and by name.
  • The asset is the specification, the tests, the context and the decisions — not the generated code.
  • The Delete Test makes “we own the spec” falsifiable, and a partial pass is more useful than a clean one.
  • Accelerators are opex and die with the process. Compiled assets are capital and survive the redesign.
  • Denominate the case in capability retained. Hours saved is the case whose beneficiaries are its opponents.
07
Part Two · The Instrument

Route by Cognition, Not by Category

Where a build pays is not decided by what kind of software it is. It is decided by how much time the work is allowed to think.

Before I had a framework for this, I had a note. It was four lines long and it was better than most of what I have read since:

The target test

  • It has to be high value or it isn’t worth it.
  • Don’t compete with humans at things they’re already good at.
  • Help them do a better job.
  • Do the things they can’t do — too hard, too slow, not enough people.

That is the Cognition Ladder written in shorthand by someone who had not yet reached for the name. It arrived from looking at real targets, not from reading a model, which is the reason I trust it. Now let me give it the name and the structure, because the structure is what makes it operable.

Three rungs, sorted by time to think

The organising variable is how much time the work is allowed for cognition. Not the software category. Not the industry. Not the department. Time.

Rung 1 — Don’t Compete. Seconds; real time.

What it looks like: live chat, conversational interfaces, immediate decisions with a person waiting.

Why it fails most often: you are asking AI to beat humans in environments humans designed for themselves. The workflows were built around human strengths. The tools reflect human mental models. The edge cases were handled by years of accumulated judgement that nobody wrote down.

Build implication: this is where vendors are strongest and where their operational depth is genuinely the product.

Rung 2 — Augment. Minutes to hours; batch.

What it looks like: far more analysis, checking, synthesis or exploration applied to problems you already have, with humans keeping judgement and control.

The formulation worth keeping: instead of sampling, you check everything. One analyst reviewing fifty transactions becomes every transaction reviewed, with the exceptions surfaced for human attention.

Build implication: this is where most first builds should live. The work is reviewable, the errors are catchable, and the vendor’s averaging is the constraint you are removing.

Rung 3 — Transcend. Overnight.

What it looks like: work that was never rational to attempt, because the coordination overhead, the calendar time, or the political-acceptability filter killed it before anyone tried.

The distinction that matters: not the same work faster. Work that did not previously exist as an option.

Build implication: no vendor sells this, because no vendor can amortise it. If it existed as a product, it would not be your advantage.

The movement of the question is the useful part. It goes from “where can we put a chatbot?” — which is the question that produces failure — to “where is human thinking being wasted on repetitive or queued work?”, and finally to “what have we never attempted because the coordination or cognitive overhead was prohibitive?”

Most organisations never leave the first question. Most vendors are built to answer it.

A note on numbers, deliberately

This book uses the ladder’s structure and none of the magnitudes that usually travel with it. There is a failure-rate band and a return-per-dollar band that circulate alongside this framework in our own earlier work and in the market generally. Neither is independently sourced, so neither appears here. That costs me a big number in an appealing place, and it is exactly the contract I set in Chapter 1.

Tiers meet rungs

Chapter 4 gave you a taxonomy of what you buy. This chapter gives you the routing. Put them together:

Tier Where the value sits Typical rung Default
Tier 1
Commodity plumbing
Operational depth, real-time reliability Rung 1 Keep buying
Tier 2
Horizontal platforms
Your workflow, your rules, your fit Rungs 2–3 Examine seriously
Tier 3
Industry verticals
Your deviation from the industry average Rungs 2–3 Examine seriously

The reasoning is short. Tier 1 mostly lives where AI is weakest and the vendor’s operational depth is the product. Tier 2 and Tier 3 mostly sit where batch and previously-infeasible cognition pay — where a build multiplies your capability rather than the vendor’s.

But the table is a default, not a rule, and the corollary is what makes it useful:

Key Insight

The rung, not the category, is the decision. A “reporting” build that has to answer in 200 milliseconds on a customer-facing page is a rung-1 build no matter what the category table says.

What has to be true for a rung-2 build to work

Three conditions. They are conjunctive, and the failure diagnosis is the useful part.

  1. High volume. Infrastructure overhead needs enough tasks to amortise against. Fifty items a month does not justify the platform, regardless of how annoying those fifty items are.
  2. Latency flexibility. Batch, queued, or overnight. Removing real-time pressure is the specific thing that lets depth in — multiple passes, self-correction, adversarial checking, exhaustive rather than sampled coverage.
  3. Errors reversible or caught by humans. Either the work is genuinely low stakes, or there is a verification step, or mistakes surface downstream before they compound.

Absent all three, you have a rung-1 build wearing a rung-2 label, and it will produce the rung-1 outcome regardless of how good the specification is. That is not a warning about ambition. It is a warning about mislabelling, which is much more common and much harder to spot from inside.

The mistake that costs the most

There is a specific expensive error waiting at this point in the process, and it deserves its own section.

A construction company once sent us their tender workflow. Fourteen steps, documented. Three different software systems. Two manual handoffs where humans copied data between spreadsheets. Their question was: which two or three steps can we speed up with AI?

Wrong question entirely. The right question was: in a world where AI can read documents, understand context and coordinate work, why do you have fourteen steps?

They weren’t asking to win more tenders. They were asking to lose tenders faster.

Here is the build-versus-buy version of the same error, and it is the reason I am telling you a story that belongs to a different book. Rebuilding your vendor’s workflow in your own code is the same mistake with a bigger bill and your name on the maintenance.

You will have reproduced the platform’s shape faithfully — including every step that exists only because the platform needed it, every field that exists only because a consultant added it in 2019, every approval that exists only because the system could not model a condition. And now you own it. Forever. With no vendor to blame and no upgrade path that fixes it for you.

The correct sequence is: extract the specification first (Chapter 8), then interrogate it. Extraction gives you the requirements. It does not oblige you to keep them. The difference between assisting an existing workflow and redesigning the work around what is now possible is the difference between a cheaper platform and a different capability — and only one of those is worth the risk you are about to take.

Shadow IT is your demand signal

The cheapest source of build candidates is already in the building, and most organisations treat it as a discipline problem.

Retool’s February 2026 survey found 60% of builders had shipped something outside IT oversight in the preceding year, and a quarter did it routinely.3

The conventional read is a governance failure. The useful read is this: where people route around the sanctioned stack, the unmet need has already been validated at their own expense. They did the demand research for you and paid for it with their evenings. Nobody builds a tool in their own time to solve a problem they do not have.

So before you run any portfolio audit, ask what people have built for themselves in the last year. That list is a map of rung-2 and rung-3 candidates that have already survived a real cost-benefit test — the most honest one there is, because the person doing the cost-benefit was also the person paying the cost.

One thing this is not, and I will say it once rather than moralise about it: this is an argument for bringing those builds inside the harness, not for tolerating unverified systems in production. Chapter 9 is where that gets resolved.

What Part Two leaves you with

You now have a trigger (Chapter 5), an asset line (Chapter 6) and a routing (Chapter 7). That is the instrument, and it is enough to decide with.

What remains is running it — which starts with the most valuable artefact in your building, the one your vendor has been maintaining on your behalf for years without either of you noticing.

Key Takeaways

  • Route by time-to-think, not by software category. The rung decides, not the tier.
  • Rung 2 needs volume, latency tolerance and catchable errors. Absent all three, it is a rung-1 build with a rung-2 label.
  • Rebuilding the vendor’s workflow in your own code is losing tenders faster, with your name on the maintenance.
  • Shadow IT is a validated demand signal, not a discipline problem — and it was paid for by the people who needed it.
08
Part Three · Running the Decision

Your Configuration Is Your Exit Specification

The thing everybody files under lock-in is the cheapest requirements document your organisation will ever own — and somebody else has been maintaining it.

The vendor spent twenty years encoding every customer’s requirements into their platform — and it turns out they were maintaining everyone’s exit specification, at their own expense.

In the first edition of this book, vendor lock-in appeared under costs. The trap. The thing that keeps you paying.

It belongs under assets. That is not a softening of my earlier position — it is a considerably stronger claim than the one I originally made, and I was too cautious. The information was available in January. I filed it in the wrong column.

The moat was never the data export

Exporting a database has been easy forever. Whatever kept organisations in place, it was not the technical difficulty of getting the rows out.

The real switching cost was that re-specifying your particular slice was prohibitively expensive.

Look at what a platform vendor actually built. Thousands of features. Each individual feature used by a small minority of customers — but every customer uses a different minority. From the outside that looks like bloat, and it is regularly mocked as bloat by people who have never had to keep a platform sold.

It is the moat. The thousands of features were never waste; they were the encoded requirements of ten thousand different businesses, accumulated over two decades of implementations. Unpicking yours from the pile was the thing you could not afford to do, and so you stayed, and the annual increase went through, and everybody carried on.

Now run an agent at it.

The unpicking is exactly the kind of work that got cheap: reading structure and emitting structure. Your slice can be lifted straight out of the configuration, and the ninety-odd per cent you never touched can be left where it is. What was prohibitively expensive is now an afternoon and a metadata export.

Key Insight

Value migration with a punchline: the moat’s own artefacts are the extraction target.

Sunk cost is captured requirements

The standard view of sunk cost is that it is wasted money and you must not let it trap you. Correct as far as it goes, and useless here.

The reframe

Old thinking

“We can’t leave. We’ve spent two million dollars configuring this.”

The sunk cost traps you.

New thinking

“We have two million dollars of requirements documentation sitting in that platform.”

The sunk cost becomes your specification.

The money is spent either way. The only remaining question is what you extract from it.

And the extraction is not reverse engineering, which is the misconception that stops people starting. You are not inferring requirements from behaviour. You are collecting requirements that were already captured, in a structured form, by people who were being paid to get them right at the time.

What each piece of configuration actually documents

Configuration element What it actually documents
Custom fields Data requirements — the name, the type, the validation, and the fact that somebody needed to record it badly enough to ask
Workflow rules Business rules — a trigger, a condition and an action that somebody decided on and defended
Reports and dashboards Output specifications — groupings, filters, calculations; the questions the business actually asks
Integrations Interface contracts — field mappings, sync frequency, error handling
Validation rules Data-quality constraints, and the hardest-won knowledge in the platform — most were written after something went wrong
Page layouts and permissions Information architecture and access policy — who sees what, who may do what

Dwell on the validation-rule row. Those constraints are the most expensive knowledge in the platform and the least likely to exist anywhere else in the organisation. “Discount cannot exceed forty per cent without approval” is not a field setting. It is a pricing policy, an approval workflow, an audit position and a piece of institutional memory about something that went wrong once — compressed into a rule that nobody ever documented as a policy.

Lose that on the way out and you will rediscover it in production, expensively, in front of a customer.

The extraction, in four moves

The extraction protocol

  1. Configuration audit. Export the metadata — objects, fields, workflows, validation rules, reports, permissions, integrations. Most platforms have a metadata API or an export path, and if yours does not, that fact is itself an input to the decision. Goal: a complete picture of what has been built.
  2. Implicit to explicit. Configuration is implicit specification. Translate it into portable statements. “Custom field” becomes “data requirement with name, type and rules.” “Workflow rule” becomes “business rule with trigger and action.”
  3. Gap identification. Which configuration serves no purpose — the dead-code equivalent? Which requirements are missing because they live as tribal knowledge? Which rules conflict with each other? This audit routinely reveals how much of the complexity was never necessary, which is worth the exercise on its own.
  4. Specification document. Five sections: data model; business rules; integrations; views and reports; permissions and access. Platform-agnostic, machine-consumable, human-readable, and under version control from day one.

The division of labour matters, because it manages expectations about what the agent is doing here. Structure-to-structure translation, pattern recognition across configuration idioms, consistency checking and gap detection: genuinely a sweet spot, and fast. Intent clarification — why was it configured this way, what problem was it solving — priority setting, and future direction: still entirely human, and the reason this is a two-week exercise rather than a two-hour one.

The shadow spec — where extraction projects quietly fail

Now the honest limit, and it is where these projects go wrong in a way nobody notices until production.

The configuration does not capture everything. Tribal knowledge. Workarounds. “We always do it this way.” The exceptions the administrator handles by hand, and the ones they handle from memory. These live in people, in spreadsheets, and in message threads from 2023.

The failure is quiet because everything looks fine. The extracted specification appears complete. The regenerated system passes every test derived from the configuration. And then the first month of production surfaces twenty behaviours nobody ever wrote down, each of which somebody depends on.

Run those interviews before you finalise anything. They take a day per team and they are the difference between a specification and an inventory.

There is a second reason they matter, which connects forward. The specification tells you what was intended. The running system tells you what actually happens. Where those two disagree, the disagreement is usually the shadow spec — and Chapter 9’s harness has to be derived from observed behaviour, not only from the configuration export, precisely because of it.

Existing behaviour is your acceptance criteria

Here is a principle that sounds lazy and is in fact the strongest position available: your current system, quirks included, is the acceptance test. If the new system does what the old one does, it is correct — plus documented improvements, minus identified unnecessary complexity.

Why that is defensible rather than unambitious:

  • Users already know the current system. Retraining is the largest hidden cost of any replacement and this reduces it to near zero for the parts that stay the same.
  • The edge cases are already handled, or you would have heard about them by now.
  • The reference implementation is running in production right now. That is the strongest oracle available in any software project, and most projects do not have one.

This is the hinge into the next chapter, and it is worth stating explicitly: the extracted specification and the characterisation harness are two views of the same artefact. One says what the system should do. The other proves what it does. Build them together or you will build neither properly.

Do this even if you stay

The last point is the practical one, and it removes essentially all of the risk from starting.

Extract the specification even if you have already decided not to leave. Three things come out of the exercise regardless of the decision:

What extraction produces, in every branch

A portable specification

A description of how part of your business actually works, which you did not previously own in any single place, and which no longer depends on a vendor to interpret.

A dead-configuration map

Everything nobody uses, which can be retired. Immediate reduction in cost, risk and the surface area of future change requests.

A costed alternative

The only thing that has ever changed a renewal conversation. Vendors do not respond to dissatisfaction; they respond to specificity.

Say it plainly, because it is the most likely good outcome of this entire book: the highest-frequency win here is not an exit. It is an organisation that owns its own requirements for the first time and negotiates from a completely different position.

One caution before you start. Extraction is a discovery activity, not a commitment. Do not announce a migration on the strength of a metadata export. The whole point of doing it first is that it is cheap, reversible and informative — three properties that evaporate the moment somebody tells the vendor.

Key Takeaways

  • The switching cost was never the data export. It was re-specifying your slice — and that just got cheap.
  • Configuration is captured requirements: fields are data, workflows are rules, reports are outputs, validations are policy.
  • The shadow spec lives in people, not in the platform. Interview for it or discover it in production.
  • Existing behaviour is valid acceptance criteria — the strongest oracle most projects will ever have.
  • Extract even if you stay. A costed alternative is the only thing that moves a renewal.
09
Part Three · Running the Decision

The Harness: How a Non-Developer Verifies

The instrument that replaced the crowd — and the two things it provably does not cover, stated at full strength rather than in a footnote.

On a recent consulting engagement — I won’t name the firm, but it was one of the big four — I could not believe what the work actually was.

Their entire delivery was pulling numbers out of the client’s systems into Excel. And every meeting was about what the spreadsheets said: how to resync them, whether they were accurate, why last week’s numbers didn’t match this week’s.

Look at what that is. They were doing custom software development — schemas, joins, versioning, a sync pipeline — but by calling it Excel they exempted it from every engineering discipline. No tests. No source control. No review. No specification. It is the shadow-spreadsheet phenomenon from Chapter 2, except performed by the people billed as the adults in the room, at partner rates.

Now look at what the meetings were. “How do we resync it, is it accurate?” — that is manual verification. Humans doing by hand, in a conference room, week after week, the exact checking a test harness does structurally and for free. They had rebuilt the 1995 verification problem inside a modern engagement and were hand-cranking their way through it.

And it strands. The deliverable dies at the end of the statement of work. Every insight assembled in those spreadsheets evaporates when the engagement closes. All of the custom-build cost, none of the compounding asset.

A manual reporting factory.

That is what happens when you do the building without building the instrument. It is the most common failure shape in this entire territory, and it is almost never recognised, because everyone involved is senior and the artefact looks like a report.

The instrument

Here is what replaces the crowd of ten thousand prior customers.

Characterisation testing, in four steps

  1. Treat the running system — including the platform you are leaving — as its own oracle.
  2. Capture what it actually does for known inputs. Real inputs, from real production traffic, across the range that matters, including the ugly ones.
  3. Turn those input/output pairs into executable, falsifiable assertions.
  4. Run them against the new system.

Pass or fail is binary. It is not a matter of opinion, and it is not a matter of reading code. That last clause is the point of the whole exercise: the buyer never has to read code.

Why this is a procurement fact rather than an engineering practice is worth being precise about, because it is the difference between this book and a testing manual.

Tests written by the builder are the builder’s claim about the builder’s own work. Structurally, that is the same object as the status report from Chapter 2 — a statement by an interested party which the buyer cannot independently check. Useful, but not verification.

A buyer-owned behavioural harness, derived from the incumbent system’s actual behaviour, is a different thing entirely. It is the structural replacement for the crowd. Whoever commissions and owns the harness holds the verification, and that is a decision about accountability, not about tooling.

Key Insight

The status report that cannot be faked is a percentage climbing toward 100 — and it is the direct replacement for “six months of progress versus six months of invoices.”

That contrast is the emotional core of this entire book. My friend from Chapter 2 could not distinguish progress from invoices. A buyer with a harness can look at one number, know exactly what it means, and know that nobody in the room can talk it upward. That is the thing that makes it reasonable for a non-technical organisation to commission bespoke software again.

Where the harness comes from

Two sources, and you need both.

  • The extracted specification from Chapter 8 — what the system was intended to do.
  • Observed production behaviour — what it actually does, including the shadow-spec behaviours nobody wrote down.

Where they disagree, the running system wins for the purposes of acceptance, and the disagreement gets logged as a decision: keep it, fix it, or drop it deliberately. Undocumented behaviour that the business quietly depends on is the single largest source of post-cutover pain, and the only way to catch it is to derive the harness from traffic rather than from documentation.

One honest limitation belongs here rather than later. If there is no incumbent, there is no oracle. Novel capability — the rung-3 work from Chapter 7 — cannot be characterised against a predecessor that does not exist. For that work you are back to acceptance criteria written in advance, which is a weaker instrument and demands more senior attention. Which is a reason to start where an oracle exists, and to treat the first build as the place your organisation learns to hold the instrument at all.

Limit one: the harness proves behaviour, not safety

This is not a caveat. It is a gate, and it belongs in the chapter that introduces the instrument.

Generated code, tested for security

55%

of generation tasks result in secure code

45–55%

the band that rate has stayed inside since 2023

95%+

syntax correctness over the same period

Veracode, “Spring 2026 GenAI Code Security Update,” 24 March 2026; longitudinal study across more than 150 large language models.9

Veracode’s finding is worth stating carefully, because the structural reading matters more than any single percentage. Across three years and more than 150 models, the ability to produce code that works improved dramatically. The ability to produce code that is safe did not improve at all. Those two curves have been diverging since 2023, and the report notes that recent flagship releases produced no statistically significant improvement on the security axis despite everything else that got better.

Which gives the sentence that should govern your gate design:

Your characterisation tests will pass happily on an injectable query.

Behavioural equivalence and security are orthogonal. A system can do exactly what the old one did, to four decimal places of fidelity, and be trivially exploitable. The harness has no opinion about that, because you did not build it to have one.

So security is a separate gate, with a separate owner, and it does not arrive with the model: review of generated code, dependency and supply-chain checks, secrets handling, permissions and data-scoping review, and an explicit written position on what the system is allowed to reach.

Limit two: the harness does not stop debt accumulating

What accumulates underneath a passing test suite

15%+

of commits from every assistant studied introduce at least one issue

22.7%

of tracked AI-introduced issues still alive at the repository’s latest version

100k+

surviving issues cumulatively, by February 2026

Liu, Widyasari, Zhao, Irsan, Chen & Lo, “Debt Behind the AI Boom,” arXiv, 26 April 2026. 302,600 verified AI-authored commits across 6,299 GitHub repositories and five coding assistants; 484,366 distinct issues identified, of which code smells were 89.3%.10

The reading is simple and unwelcome: cheap generation produces volume, and volume produces residue. Issue rates across the studied assistants ranged from around 17% to around 29% of commits, so this is not a story about one bad tool. It is a property of the mode of production.

And a behavioural harness says nothing about it. The tests keep passing while the smell count climbs, which is precisely why this is dangerous — the instrument you trust is reporting green on the axis it measures, and silent on the axis that will hurt you in year three.

This connects directly back to Chapter 6, and it is the strongest practical argument for the position taken there. If the specification and the harness are the asset, surviving issues are clearable: you regenerate from corrected upstream source and the residue goes with the old rendering. If the code is the asset, they are permanent — and you have taken a vendor’s maintenance liability in-house and put your own name on it, which is the worst available outcome of this entire decision.

I will say something uncomfortable here rather than save it for the counter-case chapter: this is the strongest evidence in this book, and it points against the naive version of the book’s own thesis. Chapter 14 grades it properly. It belongs here because you are deciding here.

Three gates, not one

Gate Question it answers Instrument Who owns it
Behaviour Does it do what the old system did? Characterisation harness derived from production traffic The buyer
Safety Is it exploitable, over-permissioned, or leaking? Security review, dependency and supply-chain checks, permissions and data-scope review Security, independently
Durability Can we still change this in three years? Specification as asset, the Delete Test, regeneration as the maintenance strategy Whoever owns the specification

Pitfall — what the harness cannot see

  • Injectable queries, weak cryptography, and everything else in the security gate.
  • Dependency and supply-chain risk introduced by whatever the agent reached for.
  • Accumulating code smells that never break a test but make change progressively more expensive.
  • Anything with no incumbent to characterise — novel capability has no oracle.
  • Licence obligations attached to whatever the generated code was derived from.

The rule that table encodes, stated as bluntly as I can: a build that clears one gate and calls itself verified is the 1995 failure with a faster loop.

The non-negotiables

Four things, none of which are optional, and all of which get cut first when a timeline slips.

  • Parallel run. The old system stays up. Both run. Differences are investigated, not explained away by whoever is most confident in the room.
  • Reversible cutover. A path back that has been tested, not assumed. An untested rollback is a plan for a meeting, not a control.
  • A measured baseline recorded before anything is built — error rates, cycle times, volumes, cost. Without it you cannot demonstrate improvement afterwards, and you will spend the following year arguing about anecdotes with people who preferred the old system.
  • A named accountable human, for the next several years. Not a project sponsor. An owner, who is still there when the project has been forgotten.

Nobody bets the quarter on a demo.

What comes next

You can now decide, extract and prove. What Part Three has not yet addressed is the option sitting between buying and building from nothing — the mature codebase that already does nearly what you want, sitting there, free, with connectors and a user interface and a deployment story.

It looks like the obvious shortcut. Chapter 10 argues it is frequently the most expensive path on the table.

Key Takeaways

  • Characterisation testing turns the running system into its own oracle, so the buyer never has to read code.
  • The harness must be buyer-owned. Builder-written tests are the builder’s claim, which is the old problem in new clothes.
  • Security has been flat at 45–55% for three years while correctness climbed past 95%. Your tests will pass on an injectable query.
  • More than a fifth of AI-introduced issues survive to the current version. Passing tests are silent on accumulating debt.
  • Three gates: behaviour, safety, durability. One gate cleared is not verification.
10
Part Three · Running the Decision

The Third Option: Regenerate, Don’t Fork

There is a mature project that already does nearly what you want. It is free. It has connectors, auth and a deployment story. Under cheap regeneration, taking it is frequently the most expensive move on the table.

Everybody treats this as a binary. Buy or build. It has not been a binary for some time, and the first edition of this book missed the third option entirely.

  1. Buy the vendor’s product and configure it.
  2. Fork something mature that nearly does it, and bend it.
  3. Generate something shaped by your own mission, mining the others for what they learned.

The first edition had two of those, and treated the second as an obvious saving. Under conditions where generation is cheap, that assumption inverts more often than anyone expects — and the reason it inverts is not about code quality at all.

Code is not neutral

A mature repository embodies thousands of design decisions: its unit of work, its state model, its data grain, its source assumptions, its audience, its latency expectations, its interface, its storage choices, what it treats as success, and what it deliberately ignores.

That bundle is its compiled North Star — the mission it was built to serve, crystallised into structure.

You can delete features. The deeper assumptions remain, and they remain in the places you cannot easily reach: schemas, module boundaries, event flows, names, tests, abstractions, dependencies, control paths. Every one of those encodes a decision somebody made about a problem that was not yours.

So adopting an existing project usually means:

What you actually inherit

their North Star
+ your desired behaviour
+ adapters between them
+ growing explanation debt

And here is the diagnosis that separates this from ordinary technical-debt talk, because it is a genuinely different animal: the debt is not necessarily poor code. It is good code organised around a different purpose.

Practitioners have started separating technical debt — local, visible, fixable in a pull request — from architecture debt, which is systemic misalignment that never surfaces as a failing unit test. Clean modules can sit inside a dysfunctional city plan.11

Key Insight

Forking a mature application is how you import architecture debt wholesale while congratulating yourself for avoiding code debt.

The shape it takes

The fork-gone-wrong timeline

Week 0

You clone a mature repo because it already has connectors, a UI, auth and a deployment story. You strip features. You rename a few entities. You write adapters so your concepts map onto theirs. It feels like you have saved four months.

Week 2

The adapters have opinions.

Week 4

Your team is explaining their domain model in every design review. New people are onboarded into two mental models: yours and the repo’s.

Week 8

The product still behaves like their product with a skin, and every new requirement arrives as a negotiation with foreign assumptions.

Nothing in this sequence involves anyone writing bad code.

And that is the part worth being scrupulous about, because this chapter is one careless paraphrase away from being a swipe at open source, which it is not.

Nothing about the original code was bad. The project may be excellent at being the system it intended to be. Its maintainers made good decisions for their mission and defended them under pressure. Fitness is relative to intent. That is precisely why it is a poor starting point for yours.

The calculus flipped

Old maths, new maths

When code was expensive
  • • Code expensive.
  • • Adaptation cheaper than regeneration.
  • • Accepting a North Star mismatch was rational — existing code saved months, and months were the scarce thing.
When generation is cheap
  • • Code cheap.
  • • Understanding expensive.
  • • Architectural mismatch expensive.
  • • Adaptation sometimes harder than regeneration.
The free code may now be the cheapest part of what you inherited.

Notice where the cost actually sits now. Not in typing — typing is the thing that got solved. It sits in comprehension, and in negotiating with somebody else’s assumptions. Both of those are human costs and neither of them has fallen at all. The line item that collapsed is not the line item the fork was saving you.

Reuse was already mostly a story

Return to the HFS Research finding from Chapter 4, because it does different work here.

Average code reuse across Global-2000-scale organisations sits at just 33% — teams rebuild about two-thirds of the functionality anyway.8

Read that as a comment on forking rather than on procurement and it becomes uncomfortable. Two-thirds was being rebuilt regardless — at adaptation prices, inside somebody else’s architecture, with the adapters accumulating opinions the whole time. The reuse everyone assumed they were getting was mostly not there. It was a story we told ourselves at the point of the decision and never audited afterwards.

And Fersht’s framing from the same study lands squarely on this chapter: AI is an amplifier of whatever already exists in your stack. Amplify a mismatched architecture and you get a faster mismatch, arriving sooner, with more of it.

Mine the discoveries, not the architecture

Here is the move that replaces the fork, and it captures most of the value people were actually after.

A mature repository is a record of everything its authors learned the hard way, in production, across years. That record is the expensive payload. The code is the cheap part.

What to take from a repository you are not forking

  • Domain distinctions — the entities and states they found they needed, and the ones they merged after discovering they were the same thing.
  • Edge-case handling — the conditions they hit in production that you have not met yet, and will.
  • Schema decisions and the migrations behind them — a migration history is a map of every place the model turned out to be wrong.
  • The test corpus — often the single most valuable artefact in the repository, and the one nobody forks for.
  • Failure shapes — the issue tracker, the postmortems, the “known limitations” section that everybody skips.
  • Protocol and format conformance — the fiddly compliance details you would otherwise rediscover one bug report at a time.

Then generate around your own North Star: your unit of work, your data grain, your definition of success — with the specification from Chapter 8 and the harness from Chapter 9 as the controlling artefacts.

One connection back that has to be made explicitly, because without it this chapter is dangerous advice: regeneration without a specification is not regeneration. It is starting again and hoping. The whole argument depends on Chapter 6 being true in your organisation. If the specification is not the asset, forking a good project is genuinely the safer option, and you should take it.

Where this argument stops

Pitfall — do not over-read this chapter

  • Still obvious reuse: languages, runtimes, libraries, frameworks, protocols, formats, cryptography, databases, infrastructure. Anything where the interface is narrow and the semantics are shared. Nobody should be generating their own TLS implementation, and the fact that they now could is not an argument that they should.
  • The argument is narrow: it is about adopting a whole application built for a different mission and bending it into yours.
  • Licence obligations do not disappear because an agent did the copying. Generated code derived from licensed sources carries real legal exposure, and it is a live question rather than a settled one.
  • Security posture cuts the other way. A widely used project has had thousands of eyes on it. Yours has had none. That is Chapter 9’s second gate arriving from a different direction, and it is a genuine argument for reuse where reuse is appropriate.

The honest summary: this is not an argument against reuse. It is an argument against inheriting a mission and calling it a head start.

What’s left of the slogan

Part Three has covered extraction, verification, and the shape of the build itself. One piece of the original provocation is still outstanding — the second half, about experts — and it is the part most likely to be over-claimed by everybody including me.

Chapter 11 takes it seriously and draws the line the first edition did not.

Key Takeaways

  • Build is not one decision. Buy, fork, and generate are three different moves with three different cost structures.
  • A mature repo carries a compiled North Star. Forking imports architecture debt that never shows up as a failing test.
  • The cost moved from typing to comprehension and negotiation — and neither of those got cheaper.
  • Mine the discoveries: domain distinctions, edge cases, migrations, the test corpus, failure shapes.
  • Regeneration without a specification is starting again and hoping. If you do not have the spec, fork.
11
Part Three · Running the Decision

Don’t Hire the Expert — Encode the Judgement

The second half of the slogan is the bigger claim and the less examined one. It is also true in a narrower way than anybody, including me, has been saying.

Start with a real target, because the abstraction is easier to believe once you have watched the test run.

An external provider comes in and trains your staff on ISO standards. Recurring fee, scheduled sessions, materials you do not own. They hold the content. In a real sense they hold your compliance knowledge, and every year they hold slightly more of it than you do.

Run the target test on it, in the open.

The four questions, applied

  • Is it valuable? Yes. Compliance is not optional, and failures are expensive in money and in licence to operate.
  • Does it compete with what current staff are good at? No. Nobody’s job title is “remembering the standard.” There is no one to threaten.
  • Does it help staff do a better job? Yes. Better-trained people produce better outcomes — and training available at the moment of need beats training delivered in March.
  • Does it do something you can’t do now? Yes. Personalised to the role, on demand, always current when the standard changes, and tied to your own procedures rather than a generic curriculum.

All four pass, which makes it a build candidate. And notice that the test is doing real work rather than rubber-stamping: most candidates fail one of the four, and the one they usually fail is the second.

Where the line actually falls

Hire the expert Encode the judgement
Real-time judgement under uncertaintyStable knowledge applied repeatedly at volume
Novel situations with no precedentSituations that recur with variation
Relationship-dependent workWork where consistency is a feature, not a compromise
Accountability that must be signedPreparation that precedes a signature
Rung 1 (Chapter 7)Rungs 2 and 3 (Chapter 7)

What encoding does to the expert’s role is the part that determines whether the project survives contact with the organisation. It moves them from applying the knowledge to improving the system and handling the edge cases. The expert stops being the throughput constraint and becomes the quality function.

That is a promotion in substance, and it should be sold as one, honestly, in the first conversation. The alternative framing — “we are automating what you do” — produces exactly the resistance it deserves, and the resistance is rational rather than reactionary.

The distinction I missed in January

This is the correction that makes the chapter worth writing, and it is the difference between a system that works and a system that is confidently useless.

The first edition wrote, approvingly, that AI “knows the expert advice — the training, the ISO, whatever stuff.” That claim is true and incomplete, and the incompleteness is precisely where these projects fail.

Key Insight

A model knowing a published standard is a claim about retrieval of public doctrine. It is not a claim about your obligations under that standard. Retrieve the first. Encode the second. Never confuse them.

The text of a standard is public, widely discussed and well represented in training data. Retrieval of it is genuinely solved, and it was solved a while ago.

Your obligations under it live somewhere else entirely. In your scope statement. Your certification boundary. Your documented procedures. Your risk register. Your auditor’s interpretations, which are not the same as anyone else’s auditor’s interpretations. Your prior non-conformances and the corrective actions you committed to and are being measured against. None of that is public. Most of it is not written down in one place. Some of it exists only as a finding letter sitting in somebody’s inbox from eighteen months ago.

A system that recites the standard is a study aid. A system that knows how the standard applies to your operation is a capability — and it is only the second one that passes the fourth question of the target test.

Which is exactly why the ISO example passes so cleanly, and why it is a better example than it first appears. The valuable part is the join between the published standard and your actual operation. That join is precisely what the external provider was least able to do, because they arrive with the standard, spend two days in a room, and leave without your context. You were never really buying their knowledge of the standard. You were buying their attendance.

The encodable asset is rarely the expertise. It is the expertise applied to your particulars — and the particulars are the part nobody has written down.

That generalises well beyond compliance. Legal review, underwriting guidelines, clinical protocols, procurement policy, engineering standards, safety cases: in every one of them the published body of knowledge is the cheap half and the organisation-specific application is the expensive half. The build target is always the second half, and the first half is a retrieval problem you should stop paying consultants to solve.

Where encoding fails

Pitfall — four ways this goes wrong, and the tell for each

  • Accountability that cannot be delegated. A signature is not a workflow step. If a named person must attest, the system prepares and the person signs — which is a legitimate and valuable design, but it is not replacement. Tell: somebody asks who is legally responsible and the room goes quiet.
  • Expertise that is really relationship. A great deal of senior expertise is trust, access and the ability to be believed by a particular person on a particular committee. That does not encode. Pretending otherwise produces a technically correct system that nobody acts on. Tell: the expert’s calendar is meetings, not analysis.
  • Judgement that is genuinely novel each time. If every instance requires a fresh frame, there is no stable knowledge to encode — you have a research problem wearing a process costume. Tell: no two cases in the last year had the same shape.
  • Confidently wrong with no reviewer. Where the failure mode is plausible incorrectness and there is no verification step, encoding converts a slow correct process into a fast wrong one. Tell: you cannot name who checks the output.

The rule underneath all four: encode the preparation, keep the judgement, and never remove the reviewer before you have evidence the reviewer is redundant. The evidence is a measured baseline and a period of parallel running, which is the same discipline Chapter 9 asked for and the same one that gets cut first when a timeline slips.

What this does to hiring

Say the consequence plainly, because it does not point where the slogan implies.

You need fewer people who write and more people who can specify and verify.

That is a harder hiring problem, not an easier one. Specification is not a junior skill. It requires domain knowledge, precision, and a willingness to be pinned down — which is temperamental as much as technical, and which a great many experienced people actively avoid because vagueness is comfortable and defensible. Verification requires someone who will read a failing test rather than an optimistic summary, and who will hold a release when the number is not where it should be.

The first build is where an organisation learns to hold that instrument, and it should be staffed accordingly — senior, deliberate, and small. Chapter 13 asks who is actually accountable for the last mile, and the answer is not something you can acquire as a licence or as a headcount requisition.

What this chapter does not own

Getting business, people and measurement to agree on what success means before the build is a real gate, and a project that skips it will fail for reasons that have nothing to do with anything in this book. It belongs to the Three-Lens work, and it is one clause here rather than a chapter.

Which is itself a correction. The first edition ran an entire chapter of alignment machinery in this position, plus a change-management timeline counting days before and after go-live. Both were doing another framework’s job. Both carried that edition’s least defensible numbers. Both are cut, and naming what was removed and why is part of the accounting.

What’s left

Part Three is complete. You can extract the specification, build the instrument, choose the right shape of build, and decide what to encode rather than hire.

What remains are the two questions that sit above the decision rather than inside it. Is this decision even strategy? And what is moving underneath it while you run the analysis?

Key Takeaways

  • Four questions decide a target: high value, doesn’t compete, helps people, does what couldn’t be done. Most candidates fail the second.
  • Hire for real-time judgement, novelty and relationship. Encode stable knowledge applied at volume.
  • Knowing the standard is retrieval. Knowing your obligations under it is the capability — and it is not public.
  • A signature is not a workflow step. Encode the preparation; keep the judgement.
  • You need fewer writers and more specifiers. That is a harder hiring problem, not an easier one.
12
Part Four · Altitude and the Moving Market

Horses, Cars, and the Board’s Version of This Question

Your build project is probably not strategy. That is not a reason to stop — it is a reason to put it in the right budget line and stop calling it something it is not.

A board meeting. The AI section of the agenda has just closed: the portfolio reviewed, the traffic lights admired, the next gate review confirmed for the following quarter. The chair speaks before the agenda moves on, walks to the whiteboard, and writes:

What will still make this company valuable when software, analysis, coding, content, reporting and basic advice become cheap?

The room goes quiet. Not because nobody understands the question — everyone understands it immediately — but because nobody knows whose job it is to answer it. The company has an AI portfolio, an AI committee, an AI lead, a rollout schedule, three vendor partnerships and a board-pack template. It does not have a single role whose mandate is the sentence on the whiteboard.

That silence is why this chapter exists in a book about buying software. The chair has asked the build-versus-buy question one floor up — and one floor up, it gets a different answer.

Four classes of project

Class What it does Verdict Resourcing
Horse Optimisation Makes the current process faster, cheaper or more accurate. The process, the org chart and the product all stay the same. Tolerate, don’t celebrate Operational budget only. Never strategic capital.
Horse Replacement Replaces a current human task directly. Avoid as flagship Re-scope so AI prepares and the human delivers.
Car Discovery Identifies the future model. The output is knowledge, not revenue. Sponsor at board level Board-sponsored, dedicated, multi-quarter.
Car Construction Builds the AI-native successor. Double down Separate-but-connected venture; multi-year capital; protected from existing-P&L gravity.

Horse Optimisation persists because it is the path of least resistance in every dimension at once: easiest to fund from an operating budget, easiest to demo, requires no organisational change, generates no political resistance and creates no regulatory complications. The wrong portfolio shape assembles itself with nobody doing anything obviously wrong.

Horse Replacement is seductive because the headcount saving fits neatly on a slide. It is mostly wrong as a flagship, because the lanes it lands in — real-time, customer-facing, regulated — stack constraints on top of each other, and because every visible failure becomes “AI isn’t ready” even when the system is quietly outperforming the human baseline it replaced.

Car Discovery produces structural maps and construction targets rather than revenue, which is precisely why no operating team can fund it from a single-year profit and loss statement. It has to be sponsored from above or it does not happen.

Car Construction is where compounding actually starts, and it needs protecting from the gravity of the existing business, whose metrics will otherwise subordinate it within two quarters.

Now apply it to your build project

Here is the least commercially convenient claim in this book, and I would rather make it than have you discover it in a board pack.

Bottom Line

Rebuilding a platform you already have — the same workflow, the same outputs, cheaper and more yours — is Horse Optimisation whichever way the decision goes.

Keep the vendor and negotiate harder: horse. Replace them with your own system that does substantially the same thing: still horse. The process is unchanged, the org chart is unchanged, the product your customers experience is unchanged. You have made the current operating model run better, which is a real and worthwhile thing to do and is not a strategy.

That is not a reason to stop. It is a reason to file it correctly, brief it to the right audience, and stop calling it AI strategy in front of a board that has a harder question on the wall.

The governing rule is worth memorising, because it settles arguments that would otherwise run for quarters: the altitude beats the metric. A project with an excellent return on investment, classified as Horse Optimisation and consuming strategic capital, is a kill candidate regardless of how good the number is. The number is not wrong. It is simply not the criterion.

What the build decision looks like at the right altitude

Not the replacement of a platform. The construction of a capability that the vendor category cannot contain — because it depends on your data, your judgement, your customers and your obligations, and no vendor can amortise it across a thousand buyers.

There is a clean test for this, and it takes about ten seconds:

The amortisation test

Could a vendor sell this to your three closest competitors without changing it?

If yes — you are building a commodity, somebody will eventually sell it to you for less than it costs you to maintain, and the buy decision deserves another look.

If no — if the thing only makes sense inside your operating reality — then you are in Car Construction territory, and the build-versus-buy question was never really the question you were answering.

This is where Chapter 7’s routing and this chapter’s altitude converge, and it is the only place they do. Rung-3 work — the previously infeasible category — is by definition work no vendor has a product for, because if there were a market for it as a product it would already exist and it would not be your advantage. That convergence is the signal that you are looking at the right kind of build.

It also determines how you should measure it, which is where most of these projects die. Horse Optimisation is legitimately measured in cost and cycle time; those are the right numbers for that class. Car Construction is not, and measuring it that way will kill it inside a year, because a capability that does not exist yet has no baseline to improve on. Chapter 6’s capability denominator was written for exactly this case.

The honest concession

Most readers who run the analysis in this book will end up with: a better renewal, a cleaner specification, one or two replaced systems, and a team that now knows how to hold a verification harness.

That is a good year. It is not a transformation, and claiming otherwise is how AI programmes lose credibility with the people who sign for them — usually in the second year, when the promised step change fails to appear and the person who promised it has moved on.

But there is a second-order value here that is genuinely worth naming, and it is easy to miss because it does not appear in the business case.

The organisation that has done this once has a specification, a harness, a baseline discipline, a security gate and a named accountable owner. Almost nobody has those. They are the preconditions for Car Construction, and they cannot be bought — they are only acquired by running a real build with real consequences.

Build one is where you buy them. That is the correct way to justify it, and it happens to also be true.

On the genre

Build-versus-buy content reliably over-promises, because the promise is what sells the engagement that follows it. Under-promising here is both the more useful position and the more honest one — and if this chapter has just talked you out of describing your project as strategy, it has done its job.

What this chapter borrows, and what it does not

The doctrine itself, the asset classes underneath it, the discovery engine and the full portfolio re-classification method belong to the Terminal Value work and are not re-argued here. What this chapter takes is three things: the four classes, the filter question, and the altitude rule. That is all it needs.

The altitude question asks whether the decision matters. The next one asks something harder: whether the decision is stable — because two things are moving underneath it that have nothing to do with price.

Key Takeaways

  • Four classes, four verdicts: tolerate, avoid as flagship, board-sponsor, double down.
  • Rebuilding a platform you already have is Horse Optimisation in either direction. Worth doing; not strategy.
  • The altitude beats the metric. A great return on a mis-classified project is still a kill candidate.
  • The amortisation test: could a vendor sell this to your three closest competitors unchanged?
  • A better renewal, a cleaner spec and two replaced systems is a good year. Build one buys you the preconditions for everything above it.
13
Part Four · Altitude and the Moving Market

Two Things Moving Under You

Who operates the software, and who is accountable for the last mile. Neither appears on a build-versus-buy scorecard, and both can invalidate your answer inside a renewal cycle.

Start with a question that is usually asked in one direction only.

Why is the user’s AI unable to use our business?

Now turn it around and point it at your own portfolio: why is our AI unable to use our vendor’s product?

Both versions are the same structural question. Neither appears anywhere on a conventional build-versus-buy scorecard, and one of them is quietly repricing half your stack while you run the analysis.

Shift one: the operator is changing

The app era had the application as custodian of the interaction. You went to the software. The software decided what a session looked like, what the funnel was, and how many clicks stood between you and the outcome.

The agent era has the personal or organisational agent as custodian of the intent. The human interface becomes the showroom and the exception desk. The machine-operable surface becomes the loading dock.

Key Insight

When the agent is the customer, UI polish stops being a moat.

What a genuinely agent-operable service exposes is different in kind from what a website exposes: state, available actions, delegated authority, the consequences of acting, and a subscription to change. Not pixels, navigation, forms and clicks — those were designed for people. And a private API built for the vendor’s own front end is not sufficient either, any more than an agent driving a user interface by impersonating human fingers counts as delegation.

This belongs in a book about buying software for two reasons, both material and both financial.

What this does to the decision, in both directions

On the buy side

A vendor your agents cannot operate is depreciating independently of price. Everything you intended to automate around it now has to be done by a person clicking. The subscription cost is flat; the effective cost is rising; and no renewal negotiation will ever surface that, because it does not appear on the invoice.

On the build side

“Build” now frequently means building a thin delegation layer over the things you keep buying. It is the smallest useful build on the table and the one most organisations skip, because it does not look like a project and nobody gets promoted for it.

Chapter 4 said keep buying Tier 1. This is the corollary: keep buying it, and make it addressable to your own agents. That is not a contradiction of the tier logic; it is the tier logic applied to a world where the operator changed.

There is a portfolio smell test that follows, and it is worth running on your own roadmap this week: if every AI initiative in the portfolio is an in-app assistant, you are funding a better concierge while the transport layer is built somewhere else.

And a procurement question that belongs in your next renewal: can an authorised agent of ours complete a durable intent against your service without impersonating fingers on your UI? The answer is diagnostic. So, frankly, is the reaction to being asked.

Shift two: the unit of purchase is changing

Start from a premise that is uncomfortable for everybody currently selling AI capability, including us.

Frontier capability is symmetric. Every serious competitor rents the same models, from the same handful of labs, on the same day they are released. Whatever advantage a new frontier model confers, it confers on everyone at once — which means it confers durable advantage on no one.

Symmetry is the exact opposite of a moat.

Run the elimination that follows. If capability is rented, then whatever remains scarce must be something that cannot be rented. Three things survive, and all three are stubbornly local:

  • Knowing how the work is actually performed, exceptions included.
  • Deciding, step by step, where intelligence belongs — and being permitted to put it into production.
  • Accountability when it misbehaves, and a form of the learning that the next project can load.

The name for that territory is the last mile: everything between the model can do this and this runs in our business and someone is accountable for it.

The correction this forces

The first edition of this book ended on a note I now find embarrassing: even one developer with AI tools can shift your make-versus-buy decisions.

That is true about capability and false about delivery, and the gap between those two things is where build projects die.

What a buyer actually needs is not capability but accountability for placement and production. That is a genuinely different unit of purchase, and most procurement functions do not have a line item for it. They can buy a licence. They can buy a headcount. They find it much harder to buy a spine.

Apply that internally, which is where most readers of this book actually sit. Your build does not need a developer. It needs someone who will own the specification, the harness, the security gate and the behaviour of the system in production for several years — and who has the standing to say no when a release is not ready.

Most organisations discover they do not have that person somewhere around month three. Not because nobody is capable, but because the role was never defined, funded or given authority, and the person doing it by default has three other jobs.

What to demand instead of a capability list

Buyers already understand the fragmented chain, because they have paid for every link in it separately: strategy consultant → analyst → architect → security → engineering → test → deploy → operations. What they cannot believe, from a list of skills, is that one spine holds that chain without the telephone-game losses at every boundary.

What they can believe, if you show it: one problem moving from an executive concern to a governed production system as a single joined pass, with inspectable artefacts at each stage.

The internal version of that test

Before an internal team earns a build mandate, ask them to show one problem making the full trip end to end — small, real, and with the artefacts visible: the specification, the harness, the baseline, the security review, the cutover plan.

Not a demo. A pass. A demo shows you the good part; a pass shows you the seams, which is the only part that predicts anything.

The same instrument works on external suppliers, and Chapter 14 supplies the ladder for grading their answer: ask which class of evidence their strongest claim belongs to, and watch what happens next.

What compounds after build one

The organisational version of a question I have written about at the individual level: what you own after build one determines whether build two is cheap.

Not the code. The compiled judgement — the specification, the harness, the approaches you rejected and why you rejected them, the baseline discipline, the security position, and the exceptions you discovered the hard way and will never discover again.

That gives a test you can run on your own numbers rather than taking a book’s word for it, and it is the falsifiable version of the “second build is cheaper” shape claim from Chapter 6:

The build-two test

The second project should start from a materially higher baseline and be led by someone more junior than the first — with the difference attributable to shared infrastructure rather than to the same people working late again.

If build two costs what build one cost, you did not build an asset. You bought a project.

That is a hard test and most organisations will fail it the first time. Failing it is useful information — it tells you precisely which artefacts did not survive contact with the second problem, and those are the ones to promote upstream before build three.

Where both shifts point

They point the same way, and it is the direction this entire book has been walking.

The durable thing is never the software. Not the vendor’s and not yours. It is the specification, the proof, the accountability and the learning — and every one of those survives a model release, a vendor change, a re-platform and a reorganisation.

Which leaves one honest job outstanding: making the strongest available case against everything above, and grading this book’s own evidence rather than yours.

Key Takeaways

  • A vendor your agents cannot operate is depreciating independently of price, and the invoice will never tell you.
  • “Build” increasingly means a thin delegation layer over what you keep buying.
  • Frontier capability is symmetric. What stays scarce is local: placement, permission, accountability and retained learning.
  • You can buy a licence and you can buy a headcount. You cannot buy a spine — internally or externally.
  • Build two should be led by someone more junior. If it costs what build one cost, you bought a project, not an asset.
14
Part Five · Honest Accounting

The Counter-Case, at Full Strength

The strongest argument against this book, and a grading of its own evidence rather than yours. The weakest link is a hole, not a caveat.

There is a rule I apply to other people’s claims, so it should apply to mine:

A claim may only be defended at the class of its weakest supporting evidence.

A grading instrument that only ever grades other people is a rhetorical device. So this chapter runs it on the book.

The evidence class ladder

Class What it is What it can honestly prove
E0An announcement or a stated intentionIntent; capital allocation
E1A role, architecture or method definitionWhat the category claims to be
E2Dated capability investment — money, headcount, programmesThat the market is betting
E3Seller-side result — the supplier’s own revenue, growth or surveyThat somebody is being paid
E4Buyer-side production outcome — named, dated, measured against a baselineThat a deployment produced a result
E5E4 plus third-party verificationThe strongest available claim

The rule that makes it usable: you cannot chain an E2 into an E4 by adding adjectives. And note what the ladder is not ranking — the classes rise in cost of production, not in importance. An announcement is not dishonest. It is simply cheap, and cheap evidence should be priced accordingly.

This book, graded

Where each load-bearing claim actually sits

Retool’s build-versus-buy numbers — E3

A vendor’s survey of its own build-inclined population. Strong evidence that the behaviour exists and is not marginal; weak evidence of a population rate. It is the single most load-bearing external source in this book, and it is seller-side. That is worth saying twice.

HFS Research’s implementation multiple and reuse figure — E2/E3

An analyst survey with a named sample and date. Good for the shape of enterprise spend. Not a controlled measurement of any particular platform, and certainly not of yours.

The SaaS pricing figures — E2/E3

Trade and vendor-adjacent surveys with dates and named publishers. Directionally consistent across three independent sources, which is worth something, and is not the same as verified.

Veracode and the surviving-issue study — the strongest evidence here

A longitudinal study across 150+ models, and a large-scale empirical paper over 302,600 commits. Note what that means: the best-supported evidence in this book is evidence against the naive version of its own thesis. That is uncomfortable, and it is the right way round.

The verification thesis itself — E1

A structural and historical argument, not a measurement. It is falsifiable, which is its virtue: had custom software thrived in periods when verification was hard and build cost was low, the argument would be dead. It did not, and the spreadsheet’s survival is the strongest support available. It remains an argument.

That AI-built replacements are durable at five years — no evidence at all

There is no E4 cohort. None. Not enough time has passed. Every statement in this book about regeneration economics and the second-build discount is structural reasoning, and the one large empirical study that touches it is pointing the other way.

Important

The weakest link in this book is the maintenance horizon, and it is a hole, not a caveat.

Objection 1 — the maintenance horizon is unproven

State the case properly, because it is the strongest one available and it deserves better than a strawman.

Nobody has a five-year cohort of AI-built SaaS replacements still running healthily. Regeneration economics is an argument, not a measurement. The measured trend in AI-authored code is accumulating surviving issues, not clearing them. It is entirely possible that the dominant story of 2029 is organisations who replaced their platforms in 2026 and now maintain something worse than what they left, with nobody to escalate to and no vendor to blame.

My response is not confidence. It is instrumentation: a measured baseline before you start, a harness that stays green, and the specification as the asset so that a bad build is recoverable rather than terminal. Those three make the failure survivable. They do not make it unlikely.

What would change my mind

  • (a) Surviving-issue rates in AI-authored code continuing to rise through 2027 while regeneration practice matures — that would suggest regeneration does not clear the residue the way the theory says it should.
  • (b) A cohort of documented replacements showing higher five-year total cost than the platforms they replaced.
  • (c) Evidence that characterisation harnesses systematically fail to catch the class of defect that actually hurts in production.

Any of those three would move this book’s recommendation materially toward “keep buying,” and I would rather write that down now than reinterpret my position later.

Objection 2 — you are asking non-technical buyers to hold an engineering instrument

The case: “read the percentage climbing toward 100” is elegant on a page and hard in an organisation. Someone has to build the harness, maintain it, keep it derived from current production traffic, and refuse to weaken it when a release is late and a director is impatient.

The response: yes — and there is a floor, which the book has stated three times and will state again. If nobody in the organisation can own the harness, the decision is buy. That is not the method failing. It is the method returning an answer, and it is the same answer as the stopping rule from Chapter 3.

There is an uncomfortable corollary worth stating rather than skating past. This makes the build option less available to exactly the organisations that would benefit most from it — the ones with the heaviest customisation burden and the least engineering capability. That is a real inequity in the argument. It is not resolved by anything in this book, and pretending it is would be dishonest.

Objection 3 — the vendor does things you are not counting

Availability engineering. Twenty-four-hour incident response. Regulatory tracking. Penetration testing. Disaster recovery rehearsal. A support line at three in the morning. An insurance position when something goes badly wrong.

The response: those are the first two rows of Chapter 4’s table, and the decomposition exists precisely so that they are counted rather than assumed away. But this objection deserves more than a pointer, because it names the most common way builds actually fail.

Key Insight

Most builds that fail do not fail on features. They fail on the operational surface nobody costed — discovered at 2am in month seven.

So price the operational surface explicitly before deciding: on-call rota and who is on it; backup and restore rehearsed, not merely configured; dependency patching; certificate management; capacity headroom; and who answers when it breaks during somebody’s annual leave. Write the names down. If you cannot, that is the answer.

Objection 4 — this is survivorship bias dressed as doctrine

The case: the people who replaced a SaaS tool and regretted it do not write case studies. Retool’s 35% counts replacements, not outcomes. Every visible example is self-selected, and the failures are invisible by construction.

My response is to concede it. Without a counter-statistic, because none exists.

We can see the behaviour. We cannot yet see the distribution of results. That is the honest position, and any author telling you otherwise is either selling something or has not looked.

What follows from conceding it is practical rather than rhetorical: prefer reversible moves; start where an oracle exists; and treat the first build as an investment in the instrument rather than as a bet on a system. All three of those are cheaper if you assume the failure distribution is worse than the visible examples suggest — which, given the selection effect, it almost certainly is.

Where building is simply wrong

Do not build here

  • Real-time, high-stakes, externally facing work where a failure is public and immediate.
  • Anything whose value is operational scale — the Tier 1 row from Chapter 4. Do not rebuild payments.
  • Anything you cannot specify. If the specification cannot be written, the build has already failed; the failure simply has not surfaced yet.
  • Anything with no oracle to characterise, unless you accept that acceptance criteria written in advance are a weaker instrument, and staff the work accordingly.
  • Anywhere you cannot name an accountable human for the next several years. Not a project sponsor. An owner.
  • Where the organisation is already in the middle of something else. Capacity is a hard constraint and it appears in none of the frameworks in this book, including mine.

Failure shapes to recognise early

Pitfall — five shapes, and the tell for each

  • The spec nobody can write. Months of workshops producing a document that describes an aspiration rather than a behaviour. Tell: there are no acceptance criteria in it, anywhere.
  • The build with no baseline. Nobody measured the old system, so improvement becomes a matter of opinion and the loudest opinion wins. Tell: the business case is written in adjectives.
  • The key-person build. One brilliant person, no specification, no harness, no second reader. Tell: the project’s status is whatever they say it is — which is Chapter 2’s principal-agent problem, rebuilt in-house with a payroll number attached.
  • The harness that was going to be added later. Tell: the tests are on the plan for after go-live.
  • The manual reporting factory. Chapter 9’s specimen. The most common shape and the least often recognised, because everyone involved is senior and the artefact looks like a report.

What survives

A narrower, better-instrumented version of the thesis.

Build where you can specify and prove. Buy where you cannot — and buy without embarrassment, because a buy decision reached through this analysis is a much better buy decision than the one you would have made by default.

The last chapter consolidates the corrections and hands over the instrument.

Key Takeaways

  • Grade the class of evidence before you count it. Cheap evidence is not dishonest; it is cheap.
  • The best-supported evidence in this book argues against the naive version of its own thesis.
  • There is no five-year cohort for AI-built replacements. That is a hole and it is stated as one.
  • If nobody can own the harness, buy. If nobody can be named as the owner for five years, buy.
  • Most builds fail on the operational surface nobody costed, not on features.
15
Part Five · Honest Accounting

What Aged, and the Decision Sheet

This is the second edition of an argument I published in January. The conclusion survived. The reasoning did not. Here is the accounting, and here is the instrument.

The corrections ledger

First edition, January 2026 This edition Why it changed
Cause. AI collapsed build costs, so the make-versus-buy calculus inverted. Custom software never died of build cost. It died of verification cost — and what changed is that verification was restructured into a harness the buyer owns. The cost story cannot explain the thirty years it describes. Most obviously, it cannot explain the spreadsheet. Ch 2, Ch 3.
Evidence. Led with developer-productivity statistics: percentage faster with an assistant, percentage of code AI-generated. Developer speed does not decide a procurement question, in either direction. It appears once, with the task named, as an exhibit in that argument. Category error, not a rounding error. And the most-cited finding on the sceptical side has since been revisited by its own authors. See below.
Lock-in. Filed under costs — the vendor trap, the thing that keeps you paying. Filed under assets. Installed configuration is your exit specification, maintained at the vendor’s expense. Too cautious. The stronger claim was available in January and I did not make it. Ch 8.
The asset. Implied that what you own after building is the software. The specification plus the harness. Generated code is regenerable output. Software you cannot regenerate is not ownership. It is a maintenance liability with your name on it instead of a vendor’s. Ch 6.
Numbers. A build-cost collapse cited to our own site; a three-step amortisation ladder; a win rate moving by twenty-two points; a named annual vendor fee. Retracted. Replaced by one sourced multiple — 2–7× licence cost on implementation and integration, HFS Research, November 2025 — and four explicitly declared shape claims. None of them had a source. Several were self-cited, which dresses an assertion as research. A sourced qualitative claim beats an invented quantitative one. Ch 4, Ch 6.

Four of those five point the same direction, and the pattern is more useful than the individual items. The first edition reasoned from inputs — how cheap is the code, how fast is the developer, how big is the licence. The decision is governed by proof: can you say what you want, and can you show that you got it.

On the evidence row, one specific example is worth the space, because it demonstrates why the whole category was the wrong place to be arguing.

The most-cited finding on the sceptical side was a randomised controlled trial showing experienced developers taking 19% longer with AI tools while believing they had been 20% faster.12 In February 2026 METR revisited it themselves. They now believe “it is likely that developers are more sped up from AI tools now — in early 2026 — compared to our estimates from early 2025.” They also disclosed a selection effect that cuts against their original number: a significant increase in developers declining to participate “because they do not wish to work without AI.” And they graded their own new data honestly: “because of the selection effects in our experiment, our data is only very weak evidence for the size of this increase.”13

Both camps were arguing about a variable that does not decide this. That is the correction, and it applies to my side of the argument at least as much as to theirs.

Also aged — in the other direction

Honesty cuts both ways, and this half is more interesting.

The first edition’s caution about AI-generated code quality turned out to be insufficient, not excessive. In January the responsible position was “AI code needs review.” The 2026 evidence is considerably worse than that: security pass rates flat between 45% and 55% for three years while syntax correctness climbed past 95%, and more than a fifth of AI-introduced issues still alive at the current version across three hundred thousand commits.

And here is why that strengthens this book rather than damaging it.

Key Insight

Evidence that cheap generation produces unsafe, debt-accumulating output is evidence for verification — which is what this edition’s thesis now is. A first edition arguing from cost would have been damaged by it. An edition arguing from proof is confirmed by it.

That is what a correct correction looks like. The new evidence fits the new thesis better than it fits the old one. If it had fitted the old one better, I would have had a different problem.

One more thing aged usefully. The first edition’s list of “what to stop buying” was too wide — reporting, analytics, customer-data enrichment, document processing, knowledge management, compliance monitoring, all in one sweep. This edition narrows it to a diagnostic you run on one platform at a time, with Tier 1 explicitly protected and the tier assignment made by three questions rather than by category. Being right about a smaller area is worth considerably more than being vaguely right about a large one.

The decision sheet

One platform. Not the portfolio. Each move points at the chapter that owns it.

Eight moves

  1. Decompose the subscription. (Ch 4) Which of the four products are you actually buying — operational depth, liability transfer, verification by crowd, or the software itself? If it is mostly the last two, continue. If it is mostly the first two, stop here and go and negotiate.
  2. Measure customisation burden. (Ch 5) Licence versus everything else: implementation amortised, consultants, internal administrators, integration maintenance, change overhead. Over 50%, keep going. Over 70%, you are already building custom software without the discipline.
  3. Run the Leverage Test. (Ch 5) Where does the customisation dollar compound? If it lives in a proprietary configuration language, every dollar spent there is a dollar next year’s model will not improve.
  4. Extract the specification before deciding. (Ch 8) Export the configuration; translate fields, rules, reports, integrations and permissions into a portable specification; interview for the shadow spec. This artefact is valuable in every branch, including “we stayed.”
  5. Write the characterisation tests first. (Ch 9) Capture the current system’s actual behaviour as executable assertions, derived from production traffic rather than documentation.
  6. Add the gates the harness does not cover. (Ch 9) Security review of generated code, dependency and supply-chain checks, permissions and data scope. A named accountable human. A measured baseline recorded before anything is built.
  7. Parallel run, then a reversible cutover. (Ch 9) The old system stays up. The percentage climbs toward 100. The rollback is tested, not assumed.
  8. Check the altitude. (Ch 12) Rebuilding the same workflow in your own code is Horse Optimisation whichever way it goes. Worth doing; not strategy. Put it in the right budget line — then ask separately what you would build that no vendor could sell to your three closest competitors unchanged.

The stopping rule

If you cannot write the tests, keep buying. Not a failure of nerve. The instrument working — and the single most useful sentence in this book.

On sequencing: moves one to four are analysis and cost a week or two of somebody’s attention. Moves five to eight are commitment. Most organisations should run one to four on their most expensive platform this quarter and decide precisely nothing until they have seen the output.

You will be surprised by what comes back. Almost everyone is. The dead configuration alone usually pays for the exercise.

Where the constraint went

One last thing, because it is the arc that has been running underneath every chapter.

Every era’s binding constraint climbed the stack. Hardware, when machines were the scarce thing. Then developers, when machines got cheap and people did not. Then licences, when the industry organised itself around selling seats. Then integration, when everybody had bought everything and none of it talked to anything else.

It has finally arrived at intent.

Which is why the risk did not disappear when code got cheap — it migrated up the stack, into the specification. My friend from Chapter 2 would still fail today if he could not articulate what his manufacturing system was supposed to do. He would just fail in a night instead of a year, and that speed is genuinely most of the improvement. But the failure mode itself relocated. It was never solved.

The gating skill moved with it: from can you manage developers to can you specify what you want. You need fewer people who write, and more who can say precisely what should exist and prove that it does.

So: stop pricing the build. Price the specification and the proof.

The bespoke era is back. This time the hard part is you.
REF
Sources & Evidence

References & Sources

The evidence base behind every claim — primary research, industry analysis, and technical specifications

Research Methodology

This ebook draws on primary research from standards bodies, independent research firms, enterprise technology vendors, and consulting firms. Statistics cited throughout have been cross-referenced against primary sources.

Frameworks and interpretive analysis developed by Scott Farrell / LeverageAI are listed separately below — these represent the practitioner lens through which external research is interpreted, and are not cited inline to avoid self-promotional appearance.

LinkedIn Commentary

Andrej Karpathy (X, 27 December 2025) — I've never felt this much behind as a programmer [1]

The 10X admission and the "skill issue" framing that set the tone for 2026

https://x.com/karpathy/status/2004607146781278521

Case Studies

Theo Browne (YouTube) — You're falling behind. It's time to catch up [2]

90% of personal code AI-generated, 70%+ across teams; "getting into it now is getting into it late"

https://www.youtube.com/watch?v=Z9UxjmNF7b0

Primary Research & Standards Bodies

Retool (17 February 2026, survey of 817 builders) — The Build vs. Buy Shift: How Vibe Coding and Shadow IT Have Reshaped Enterprise Software [3]

35% have already replaced at least one SaaS tool with a custom build; 78% expect to build more in 2026; 60% built something outside IT oversight

https://retool.com/blog/ai-build-vs-buy-report-2026

Wikipedia — VisiCalc [7]

"VisiCalc ... was the first spreadsheet computer program for personal computers ... often considered the application that turned the microcomputer from a hobby for computer enthusiasts into a serious business tool," released 1979

https://en.wikipedia.org/wiki/VisiCalc

HFS Research with Unqork (18 November 2025, n=123 Global 2000-scale organisations) — AI Won't Save Enterprises from Tech Debt Unless They Change the Architecture First [8]

Organisations spend "2–7x their license cost on implementation and integration"; "average code reuse is just 33%, meaning teams rebuild about two-thirds of functionality"

https://www.hfsresearch.com/press-release/ai-wont-save-enterprises-from-tech-debt-unless-they-change-the-architecture-first/

Veracode (24 March 2026, 150+ large language models) — Spring 2026 GenAI Code Security Update [9]

"only 55% of generation tasks result in secure code"; security performance has "remained essentially flat, hovering between 45% and 55%" since 2023 while syntax correctness passed 95%

https://www.veracode.com/blog/spring-2026-genai-code-security/

Liu, Y., Widyasari, R., Zhao, Y., Irsan, I. C., Chen, J. & Lo, D. — arXiv:2603.28592v2, 26 April 2026 — Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild [10]

302,600 verified AI-authored commits across 6,299 repositories; 484,366 distinct issues; "more than 15% of commits from every AI coding assistant introduce at least one issue"; "22.7% of tracked AI-introduced issues still survive at the latest version of the repository"; over 100,000 surviving issues by February 2026

https://arxiv.org/html/2603.28592v2

METR (10 July 2025) — Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity [12]

Randomised controlled trial: developers took 19% longer with AI tools while estimating afterwards that AI had made them 20% faster

https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/

METR (24 February 2026) — We are Changing our Developer Productivity Experiment Design [13]

"we believe it is likely that developers are more sped up from AI tools now — in early 2026 — compared to our estimates from early 2025"; a significant increase in developers declining to participate "because they do not wish to work without AI"; "because of the selection effects in our experiment, our data is only very weak evidence for the size of this increase"

https://metr.org/blog/2026-02-24-uplift-update/

Industry Analysis & Vendor Research

SaaStr — The Great Price Surge of 2025 [4]

"SaaS pricing is up by approximately 11.4% compared to the same time in 2024—a stark difference from the 2.7% average market inflation rate of G7 countries"

https://www.saastr.com/the-great-price-surge-of-2025-a-comprehensive-breakdown-of-pricing-increases-and-the-issues-they-have-created-for-all-of-us/

Zylo (page dated 9 June 2026) — 2026 SaaS Pricing Trends (reporting the 2026 SaaS Management Index) [5]

79% encountered price increases at renewal; 78% experienced unexpected consumption or AI charges; 61% cut projects because of unplanned SaaS cost increases

https://zylo.com/blog/saas-pricing-trends

Gartner / Campus Technology — Gartner 2026 worldwide IT spending forecast, as reported by Campus Technology, 20 May 2026 [6]

Total IT spending $6.31 trillion for 2026; software segment growth 15.1%

https://campustechnology.com/articles/2026/05/20/gartner-estimates-worldwide-it-spending-at-6-31t-for-2026.aspx

Stalbouskaya, N. — The New Stack, 22 October 2025 — Technical Debt vs. Architecture Debt: Don't Confuse Them [11]

Architecture debt is systemic misalignment rather than a local code smell; clean modules can sit inside a dysfunctional "city plan"

https://thenewstack.io/technical-debt-vs-architecture-debt-dont-confuse-them/

LeverageAI / Scott Farrell — Practitioner Frameworks

The interpretive frameworks, architectural patterns, and practitioner analysis in this ebook were developed through enterprise AI transformation consulting. The articles below are the underlying thinking behind those frameworks. They are listed here for transparency and further exploration — not cited inline, as this is the author's own analytical voice.

Scott Farrell — Custom Software Didn't Die of Cost. It Died of Verification.

The verification-cost thesis that this edition takes as its spine

https://leverageai.com.au/wp-content/media/articles/105-custom-software-verification.html

Scott Farrell — AI Legacy Takeover

"Trust the tests, not the AI" and characterisation testing as the trust model for governed legacy replacement

https://leverageai.com.au/wp-content/media/articles/48-ai-legacy-takeover.html

Scott Farrell — Start Leveraging AI: The Platform Trap and the Escape Path

The three tiers, the plugging-in-versus-fitting-in test, and the customisation-burden framing

https://leverageai.com.au/wp-content/media/articles/46-start-leveraging-ai.html

Scott Farrell — The Prompt Is Source

Source before the source code; the durable upstream package; the Delete Test

https://leverageai.com.au/wp-content/media/articles/154-the-prompt-is-source.html

Scott Farrell — Wiki Is CapEx

Opex dies with the process, CapEx carries over; capability as the denominator instead of labour hours

https://leverageai.com.au/wp-content/media/articles/112-wiki-is-capex.html

Scott Farrell — Maximising AI Cognition and AI Value Creation

The Cognition Ladder: don't compete, augment, transcend — routed by time allowed for cognition

https://leverageai.com.au/wp-content/media/articles/27-maximising-ai-cognition.html

Scott Farrell — Stop Automating. Start Replacing.

The Spock Question, the fourteen-step tender workflow, and the assistive/layered/reimagined levels

https://leverageai.com.au/wp-content/media/articles/24-stop-automating-start-replacing.html

Scott Farrell — The Open Source Shortcut Trap

The compiled North Star, architecture debt, the fork-gone-wrong shape and the regeneration calculus

https://leverageai.com.au/wp-content/media/articles/148-open-source-shortcut-trap.html

Scott Farrell — The Three-Lens Framework

Pre-build alignment across business value, role impact and measurement — the gate this book deliberately does not own

https://leverageai.com.au/wp-content/media/articles/19-three-lens-framework.html

Scott Farrell — The Terminal Value Doctrine

The board question, the four project classes, the classification filter and "the altitude beats the metric"

https://leverageai.com.au/wp-content/media/articles/61-terminal-value-doctrine.html

Scott Farrell — Agent Addressability

When the agent is the customer, UI polish stops being a moat; state, actions, authority, consequences and change feeds as the delegation surface

https://leverageai.com.au/wp-content/media/articles/111-agent-addressability.html

Scott Farrell — Forward Deployed Engineering

Capability symmetry, the last mile, and accountability for placement and production as the real unit of purchase

https://leverageai.com.au/wp-content/media/articles/175-forward-deployed-engineering.html

Scott Farrell — Sell the Compression, Not the Components

The joined pass from executive problem to governed production system as the unit a buyer can believe

https://leverageai.com.au/wp-content/media/articles/166-sell-the-compression-not-the-components.html

Scott Farrell — Compounded Execution Capital

What compounds after an engagement: compiled judgement, rejected approaches, tests and evidence rather than the artefact

https://leverageai.com.au/wp-content/media/articles/165-compounded-execution-capital.html

About This Reference List

Compiled July 2026. All URLs verified at time of compilation. Regulatory documents and standards specifications are subject to revision — check primary sources for the most current versions.

Some links to academic papers and vendor research may require free registration. Government and standards body publications are freely accessible.