Leverage AI
LeverageAI · Build vs Buy

Don't Buy Software, Build AI Instead: Cheap Code Was Never the Argument

📖 This article has an expanded ebook edition — read the full ebook.

The build-vs-buy case has aged in both directions. Construction cost was never the binding constraint — verification was — and that changes which software you should still be renting.

Scott Farrell · LeverageAI · Second edition, July 2026

TL;DR

The slogan went mainstream. The reasoning didn't.

In December 2025, Andrej Karpathy posted that he had never felt this far behind as a programmer, that he had "a sense that I could be 10X more powerful if I just properly string together what has become available over the last ~year," and that failing to claim the boost "feels decidedly like skill issue."11 Theo Browne, running production teams, put numbers on the same shift: roughly 90% of his own code AI-generated, at least 70% across the teams he runs.12

Six months later, the organisational conclusion everyone drew from that has stopped being a forecast. In Retool's February 2026 survey of 817 builders, 35% had already replaced at least one SaaS tool with something they built themselves, 78% expected to build more in 2026, and 60% had shipped something outside IT oversight in the preceding year.1 That is a build-inclined sample — Retool sells a build tool — so read it as evidence of direction among people equipped to act, not as a population rate. Directionally, it is unambiguous.

Meanwhile the other side of the ledger got worse in ways with dates attached. SaaS pricing rose about 11.4% year on year against 2.7% average G7 inflation.3 In a 2026 survey of IT leaders, 79% met a price increase at renewal, 78% met unexpected charges tied to consumption or AI features, and 61% cut projects because of unplanned SaaS cost increases.4 Gartner has software as the fastest-growing IT segment of 2026, up 15.1%.5 Software is supposedly becoming free, and the software line keeps going up.

So the conclusion is right. Building more of your own software is now rational for a widening class of work. Here is my problem with that: the reasoning underneath it is wrong, and a right answer held for a wrong reason fails silently at the boundary. It will tell you to build things you have no way to verify, and to keep paying for things whose only remaining value was verification you can now produce yourself.

It died of verification, not cost

The standard story is that custom software was too slow and too expensive, so packaged software won on price and speed. That story cannot explain the thirty years it is meant to describe.

A friend of mine tried to build his own manufacturing system. He was not a developer. He could not lead the project. He got comprehensively ripped off — which is roughly what happened to most organisations that tried the same thing. His problem was never that code was expensive. His problem was that he could not tell whether he was getting what he paid for. Every status update was a claim he had no way to check; by the time the gap between the claim and the reality was visible, the money was gone.

That is not a cost failure. It is a trust failure, and a structural one: it does not matter how honest your developer is if you have no independent way to verify the honesty.

Which makes the winner legible in a way the cost story never manages. Packaged software won as a trust technology. You could see the product before paying. Ten thousand prior customers had already de-risked it — they had done your verification, collectively, at their own expense.

"Nobody got fired for buying SAP" was never about SAP being good. It was about the buyer's inability to verify anything else.

Read that way, the old line stops being a joke about corporate cowardice and becomes a precise description of how buyers manage risk when they cannot inspect. If you cannot check the work yourself, you outsource the checking to the crowd. "Everyone else runs it" is the verification.

There is one more piece of evidence, and it is the one that settles the argument. Custom software supposedly died for thirty years — during which the most widely used custom-software platform in history was flourishing in every office in the world. VisiCalc, Lotus 1-2-3, Excel. Corporate IT spent three decades trying to stamp out shadow spreadsheets and never could, because the person with the problem kept building the tool anyway. The spreadsheet was the one place ordinary people got to be build-first the entire time — precisely because it was the one path that did not route through a developer they could not verify. Only the verification story explains both facts at once.

A subscription contains four products. AI repriced one of them.

Once you see the death as a verification problem, the buy side decomposes. What you are actually paying a platform vendor for is at least four separable things:

What's in the subscriptionWhat it really isRepriced by AI?
Operational depthGlobal infrastructure, 24/7 incident response, fraud detection at scale, carrier and regulator relationshipsNo
Liability transferCompliance certifications, audit posture, someone else's name on the failureNo
Verification by crowdTen thousand prior customers who established that it worksYes — replaceable
The software itselfFields, workflows, rules, reports, screensYes — collapsed

This is why "anything you're paying for, replace it with your own software" is a good provocation and a bad rule. Auth, payments, messaging, CDN and storage stay bought — not because their code is hard, but because the value is operational muscle you would have to build a company to replicate. That is Tier 1, and it is not moving.

The flip zone is the middle: horizontal platforms — CRM, ITSM, HCM, ERP, marketing automation — and industry verticals, which promise specificity and deliver an average. The vendor optimises for their 80th-percentile customer; everything that makes you distinctive is, by definition, customisation work. In that zone the licence fee is the small number. HFS Research's November 2025 study of 123 Global-2000-scale organisations found enterprises spend two to seven times licence cost on implementation and integration, with average code reuse of just 33% — teams rebuild about two-thirds of the functionality anyway.2

The first edition of this argument had a table of dollar figures here. Every number in it was invented. I have replaced them with the one sourced magnitude I can defend, because the shape is the point: for configured platforms, the subscription is not the cost.

The instrument: where does the customisation dollar get leverage?

The old question was "is it cheaper to build or buy?" That question is now aimed at the wrong variable. The better one, which I have written about at length elsewhere as the Leverage Test, is: where does your customisation dollar get AI leverage?

Money spent configuring a proprietary platform gets almost none. It lives in closed configuration languages that agents can barely touch, cannot be refactored, cannot be tested, cannot be meaningfully version-controlled, and — this is the part people miss — gains nothing from model improvements. Money spent on spec-driven code gets most of the leverage available, compounds, and treats every model release as a free upgrade. Same dollar, wildly different trajectory.

That asymmetry also explains the debt problem. Platform configuration debt and codebase debt both compound. Only one has fifty years of engineering discipline attached to it. Configuration has the complexity of software with none of the tooling: no refactoring, no automated tests, no diffable history, and no specification to regenerate from. Its five-year trajectory is archaeology — "we can't touch that, no one knows what it does" — and its seven-year trajectory is ossification. A codebase with a spec and a test harness has the opposite trajectory: it can get cleaner at year five, because regeneration keeps getting cheaper.

The 50% diagnostic. If customisation burden exceeds half your total cost of running a platform, run serious build analysis. Under 30%, keep buying. Over 70%, you are already building custom software — you are just doing it without version control, tests, or the word "development" anywhere in the budget line. This is a prompt for analysis, not a verdict.

Your configuration is your exit specification

Here is the part that inverts the usual lock-in conversation, and the place where my first edition was not wrong so much as too timid. I filed vendor lock-in under costs. It belongs under assets.

SaaS vendors write thousands of features. Each feature is used by a small fraction of customers — but everyone uses a different fraction. That looks like bloat. It is the moat. And the switching cost was never the data export; exporting a database has been easy forever. The switching cost was that re-specifying your particular slice was prohibitively expensive, so you stayed.

Now run an agent at it. Your custom fields are data requirements. Your workflow rules are business rules. Your reports are output specifications. Your integrations are interface contracts. Your validation rules are the constraints that are hardest to discover and are already written down. That configuration is not technical debt; it is an executable record of what you actually need, and AI is extremely good at exactly this kind of structure-to-structure translation.

The vendor spent twenty years encoding every customer's requirements into their platform — and it turns out they were maintaining everyone's exit specification, at their own expense.

Two honest caveats. First, the configuration does not contain everything: the workarounds, the "we always do it this way", the tribal knowledge in the head of the admin who is one resignation from being a single point of failure. Capture that by interview and observation before you touch anything. Second, extraction is worth doing even if you decide to keep buying. A costed, credible build alternative is the only thing that has ever changed a renewal conversation. The analysis pays for itself before the decision is made.

The harness — and the two things it does not cover

So what replaced the crowd of ten thousand prior customers? A test suite you own and can read as a percentage climbing toward 100.

The mechanism is characterisation testing: treat the running system — including the platform you are leaving — as its own oracle. Capture what it actually does for known inputs. Turn those input/output pairs into executable, falsifiable assertions. Pass or fail is binary. It is not a matter of opinion and it is not a matter of reading code. That is the instrument my friend never had: a way for a non-developer to verify a system's behaviour without inspecting a single line of it. You don't trust the AI. You trust the tests.

This is why the return of custom software is a spiral rather than a circle. Two things changed together — build cost collapsed and verification restructured — and only the second one explains why the outcome should be different this time.

Now the limits, at full strength, because a piece that skips them is not worth reading.

A behavioural harness proves behaviour. It does not prove safety. Veracode's March 2026 longitudinal study across more than 150 models found that only about 55% of generation tasks produce secure code, and that this rate has stayed essentially flat between 45% and 55% since 2023 — while syntax correctness climbed past 95%.6 Recent flagship releases produced no statistically significant improvement. Your characterisation tests will pass happily on an injectable query. Security needs its own gate, and it does not arrive with the model.

A harness does not stop debt accumulating. A large-scale April 2026 study of 302,600 verified AI-authored commits across 6,299 repositories found 484,366 distinct issues introduced, more than 15% of commits from every assistant carrying at least one, and — the number that should govern your maintenance planning — 22.7% of tracked AI-introduced issues still alive at the repository's latest version, with the cumulative surviving count passing 100,000 by February 2026.7 Cheap generation produces volume, and volume produces residue.

Which is exactly why the durable asset cannot be the code. Keep the specification, the acceptance tests, the context and the decisions that cannot be safely re-inferred; treat the generated implementation as regenerable output. Otherwise you have not escaped a vendor's maintenance liability. You have taken it in-house and put your own name on it.

What I got wrong in January

The first edition of this piece reached the right conclusion by the wrong route. Five corrections, in the open.

  1. I argued from cost. The variable is verification. "AI made code cheap, so build" is the popular version and it is not sufficient. Build cost collapsing tells you building is possible; verification restructuring tells you it is safe. Only the second one is a decision rule.
  2. I led with developer-productivity statistics. Task-completion speed was never capable of settling a procurement question. The famous Copilot result — 55.8% faster — was measured on implementing an HTTP server in JavaScript.10 That is a real finding about a real task and it tells you nothing about whether to replace your ITSM platform. The same applies in the other direction: METR's widely cited finding that experienced developers were 19% slower with AI8 was revisited by METR itself in February 2026, which reported that developers are now likely more sped up than the 2025 estimate, that participants were increasingly refusing to work without AI, and that the new data is "only very weak evidence" either way.9 Both camps were arguing about a variable that does not decide this.
  3. I filed lock-in as a cost when it is an asset. Too cautious. Installed configuration is the exit specification, maintained at the vendor's expense.
  4. I implied the durable asset was the software. It is not. It is the specification plus the harness. Software that cannot be regenerated is not ownership; it is a liability with your name on it.
  5. I invented numbers. A $500K-to-$50K build-cost collapse cited to my own site. A platform amortisation ladder of $200K, then $80K, then $40K. Win rates moving from 12% to 34%. None of those had sources. They are gone, and where I can only defend a shape — the second build costs less than the first because the spec and harness carry over — this edition says shape and stops.

The decision sheet

Run this on one platform, not on your portfolio. The portfolio version is a strategy offsite; the single-platform version is a decision.

  1. Decompose the subscription. Which of the four products are you actually buying — depth, liability, verification, software? If the answer is mostly the last two, continue.
  2. Measure customisation burden. Licence versus everything else: implementation amortised, consultants, internal admins, integration maintenance, change overhead. Over 50%, keep going.
  3. Run the Leverage Test. Can AI help with this customisation at all? If your customisation lives in a proprietary configuration language, every dollar you spend there is a dollar that will not compound.
  4. Extract the spec before deciding. Export the configuration. Translate fields, rules, reports, integrations and permissions into a platform-independent specification. Interview for the shadow spec. This artefact is valuable in every branch, including "we stayed."
  5. Write the characterisation tests first. Capture the current system's actual behaviour as executable assertions. If you cannot write the tests, the decision is already made: keep buying. That is not a failure; it is the instrument working.
  6. Add the gates the harness does not cover. Security review, dependency and supply-chain checks, a named accountable human, and a measured baseline recorded before anything is built.
  7. Parallel run, then reverse-able cutover. The old system stays up. The percentage climbs toward 100. Nobody bets the quarter on a demo.
  8. Check the altitude. Rebuilding the same workflow in your own code, faster, is still a decision about making the current operating model run better. Worth doing; not a strategy. The board's version of this question is what the company is worth when software, analysis and basic advice are cheap — and that question is answered by what you build next, not by what you stop renting.

The last thing worth saying is the thing that has aged best. The risk did not disappear when code got cheap; it migrated up the stack, into the specification. My friend would still fail today if he could not articulate what his manufacturing system was supposed to do — he would just fail in a night instead of a year, which is genuinely most of the improvement. Every era's binding constraint has climbed: hardware, then developers, then licences, then integration. It has finally arrived at intent.

So stop pricing the build. Price the specification and the proof. The bespoke era is back, and this time the hard part is you.

If you are weighing a build-versus-buy call, the question to sit with isn't "can we build it?" It's "can we specify it, and can we prove it?" LeverageAI works with mid-market organisations on exactly that decision — including the analysis that ends with "keep buying, and here is the costed alternative that changes your renewal."

References

  1. Retool. "The Build vs. Buy Shift: How Vibe Coding and Shadow IT Have Reshaped Enterprise Software." Published 17 February 2026; survey of 817 Retool builders and customers fielded late 2025. https://retool.com/blog/ai-build-vs-buy-report-2026 — "35% of them have already replaced at least one SaaS tool with a custom build"; "78% expect to build more of their own tools in 2026"; "60% of builders across levels of seniority have built something outside IT oversight."
  2. HFS Research (with Unqork). "AI Won't Save Enterprises from Tech Debt Unless They Change the Architecture First." 18 November 2025; 123 respondents at Global 2000–scale organisations, surveyed September 2025. https://www.hfsresearch.com/press-release/ai-wont-save-enterprises-from-tech-debt-unless-they-change-the-architecture-first/ — organisations spend "2–7x their license cost on implementation and integration"; "average code reuse is just 33%, meaning teams rebuild about two-thirds of functionality."
  3. SaaStr. "The Great Price Surge of 2025." https://www.saastr.com/the-great-price-surge-of-2025-a-comprehensive-breakdown-of-pricing-increases-and-the-issues-they-have-created-for-all-of-us/ — "SaaS pricing is up by approximately 11.4% compared to the same time in 2024—a stark difference from the 2.7% average market inflation rate of G7 countries."
  4. Zylo. "2026 SaaS Pricing Trends" (reporting Zylo's 2026 SaaS Management Index), page dated 9 June 2026. https://zylo.com/blog/saas-pricing-trends — "79% of IT leaders encountered price increases at renewal in the past 12 months"; "78% of IT leaders experienced unexpected charges tied to consumption or AI features during the past year"; "61% of organizations cut projects or initiatives because of unplanned SaaS cost increases in the past 12 months."
  5. Gartner, 2026 worldwide IT spending forecast, as reported by Campus Technology, 20 May 2026. https://campustechnology.com/articles/2026/05/20/gartner-estimates-worldwide-it-spending-at-6-31t-for-2026.aspx — total IT spending forecast $6.31 trillion for 2026; "the software vertical is forecast at $1.44 [trillion] in spending, a 2026 growth of 15.1%." (Secondary reports of the software dollar figure vary between $1.43T and $1.44T; the growth rate is consistent.)
  6. Veracode. "Spring 2026 GenAI Code Security Update." 24 March 2026; longitudinal study across 150+ large language models. https://www.veracode.com/blog/spring-2026-genai-code-security/ — "only 55% of generation tasks result in secure code"; security performance has "remained essentially flat, hovering between 45% and 55%" since 2023.
  7. Liu, Y., Widyasari, R., Zhao, Y., Irsan, I. C., Chen, J. & Lo, D. "Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild." arXiv:2603.28592v2, 26 April 2026; 302,600 verified AI-authored commits across 6,299 GitHub repositories. https://arxiv.org/html/2603.28592v2 — "484,366 distinct issues"; "more than 15% of commits from every AI coding assistant introduce at least one issue"; "22.7% of tracked AI-introduced issues still survive at the latest version of the repository."
  8. METR. "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity." 10 July 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ — randomised controlled trial; developers took 19% longer with AI tools while estimating they had been 20% faster.
  9. METR. "We are Changing our Developer Productivity Experiment Design." 24 February 2026. https://metr.org/blog/2026-02-24-uplift-update/ — "we believe it is likely that developers are more sped up from AI tools now — in early 2026 — compared to our estimates from early 2025"; "because of the selection effects in our experiment, our data is only very weak evidence for the size of this increase."
  10. Peng, S., Kalliamvakou, E., Cihon, P. & Demirer, M. "The Impact of AI on Developer Productivity: Evidence from GitHub Copilot." arXiv:2302.06590, 13 February 2023. https://arxiv.org/abs/2302.06590 — developers using GitHub Copilot "completed the task 55.8% faster than the control group" on an HTTP-server implementation task.
  11. Andrej Karpathy, post on X, 27 December 2025. https://x.com/karpathy/status/2004607146781278521 — "I've never felt this much behind as a programmer… I have a sense that I could be 10X more powerful if I just properly string together what has become available over the last ~year and a failure to claim the boost feels decidedly like skill issue."
  12. Theo Browne. "You're falling behind. It's time to catch up." YouTube. https://www.youtube.com/watch?v=Z9UxjmNF7b0 — "I am writing the majority of my code with AI now. I would say more than the majority, like 90%. And for the teams that I run, we're at at least 70% AI generated code."
  13. Rahul, head of applied AI at Ramp, post on X, 31 December 2025. https://x.com/rahulgs/status/2006090208823910573 — "you are guaranteed to lose if you fall behind"; "If agents are being held back because of lack of context that's your fault."
  14. Scott Farrell / LeverageAI. "Custom Software Didn't Die of Cost. It Died of Verification." https://leverageai.com.au/wp-content/media/articles/105-custom-software-verification.html — verification cost, packaged software as trust technology, the spiral, and the risk migrating into the spec.
  15. Scott Farrell / LeverageAI. "Start Leveraging AI: The Platform Trap and the Escape Path." https://leverageai.com.au/wp-content/media/articles/46-start-leveraging-ai.html — the three tiers, the Leverage Test, the 50% diagnostic, platform debt versus code debt, sunk cost as spec.
  16. Scott Farrell / LeverageAI. "Maximising AI Cognition and AI Value Creation." https://leverageai.com.au/wp-content/media/articles/27-maximising-ai-cognition.html — the Cognition Ladder: don't compete, augment, transcend.
  17. Scott Farrell / LeverageAI. "The Terminal Value Doctrine." https://leverageai.com.au/wp-content/media/articles/61-terminal-value-doctrine.html — the altitude check and the four project classes.
  18. Scott Farrell / LeverageAI. "Stop Automating. Start Replacing." https://leverageai.com.au/wp-content/media/articles/24-stop-automating-start-replacing.html — the Spock Question and why speeding up a bad process makes it fail faster.
  19. Scott Farrell / LeverageAI. "The Open Source Shortcut Trap." https://leverageai.com.au/wp-content/media/articles/148-open-source-shortcut-trap.html — the compiled North Star, architecture debt, and regeneration versus forking.
  20. Scott Farrell / LeverageAI. "The Prompt Is Source." https://leverageai.com.au/wp-content/media/articles/154-the-prompt-is-source.html — the upstream source package and the Delete Test.
  21. Scott Farrell / LeverageAI. "Wiki Is CapEx." https://leverageai.com.au/wp-content/media/articles/112-wiki-is-capex.html — opex dies with the process; the compounding asset carries over.
  22. Scott Farrell / LeverageAI. "Agent Addressability." https://leverageai.com.au/wp-content/media/articles/111-agent-addressability.html — when the agent is the customer, UI polish stops being a moat.
  23. Scott Farrell / LeverageAI. "Forward Deployed Engineering." https://leverageai.com.au/wp-content/media/articles/175-forward-deployed-engineering.html — capability symmetry, the last mile, and the evidence class ladder.
  24. Scott Farrell / LeverageAI. "Compounded Execution Capital." https://leverageai.com.au/wp-content/media/articles/165-compounded-execution-capital.html — what an organisation actually owns after a build.
  25. Scott Farrell / LeverageAI. "Sell the Compression, Not the Components." https://leverageai.com.au/wp-content/media/articles/166-sell-the-compression-not-the-components.html — the unit a buyer can believe.
  26. Scott Farrell / LeverageAI. "The Proposal Compiler." https://leverageai.com.au/wp-content/media/articles/32-proposal-compiler.html — the same inversion of customisation economics, applied to selling.