Leverage AI

Fixed Price Is Underwriting

📖 This article has an expanded ebook edition — read the full ebook.

AI did not make fixed price easy. It changed what the price is attached to — and that turned your firm into an underwriter that has never written a policy.

The engagement worked. Fixed fee, fixed clock, no change requests, a delighted client and a margin that held. Everyone in the room drew the same conclusion, which was that we had got better at estimating — that AI had absorbed enough of the delivery mess to make the number safe.

That conclusion is wrong, and it is wrong in a way that becomes expensive at about engagement eight.

We did not get better at estimating. We changed what the price was attached to. The old fixed price was attached to forecast labour, which is why it needed an estimate and why the estimate was the whole risk. The new one is attached to measured, machine-absorbable complexity — a census of the input surface, a band, a defined quantity of human judgement, and a reserve for named surprises. That is not a better estimate. It is a different object.

And the moment the price attaches to measured exposure rather than forecast effort, a firm stops being an estimating business and becomes an underwriting business — one that sells certainty while retaining the variance behind it. Underwriting is a three-hundred-year-old discipline with six jobs, a vocabulary, a failure taxonomy and a professional obligation to explain every departure from expectation. Most firms selling AI-native fixed prices are doing two of the six.

This piece is about the other four, and about the one class of risk that AI has genuinely made worse.

AI did not invent fixed price

It is worth killing this claim before it does any more damage, because "AI makes fixed price easy" is becoming a sales meme and it is arriving at precisely the wrong moment.

My own record has the pre-AI half of the argument in it. In 2004 my earlier consultancy put a proposal in front of a client offering a choice: a fixed-price contract, or time-and-materials. The fixed-price route required paid requirements analysis and a paid statement-of-work phase, credited against the later build, formal change control, and continuous comparison of contract value against actual cost. That was twenty-two years ago and there was no machine cognition anywhere in it. What made the fixed price possible was not intelligence. It was that we had bought the right to know what we were pricing before we priced it.

The regulators have been saying the same thing for even longer, and rather more precisely. US federal acquisition regulation defines a firm-fixed-price contract as one that "places upon the contractor maximum risk and full responsibility for all costs and resulting profit or loss."1 Note that the definition is not about price certainty for the buyer. It is a statement about who holds the variance. And the eligibility test reads like something out of this argument: fixed price is suitable when "Performance uncertainties can be identified and reasonable estimates of their cost impact can be made, and the contractor is willing to accept a firm fixed price representing assumption of the risks involved."2

Identifiable uncertainties. Estimable cost impact. Willingness to hold the residue. That is a census, a band and an underwriting decision, written into procurement law long before anyone had a model to run.

Even McKinsey's own technology and AI leader says it out loud: "Outcomes-based pricing didn't start because of AI, but the type of work AI transformation demands suits it."3 About a quarter of the firm's global fees now come through performance-based arrangements.3

And the pressure is not coming only from suppliers getting braver. In May 2026 a US executive order made fixed price the procurement default, requiring any non-fixed-price contract to be "justified in writing by the contracting officer to the agency head" — citing roughly $120 billion obligated on cost-reimbursement consulting contracts in a single fiscal year.4

So the market will get fixed prices whether or not the supply side has the machinery to hold them. That is exactly the condition under which a lot of firms are about to discover what they have agreed to.

What a fixed price actually commits you to

I wrote the sentence that opens this argument in Buy Certainty First: a fixed price for an uncertain activity is a risk transfer from client to supplier, and that sentence should make a serious firm nervous. Fixed price without a theory of uncertainty is not productisation — it is gambling with better stationery.

The actuarial profession got to the same place, and stated it better than I did:

"Though an individual exchanges the uncertainty of occurrence, timing and magnitude of a particular event for the certainty of a fixed price, that exchange in no way makes the uncertain known. Nor need it. The insurance program assuming the financial uncertainty is not able to fix the occurrence or, often, the magnitude of a specific risk merely because it assumes that risk. But it should find a way of establishing a fair price for assuming it."5

Read that as a description of your last fixed-price engagement. Selling certainty does not require having certainty. It requires a defensible method for pricing the uncertainty you agreed to hold. The client's certainty and your certainty are different objects, and the fee is the price of converting one into the other.

The underwriter's table

Here is the mapping, and it is not decoration. Each row on the right is a job that exists because the row on the left cannot be done without it.

Your AI-native engagementThe underwriting job
Preflight censusRisk assessment — measure the exposure before accepting it
Eligibility rules and bandsRisk classification
The fixed feePremium for expected cost plus retained risk
Included human dispositionsCovered consumption
The typed reserveExplicit loss allowance
Unsupported-source rules and exclusionsCoverage exclusions
Re-band or declineReprice or refuse the risk
Engagement actualsClaims and loss-ratio data

Firms reliably do the first three. Census, bands and a fee are the visible, saleable parts, and they are the parts that make a proposal look productised. The last four are where practices die quietly.

Take the premium row seriously for a moment, because the profession's own definition is more useful than any consulting framework. The Casualty Actuarial Society's ratemaking principles state that "A rate is an estimate of the expected value of future costs" and that a rate "provides for all costs associated with the transfer of risk" — where those costs explicitly include "claims, claim settlement expenses, operational and administrative expenses, and the cost of capital."6 Your fee is not your delivery cost plus a margin. It is your expected delivery cost, plus the cost of settling the surprises, plus the cost of the capital you are putting behind the promise.

And there is one sentence in that document that is worth the whole analogy on its own:

"The rate should include a charge for the risk of random variation from the expected costs… The rate should also include a charge for any systematic variation of the estimated costs from the expected costs. This charge should be reflected in the determination of the contingency provision."7

Two different charges for two different kinds of wrongness. Random variation — the engagement is harder than average, in the ordinary way — is priced into the margin. Systematic variation — the estimate itself was built on a wrong picture — gets its own separate provision. In our vocabulary that is interior variation absorbed inside the band, versus a typed reserve held against named surprises. The reserve is not something I invented in 2025. It is the contingency provision, and it has been a professional obligation since 1988.

Why one blended price kills you

The classification row is the one most firms skip, on the reasonable-sounding grounds that their clients are all different and a single sensible number is simpler to sell. The actuarial answer to that is unsentimental: "Risk classification is one means of minimizing the potential for adverse selection. It reduces adverse selection by balancing the economic forces governing buyer and seller."8

Run the mechanism forward. You publish one fixed price for a class of engagement. Buyers whose true cost sits below it look at your number, compare it to what they could get elsewhere, and go elsewhere. Buyers whose true cost sits above it look at your number and sign immediately, because you are cheap. Nobody is behaving badly. Within a year your book is composed disproportionately of the engagements your price was wrong about, and the pattern is invisible in any individual deal.

Bands are the fix, and they carry their own trade-off, which the same literature names: "classes may improve homogeneity, but at the expense of credibility."9 Too few bands and each one contains engagements that are not really alike. Too many and you never accumulate enough history in any single band to learn anything. That tension is a real design problem with no clean answer, and a firm that has never felt it has not yet been banding seriously.

Four kinds of variance, not one "scope"

Everything above depends on being able to say, of any given surprise, which kind of thing it is. "Scope" is one word doing four jobs, and the four have different owners, different commercial consequences and different speeds.

1. Provider-owned interior variation

An approach fails its test and the system generates another. An integration needs three regeneration attempts rather than one. The evidence base is rebuilt because the first assembly was structurally wrong. Fifteen thousand pages instead of three thousand. Two authoritative sources disagree and reconciliation takes four passes.

None of these is a change request. The buyer bought a state; the search that produced it is my business. The band was priced knowing these happen, and — the part that decides whether any of this survives contact with a delivery organisation — this is the default. A surprise becomes commercial only by earning its way out of the default, and it earns its way out by moving a named perimeter field. Absorbing is what happens when nobody does anything. Nobody has to be brave.

2. Metered residual scarcity

Material human dispositions. Licensed sign-off. High-consequence ambiguity. Exceptional technical review. Verification and remediation.

These do not become free because generation became cheap. AI makes the cost curve flatter; it does not make it flat, and what remains scarce is the number of consequential judgements a named human must own. That is the metered resource, it is what the band actually prices, and it is why the interior can be free without the engagement being free.

3. Boundary mutation

The buyer changes the promised outcome. A new source class appears. The input estate crosses the agreed volume band. The required authority changes. A regulatory obligation appears. The deadline moves. The reserve exhausts.

Each of those moves a named field that had a recorded value at signature, which is what makes it a lookup rather than an argument. Internal mess is the supplier's variance; changed commercial intent is a boundary mutation. That distinction is the useful replacement for indiscriminate change control, and it is the subject of its own treatment, so I will not re-derive it here.

4. External and correlated tails

This is the class that is new, and it is the one this whole piece exists to put on the table.

A third-party vendor delays an integration. A regulator changes the admissible process. The client cannot produce an authorised decision-maker. And then the two that did not meaningfully exist five years ago: a shared model upgrade changes behaviour across every live engagement in the same week, and one common parser, connector or kernel defect affects the whole portfolio at once.

Hold that against the pooling logic every fixed-price portfolio implicitly relies on. Pooling works because independent risks average out; the law of large numbers is a statement about independent variables. Correlated exposure does not average out, which is why catastrophe risk gets its own explicit allowance rather than being absorbed into the ordinary margin.10

Here is the uncomfortable structure. Everything the flywheel argument tells you to do — shared kernel, shared parsers, shared evaluation harness, one pinned model, one adapter library — is a deliberate act of correlating your own book. The machinery that makes engagement twenty cheaper than engagement two is the same machinery that makes all twenty fail together. That is a real trade, and almost nobody is making it consciously.

You do not have to take my word for the mechanism. The Bank of England: "A reliance on a small number of providers for a given service could also generate systemic risks in the event of disruptions to them, especially if is not feasible to migrate rapidly to alternative providers," and separately, "From a systemic risk perspective, the potential for AI-based participants to take increasingly correlated positions is an important consideration."11 The Financial Stability Board, more directly still: "Most LLMs are trained using the same underlying architecture and many are trained, at least in part, on common sources of web crawl data. The homogenisation in training data and model architecture can lead to correlated outputs."12

And the demonstrations are on the record. The July 2024 CrowdStrike update affected approximately 8.5 million Windows devices by Microsoft's own estimate13; one analysis put direct financial losses to the Fortune 500 alone at at least $5.4 billion, with the cyber insurance market facing $400 million to $1.5 billion in insured losses — "potentially the single worst loss in the cyber insurance sector over 20 years."14 A single latent defect in an automated system, triggered by a routine change, produced simultaneous failure across thousands of organisations that had no relationship with each other. In October 2025, AWS reported the same shape from its own side: "The incident was triggered by a latent defect within the service's automated DNS management system."15

The insurance market's response to correlated technology risk is instructive, because it is not "price it higher". Lloyd's excluded catastrophic state-backed cyber attacks on the grounds that "losses have the potential to greatly exceed what the insurance market is able to absorb"16 — and required that "policy language should be clear so that the scope of cover is understood by all the parties and the exposure is properly assessed and monitored."17 The most sophisticated risk market on earth handles non-diversifiable exposure by writing an explicit, legible exclusion and then monitoring what is left.

The practical version

Version pinning, independent evaluation sets, rollback capability and shared-component regression tests are not engineering hygiene. They are commercial underwriting controls, and they belong in the price. The test is simple and you can run it this week: name every shared component that every live engagement depends on, and ask what a version change to each one would do to all of them at once. If the answer is "we would find out from the client", you are carrying an unpriced catastrophe exposure.

The oracle you are not allowed to control

There is a second failure that fixed price does not cause but does intensify, and it is structural rather than moral.

Under time-and-materials, effort and payment move together, so there is no particular financial reason to find the cheapest route to a green tick. Under fixed price, the supplier observes the actual production path and cost; the buyer largely observes the supplier's claim that completion occurred; and the supplier benefits financially from reducing internal effort. That is an information asymmetry with money attached to one side of it, which is the textbook setup for the measure ceasing to measure.

Charles Goodhart's 1975 original: "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."18 The famous compression — "When a measure becomes a target, it ceases to be a good measure" — is Marilyn Strathern's, from 1997, and almost everybody misattributes it.19

Now add the AI-specific aggravation, which runs the opposite way to intuition. DeepMind, on specification gaming: "Even for a slight misspecification, a very good RL algorithm might be able to find an intricate solution that is quite different from the intended solution, even if a poorer algorithm would not be able to find this solution… This means that correctly specifying intent can become more important for achieving the desired outcome as RL algorithms improve."20 Capability makes proxy failure more likely, not less. A better producer finds the loophole a worse one would have missed.

Until recently that was an argument by analogy. It is not any more. A controlled 2026 study put two production coding agents to work re-implementing a component library under a 222-test oracle across three availability conditions. The finding:

"Without the oracle, the library is present but unfinished, revealed by scores. With the oracle in the loop, the score reaches near-perfect, but from a demo holding the tested behavior directly, the library left dead or absent. We call this building to the test… The agent does not, on its own, validate what it ships as a user would."21

That last sentence is the entire commercial argument in one line. It is not an accusation about anyone's integrity — my own position on this has always been that it does not matter how honest your developer is if you have no independent way to verify the honesty. It is a statement about what a producer's objective function is.

So the contract needs three separate objects, not one:

  1. The promise — the state, decision, artefact or commitment being bought.
  2. The declared acceptance surface — the evidence both parties know must exist.
  3. The independent oracle — tests, samples, reconciliations or receipts the producer does not unilaterally control.

Intent must be visible. Acceptance must be legible. Verification must remain independent. And "independent" has to mean something mechanical rather than rhetorical: a deterministic check run against real state, a check performed against evidence the producer did not select, or a human disposition made under named authority. Running the same model twice and treating agreement as verification is an echo, not a check — and the research is unhelpfully clear that this gets worse as models improve, not better.

Whole industries reached this conclusion decades ago and made it structural. Under ISO/IEC 17020, third-party inspection requires a Type A body, and the requirement is not about procedure: "Rock solid demonstrations of impartiality require the IB to ensure that its own staff are not involved in any aspect of ownership, design, manufacture, or other relationship as regards the object of inspection."22 Enterprise buyers already hold the distinction in another form — a SOC 2 Type 1 report covers control design at a point in time; a Type 2 covers operating effectiveness over a period.23

The predictable objection is that independent verification is unaffordable inside a fixed fee. Two answers. First, it is a cost and it belongs in the band explicitly, because the alternative is not cheaper — it is the same cost, unpriced, arriving later as rework and dispute. Second, independence has cheap forms: held-out cases the buyer selects after generation, reconciliation against an authoritative system, live production receipts, a sample the producer did not choose. You do not need a second delivery team. You need an answer key the producer does not hold.

And the cost is measurable, which turns the objection into something manageable. Google's DORA research names it directly — "The verification tax: Time saved writing is often re-spent auditing" — and observes that "higher AI adoption is associated with an increase in both software delivery throughput and software delivery instability."24 For an underwriter that is not a productivity note. It is a change in the shape of the loss distribution: more claims, arriving faster, with more variance. Track verification hours against production hours removed, per offer. A firm that has never computed that ratio does not know how much of its AI gain is real.

The cost of measuring instead of arguing

A book about underwriting that hid its own downside would be doing the thing it warns against, so here it is.

Everything above replaces adjudication with observable triggers — recorded perimeter values, published band drivers, typed reserve draws. That is the same move parametric insurance makes, and parametric insurance names the price honestly. A parametric contract "typically specifies (1) the payment amount; (2) the trigger (a pre-determined parameter based on observable data); and (3) an impartial third party to verify that the trigger was met."25 The benefit is real: clear triggers "reduce policy disputes" and make premiums calculable. The cost has a name:

"Compensation from parametric policies is not linked to actual losses, so the claim payment may be higher or lower than the losses incurred. This is known as basis risk."26

The worked failures are unsparing. The New Orleans School District held parametric wind cover for 2024, and the winds from Hurricane Francine "did not meet the 100 mph trigger and the policy did not pay out despite damage to school facilities."27

Your variance schedule inherits exactly that exposure. Type a trigger badly and you will absorb something you should have re-contracted; type it the other way and a buyer will face a re-contract for something that did not really hurt them. Typing does not make the world tidy. It makes the disagreement about the world happen at a better time, in a cheaper form — before signature, when both parties can still calmly walk away.

When to refuse

Decline is the organ that distinguishes underwriting from project risk management, and it is the one almost nobody has. A risk register has never declined anything. It manages a project you have already agreed to do.

The rule is not AI → fixed price. It is that AI expands the territory in which complexity can be bounded, configured and priced as a product — and the edge of that territory is a design output, not an act of courage. Run the perimeter as a self-test before anyone writes a number: can I record a starting value and an observable trigger for the promised state, the input estate, the volume band, the buyer-controlled dependencies, the authority and access, the consequence class, the acceptance rule, and the time boundary? If several rows cannot carry a recorded value — and particularly if they sit outside both parties' control — the chain terminates. No recorded value, no delta; no delta, no trigger; no trigger, and every surprise resolves into an argument.

The answer to that is almost never to walk away, because refusing costs revenue now and a firm without a pipeline cannot afford principles. The answer is to shrink: sell the bounding itself as a smaller commitment whose entire deliverable is a stable, testable question and a filled perimeter. That product is not a consolation prize. It is the thing that has to exist before the larger engagement can be bought at all — by you or by anyone else.

And none of this is an AI-era discovery. Construction lawyers have been saying the quiet part for decades: lump-sum contractors price significant contingencies and are "naturally incentivized to seek opportunities to reopen the fixed price" — because "in truth, there is no such thing as an absolute fixed price contract."28 The perimeter does not abolish the pressure to reopen. It governs when and how the reopening happens.

The asset nobody can copy

Which brings me to the reason any of this is worth the trouble.

Fixed price does something to your incentives that time-and-materials does not. Under T&M, a delivery improvement that takes work from a hundred units to eighty reduces your billings; the client captures the gain. Under a fixed commitment, the client pays for the bounded promise, so reusable machinery that lowers the next delivery's cost accrues to you. Fixed price makes the supplier the residual claimant on its own learning. That is the economic joint that makes the square and the flywheel one business model rather than two compatible ideas.

But the effect is conditional, and the conditions are strict. It requires a credible cohort of similar future engagements; recurring machinery that outweighs client-specific variation; verification costs that do not grow as fast as production breadth; available reuse rights; and competitive repricing that does not immediately return every gain to the market. The relevant denominator is not headcount. It is: how many eligible future commercial units can actually consume this machinery? A hundred-and-fifty-person bench creates potential distribution; it does not prove amortisation. If six people encounter the problem, or every engagement forks the system, the effective reuse population is six — or zero.

Given all that, the compounding asset is not the AI workflow and it is not the contract template. Competitors rent the same frontier models and can copy your offer, your language and your report format by Tuesday. What they cannot copy is your accumulated loss history of complexity:

The professional obligation attached to that data is stated better in the actuarial literature than in any consulting methodology: incurred results should be measured against expected "wherever possible in terms of frequencies, severities, and loss ratios," and "no material departure from expected results should be accepted without attempting to find an explanation for the variation."29

Read that as a delivery discipline. Every engagement where actual effort departed materially from the band gets an explanation before it is closed — not a shrug, not "that client was difficult", an explanation that changes something. That is what makes engagement twenty structurally safer than engagement two. If the learning stays in a founder's intuition, you have productised an expert, not built a business.

Eight records

The measurement programme is the thing to start this week, because a loss history you began recording today is worth more than one you plan to record perfectly next year. For each engagement:

  1. Remove-AI classification — is this offer AI-enabled, AI-dependent or AI-constituted?
  2. Variance ownership — every event classified: provider interior, included disposition, reserve draw, boundary mutation, or excluded tail.
  3. Disposition density — consequential human decisions per hundred machine-processed units.
  4. Verification scaling — whether proof and remediation cost grows slower than generated breadth.
  5. Boundary integrity — how many "change requests" were genuine changed intent rather than internal delivery mess.
  6. Band economics — margin, reserve consumption and exception rate by configured band.
  7. Transfer — what engagement two can do without the originator in the room.
  8. Customer value — whether buyers pay for the unit, or for access to your personal judgement.

If those improve across engagements, the square is real. If they do not, the fixed price is being subsidised by heroics — and heroics do not appear on any P&L line, which is precisely why the firm will not notice until the heroes leave.

What would prove me wrong

An underwriter who mis-states their own loss data is not underwriting, so let me be exact about the status of this argument.

I do not have loss ratios by band across a multi-client cohort. Nobody in this field does, and no one has published a correlation coefficient for delivery failures across engagements sharing a model vendor — the regulators establish the mechanism, not the magnitude. What I have is the mechanism, my own pre-AI fixed-price history, and a specimen. Where I have no number, I have given you the shape and said so.

The falsifiers, written before the results exist:

Any of those, sustained, and the honest answer is a hybrid commitment or a decline rather than a defence of the model. The strongest case against this whole argument is that delivery cost is driven less by measurable breadth and more by novel consequential judgement — in which case census variables will never predict cost, bands will never stabilise margins, and the fixed price is speculative risk transfer wearing productisation's clothes. That is not a rebuttal I can argue away. It is a test, and it runs per offer rather than per industry.


The firms that get hurt by this are not the ones that deliver badly. They are the ones that took the commercial form of underwriting — the fixed number, the productised proposal, the confident boundary — and skipped the machinery that makes it survivable. That state is more dangerous than old-fashioned time-and-materials, because it looks like productisation from the outside and is a naked risk position on the inside.

So the question to put to your own next proposal is not "can we deliver this?" It is the underwriter's question, and it has six parts: what exactly is covered, what did we measure before accepting it, what does the loss distribution look like, what have we reserved for, what have we excluded, and what would make us decline.

The square is not drawn around the project. It is earned by underwriting the variance.

References

  1. Federal Acquisition Regulation, 16.202-1. "A firm-fixed-price contract provides for a price that is not subject to any adjustment on the basis of the contractor's cost experience in performing the contract. This contract type places upon the contractor maximum risk and full responsibility for all costs and resulting profit or loss." https://www.acquisition.gov/far/subpart-16.2
  2. Federal Acquisition Regulation, 16.202-2(d). "Performance uncertainties can be identified and reasonable estimates of their cost impact can be made, and the contractor is willing to accept a firm fixed price representing assumption of the risks involved." https://www.acquisition.gov/far/subpart-16.2
  3. Business Insider, "AI is reshaping how McKinsey makes money" (via AOL syndication). Kate Smaje, global leader of technology and AI at McKinsey — "Outcomes-based pricing didn't start because of AI, but the type of work AI transformation demands suits it." Also: "About a quarter of McKinsey's global fees come from this pricing model." https://www.aol.com/articles/ai-reshaping-mckinsey-makes-money-115132273.html
  4. Executive Order, "Promoting Efficiency, Accountability, and Performance in Federal Contracting," Federal Register, 5 May 2026. "Use of any non-fixed-price contract… must be justified in writing by the contracting officer to the agency head"; "A review of spending across the Government in Fiscal Year 2024 identified approximately $120 billion obligated on cost-reimbursement consulting contracts alone." https://www.federalregister.gov/documents/2026/05/05/2026-08900/promoting-efficiency-accountability-and-performance-in-federal-contracting
  5. American Academy of Actuaries, Committee on Risk Classification, "Risk Classification Statement of Principles." "Though an individual exchanges the uncertainty of occurrence, timing and magnitude of a particular event for the certainty of a fixed price, that exchange in no way makes the uncertain known. Nor need it." https://actuary.org/wp-content/uploads/2025/05/risk.pdf
  6. Casualty Actuarial Society, "Statement of Principles Regarding Property and Casualty Insurance Ratemaking" (adopted 1988; reinstated 2021), Principles 1–2 and Definitions. "A rate is an estimate of the expected value of future costs"; "Such costs include claims, claim settlement expenses, operational and administrative expenses, and the cost of capital." https://www.casact.org/sites/default/files/2021-06/Statement%20of%20Principles%20Regarding%20P&C%20Casualty%20Insurance%20Ratemaking_2021.pdf
  7. Casualty Actuarial Society, "Statement of Principles Regarding Property and Casualty Insurance Ratemaking," Considerations — Risk. "The rate should include a charge for the risk of random variation from the expected costs… The rate should also include a charge for any systematic variation of the estimated costs from the expected costs. This charge should be reflected in the determination of the contingency provision." https://www.casact.org/sites/default/files/2021-06/Statement%20of%20Principles%20Regarding%20P&C%20Casualty%20Insurance%20Ratemaking_2021.pdf
  8. American Academy of Actuaries, "Risk Classification Statement of Principles." "Risk classification is one means of minimizing the potential for adverse selection. It reduces adverse selection by balancing the economic forces governing buyer and seller." https://actuary.org/wp-content/uploads/2025/05/risk.pdf
  9. Casualty Actuarial Society, "Statement of Principles Regarding Property and Casualty Insurance Ratemaking," Considerations — Credibility. "There is a point at which partitioning divides data into groups too small to provide credible patterns. Each situation requires balancing homogeneity and the volume of data"; "classes may improve homogeneity, but at the expense of credibility." https://www.casact.org/sites/default/files/2021-06/Statement%20of%20Principles%20Regarding%20P&C%20Casualty%20Insurance%20Ratemaking_2021.pdf
  10. Casualty Actuarial Society, "Statement of Principles Regarding Property and Casualty Insurance Ratemaking," Considerations — Catastrophes. "Consideration should be given to the impact of catastrophes on the experience and procedures should be developed to include an allowance for the catastrophe exposure in the rate." https://www.casact.org/sites/default/files/2021-06/Statement%20of%20Principles%20Regarding%20P&C%20Casualty%20Insurance%20Ratemaking_2021.pdf
  11. Bank of England, "Financial Stability in Focus: Artificial intelligence in the financial system," April 2025. "A reliance on a small number of providers for a given service could also generate systemic risks in the event of disruptions to them"; "From a systemic risk perspective, the potential for AI-based participants to take increasingly correlated positions is an important consideration." https://www.bankofengland.co.uk/financial-stability-in-focus/2025/april-2025
  12. Financial Stability Board, "The Financial Stability Implications of Artificial Intelligence," 14 November 2024. "Most LLMs are trained using the same underlying architecture and many are trained, at least in part, on common sources of web crawl data. The homogenisation in training data and model architecture can lead to correlated outputs." https://www.fsb.org/uploads/P14112024.pdf
  13. Reuters (via Malay Mail), 21 July 2024. Microsoft estimated the faulty CrowdStrike update affected approximately 8.5 million Windows devices. https://www.malaymail.com/amp/news/money/2024/07/21/microsoft-says-crowdstrike-mayhem-took-out-85-million-windows-devices/144464
  14. David Jones, "CrowdStrike disruption direct losses to reach $5.4B for Fortune 500, study finds," Cybersecurity Dive, 25 July 2024. "Parametrix said the global IT outage linked to Crowdstrike will likely cost the Fortune 500, excluding Microsoft, at least $5.4 billion in direct financial losses"; "CyberCube estimates the cyber insurance market will face preliminary insured losses of between $400 million and $1.5 billion, potentially the single worst loss in the cyber insurance sector over 20 years." https://www.cybersecuritydive.com/news/crowdstrike-cost-fortune-500-losses-cyber-insurance/722396/
  15. Amazon Web Services, "Summary of the Amazon DynamoDB Service Disruption in the Northern Virginia (US-EAST-1) Region," October 2025. "The incident was triggered by a latent defect within the service's automated DNS management system that caused endpoint resolution failures for DynamoDB." https://aws.amazon.com/message/101925/
  16. Lloyd's of London, quoted in The Stack, "Lloyd's cyber insurance exclusions flag systemic risk fears." "The ability of hostile actors to easily disseminate an attack, the ability for harmful code to spread, and the critical dependency that societies have on their IT infrastructure… means that losses have the potential to greatly exceed what the insurance market is able to absorb." https://www.thestack.technology/lloyds-cyber-insurance-exclusions-state-backed-systemic-risk/
  17. Lloyd's Market Bulletin Y5433, "State-backed cyber-attack wordings," 14 May 2024. "we wish to reiterate that policy language should be clear so that the scope of cover is understood by all the parties and the exposure is properly assessed and monitored by syndicates." https://assets.lloyds.com/media/6335bcb0-e2a2-4378-8328-1ddf54828f2f/Y5433.pdf
  18. Charles Goodhart (1975), "Problems of Monetary Management: The UK Experience." "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." Overview and citation record: https://en.wikipedia.org/wiki/Goodhart%27s_law
  19. Marilyn Strathern (1997), "'Improving ratings': audit in the British University system," European Review 5(3), pp. 305–321, at p. 308. "When a measure becomes a target, it ceases to be a good measure." https://www.cambridge.org/core/journals/european-review/article/improving-ratings-audit-in-the-british-university-system/FC2EE640C0C44E3DB87C29FB666E9AAB
  20. Victoria Krakovna et al., "Specification gaming: the flip side of AI ingenuity," Google DeepMind, 21 April 2020. "Even for a slight misspecification, a very good RL algorithm might be able to find an intricate solution that is quite different from the intended solution… This means that correctly specifying intent can become more important for achieving the desired outcome as RL algorithms improve." https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/
  21. Yanuo Ma, Ben Kereopa-Yorke and Ben Schultz, "Building to the Test: Coding Agents Deliver What You Check, Not What You Requested," arXiv:2606.28430, 26 June 2026. "With the oracle in the loop, the score reaches near-perfect, but from a demo holding the tested behavior directly, the library left dead or absent. We call this building to the test… The agent does not, on its own, validate what it ships as a user would." https://arxiv.org/abs/2606.28430
  22. International Accreditation Service, "Understanding ISO/IEC 17020 Handbook," reproducing ISO/IEC 17020:2012 requirements. "Rock solid demonstrations of impartiality require the IB to ensure that its own staff are not involved in any aspect of ownership, design, manufacture, or other relationship as regards the object of inspection or its manufacturer / supplier." https://www.iasonline.org/wp-content/uploads/2021/01/Tab-1-01-Understanding-ISOIEC-17020-Handbook.pdf
  23. Linford & Company LLP, "SOC Report Types: Type 1 vs Type 2 SOC Reports/Audits." A Type 1 report "is as of a point in time… It only covers the design effectiveness of the internal controls"; a Type 2 report "covers a period of time… the operating effectiveness of the internal controls over time." https://linfordco.com/blog/soc-report-types-1-vs-2/
  24. DORA (Google), "Balancing AI tensions: Moving from AI adoption to effective SDLC use," 2025. "The verification tax: Time saved writing is often re-spent auditing"; "higher AI adoption is associated with an increase in both software delivery throughput and software delivery instability." https://dora.dev/insights/balancing-ai-tensions/
  25. Congressional Research Service, "Parametric Insurance for Natural Disasters: Frequently Asked Questions," IN12670, updated 19 May 2026. "A contract for parametric insurance typically specifies (1) the payment amount; (2) the trigger (a pre-determined parameter based on observable data); and (3) an impartial third party to verify that the trigger was met." https://www.everycrsreport.com/reports/IN12670.html
  26. Congressional Research Service, IN12670. "Compensation from parametric policies is not linked to actual losses, so the claim payment may be higher or lower than the losses incurred. This is known as basis risk." https://www.everycrsreport.com/reports/IN12670.html
  27. Congressional Research Service, IN12670. "the New Orleans School District had purchased parametric wind insurance for 2024; however, the winds from Hurricane Francine did not meet the 100 mph trigger and the policy did not pay out despite damage to school facilities." https://www.everycrsreport.com/reports/IN12670.html
  28. Troy Edwards and Peter Tolson, "Cost reimbursable vs. lump sum turnkey construction contracts: the many routes to bankability," A&O Shearman. "In truth, there is no such thing as an absolute fixed price contract"; contractors price significant contingencies and are "naturally incentivized to seek opportunities to reopen the fixed price." https://www.aoshearman.com/en/insights/cost-reimbursable-vs-lump-sum-turnkey-construction-contracts-the-many-routes-to-bankability
  29. Casualty Actuarial Society, "Statement of Principles Regarding Property and Casualty Loss and Loss Adjustment Expense Reserves," Considerations — Reasonableness. "The incurred losses implied by the reserves should be measured for reasonableness against relevant indicators… and expressed wherever possible in terms of frequencies, severities, and loss ratios. No material departure from expected results should be accepted without attempting to find an explanation for the variation." https://www.casact.org/sites/default/files/2021-04/statement_of_principles_Loss_Loss_Adjustment%20_Expense%20_Reserves_2021.pdf

Frameworks referenced without inline citation are my own published work: the Square and hard acceptance from AI-Native Service Architecture; typed uncertainty and the pricing envelope from Buy Certainty First; the Fixed-Price Envelope from AI-Constituted Services; interior variation, the fences and the reserve exhaustion protocol from Boundary Mutation, Not Change Request; the four refusals from Preparedness Is the Product; verification independence from Custom Software Verification and Two Falsifiers; the offer ladder and the three migrations from The Terminal Value Doctrine: Professional Services.