Two Falsifiers
Verification semantics for AI-native services

Two Falsifiers

Should this exist, and was the promise kept?

One acceptance test cannot answer two questions. Every bounded offer needs an entry falsifier that can stop the work before it starts, and a distinct exit falsifier that can refuse it at the end.

Tests tell you whether the machine worked. Falsifiers tell you whether the commitment survived.

By the end of this book you can

  • ✓ Write two different disconfirmation conditions before work starts
  • ✓ Build a Falsifiability Spine for any bounded offer — six joints, each with an owner
  • ✓ Tell an implementation test from outcome evidence, and stop asking one to do the other’s job
  • ✓ Map every claim to a deterministic check, bounded AI judgment, or an accountable human
  • ✓ Make “stop” a valid accepted result — and get paid for it
  • ✓ Draft both ends into a contract, with the consequence attached
01
Part I: Two Questions, One Instrument

What Was Actually Accepted

The invoice goes out and nobody argues. Ask the awkward question anyway.

The closing meeting runs the way closing meetings run. The work landed. The client is pleased. The partner is pleased. Somebody says “genuinely excellent job” and means it. The invoice goes out and nobody argues about it.

Now ask the awkward question. Not what was delivered — that part is easy, there is a list, and everything on it exists. What was accepted? Put more usefully: what observation, had it come out the other way, would have made a named person on the buyer’s side say no, that isn’t the thing we bought?

In most engagements the honest answer is a feeling, held by several senior people at roughly the same time. That is not a defect of those people. They were paying attention and would justify the call well if you asked. What they did not have was an instrument.

And it is worse than one missing instrument, which is what this book is about. It is one instrument being asked to do two entirely different jobs.

Three objects, not one

The distinction that starts to fix this is already published, and worth crediting before extending.

Definition

Outcome
What was sold. The bounded state the buyer is paying to have exist.
Test
The instrument by which we can legitimately say the outcome occurred.
Falsifier
A named observation, written in advance, that would show the outcome did not occur — or, at the other end, that it should never have been attempted.

One sentence of illustration: if I sell you a verified current-state architecture, the architecture is the outcome; the evidence coverage, the reconciliation checks and the human dispositions are the tests by which I can prove it.

Collapse the two and “outcome-based” reverts to vague consulting language, for a precise reason: there is nothing left that could fail.

That book placed the falsifier in the acceptance sequence and then deliberately stopped, saying so in terms: the semantics of falsification — what makes an engagement wrong to start versus wrong to accept — is a distinct piece of doctrine with its own treatment. This is that treatment.

The second collapse

Separating outcome from test gets you halfway. The test is still answering two questions, and they live at opposite ends of the engagement.

Should this commercial object exist at all? Is this the right problem, the right boundary, the right placement of machine cognition, the right buyer, an economics that closes? That question is answerable in the first fortnight and unanswerable after the fact.

Was the promise, as written, actually kept? Does the bounded state exist, and is there evidence a stranger could check? That question is meaningless before the work and decisive at the end of it.

They fail independently, which is what makes one instrument insufficient rather than merely inelegant. The premise can be sound and the delivery broken. Or the delivery can be immaculate and the premise dead — the beautifully executed irrelevance. Everything works. Nothing moves.

Key Insight

Any single instrument covering both moments is guaranteed to be the wrong shape for at least one of them. Not likely to be. Guaranteed.

The fields that separated them thirty years ago

None of this is our idea, and the fastest way to make it credible is to say whose it is. Systems engineering has held the distinction since before most of us were working.

Verification is “the confirmation, through the provision of objective evidence, that specified requirements have been fulfilled.”1 Validation is “the confirmation, through the provision of objective evidence, that the requirements for a specific intended use or application have been fulfilled.”2

The two sentences look almost identical and the difference between them is the whole book. One is measured against the specification you wrote. The other against what the buyer actually needed. Different processes, different references. An implementation test is a verification instrument. An exit falsifier is a validation instrument. Asking the first to do the second’s job is not a stretch; it is a category error with a standards-body vocabulary available to describe it.

Regulators are blunter about this than consultancies are. In the FDA’s software validation guidance, “software testing is one of many verification activities intended to confirm that software development output meets its input requirements” — the test demoted to a component of a component — and the honest reason a promise needs a falsifier rather than exhaustive proof follows immediately: “a developer cannot test forever, and it is hard to know how much evidence is enough.”3

The sentence that should unsettle you

Here is a man writing about the most rigorous software assurance regime humanity has built, explaining what it cannot do. Modern aircraft are extraordinarily safe and no serious airplane incident has been traced to faulty software — “some have been traced to faulty requirements for systems implemented in software, but that is outside the remit of DO-178C.”4

A world-class verification standard, applied perfectly, still leaves failures traced to faulty requirements — because requirement correctness is outside its remit.

Read that against your own acceptance criteria. Whatever sits at the end of your engagement, however rigorous, it is testing the thing you specified and has nothing to say about whether you specified the right thing. That gap has a name in this book — the entry falsifier — and nothing on the right-hand side of an engagement can fill it.

Why the collapse is newly dangerous

The two questions were always distinct. What changed is that you can no longer get away with conflating them, and the reason is structural rather than nostalgic. When production was expensive, activity was a serviceable proxy for progress: you could watch the workshops happening and reasonably believe the work was happening, because the workshops were most of the cost. Cheap probabilistic production breaks that proxy in both directions at once. The interior can now take an unexpected route and still be exactly right — and it can produce a great deal of confident, well-formatted, internally consistent output that is wrong, faster than anyone can read it.

Our own prior work put the general form of this plainly: control that lives only as reading the middle will always lose to the volume of the middle. That is the parent doctrine and this book does not re-teach it. What matters here is the consequence: when process conformance stops being a control, the two ends carry the entire load — and they currently carry one instrument between them.

The enemy, named

There is a second antagonist in this book, softer than the category error and more common. It is “outcome-based” used as a mood: a promise phrased so warmly that no observation could contradict it.

Pitfall: the criterion that cannot fail

“Stakeholders confirm the recommendations are actionable.”

Which stakeholders, by name? What would non-confirmation look like as an observation? What happens commercially if they don’t? Three questions, no answers — and no result of the engagement that would make the sentence false.

That is not a weak test. It is a sentence wearing the costume of a test. A criterion with no false state is not a low bar; it is not a bar. Chapter 4 takes it apart properly and repairs it.

The line the rest of this book is downstream of

There is a sentence underneath everything that follows, and it belongs here rather than at the end.

Tests tell you whether the machine worked. Falsifiers tell you whether the commitment survived.

Two proofs are coming. In Chapter 9 a migration passes every implementation test it has and does not close its promise. In Chapter 11 an engagement stops in week two and the supplier is paid in full, because stopping was the thing that was sold.

Both are downstream of a structural question, which is where the next chapter starts: if one instrument cannot do both jobs, what does an engagement actually look like when it carries two?

02
Part I: Two Questions, One Instrument

The Falsifiability Spine

Six joints. Remove any one and the failure it produces is specific.

An engagement that carries two instruments has a shape, and the shape is already half-published. The acceptance sequence for a bounded AI-native engagement runs:

published sequence
Promise → Falsifier → Latitude → Evidence → Acceptance

That is the inherited structure. I want to make one change to it, in the open, so you can see exactly what was taken and what was added: the second joint is doing two jobs, so it becomes two joints.

the Falsifiability Spine
Promise → Entry falsifier → Method latitude → Evidence chain → Exit falsifier → Acceptance

Six joints. This is the artefact the rest of the book fills in, and it is not re-derived after this chapter — later chapters point back at it.

Joint Question it answers Who owns it, and when it is written
PromiseWhat bounded state will exist when this is over?Both parties, fixed at signature
Entry falsifierWhat observed fact would show this should not proceed?Written by the supplier before signature; callable by a named person on each side
Method latitudeWhat is the interior free to do without asking?Supplier, silently, throughout
Evidence chainWhat must the interior produce as it goes, rather than reconstruct afterwards?Supplier, continuously; specified before work starts
Exit falsifierWhat evidence would show the promise was not kept?Written before signature; evaluated by the buyer’s named acceptor
AcceptanceDid the falsifier stay silent, and is the chain complete?Named humans, both sides, one event

Two things in that table are easy to skim past. The entry falsifier is callable by a named person on each side — the buyer can invoke it too, which stops it being a supplier’s escape hatch. And the evidence chain is produced as the work goes: a chain assembled in the final week is a report about the work rather than evidence of it, and a falsifier cannot fire on a report.

Why this is a spine and not a checklist

A checklist is a set of items that could be reordered or shortened without changing what it is. A spine is a structure where removing any element produces a specific, predictable failure. Take them out one at a time.

Remove the promise and the other five joints have no referent. This is the one people forget, because a vague promise still feels like a promise. But a falsifier is always of something; if the promise is “we will help you understand your estate”, no quantity of rigour downstream will produce a sharp falsifier, because there is nothing sharp to falsify.

Remove the entry falsifier and production absorbs a dead premise. The engagement gets executed beautifully and should never have started. Nobody can point at a failure, because there isn’t one — there is an absence, and absences do not appear in status reports.

Remove method latitude and you are buying activity again. Every economic gain the architecture offers gets paid straight back out in supervision, and the cheap interior buys the buyer nothing at all.

Remove the evidence chain and the exit falsifier has nothing to fire on. It degrades into a retrospective argument about what everyone remembers, which is the condition the instrument was built to escape.

Remove the exit falsifier and acceptance resolves to seniority. Whoever is more senior, or simply more determined, decides whether the promise was kept.

Five removals, five different failures. That is the test of whether a structure is real rather than decorative: if two joints produced the same failure when removed, they would be one joint, and the boundary between them would be wrong.

Why falsifier and not criterion

This is the mechanism underneath the whole design, and it is worth slowing down for.

A test constructed to demonstrate that a thing happened can nearly always succeed. The space of confirming observations is large, and the producer chooses which observations to make. That is not cheating — it is what a competent supplier does: you build the evidence you know how to build, from the work you actually did.

A condition constructed to demonstrate that a thing did not happen works differently. It has to name, in advance, an observation that would be uncomfortable to find. Once that observation is named, whether it occurred is a matter of looking rather than of arguing.

Key Insight

A confirming test lets the producer choose the evidence. A disconfirming observation, named in advance, does not.

Which gives the timing rule everything else depends on. Both falsifiers must be written before the work. Written afterwards, a falsifier is a rationalisation with a footnote — you will select the observation you already know did not occur, and you will do it sincerely.

Popper, once

The logical shape here is old, and I want to spend exactly one section on it. Popper’s demarcation rests on an asymmetry: “it is logically impossible to verify a universal proposition by reference to experience… but a single genuine counter-instance falsifies the corresponding universal law.”5 Corroboration counts “only if it is the positive result of a genuinely ‘risky’ prediction, which might conceivably have been false.”5

A test that could not have come out badly is not evidence. It is a ritual.

Take the caveat too, because without it this book would be making a naive claim it cannot support. Popper explicitly allowed that in practice a single conflicting instance is never methodologically sufficient for falsification on its own, and that theories are often retained in the face of anomalous evidence.5 That matters commercially rather than philosophically: a falsifier firing produces a disposition, not an automatic termination. The vocabulary for that arrives in the next chapter.

That is the whole of the philosophy. This is a book about commercial semantics.

“You’ve just moved the argument earlier”

Yes. That is exactly what this does, and it is the strongest thing about it.

Consider when acceptance disputes are currently settled. By the time a surprise arrives, both parties have a financial interest in the answer, at least one is embarrassed, and the relationship is the loudest thing in the room. Under those conditions the answer is not determined. It is negotiated — and whoever negotiates better, or has more to lose, wins.

Move the decision to a point where nobody has an interest yet and the character of the question changes. It stops being who is right about what we meant and becomes did the thing we wrote down happen. Arriving at the answer becomes a lookup. The same reasoning underlies classifying commercial surprises before an engagement starts rather than after one arrives.

The honest residue: this does not make the conversation easy. It makes it early, which is a different and much more achievable thing. Nobody has ever complained that a difficult conversation happened while it was still cheap.

Where the spine sits

One boundary, so nothing is misread. The spine sits inside a bounded commercial perimeter this book does not redraw — promise, unit and band, time boundary, authoritative inputs, terminal states, exclusions, acceptance tests, authority. That topology is published; it is the parent architecture here, not the contribution. What this book adds is what happens at the two ends of that perimeter.

One last thing before we build the joints, because it shapes both of the next two chapters. The two falsifiers look symmetrical on the page and are not symmetrical in practice. One of them is written by the party who loses money when it fires. The other is evaluated by the party who paid.

03
Part I: Two Questions, One Instrument

Should This Exist? The Entry Falsifier

The sentence a supplier writes, before signature, that would mean they should not be hired.

Start with the uncomfortable half of that asymmetry. The entry falsifier is a supplier writing down, before signature, the observation that would mean they should not be hired.

Every incentive in professional services pushes against writing that honestly, and hardest exactly when the pipeline is thin — which is to say, when the temptation to take the wrong engagement is highest. Any doctrine that relies on people being good at that moment is not a doctrine; it is a hope. So the resolution is not virtue. It is that the disconfirmation has to have a paid form.

The note in the source is a claim about placement rather than enthusiasm: “Falsifiable keeps coming up. I think that’s probably more important at the outset.” The move is not “falsifiability is good”. It is falsifiability belongs here.

Seven ways an engagement is wrong before it starts

The question has a fixed form — what evidence would make us say… — and seven answers. These are not risk categories. Each is a class of engagement that should not proceed, with a characteristic observation attached.

1. Wrong problem

The thing being solved is not the binding constraint. The observation is not “we disagree about priorities” — it is a measurement showing the reachable portion is too small to move the aggregate. Chapter 11 runs this one end to end.

2. Wrong client

The organisation cannot receive the outcome. No decision owner; or no authority to act within the window; or the decision has already effectively been made and the engagement is being bought as cover.

3. Wrong promise

The state named in the promise is not the state the buyer needs. Usually discovered by asking what they would do with it — the promise that survives no contact with an intended action is the wrong promise.

4. Wrong commercial boundary

The perimeter as drawn either excludes something the outcome depends on, or includes something nobody can evidence. Both are fatal and neither is visible without drawing the perimeter first.

5. Impossible economics

The consequential human dispositions required to keep the promise are a large fraction of the unit count — so the offer is human-constituted whatever the technology says. This one is measurable, which makes it the most useful class in the list.

6. Inappropriate AI placement

The work in question is authority, consequence or physical capacity — not cognition. Cheap cognition does not reach it, and no amount of cheaper cognition ever will.

7. Stop or redesign

The residual class. The premise survives, but not in this shape. This is the reshape result, and it is the most common one in practice.

Notice which are observable in week one. Classes 2, 4 and 6 usually are; class 5 is checked before quoting; classes 1 and 3 depend on what the promise is about. An entry falsifier can only be built from the observable ones — a kill class you cannot see inside the window is a design constraint, not a falsifier.

The four parts, all mandatory

Anatomy of an entry falsifier

  • The observation. A fact, not a judgement. “The sponsor seems disengaged” is a judgement. “No named individual with authority to make this decision inside the window has been identified” is an observation. Omit it and you have a worry.
  • The window. A date, not “early”. Omit it and the falsifier is never called.
  • The caller. A named person on each side. Not a committee. Omit it and the falsifier becomes a discussion.
  • The accepted results. Stop, reshape, or proceed — written before, so that firing is a disposition rather than a crisis. Omit it and firing becomes a negotiation.

All four, or it is decoration. This is the commercial version of a discipline our pre-thinking work already applies to analysis — the first-signal test: name the observable evidence that could change the frame, and make it falsifiable inside a short horizon rather than eventually.

Why the clock is load-bearing

The window looks like administrative tidiness and is not. Without it the entry falsifier competes with sunk effort, and it loses on a curve that steepens weekly. Week one: nothing spent, nobody told, and calling the falsifier costs an awkward email. Week nine: three people staffed, two steering updates delivered, and calling it costs a reputation. The observation is most available early and least expensive early, and those two facts point the same way.

There is a second effect. An early window makes the falsifier easier to write honestly, because the supplier commits to observe something before investing in the answer. Ask a team in week eight what would show the engagement was wrong and you get a careful, defensible, useless sentence.

Stop and reshape are accepted results

Here is the claim that gets softened, so I will state it without softening: a responsible bounded service must permit a finding that the intervention should not proceed — and must be able to be paid for it.

The mechanism is structural rather than moral: the disconfirmation is the deliverable. The buyer purchased the answer, not the build — which makes the entry falsifier partly a product design instrument rather than only a governance one. A supplier who sold the build cannot stop and be paid; they can only fail gracefully.

Our own certainty-product doctrine already states the commercial position:

“Sometimes the right outcome of a certainty purchase is that the transformation should not proceed now. That is not a failed sale. It is the option resolving to zero exercise, which can be the highest-value resolution for both parties.”

And the sentence that separates firms who mean it from firms who say it: “Firms that only celebrate exercised options will pressure consultants to recommend builds. Firms that celebrate correct non-exercise will keep the mirror honest. Your incentive design will show which one you are.”

What is delivered, what is invoiced and what is signed when this actually happens is Chapter 11, worked all the way down. It is not a footnote and it does not belong here.

An institution built for exactly this

Clinical trials face this problem where stopping late kills people. A Data Monitoring Committee “reviews accumulating data on a regular basis… and recommends to the sponsor whether to continue, modify, or stop a trial or trials”, and it “is established by the sponsor but should be independent of the sponsor and the trial conduct”.6

The payload is what a DMC is permitted to say. It can recommend the sponsor stop the trial because the investigational product “is not effective”.6 That sentence is pre-authorised. Nobody has to be brave enough to invent it in the room.

Four design rules, transferable as they stand

  1. The stop authority is named in advance, in a charter written before the work.
  2. It is independent of the party doing the work.
  3. “Is not effective” is an explicitly listed legitimate reason to stop.
  4. The stopping analysis uses planned procedures, not improvised judgement.

A ten-week commercial engagement is not a clinical trial. What transfers is those four rules, not the ceremony.

The standard already contains the determination

If that feels exotic, look at the most-cited AI governance framework in the world. NIST’s AI Risk Management Framework asks, under MANAGE 1.1, for “a determination… as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed.”7 Under MANAGE 2.1 it requires “viable non-AI alternative systems, approaches, or methods” to be taken into account.7 Building nothing is on the table in the standard.

So the determination exists. What does not exist anywhere in the framework is an attachment: no price, no contract, no named acceptance owner, no consequence — a governance activity floating free of the commercial object it would have to govern.

Key Insight

This book is extending a standard, not rivalling one. The determination is already written down; what it lacks is a window, a caller, and an invoice.

When it fires

A falsifier firing produces a disposition, not an automatic termination. The vocabulary is already ours: Kill when continuation is not justified; Fix when the underlying opportunity may be real but a remediable constraint is wrong; Double-Down when the evidence supports concentrated commitment. Fix is the reshape result.

And the reciprocal that almost everyone forgets: record the reopen condition at the same moment. A stop without a reopen condition is an opinion. Chapter 11 writes one out.

“No client will let me write a kill condition into a proposal”

Three answers, in order of how much they help.

Some won’t, and that is information. A buyer who cannot tolerate a named observation that would end the engagement is telling you how the other end will go. The conversation you cannot have in week zero is the one you will certainly have in week twelve, with more money on the table.

The paid form changes what you are asking. You are not requesting permission to walk away; you are describing what they receive if the answer turns out to be no — a deliverable, with a price against it.

And the reframe, which is what lands in the room. A supplier who can name what would make this the wrong engagement has demonstrably thought about the engagement rather than the sale. That reads as competence, not hedging — especially to a buyer burned once by a confident supplier.

Myth vs reality

Myth: a kill condition in a proposal signals doubt and loses deals. Reality: the buyer’s reaction to it is the cheapest qualification signal available, and it arrives before you have spent anything.

Some buyers will refuse the conversation outright. What to do then is a method question, and Chapter 12 answers it with three specific cases rather than a platitude.

One boundary, before we move

There is an adjacent instrument that is easy to conflate with this one. Product-level kill conditions ask whether an offer should keep existing, reviewed on a fixed date against evidence — a product that cannot fail its own test is only an attractive story. The entry falsifier asks whether this instance should proceed. Same logic, different scope. Keep them apart or you will end up cancelling a product because one engagement was wrong.

That is the instrument at the front of the engagement. The one at the back asks the opposite question — not whether there was a promise worth making, but whether the promise that was made was kept — and it turns out to be a differently shaped instrument rather than a stricter one.

04
Part I: Two Questions, One Instrument

Was the Promise Kept? The Exit Falsifier

A stated observation, plus a stated consequence. Anything less is a metric.

Differently shaped, not stricter. That distinction is easiest to see by starting with a criterion that has the wrong shape, and Chapter 1 already put one on the table:

Stakeholders confirm the recommendations are actionable.

The instinct is to call that soft and tighten it — add a stakeholder list, define “actionable”, date the confirmation. That instinct is wrong, and understanding why is most of this chapter. The problem is not vagueness. It is that no result of the engagement would make the sentence false. Tighten it all you like and it still cannot come back negative.

What the instrument actually is

The exit falsifier states what evidence would show the bounded promise was not kept — regardless of activity, effort or technical elegance. And acceptance follows from it, in a definition worth reading twice:

Definition

Acceptance

Acceptance means exactly two things: the exit falsifier was not triggered, and the promised evidence chain is complete. Not that the buyer liked the answer. Not that the calendar ended.

People skip the second clause. A silent falsifier proves nothing if the evidence that would have made it speak was never produced — and without both conditions the cheapest way to pass is to generate less evidence.

“This is just acceptance criteria with extra steps”

Worth meeting immediately, because a reader holding this objection is not reading the rest of the chapter. Three differences, each of which changes an outcome.

Acceptance criterion vs exit falsifier

Acceptance criterion
  • • A positive threshold — can be met in ways nobody intended
  • • Evaluated at the end, usually by the party that produced the work
  • • Describes a desired state; the consequence of missing it is negotiated later
Exit falsifier
  • • A negative observation named in advance — it occurred or it did not
  • • Fires on evidence the producer did not author
  • • States what happens commercially when the observation is true

Concede the honest part: a good acceptance criterion written by a careful person is already halfway here. The falsifier makes “careful person” unnecessary — which matters, because the careful person is not always in the room at the close.

The register already exists

If writing acceptance as a negative sounds strange, an entire profession signs opinions in exactly that form. A limited assurance conclusion is “framed in a negative sense: ‘Based on the procedures performed, nothing came to our attention to indicate that the management assertion on XYZ is materially misstated.’”8

Read that back as a spine joint: the falsifier did not fire. A professional opinion drafted as a disconfirmation, relied upon commercially every day.

Take the honest ceiling from the same source, because this book should not oversell its instrument: assurance risk “is never reduced to nil and therefore, there can never be absolute assurance.”8 A falsifier-based acceptance is not a guarantee. The auditing profession says so in its own standards; so should we.

The drafting lesson: a threshold is not a falsifier

An industry that has been writing bounded promises for a century still fights about the acceptance condition, and one decided case explains precisely why. In Mears Ltd v Costplan Services (South East) Ltd, the parties had contracted what most of us would call a model acceptance criterion: a measurable, unambiguous threshold — a room more than three per cent smaller than drawn. Exactly the kind of clause that survives legal review and looks like rigour.

The Court of Appeal held that the clause “simply provides a mechanism by which a breach of contract can be indisputably identified”, and that materiality had been introduced “only in relation to room size… and not in relation to the resulting breach.”9 The commentary draws the moral for drafters: parties may deem a clause a condition, or deem that its breach precludes practical completion, “but if that option is taken, they should do so clearly so that there is less risk of a dispute as to the proper construction of the clause.”10

Key Insight

An exit falsifier is a stated observation plus a stated consequence. A measurable threshold on its own identifies a breach; it does not close or refuse a promise.

Most acceptance criteria in circulation are thresholds. They say when something is wrong and nothing about what happens next, which is how a clear measurement produces an unclear argument.

One further constraint from the same judgment shapes the design: a latent defect cannot prevent completion. A falsifier can only fire on evidence that exists at the acceptance moment, which is why the evidence chain must be specified in advance rather than assembled in hindsight.

Four diagnostics, and three repairs

Three of these are published; the fourth is what the Mears lesson adds.

  1. Could a competent third party run it? If it requires the author to interpret it, it is a preference, not a test.
  2. Could it come back negative? Write the failing result as a sentence. If you cannot, there is no test.
  3. Would a negative result change what the buyer does? If not, you are measuring something nobody cares about.
  4. Is the consequence written down? The first three tell you whether the observation is real. The fourth tells you whether it is commercial.

Run them against three criteria you have almost certainly signed.

Pitfall: three criteria, and their repairs

“The final report is delivered and presented to the steering committee.”
Delivery is an event, not a state. Repair: name the state the report must establish, the coverage threshold it must reach, and what happens if it is not reached.

“Stakeholders confirm the recommendations are actionable.”
Unfalsifiable by construction. Repair: name the decision the recommendation supports, the evidence class required for that decision to be defensible, and the consequence of the class not being met.

“All identified issues are addressed.”
Circular — identified by whom, against what boundary? Repair: coverage against a declared boundary, typed unknowns for whatever was not observed, and a named consequence for an untyped gap.

Notice what every repair does. It replaces an adjective with an observation, and then attaches a consequence to the observation. That is the entire craft.

Two ways an evidence chain fails

Auditing splits evidence quality from quantity. Sufficiency is “the measure of the quantity of audit evidence”; appropriateness is “the measure of the quality of audit evidence; that is, its relevance and its reliability in providing support for the conclusions”.11

An exit falsifier should be able to fire on either. “We ran a lot of tests” attacks sufficiency only, and says nothing at all about whether the tests were the right kind — which is the failure Chapter 9 walks all the way down. Which instrument is allowed to supply which kind of evidence is Chapter 6.

Typed unknowns are part of the promise

One more thing makes an otherwise impossible falsifier writable. A perimeter that can only emit success is a liar. Unknowns become typed deliverables, and “not observed” within a boundary is never “does not exist”.

The payoff is direct. If the promise is “we will understand your estate”, nothing is writable. If the promise is “every unit reaches a typed terminal state with named evidence”, the exit falsifier states itself: a unit ended in an untyped shrug, or a typed state is not supported by its cited evidence. A typed promise gives you a falsifier almost for free.

Buyers hear typed unknowns as incompleteness, right up until they have been burned once by an over-confident report — after which they become the strongest advocates for typing you will ever meet.

Do this before you read on

Take the acceptance clause of the engagement you are running now and write the failing sentence: the result that would mean the promise was not kept. If you cannot write it, you do not have an exit falsifier — you have a deliverables list, learned in ten minutes rather than at the close.

Everything in this chapter and the last concentrates precision at the two ends of the engagement. That has a cost, and it is the discipline nobody holds: keeping precision out of the middle. It does not fail by argument. It fails by drift.

05
Part I: Two Questions, One Instrument

The Middle You Must Not Supervise

Latitude is never taken away by argument. It is taken away one reasonable revision at a time.

Before defending the middle, look at one honestly, because most people have never been shown what the interior of a bounded AI-native engagement actually does. It assembles the evidence base and then rebuilds it, because the first assembly used a source that turned out not to be authoritative. It generates eight candidate framings and discards six. A line of analysis dies in week three on a constraint nobody declared at kick-off. Two sources disagree and reconciliation takes four passes.

Put that on a timesheet and it is waste. A buyer watching it would reasonably ask why they are paying for the six discarded framings. They are not — the buyer bought a state, and the search that produced it is the supplier’s business. That is not generosity. It is the deal.

What the customer is actually buying

The customer buys precision of intent and precision of outcome. They do not buy precision of internal procedure.

Follow that where it goes, because it is uncomfortable. It reclassifies most of a traditional statement of work as a description of something nobody purchased. Conduct five workshops. Interview twelve stakeholders. Develop design. Run analysis. Hold weekly status meetings. Prepare report.

Those are descriptions of the middle.

Suppliers write them for a reason that deserves respect rather than mockery: activity descriptions are a defence. They make effort visible when outcome is not provable, and for most of the history of professional services it was not. Once outcome is provable, those clauses stop protecting you and become the thing that gets argued about, because they are the only specific commitments in the document.

The parent rule here is published and this book does not re-teach it: Tight Intent, Loose Method — hold purpose constant, and treat mechanism, timing, sequence, commercial model and implementation layer as variables the review is authorised to challenge. In its canonical form the rule is about prompting: “The thing to be precise about is intent. The thing to stop over-specifying is procedure.” What this book does is promote it from a design rule to a contract one: write it into the agreement rather than into the prompt.

Where you stay prescriptive

Latitude is not a licence, and a chapter that only defends freedom will be read as one. There are three places to stay exact.

The useful observation, which saves a decision later: those are exactly the places where the exit falsifier must be deterministic rather than judgemental. The stay-prescriptive list and the deterministic column of the evidence map are the same list seen from two ends.

The skill is sorting “must be exactly X” from “must be good” — and not accidentally prescribing the second kind. Which is what goes wrong next.

How latitude is actually lost

Nobody decides to supervise the middle. There is no meeting where someone proposes abandoning the architecture. What happens is a sequence of individually reasonable revisions to one clause, each with a good local argument, over about three weeks.

The drift sequence — one clause, four revisions

1. The falsifier is written

“Fires if any claim in the pack cannot be traced to the agreed evidence.”

Clean. It names a state.

2. Someone adds an evidence requirement

“…traced to the agreed evidence, with the supporting extract attached to each claim.”

Still defensible — but the clause has started describing an artefact rather than a state.

3. Someone specifies the format

“…extract attached in the agreed traceability matrix template.”

The clause now names a tool.

4. The tool implies a method

“…populated during weekly evidence reviews.”

The acceptance criteria are now a specification of the interior, and there is a recurring meeting in the contract.

Each step had a good local argument. Collectively they reinstated the activity-shaped statement of work the architecture was built to replace.

The asymmetry is what makes this dangerous. Every addition is individually defensible, and there is no natural moment at which someone says that is enough now — because saying it requires arguing against a specific reasonable request on general grounds.

The counter-move

One question, applied to every acceptance requirement before it goes in: does this describe the state, or the route to it? Route language goes back.

Run it on the sequence above. Revision two — “with the supporting extract attached” — describes a state, just about: a claim either carries an extract or it does not. It survives. Revision three names a template, which is a route; it goes back, and the repair is to say what the extract must do (resolve to its source) rather than what it must look like. Revision four is entirely route.

The test also catches the well-meant additions from the buyer’s side that nobody feels able to refuse. It is much easier to say “that describes our route rather than your state” than to say no.

Method divergence is not failure

If the interior takes an unexpected route — a different decomposition, a different model, a different order of attack — and still produces the agreed provenance, coverage and acceptance evidence, that is the architecture working. Not tolerated. Working.

The failure mode to watch for is subtler than supervision: treating divergence as a deviation requiring explanation. It usually arrives as a reasonable request for a status update, and it ends with a supplier who has learned it is cheaper to follow the declared plan than to take the better route. Chapter 10 runs an engagement that abandons its declared approach in week three and keeps its promise anyway.

Aviation made this move on purpose

This is not an efficiency argument dressed up as a principle. The most safety-critical software domain there is made exactly this move, deliberately, for safety reasons. DO-178B was “a total re-write of DO-178 to move away from the prescriptive process approach and define a set of activities and associated objectives that a design assurance process must meet”, allowing “flexibility in the development approaches that could be followed”.12

Objectives, not procedures. Tight intent, loose method, hard verification — in aviation software, thirty years ago. The same standard also answers a question this book has not reached: how much verification is enough? Rigour scales with consequence, and the scaling is published as a table.

DO-178C objectives by Design Assurance Level12

71

DAL A — catastrophic

69

DAL B — hazardous

62

DAL C — major

26

DAL D — minor

0

DAL E — no safety effects

Zero objectives where nothing consequential can happen. Seventy-one where the aircraft can be lost. Falsifier mass scales the same way, and Chapter 7 sets that rule properly.

If the middle is genuinely free, though, everything the buyer relies on has to be produced at the ends — and an exit falsifier is only as good as the evidence it can fire on. Which raises the question the next chapter exists to answer: who, or what, is allowed to supply that evidence?

06
Part I: Two Questions, One Instrument

The Evidence Chain

Three instruments, one claim each — and one rule that costs suppliers money.

A falsifier is only worth writing if somebody can say, without arguing, whether it fired. That requires two things settled in advance: what evidence will exist, and who is allowed to produce it.

The second half is where most designs quietly fail. An exit falsifier whose supporting evidence is authored by the system whose work is being judged will not fire — not because anyone is dishonest, but because the producer’s objective function is completion. Authoring systems optimise for successful completion and green, and that is the objective function of getting the job done.

The Evidence Map

The artefact is a table, built before work starts. Every claim inside the promise is assigned to exactly one of three instruments, with the reason in the row — the reason column is what makes the map arguable later.

Instrument What it is good for How it fails alone
Deterministic check Identity, coverage, reconciliation, completeness, provenance, conformance. It passes or it fails and nobody’s seniority is involved. Silent about everything outside its schema. It can certify the completeness of a hollow process.
Bounded AI judgment Search and semantic breadth — reading everything, nominating candidates, typing exceptions, surfacing the disagreement a human should look at. Soft, correlated, and capturable by eloquent narrative. Commentary without teeth.
Accountable human disposition The consequential calls, made under named authority, each one bounded and recorded. Does not scale — and gets spent on things the first two instruments should have caught.

“Bounded” is doing heavy lifting in that middle row, so make it concrete rather than reassuring: a defined question, a scoped evidence slice, a structured output schema, citation obligations, and no write access to anything authoritative. On the third row, the boundary is equally exact: the machine may investigate, synthesise and nominate every one of those calls. It may not become a disposition, and it may not mutate an authoritative state.

The three are not redundant, and the failure column is why. Deterministic alone certifies hollow process; judgement alone is commentary. Together, the machine sees the messy reality and the machinery ensures the required checks cannot quietly disappear.

The rule

Key Insight

No single-model self-verification may be presented as independent proof.

Be precise, because the rule is both over- and under-applied. It does not forbid using a model to check work; models catch real defects. It does forbid presenting that check as the independent evidence on which an exit falsifier is evaluated, and it forbids treating agreement between two runs, two prompts, or two models of the same family as verification.

Why the rule holds

A firm’s own doctrine is a poor witness in its own defence, and this rule costs suppliers money, so it needs external grounding. Across an evaluation of more than 350 large language models, “models agree 60% of the time when both models err”, and — the part that removes the obvious escape route — “larger and more accurate models have highly correlated errors, even with distinct architectures and providers.”13 In the judge setting specifically, each judge “systematically inflates the accuracy of models that are less accurate than itself, due to correlated errors (the judge marks incorrect answers as correct if both models agree on the incorrect answer).”14

And the trend, which closes off “this will fix itself as models improve”: “model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures.”15

Two numbers that have to be held together

60%

of the time, two models agree when both are wrong — on one leaderboard dataset

80%+

agreement between a strong model judge and human preference — about the same as between humans

The second figure is not a footnote and the literature is not one-sided. Strong judges do match human preferences well, “achieving over 80% agreement, the same level of agreement between humans”.16 So the honest position is narrower than a prohibition and sharper than a preference:

Model judgment is a legitimate link in an evidence chain. It is never the closer. Ranking candidates is not the same job as closing a commercial promise.

Our own corpus circles this from three directions: verification needs mechanically different checks rather than copies of the same model voting on the same prompt; the failure mode of a loop with no external closer is not a crash but “a green tick over work that never happened”; and a check written in the session’s own terms can confirm itself. The external result is what turns those from a design preference into a rule.

Independence, then, has to be mechanical: a different property, checked a different way, against something outside the model. Not a second opinion — a different question. Three shapes qualify: a deterministic check run against real state; a check performed against evidence the producer did not select; and a human disposition made under named authority. One shape does not, and it is the one in constant use: the same model, on the same context, asked whether the work looks right.

What a link in the chain looks like

Every claim travels as four fields rather than one: claim, exhibit, resolvable pointer, and a confession of what could not be verified.

The fourth field is load-bearing rather than polite: you cannot falsify a promise whose supporting claims never declared their gaps. A chain of uniformly confident claims gives an exit falsifier nothing to catch. The confession is the surface the instrument acts on — and it tells the accountable human where to spend attention: where the exhibits are thin.

The analogy that needs no explaining

Any enterprise buyer already holds this distinction. A SOC 2 Type 1 report “is as of a point in time… It only covers the design effectiveness of the internal controls”. A Type 2 report “covers a period of time… the operating effectiveness of the internal controls over time”.17

An implementation test is Type 1. An exit falsifier is Type 2. Used with a CIO or a procurement lead, that sentence ends the conversation in thirty seconds.

A tension worth naming out loud

There is a contradiction here with our own delegation doctrine, better named than discovered. Hidden Gates says hide the rubric from the worker, because “a gate the worker can’t see is a gate it can’t game”. This book says publish the exit falsifier in the contract.

The same source draws the other boundary, and it keeps humans in the right place: gates guard completeness and process. They do not guard taste.

So who decides whether it fired?

Named humans, named in advance, on both sides. The map exists to make that judgement a lookup rather than a debate.

Some rows will still be contested, and the map earns its keep there too, because it tells you what kind of disagreement you are having. A disputed deterministic check is a defect — look at the machinery. A disputed disposition is a judgement call with an owner. A disputed bounded-judgment row almost always means the row was mis-assigned, and the fix is upstream rather than in the meeting.

All of which assumes the falsifier is pointed at something the supplier can actually keep and evidence. If it is not, none of this machinery helps — you have built a rigorous instrument aimed at something outside your control.

07
Part I: Two Questions, One Instrument

What a Falsifier May Attach To

A promise too big to keep is also too big to falsify.

Here is the sentence that undoes everything so far, usually said by the person in the room with the best intentions: “We’ll guarantee the business outcome.”

True business outcomes — realised margin, adoption, fleet uptime across weather and operator behaviour — depend on factors the supplier does not control. Sell those as unbounded guarantees and you price yourself out of the market, destroy margin, or quietly rewrite the promise later through change requests. The safer formulation is not hours versus outcomes; it is hours versus a bounded state, decision, deliverable or commitment.

For this book the danger is sharper than commercial exposure. An exit falsifier attached to a customer business outcome fires for reasons the supplier neither caused nor could have prevented, which is worse than no falsifier at all: it makes the instrument look unserious the first time anyone uses it.

The fence: keepable and evidenceable

A falsifier may attach only to a state the supplier can (a) keep and (b) evidence. Both tests. Neither on its own, and the interesting judgement lives in the cases that pass one and fail the other.

Two tests, three cases

Keepable, not evidenceable

“The client’s team now understands the architecture.”

The supplier can genuinely cause it. Nobody can produce an observation that would show it did not happen. Repair: attach to an artefact and a demonstrable use of it.

Evidenceable, not keepable

“Backlog reduced by thirty per cent within the quarter.”

Beautifully measurable, and dependent on volumes, staffing and an external step the supplier does not control. Repair: attach to the supplier’s contribution — the assessed portion, the demonstrated throughput of the mechanism under stated conditions — and let the buyer own the aggregate.

Both

A verified decision. A maintained state. An assessed estate with typed coverage. A protected period under a defined commitment.

These are commercial units, not feature names — and each of them is a state a falsifier can be pointed at without embarrassment.

One test to carry into a drafting session: if it fires, can we tell whether we caused it? If the answer is no, the falsifier is attached to the wrong thing, and no amount of measurement precision will rescue it.

The experiment has already been run

The public sector has decades of outcome contracting behind it and the findings are not ambiguous. The US Government Accountability Office examined performance-based logistics arrangements and found that of fifteen programme offices, only one had updated its business case with actual cost and performance data. In that single case, “it determined that the performance-based logistics contract did not result in expected cost savings and the weapon system did not meet established performance requirements.”18

And the evidence itself: programme officials “typically relied on cost and performance data generated by the contractors’ information systems”, and the offices “had not determined whether contractor-provided data were sufficiently reliable”.18

Three separate failures, all of them this book’s argument. Fourteen of fifteen never re-tested the premise — no exit falsifier, no owner, no date. The one that did test it, fired — which is what the instrument is for, and exactly why nobody volunteers to run one. And the outcome evidence was authored by the party being assessed — Chapter 6’s rule, breached in contract form, twenty years before anyone said “LLM-as-judge”.

Outcome contracting is not the answer either

A book arguing for falsifiable promises has to say what outcome contracting actually does in the field rather than borrow its glow. It has a documented gaming taxonomy: contracts risk “‘cherry picking’, where eligible individuals are not referred or accepted onto a service if they seem unlikely to achieve payable outcomes, ‘creaming’, in which providers focus their efforts on those individuals who are easiest to help, and ‘parking’, which is neglect of those who may be more difficult to achieve outcomes with.”19 Plus the deeper problem: “it can be difficult to set simple, measurable outcomes that align effectively with complex social problems.”19

So “outcome-based” is not a solution; it is a different set of failure modes. The first three are prevented by eligibility and exclusions at the perimeter — it is the bounded promise, not the falsifier, that stops cherry-picking, and if your perimeter admits selective intake no instrument downstream will help. The fourth is prevented by the fence: attach to what you can keep and evidence, and attribution stops being a contested inference.

Why Goodhart bites harder here

Attribute it properly, because almost everyone gets it wrong. Goodhart’s 1975 original was “any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes”. The famous compression — “when a measure becomes a target, it ceases to be a good measure” — is Marilyn Strathern’s, from 1997.20 Our own corpus already makes that correction and this book stays consistent with it.

The sentence that explains the timing of this whole book, though, is neither of theirs: “the importance of Goodhart effects depends on the amount of power directed towards optimizing the proxy, and so the increased optimization power offered by artificial intelligence makes it especially critical for that field.”21

An AI-native service is, by construction, a large amount of optimisation power aimed at whatever you wrote down as the test.

More capability aimed at a proxy makes proxy failure more likely, not less. Which is the structural reason a falsifier must attach to the state rather than the proxy — and why the problem gets worse as the technology gets better.

How much falsifier a promise deserves

A routine bounded promise earns a light instrument; an irreversible commitment earns full adversarial evidence — the same rule Chapter 5 found in external form. The practical consequence is a discipline rather than a permission: do not write four exit falsifiers for a two-week engagement. Write one that can fire.

The seam with change control

One thing belongs to neither instrument alone and should be said here, because drafters need it: a boundary that moves does not reset the falsifier. It replaces it — and the replacement is written the same way the original was, with an observation, a consequence and a named acceptor. A re-contract that leaves the old acceptance criteria in place has moved the promise and left the instrument pointing at the previous one.

What you now hold

An artefact with six joints. Two instruments with different shapes, at different ends, owned by different people. Three evidence types with one hard rule about who may author proof. And a fence that decides what any of it may be pointed at. None of which is worth anything until it has been run on a real promise — and there are exactly two ways to find that out: write one out in full, and then break it.

08
Part II: Three Spines, and Two Ways They Break

Spine One: The Decision Product

Six joints, filled in. Including the one that was written badly first.

Evidence boundary — stated once, for all of Part II

Every engagement in this part is a worked model with stated assumptions. Names, volumes and windows are chosen to make the mechanism decidable, not to report an outcome. Nothing here is client data, and no figure in Part II is a measured result. The posture is inherited: a specimen is stated as a specimen, and n = 1 is stated as n = 1.

Read on anyway, because a spine either decides things or it does not, and you can test that on a model. What a model cannot tell you is how often each branch occurs in the wild — and this book does not claim to know.

The service

A fixed-price decision product: an engagement whose promise is a decision pack for one named, dated decision. The specimen decision, chosen to be generic and structurally typical: which claims platform do we commit to for the next five years?

This service goes first because it is the purest case. The promise is cognitive, the evidence is documentary, and there is no physical or third-party dependency to muddy the fence from Chapter 7. If the spine cannot be written cleanly here, it cannot be written anywhere.

Joint 1 — the Promise

In the words that would appear in the document:

A decision pack for the named decision, containing every viable option, each traceable to the agreed evidence set and scored against the agreed decision criteria, with every residual uncertainty typed.

Then the half most promises omit, which is what makes the rest of the spine writable.

Every clause in the promise names a state, and states can be disconfirmed. “We will support the decision” cannot — which is why a weak promise cannot be rescued by a strong falsifier.

Joint 2 — the Entry falsifier

ObservationEither (a) no named individual with authority to make this decision inside the stated window has been identified; or (b) fewer than three of the four declared authoritative evidence sources are accessible inside the audit boundary.
WindowEnd of week one.
CallerThe named delivery lead and the named sponsor, jointly.
Accepted resultsStop; or reshape to a smaller square whose deliverable is a decidable question.

Now the part that matters more than the table: why those two observations and not others.

Observation (a) is kill class two — wrong client, an organisation that cannot receive the outcome. It is trivially observable in week one, and it is the failure most often discovered in week nine, when the pack lands and nobody has the authority to act on it.

Observation (b) is kill class four — wrong commercial boundary. A perimeter that excludes something the outcome depends on. The three-of-four threshold is a judgement, and I want to be honest that it is one: it is set where the pack can still carry a defensible option set. A service promising exhaustive coverage would set it at four of four. A service promising a directional read might set it at two.

Two kill classes are deliberately not in this falsifier, and the reasons are instructive. Impossible economics is checked before quoting, when the evidence map is first drafted, not in week one when the price is already agreed. Wrong problem is not observable this early for a decision product, because the buyer named the decision — there is nothing for the supplier to discover about whether it is the right one. Selecting which classes an entry falsifier can actually watch is most of the craft.

Joint 3 — Method latitude

Free, and commercially silent: framing generation, retrieval strategy, model choice, number of passes, reconciliation route, how many candidate option sets are built and discarded, whether the evidence base is torn down and rebuilt. None of it reaches the buyer and none of it is a change request.

Not free, and named explicitly — Chapter 5’s three exceptions plus one commercial one:

  1. The pack’s structure, because it feeds a downstream reader and must match the agreed shape.
  2. The typed-state vocabulary — a fixed catalogue, with no synonyms invented under pressure.
  3. The provenance format: every claim’s pointer must resolve.
  4. Authority: no authoritative state is mutated by anything other than a named human, ever.

Enumerating the exceptions protects the latitude rather than eroding it. An unbounded claim of freedom invites a buyer to bound it for you, in a meeting, using their vocabulary.

Joint 4 — the Evidence chain

Claim Instrument Reason
Every declared source was read, or explicitly typed as inaccessibleDeterministicBinary and machine-observable; no judgement involved
Each claim carries a resolvable pointer to its sourceDeterministicA pointer resolves or it does not
Where two sources bear on the same quantity, they tie within tolerance or the disagreement is typedDeterministicArithmetic, with a stated tolerance agreed before work
Every option carries an evidence position; none is implicitly unevaluatedDeterministicStructural completeness against the option register
The option set is exhaustive against the agreed criteriaBounded AI judgment (nominating)A semantic-breadth task; the machine is better at it than a human with a fortnight
This constraint is a genuine blocker rather than a preferenceBounded AI judgment, then human dispositionMachine nominates with exhibits; consequence means a human decides
This source disagreement is material to option BHuman dispositionConsequence — it changes what is recommended
This residual uncertainty is acceptable at board levelHuman dispositionAuthority; nobody else can hold it
The option scoring is correctSee belowThe row where a naive design breaks Chapter 6’s rule

Read the ordinary rows first, because they are doing real work. Notice that four of the nine are deterministic, and that each of those four is deliberately narrow: was it read, does the pointer resolve, do the numbers tie, is anything unevaluated. None of them asks whether the pack is any good. That is not an oversight — it is the division of labour. Deterministic checks buy you the right to argue about the interesting things, by removing everything that should never have been arguable.

Then the trap row.

The row that breaks the rule, and its repair

The naive design: “A second model reviews the option scoring and confirms it.” It is fast, it is cheap, and it produces an artefact that looks exactly like evidence.

Why it fails: a judge drawn from the same family shares the blind spots that produced the scoring in the first place. Agreement between them is an echo. It is worse than nothing, because it converts an open question into a closed one.

The repair — and note that it is not “add a human to everything”: decompose the row. The criteria weights become a deterministic calculation over disposed inputs. The inputs to that calculation are human dispositions, each one bounded and recorded. And the model is demoted to what it is genuinely good at — nominating candidate scorings with exhibits attached, for a human to dispose. Same work, same cost, completely different evidence.

One observation about the shape of this map, because it matters in two chapters’ time: the disposition column is long relative to the unit count. A decision product is disposition-dense by nature — the whole thing turns on a handful of consequential calls. Chapter 10’s spine has the opposite shape, and the contrast turns out to be a commercial signal rather than a curiosity.

Joint 5 — the Exit falsifier

Three observations, each with its stated consequence — the fourth diagnostic from Chapter 4.

  1. Any option in the final pack cannot be traced to the agreed evidence and the agreed decision criteria. Consequence: the pack is not accepted and does not close the engagement.
  2. Any declared authoritative source is neither read nor explicitly typed as inaccessible within the boundary. Consequence: as above.
  3. Any material ambiguity was closed by the machine rather than disposed by a named human. Consequence: as above, and the specific ambiguity returns for disposition.

Look at what those three have in common. Every one is a negative observation about the evidence chain rather than a positive statement about the pack’s quality. That is deliberate, and it is what makes them checkable by someone who did not do the work — which is the only kind of check that survives a disagreement.

Joint 6 — Acceptance

The named decision owner signs, against a one-page acceptance record that lists the three falsifier observations and states that each was checked. Not a satisfaction survey. Not a presentation.

And the clause that the promise made possible back at Joint 1: “recommend stop” is a valid accepted result. If the pack’s honest conclusion is that none of the options should be pursued, and it is traceable and complete, the promise was kept. The buyer bought a decision pack, not a decision. Chapter 11 works the commercial settlement of that all the way down.

The falsifier that nearly wasn’t written

None of the above arrived fully formed. The first draft of the exit falsifier read:

The pack is comprehensive and decision-ready.

Run Chapter 4’s four diagnostics against it. Could a competent third party run it? No — “comprehensive” requires the author. Could it come back negative? Not in any form anyone could state in advance. Would a negative result change what the buyer does? Unknowable, because there is no negative result. Is the consequence written down? No. It fails all four.

Repair it in stages and watch what happens. First, replace the adjective with an object: “every option in the pack is traceable.” Better — there is now something to look at — but traceable to what, and judged by whom? Second, name the referent: “every option is traceable to the agreed evidence and the agreed decision criteria.” Now a third party can run it, because both referents were fixed at signature. Third, invert it, so it is a disconfirmation rather than a claim: “fires if any option cannot be traced…” Fourth, attach the consequence: “…the pack is not accepted and does not close the engagement.”

Four edits, about ten minutes, and an unfalsifiable sentence became observation one. The interesting thing is that nobody was being lazy in the first draft. “Comprehensive and decision-ready” is what a good supplier honestly intends. It just cannot be checked, and the gap between intending something and being checkable about it is where acceptance disputes live.

Writing this spine took about ninety minutes, before signature, with the buyer in the room for two of the six joints.

It was also the easy case. The promise here is about evidence, which makes an evidence-shaped falsifier almost natural. What happens when the promise is about an operational state — and the test suite is genuinely, thoroughly green?

09
Part II: Three Spines, and Two Ways They Break

Spine Two: The Migration, and the Green Board

Every implementation test passes. The promise is not closed.

Here is the test report, and I want you to believe it, because the chapter does not work if you suspect the engineers.

  • Row counts match across every table.
  • Checksums tie on every extracted file.
  • Referential integrity holds; no orphaned foreign keys anywhere in the target.
  • Every field in the mapping specification is mapped. The unmapped-field count is zero.
  • Reconciliation passes within the stated tolerance on every account class, including the two awkward ones that nobody wanted to touch.
  • The suite runs in CI. It has been run forty times. It is green.

Nobody cut a corner. Nobody is lying. This is a good suite, built by people who were paying attention, and it will pass any review you point at it.

The promise is not closed. As in Chapter 1, the failure here is an absence rather than a defect — and absences do not appear in test reports. The evidence boundary stated at the top of Chapter 8 applies to this engagement as well.

The promise, as sold

The target system running as the system of record, with the legacy system retired, by the stated date.

That is an operational state — not an artefact, and not a transformation. It passes Chapter 7’s fence in both directions: the supplier can keep it, and the supplier can evidence it. The engagement did the right thing at the promise joint.

Note what it deliberately is not: “a successful migration”. That phrase names an activity and would have been unfalsifiable from the first day.

Two questions, one of which was purchased

What the tests answered: did the mechanism transform the data correctly?

What was sold: is this organisation now safely running on the new system?

Both are honest questions and only the second one was bought. The transformation is a necessary condition of the operational state, and it is not remotely a sufficient one. Everything in this chapter lives in that gap.

What is missing, walked all the way down

Missing piece one: authorised cutover. The state exists technically. It does not exist organisationally, because no named authority has signed that the target is now the system of record.

Watch what that absence does in practice, because it is not abstract. Two teams keep writing to the old system “just in case” — nobody told them not to, and nobody had the standing to. The finance close gets run twice, once in each system, because the controller will not sign a close from a system that has not been declared authoritative. An integration partner is still polling the legacy endpoint because their change request was never raised. Every one of those is downstream of a signature that does not exist. The data is correct in a system that is not yet, in the only sense that matters commercially, the system.

Missing piece two: evidenced rollback. The rollback procedure exists. It is documented, it has been reviewed, and it is genuinely well written. It has never been exercised against production-shaped data.

So “we can roll back” is a belief rather than a receipt — and it is precisely the belief that will be tested on the worst possible day, by people who have not slept, under a clock. A rollback that has never been run is a plan for a rollback.

The two share a property worth naming: both are observations about authority and demonstration, and neither is expressible in a schema built from a transformation specification.

Why the tests were silent

The mechanism is one sentence: a test can only ever speak about what is in its schema.

Follow the chain for this engagement. The test schema was derived from the transformation specification. The transformation specification described a data mapping. And the word “authorised” appears nowhere in a data mapping, because authorisation is not a property of data. There was no point at which anyone made a mistake. The schema simply could not reach the thing that was sold.

Key Insight

The exit falsifier’s job is to name the observations the schema cannot contain. If you can derive your acceptance evidence entirely from your implementation specification, you have not written an exit falsifier — you have written a second copy of the specification.

Our own corpus has two names for the general form. On the evidence ladder, tested → accepted is a four-rung jump, and “green dashboards are not acceptance”. And the reason the producer will not notice: authoring systems optimise for successful completion and green.

Medicine has the expensive version of this

There is a field that formalised this failure, graded it, and has the most costly worked example available. A surrogate endpoint “does not measure the clinical benefit of primary interest in and of itself, but rather is expected to predict that clinical benefit or harm based on epidemiologic, therapeutic, pathophysiologic, or other scientific evidence.”22

Substitute the terms and read it again: an implementation test does not measure the commercial benefit of primary interest in and of itself; it is expected to predict it. Every failure of this kind follows from forgetting that sentence.

The mechanism of failure is known too. As Fleming and Powers put it, quoting Fleming and DeMets: “A correlate does not a surrogate make.”23 The reason is off-target effects — things the proxy was never watching. Which is the same sentence as a test can only speak about what is in its schema, written by biostatisticians.

CAST

Drugs that suppressed arrhythmia after myocardial infarction were in wide clinical use on exactly that reasoning: arrhythmia predicts sudden death, so suppressing arrhythmia should prevent it.

The enrolment design is the whole argument, and it deserves its own beat. The Cardiac Arrhythmia Suppression Trial randomised only patients whose arrhythmia had already been successfully suppressed — 1,727 of the 2,309 recruited. The mechanism test passed, by construction, for every patient in the study.

CAST: the mechanism worked24

1,727

of 2,309 recruited had their arrhythmia successfully suppressed — and only those were randomised

2.5

relative risk of total mortality on active drug (95% CI 1.6–4.5)

10

months mean follow-up before the arm was discontinued

The investigators’ conclusion has the exact structure of an exit falsifier: neither drug should be used “even though these drugs may be effective initially in suppressing ventricular arrhythmia.”24 The mechanism worked. The commitment did not survive. Those are two different findings and the trial reported both.

The counter-example, so this is not an argument against proxies

A proxy can be promoted to an outcome-closing instrument. In December 2025 the FDA qualified total hip bone mineral density by DXA as a validated surrogate endpoint — but as “the percentage change from baseline at 24 months in total hip BMD assessed by DXA”, for post-menopausal women with osteoporosis at risk of fracture, after a formal qualification process.25

Count what that took: a precise measurement, a precise window, a precise population, and a formal qualification. That is the bar an implementation test must clear before it is allowed to close a commercial promise. Almost no acceptance test in an AI-native service clears any of the four — and yet they are routinely asked to do the job.

The spine, for this migration

PromiseThe target system running as the system of record, legacy retired, by the stated date.
Entry falsifierFires if, by end of week two, either the legacy system’s authoritative record set cannot be enumerated, or no named authority for the cutover decision has been identified. Caller: delivery lead and operations owner, jointly. Accepted results: stop, or reshape to a data-transformation promise without the operational state.
Method latitudeExtraction approach, transformation tooling, staging strategy, parallelisation, number of rehearsal runs. Not free: the reconciliation tolerance, the record-set definition, and any write to the legacy system.
Evidence chainThe full deterministic suite — plus two obligations no data test produces: a signed cutover authority record, and a rollback rehearsal artefact captured against production-shaped data.
Exit falsifier(1) No named authority signed the cutover inside the window → the state is not accepted, the legacy system remains the system of record, and the engagement does not close. (2) No rollback was demonstrated against production-shaped data inside the window → as above. (3) Any account class reconciles outside tolerance without the disagreement being typed and disposed → as above.
AcceptanceThe named operations authority signs that the state exists and that no falsifier fired.

Note what did not happen. The deterministic suite did not shrink by a single test. Nothing was replaced or downgraded. Two observations were added that the suite could never have produced, and that is the entire repair.

The same engagement, run again

Same reality, two regimes

✗ Without the spine

  • • Week six: rollback documented, never rehearsed. Nobody is accountable for rehearsing it.
  • • Week eight: no cutover authority exists, so nobody discovers the two teams still writing to legacy.
  • • Close: “everything’s green”, and the discovery happens in the quarter after.

✓ With the spine

  • • Week six: the rehearsal is scheduled because it is an evidence obligation, not a good intention. It fails on first attempt — as rehearsals do — and there are three weeks to fix it.
  • • Week eight: the cutover authority has a name, and that name discovers the two teams. A fortnight early instead of a quarter late.
  • • Close: the acceptance record names three observations and states that each was checked. It takes less time, not more.
Same money at stake, same people, entirely different meeting.

“Our tests already cover this”

This is the objection the chapter exists for, and it deserves an exercise rather than a rebuttal. Put your promise clauses in one column and your test list in the other, and go line by line asking one question: which test proves this clause?

Promise clause Test that proves it
Every record in the declared set exists in the targetRow-count and checksum reconciliation — clean mapping
Records are correct against the mapping specificationField-level mapping suite — clean mapping
The target is the system of record— nothing —
The legacy system can be reverted to if required— nothing —

The diagnostic generalises, and it is worth writing on a whiteboard: any promise clause with no test opposite it is either an exit falsifier waiting to be written, or a clause that should not be in the promise. Both outcomes are useful. There is no third.

What this case proves is narrow and important: a suite can be complete against its specification and silent about the promise, and the gap is closed by adding observations, not by adding tests. Adding tests would have made the board greener.

This engagement followed its plan exactly and missed its promise. The next one abandons its plan entirely in week three — and keeps it.

10
Part II: Three Spines, and Two Ways They Break

Spine Three: The Assessed Estate, by an Unexpected Route

The approach everyone assumed is abandoned in week three. Nothing commercial happens.

Three weeks in, the delivery lead gets the update every delivery lead dreads. The approach everyone assumed at kick-off has been abandoned and replaced by something nobody proposed.

Under an activity-shaped statement of work, this is a conversation — possibly an uncomfortable one, certainly one involving the phrase “as discussed at kick-off”, and quite likely one that ends with a written explanation nobody will ever read.

Under a spine, it is Tuesday.

That is the claim this chapter has to earn, and the word to hold onto is not tolerated. Nothing commercial happened, and that is correct.

The service

A fixed-price assessed estate: a generative analysis engagement whose promise is coverage of a declared boundary, with every unit in a typed terminal state, each state supported by cited evidence.

This one goes third because latitude is at its widest here and the exit falsifier is at its cleanest — and because it is the shape most people assume cannot be fixed-priced at all, which makes it the useful case rather than the easy one. Structurally it differs from both earlier spines: the promise is about coverage rather than about a decision or an operational state, and that changes the falsifier’s entire character.

The spine

PromiseCoverage of the declared boundary across the declared source classes; every unit in a typed terminal state; every state supported by cited evidence.
Entry falsifierFires if, by end of week one, the estate’s unit population cannot be enumerated from the declared source classes within a stated margin. Caller: delivery lead and estate owner, jointly. Accepted results: stop, or reshape to a bounded enumeration engagement whose deliverable is the population itself.
Method latitudeDecomposition strategy, extraction route, model choice, ordering, parallelism — and whether the whole approach is discarded and rebuilt.
Evidence chainUnit register with provenance; a typed state per unit with cited evidence; a disposition record for every ambiguous unit.
Exit falsifierThree observations — see below.
AcceptanceThe named acceptor signs coverage and typing. Not conclusions — this engagement does not sell conclusions.

The entry falsifier is worth a sentence of its own, because it is the sharpest of the three in the book. A coverage promise over a population you cannot count is simultaneously unkeepable and unfalsifiable — you can neither deliver it nor be shown to have failed. If the denominator cannot be established in week one, there is no promise here, only an activity.

The divergence

The declared route was source-by-source extraction: work through each declared source class in turn, extract its units, type them, move on. It is the obvious approach and it is what the kick-off deck described.

In week three the interior finds something. Two of the four source classes share an identifier convention that nobody documented — and that convention makes a unit-first decomposition far more reliable than a source-first one, because it collapses an entire class of duplicate-unit ambiguity that source-first extraction would otherwise have had to dispose one case at a time, by hand, in front of a senior human.

So the interior abandons the first approach entirely, rebuilds the evidence base under the new decomposition, and arrives at the promised state by a route nobody in the kick-off would have described. The cost of the abandoned work is real and it is absorbed by the supplier.

Why nothing commercial happened

Three checks, each answered against what was actually recorded.

  1. Did any recorded perimeter field move? No. The promise, the declared boundary, the source classes, the date and the named acceptor are all unchanged.
  2. Did the evidence obligations change? No. Same unit register, same typed states, same provenance requirement, same disposition record.
  3. Did the buyer’s position change? No. The state they will receive is identical, and in one respect better, since fewer units will require disposition.

Three noes, so the movement is interior variation, and interior variation is commercially silent. Classifying surprises is a separate published instrument and this book hands off to it rather than re-deriving it.

What belongs to this book is the commercial reading, and it is blunter: latitude is not a favour the buyer grants. It is the thing they are paying less for. Clawing it back mid-engagement is a price increase disguised as governance.

The exit falsifier for a coverage promise

  1. Any unit in the register is untyped. Consequence: not accepted; the engagement does not close.
  2. Any typed state is not supported by its cited evidence — the pointer does not resolve, or the evidence does not support the state claimed. Consequence: not accepted; the unit returns for retyping.
  3. “Not observed” is reported anywhere as “does not exist”, in the pack or in any downstream artefact this engagement produced. Consequence: not accepted, and the claim is corrected before acceptance.

Now notice something about those three that is not true of Chapter 8’s or Chapter 9’s: all three are deterministic checks. A register can be scanned for untyped units. A pointer either resolves or it does not. The phrase “does not exist” can be searched for.

Key Insight

A promise stated in types gives you a falsifier a machine can evaluate. A promise stated in adjectives gives you an argument.

Three units, three states

The typed catalogue is inherited rather than invented here — directly mapped, partially mapped, not observed, insufficient evidence, inaccessible within boundary, unsupported source type, ambiguous — human decision required, and excluded from this phase. What matters is using it, so here are three actual units.

Unit Typed state What the evidence looked like
A supplier contract with an obligations scheduleDirectly mappedTwo sources agree; both pointers resolve to the clause; no disposition needed. The machine typed it and nobody had to look.
A records system held by a business unit that refused access inside the timeboxInaccessible within boundaryThe source exists and is named. Access was refused, with a date and a named refuser recorded. This is a delivered state, not a gap — the buyer now knows something specific they did not know.
An entitlement whose start date differs between the operational record and the contract registerAmbiguous — human decision requiredThe machine nominated both readings with exhibits attached. A named human disposed it in favour of the operational record, with the reasoning recorded and dated.

Typed states are how a coverage promise stays keepable without pretending to omniscience. The engagement does not promise to resolve every unknown. It promises that nothing ends in a shrug — and the middle row is the proof, because an inaccessible source recorded with a date and a name is worth more to a buyer than a silence.

The ratio, which is itself a signal

Fill in the evidence map for this spine and its shape is the opposite of Chapter 8’s. Here the deterministic column is long — coverage, typing support, pointer resolution and boundary discipline are all machine-checkable — and the disposition column is short, holding only genuine material ambiguity. The decision product had a short deterministic column and a long disposition one.

That contrast is not a curiosity. When dispositions approach the number of units, the offer is human-constituted whatever the technology says, and the old cost curve has been rebuilt with extra steps. Which is kill class five from Chapter 3 — impossible economics — in measurable form rather than argued form.

So the evidence map does entry work as well as exit work, and that is the argument for building it before signature rather than at kick-off. It is the only instrument in this book that answers a question at both ends.

“If you can change the method, what did I buy?”

A fair question, and it deserves an answer in the buyer’s terms rather than reassurance about trust.

None of that requires the buyer to understand the route, and that is the point. Latitude is affordable because the evidence is continuous, not because the buyer is trusting. A buyer who can watch coverage move does not need to watch method — and, in my experience, stops asking within about a fortnight.

The honest limit: this only holds if the evidence chain is genuinely produced as the work goes. An evidence chain assembled in the final week is a report about the work rather than evidence of it, and a buyer who is shown one will correctly go back to asking about method.

Three spines, then, and the useful finding is that the joints are constant while what fills them is not. A decision product’s exit falsifier is about traceability. A migration’s is about authority. An assessed estate’s is about coverage — three genuinely different instruments from one structure.

All three assumed there was a promise worth keeping. The case none of them runs is the one where the entry falsifier fires, the work stops in week two — and the supplier is paid.

11
Part II: Three Spines, and Two Ways They Break

The Accepted Stop

The work stops in week two. The full fee is payable. Here is the paperwork.

The analysis came back on Tuesday. Everyone has read it and nobody wants to be the first to say it out loud.

The tension in the room is not analytical — the finding is clear enough. It is commercial. The engagement is signed, a team is mobilised, and the evidence says the thing everyone assembled to build will not do what it was assembled to do.

In most firms the conversation quietly becomes about how to deliver something. That is not corruption and it is not weakness. It is the absence of a pre-agreed alternative: there is no path on the table called stop, so the room invents the only path it can see.

With an entry falsifier, this is not a crisis. It is a lookup, and there is a settlement waiting for it. The evidence boundary from Chapter 8 applies here too — this is a worked model.

Six words that make the difference

The service is a bounded diagnosis, and the promise reads:

We will establish, on evidence, whether the proposed AI intervention can reduce the named operational backlog.

Stop and look at that wording, because everything in this chapter rests on it. We will establish whether — not we will build. Six words apart, and a completely different commercial object.

Key Insight

An engagement can only stop and be paid if the promise was bounded to the answer rather than to the artefact. A supplier who sold the build cannot stop and be paid; they can only fail gracefully.

Which makes the entry falsifier partly a product design instrument rather than only a governance one. If you cannot see how your current offer would survive a stop, the problem is not your governance. It is what you are selling.

The finding

The assumption at signature was straightforward and widely held inside the client: the backlog is slow because assessing each case is cognitively expensive, so an AI-native assessment step should compress it.

What the evidence shows by the end of week two is that the assessable portion of the cycle is a minority of elapsed time. The binding constraint is a single external authorisation step with a fixed turnaround that neither party controls.

The arithmetic follows without needing a number invented for it: compressing the assessable portion to near zero would leave the aggregate cycle substantially unchanged, because the constraint sits outside the portion being compressed. That is the shape of the finding, and the shape is what decides it.

How it was obtained matters, because a finding this consequential should not be magic: an elapsed-time decomposition across the declared record set. A deterministic check, not a judgement — which is why it can be argued with on its own terms rather than on the analyst’s authority.

Which kill class fired, and why it was a lookup

Kill class six from Chapter 3: inappropriate AI placement, with wrong problem as the secondary. The work in question is authority and external timetable, not cognition. Cheaper cognition does not reach it, and no future cheapness will.

Here is the wording that was actually in the contract:

Fires if the assessable portion of the declared cycle is less than the share required for the intervention to move the aggregate, measured against the declared record set, by end of week two.

Read what that wording does. It names a measurement, a denominator, a threshold and a date. Nobody in the room had to argue about whether the intervention was “worth it”, because that was never the question on the table. The only question was what the decomposition showed — and everyone was looking at the same decomposition.

Contrast the version that would have been written by a well-intentioned team without this discipline: “We will confirm the intervention is viable.” Same intention. No observation, no denominator, no date, no caller. In week two it produces a discussion about optimism, and optimism wins, because the alternative is telling a client you would like to stop.

What happens without it

The same finding, two regimes

✗ No entry falsifier

  • The engagement continues because it was sold. Nobody decides to waste the money; there is simply no mechanism for stopping, and stopping would require someone to volunteer for the conversation.
  • The supplier delivers a beautifully executed irrelevance. Every implementation test passes. The assessment step is fast, accurate, well-governed. It moves nothing.
  • The buyer discovers it in production, after paying twice — once for the build, once for the programme that follows when the backlog does not move and somebody is asked why.

✓ Entry falsifier, called in week two

  • • The observation is checked against the declared record set. It is true.
  • • The named caller on each side confirms the accepted result: stop.
  • • The settlement below runs, and the engagement closes in week three.
The same category error as Chapter 9, at the other end of the engagement: there, green tests over an unkept promise; here, a green engagement over a dead premise.

The settlement

What the buyer receives
  • The evidenced identification of the real binding constraint, with the decomposition and its provenance attached.
  • The demonstration that the proposed intervention cannot move it, stated as a bounded claim with its method attached.
  • The typed statement of what was and was not observed inside the boundary — including anything inaccessible, named as such rather than silently absent.
  • The reopen conditions.
  • The redirected question — where the constraint actually is, stated precisely enough for someone else to work on. Easy to forget, and often the most valuable line in the pack.
What the buyer pays The agreed fee for the bounded engagement. Not a discount, not a goodwill reduction. The promise was establish whether, and it was established. A discount here is not generosity — it is an admission that the promise was really the build, and it undoes the entire design for every engagement that follows.
What is signed An acceptance record stating that the entry falsifier fired on the named observation, that the accepted result is stop, and that the evidence chain supporting the finding is complete. Signed by the same named acceptor who would have signed a completion. It does not say the project failed, and it does not say anyone was wrong.
What happens next Three honest paths. Reshape — a smaller square aimed at the real constraint, priced separately. Referral — the constraint belongs to someone else’s discipline, and saying so is worth more than pretending otherwise. Or the honest nothing: the buyer takes the finding and does not buy anything else. That third path is not rare and should not be described as though it were.

Reopen conditions

A stop without reopen conditions is an opinion. Write them at the same moment, in three fields.

The pattern is inherited — a kill registry records why something was rejected, how close it was, and what would reopen it. What matters commercially rather than intellectually is this: reopen conditions are what let the buyer defend the stop internally. Without them, a stop is indistinguishable from a supplier who could not do it.

“Isn’t a stop just a failed engagement with better PR?”

Answer it only in the currency that counts, which is the table above. What was delivered: five named artefacts, one of which redirects the buyer’s programme. What was paid: the full fee. What was signed: an acceptance record, by the acceptor. If that is a failed engagement, it is a failed engagement with a signature and an invoice, which is not what the objection means.

But there is a sharper version of the test, and it is about the firm rather than the engagement. A stop is real only if it is reported in the same forum where wins are reported. If stops are absorbed quietly and wins are announced at the monthly, the incentive gradient will find its way into the next entry falsifier — not through anyone’s dishonesty, but through the wording people choose when they are drafting under a quota.

“Capacity freed by honest non-build is often more valuable than a thin win that consumes a team for a year. Measure capacity returned to the bench as an outcome… Otherwise the system will still prefer any signature over a correct stop.”

And the line that ends the argument: “Your incentive design will show which one you are.”

There is an adjacent doctrine that reaches the same place from the delivery side: contract falsifying tests as success, and stop declaring victory at presentation acceptance. Two different books, one commercial fact: a supplier that cannot be paid for a true negative will eventually stop producing them.

Institutions that mean it build a body for it

Chapter 3 drew four design rules from the clinical Data Monitoring Committee. Only one of them matters here, and it is the narrowest: “is not effective” is a pre-authorised sentence, written into a charter before enrolment, held by a body independent of the sponsor.

Pre-authorising the sentence is most of what makes it sayable. Nobody has to be brave in week two; they only have to read.

What Part II has shown

Three spines written out, one broken at the exit, one stopped at the entry. And running through all of it, one asymmetry that is worth more than any of the individual findings:

The entry falsifier costs an hour before signature and is impossible to add afterwards.

An exit falsifier can be retrofitted badly — late, thinly, with evidence that was never designed to support it — and something is salvaged. The entry falsifier cannot be retrofitted at all. By the time you want it, the observation is stale, the money is spent, and the only honest version of it is a post-mortem.

The instrument is proved. What remains is getting it written down — with the buyer, and then into the contract.

12
Part III: Writing It Down

The Pre-Commitment Workshop

Ninety minutes, five moves, one page. It fits inside the meeting you are already having.

Everything in Part II assumed both falsifiers already existed. They do not write themselves, and — this is the claim of the chapter — they cannot honestly be written alone.

A falsifier drafted by the supplier alone is a falsifier the buyer can dispute later, which defeats the design. The instrument works by moving the argument to a moment when neither party has a financial interest in the answer, and that moment only exists if both parties are in it.

The posture is one we have written about before: a falsification meeting. Arrive with a killable theory and ask the other side to test it, rather than arriving blank (“what keeps you awake at night?”) or arriving certain (“I diagnosed you from your website”). What is different here is the object. That meeting tests a diagnosis. This one tests a promise, and the buyer is not being asked to supply the insight — they are being asked to supply the conditions under which the promise would be broken. Every question in the room is about their risk, which is why they engage.

The five moves

The failure signals are what make this usable — they are what a facilitator watches for, and each has a specific repair.

1. State the bounded promise in one sentence — 10 minutes

Question: what bounded state will exist when this is over? Output: a sentence naming a state, not an activity.

Failure signal: the sentence needs a conjunction. Two states joined by “and” are either two engagements, or one promise carrying a hidden second one that will never be falsified — and that hidden clause is where the dispute will eventually go. Facilitation: write it on the board and leave it there. Every later move is checked against it.

2. Write the entry falsifier — 25 minutes

Question: what would we have to see, in the first two weeks, to conclude this should not proceed? Output: observation, window, caller, accepted results.

Failure signal: the room produces risks instead of observations. “The data might be poor” is a risk. “Fewer than three of the four declared sources are accessible by Friday of week one” is an observation. When you hear a risk, ask the same question again with the emphasis moved: what would you see?

Second failure signal: nobody will hold the caller role. Do not paper over it — that is itself a finding, and it belongs in move five.

3. Write the exit falsifier — 25 minutes

Question: what evidence would show we did not keep this promise, even if everything we did was excellent? The clause after the comma is doing the work — it separates the promise from the effort, and it is the clause that stops the room drifting into a quality discussion.

Failure signal: someone proposes a threshold with no consequence. Ask what happens commercially when it is true. If the answer is “we’d discuss it”, it is not written yet.

4. Map the evidence — 20 minutes

Question: for each claim, what will show it — a deterministic check, a bounded AI judgment, or a named human’s disposition? With the reason in the row.

Failure signal: a row whose evidence is the producing system’s own assessment. That is Chapter 6’s rule, applied live and in front of the buyer, which is the best possible place to apply it. Facilitation: do not attempt every claim. Do the six that matter and note that the schedule completes before signature.

5. Name the two humans — 10 minutes

Question: who calls the entry falsifier, and who accepts against the exit falsifier? By name and role, on both sides.

Failure signal: “the steering committee.” A committee is how a falsifier becomes a discussion. Push for a name — and if there genuinely is not one, you have just discovered an entry-falsifier observation for free, and it is the one that maps to wrong client.

When the buyer refuses

Three cases, and they behave differently enough that a single answer would be useless.

The buyer who will not allocate ninety minutes. Send two sentences by email — the promise and the proposed exit falsifier — and ask: is this what you would want us judged against? The reply is the signal, and so is its absence. A buyer who will not spend ten minutes on the thing they will later be asked to accept has told you how acceptance will go.

The buyer who will not name an acceptor. Do not treat this as an administrative gap to be chased later. It is a genuine entry-falsifier observation, and the honest move is to say so and offer the reshape: a smaller engagement whose deliverable is a decidable question with an owner attached. That is frequently the most valuable thing you could sell them, and it is a real product rather than a consolation prize.

The buyer who wants the falsifier softened after seeing it. The most informative moment in the sale, and almost never bad faith. It usually means the promise as written is broader than what they believe you can deliver — they are protecting the relationship from a clause they expect to fire. So the right response is not to blunt the instrument but to narrow the promise. The falsifier is downstream of the promise; softening it leaves an unkeepable promise with a broken instrument attached, which is worse than either problem alone.

A falsifier that survives the buyer’s first attempt to soften it is worth more than one that was never tested.

The leave-behind

One page: the promise sentence, the two falsifiers with their consequences, the six mapped claims, and the two names. Nothing else.

It has one borrowed test to pass: if they forward a single artefact internally, that page must travel without you in the room. If it needs explaining, it will be explained by whoever is nearest, in their own words, to the person who later decides whether the falsifier fired.

Worth naming what the page replaces: roughly two pages of activity description. Fewer words, more obligation — the trade the next chapter puts into a contract.

Four facilitation details

  • Write the failing sentence in the buyer’s words, not yours. If they cannot recognise it as something they would say, they will not invoke it.
  • Time-box move one hardest. Rooms will happily spend an hour polishing the promise and then rush the falsifiers, which is exactly backwards. An imperfect promise with two sharp falsifiers is recoverable. A perfect promise with none is not.
  • End with a dated owner on every open row. “We’ll finalise offline” is how falsification dissolves back into relationship maintenance, usually within four days.
  • Do not leave a softening request unrecorded. If a falsifier was softened in the room, write down what changed and why, in the leave-behind, where both parties can see it. The record is what stops the second softening.

What the room now holds is two sentences that can fail, a map of who proves what, and two names. None of which is a contract — and a falsifier that lives only in a workshop artefact will be quietly renegotiated by the first person who writes the statement of work.

13
Part III: Writing It Down

Drafting Both Ends

Five clause shapes. Fewer words than the activity description they replace, and considerably more obligation.

Illustrative drafting, not legal advice. What follows shows structure and vocabulary. It is not jurisdiction-specific and it should be reviewed by somebody whose job that is. Said once, and not repeated — a chapter that hedges every clause is worth nothing to the person who has to write one.

One thing before the clauses, because it decides where all of this has to live: types that exist only in a deck get synonyms invented under pressure. What is not in the schedule is not in the contract.

The shape is already written

The skeleton of an AI-native statement of work fits in six lines:

Here is the state we will establish.
Here are the evidence boundaries.
Here are the valid uncertainty states.
Here is what would invalidate the engagement.
Here is what constitutes completion.
How we exhaustively get there is largely our problem.

Line one is the promise. Lines two and three are the evidence chain and its typed vocabulary. Line four is the entry falsifier; line five the exit falsifier and acceptance; line six method latitude, stated as a refusal to be specified. The contract and the artefact are the same object seen from two sides. What the six lines do not name is which instrument proves which claim — the piece that most commonly goes missing, and what clause shape four adds.

Clause shape 1 — the bounded promise

promise
The Supplier will establish the following state by [date]: [state], across [declared boundary],
excluding [exclusions]. The Supplier does not promise [explicit non-promises].

Two drafting errors are worth more than the model clause, because they are the two that actually appear.

Two errors, repaired

Activity smuggled into a state. “Supplier will conduct a comprehensive assessment of the estate and deliver findings.”
Repair: name the state — “every unit in the declared population carries a typed terminal state with cited evidence” — and let the assessment be method.

A promise with a hidden second promise. “…and support the client’s decision-making process.”
Repair: delete it, or make it a second bounded promise with its own falsifier. An unfalsified clause in a falsifiable contract is where every dispute will go — precisely because it is the only clause with nothing to check it against.

Clause shape 2 — the entry falsifier

entry falsifier
If, on or before [date], [observation stated as a fact], either party’s Named Caller
([role], [role]) may invoke this clause. On invocation the parties will adopt one of:
  (a) STOP — the Supplier delivers [finding pack], the full Fee is payable, and the
      engagement closes;
  (b) RESHAPE — the parties agree a replacement bounded promise within [n] business
      days, failing which (a) applies.

The commercial consequence is what most drafts omit, and without it the entry falsifier is a statement of intent that will be argued about at the moment it is invoked. The stop branch mirrors Chapter 11’s settlement: the Supplier delivers the evidenced finding, the typed observation record and the reopen conditions, and the full fee is payable. Say “full” in the clause — a drafter who leaves the fee open on a stop has written a discount into the future without knowing it. The reshape branch needs its default at the end, or it is an open-ended negotiation held by the party with more time.

Clause shape 3 — the exit falsifier

exit falsifier — two worked examples
(1) If any option in the Decision Pack cannot be traced to the Agreed Evidence Set and the
    Agreed Decision Criteria, the Decision Pack is not accepted and does not close
    this engagement.

(2) If no Named Authority has signed the Cutover Record on or before [date], the Target System
    is not accepted as the system of record, the Legacy System remains the system
    of record, and this engagement does not close.

Both are observations with consequences attached, drawn from Part II so you can see one structure carrying different content.

The drafting note that matters most comes from a decided case. If you intend an observation to preclude acceptance, say so in terms. Otherwise a measurable threshold is read as identifying a breach rather than refusing the promise — parties may deem a clause a condition or deem that its breach precludes practical completion, “but if that option is taken, they should do so clearly.”10

One further caution: keep the number of exit falsifiers small. Three that can fire beat nine that describe quality, and a long list is usually a sign that the promise is carrying hidden clauses.

Clause shape 4 — the evidence schedule

schedule N — evidence
| Claim | Instrument | Reason | Produced by |


No acceptance evidence may be constituted solely by an assessment produced by the system that
produced the work being assessed.

Chapter 6’s rule as a term rather than a principle, plain enough that a procurement lawyer can read it without a glossary. It is a schedule rather than a playbook page for the same reason the typed catalogue is: a rule living outside the contract gets reinterpreted by whoever is under the most pressure.

Note what it buys the buyer, because a one-sided clause invites deletion: it tells them, before signature, exactly which claims will be machine-proved and which will rest on a named person’s judgement. Very few buyers have been shown that distinction, and none of them dislike it.

Clause shape 5 — acceptance

acceptance
Acceptance occurs when the Named Acceptor ([name/role], substitute [name/role]) confirms,
within [n] business days of delivery, that no Exit Falsifier has fired and the Evidence Schedule
is complete. If no confirmation or notice of non-acceptance is given within that period,
acceptance is deemed to have occurred. Notice of non-acceptance must identify the Exit Falsifier
relied upon.

Three sub-clauses make it operable. The named acceptor, by role, with a substitute — an acceptance clause with an unnamed acceptor has no acceptor. The window, plus what happens on silence: deemed acceptance after a stated period is conventional, but deemed acceptance with no evidence obligation is how a supplier passes by exhausting the buyer. And the consequence of non-acceptance.

The buyer’s reciprocal rights

If only the supplier can invoke the machinery, the machinery is marketing. Three rights, drafted in.

  • The buyer may call the entry falsifier on the same observation and within the same window.
  • The buyer may require the evidence schedule before signature, rather than receiving it as a deliverable of the engagement.
  • The buyer may refuse acceptance on a falsifier the supplier says did not fire — with a named path when the two disagree.

That third right needs its path drafted or it becomes a dispute clause in disguise, and the path is shorter than people expect because the evidence schedule already sorts the disagreement by type. A disputed deterministic row is a defect: go and look at the machinery, and one of you is simply wrong. A disputed disposition row goes to the two named humans, because that is what they are for. Different disagreements, different routes, neither of them a lawyer.

If you are on the buying side

Four questions before you sign

  1. What is the acceptance evidence, specifically, and in what form will it arrive?
  2. Who signs it — by name and role, on both sides?
  3. What result would count as a failure? Ask them to say it out loud.
  4. What happens commercially if it fails?

A supplier who cannot answer the third has not sold you an outcome. They have sold you an activity with an optimistic adjective.

With these five clauses the promise, both falsifiers, the evidence schedule and the two names all sit in one document, and none of them can be quietly renegotiated by the person who writes the status report. Which covers a new engagement — and most readers are inside an old one.

14
Part III: Writing It Down

Monday, and What Would Falsify This

Four moves on the engagement you are already inside — then the same instrument, pointed at this book.

The next proposal is months away and, more importantly, hypothetical. There is an engagement running right now with one acceptance test in it, and that test is doing two jobs. Start there.

None of the four moves below requires renegotiating anything. They are diagnostics first and contract changes second, and three of the four are done alone in under an hour.

Four moves

1. Write the failing sentence

Take your current acceptance criteria and write, in one sentence and in the buyer’s language, the result that would mean the promise was not kept.

If you cannot write it, you have a deliverables list rather than an exit falsifier — and you have learned that in ten minutes rather than at the closing meeting.

Constraint that does most of the editing: write it as something you would be willing to read aloud to the buyer.

2. Add the consequence

Beside the observation, write what happens commercially when it is true. A threshold identifies a breach; it does not close or refuse a promise.

If the honest answer is “we’d have a conversation”, write that down and read it back. It usually rewrites itself.

3. Map the evidence

Three columns — deterministic check, bounded AI judgment, accountable human disposition — and every claim in exactly one, with the reason in the row. Any row whose evidence is the producing system’s own assessment is not evidence: replace it, or mark it unproven.

Do the six claims that matter. A complete map you never finish is worth less than six honest rows.

4. Back-test the entry falsifier you should have written

Retrospectively, for the engagement you are inside: what observation, in the first fortnight, would have told you this should not proceed as scoped? Then the question that makes it useful — would it have fired?

Either answer is cheap. If it would have fired, you know what the next proposal needs. If it would not have, you have a tested falsifier to reuse.

Four refusals for the first ninety days

Refusals rather than aspirations, because a refusal is something you can hold in a meeting and an aspiration is something you can defer.

  • No acceptance criterion that cannot come back negative.
  • No acceptance evidence authored solely by the producing system.
  • No entry falsifier without a window and a named caller.
  • No falsifier attached to an outcome the supplier cannot keep and evidence.

They are cheapest to apply in the proposal, and they are still worth applying in the last week of an engagement — because the acceptance conversation is going to happen either way, and the only variable is whether anything has been written down before it starts.

Why this is arriving now

Outcome pricing is reaching professional services before the adjudication problem has been solved. About a quarter of one major firm’s global fees are now reported to come from outcome-based arrangements, with clients increasingly arriving not with a defined scope but with a result they want and asking the firm to price against delivering it.26 Be careful with that figure and I will be careful with it here: the original reporting is paywalled, and the numbers come from firms’ own characterisations at media events rather than from audited disclosures. Treat it as a direction of travel, not a measurement.

The direction is not the interesting part anyway. This is:

What the reporting does not tell you is “what baselines are set, who adjudicates whether a target was hit, and whether partner compensation changes as risk shifts from client to firm.”26

Who adjudicates whether a target was hit. That is the exit falsifier and its named acceptor, and it is unbuilt in most of the contracts being signed this quarter. The current answer, in practice, is whoever is more senior in the room at the closing meeting.

What would falsify this book

Fourteen chapters spent demanding that other people write down what would prove them wrong. Here is the same instrument, pointed the other way.

The limitation first, flatly: there is no evidence that two-falsifier governance improves engagement outcomes. No such study exists. The argument in this book is structural — it says one instrument is the wrong shape for two jobs, and shows what the right shapes look like. It does not say that a measured population of engagements did better, because there is no such population and I will not manufacture one.

Claim What would test it What a bad result would mean
An explicit entry falsifier reads as competence, not hedging Track whether proposals carrying one shorten or lose the sale, across a comparable set If it reliably loses deals, this doctrine has a commercial defect rather than a communication problem — and Chapter 3’s answer is wrong
An exit falsifier can be written before the method is chosen Ask teams to draft one at proposal stage and compare it with what they would write after selecting an approach If they genuinely cannot, the endpoints are not separable from the middle and Chapter 5’s defence of latitude is weaker than claimed
Accepted stops settle commercially Whether stop findings are paid in full outside pre-existing strong relationships If only strong relationships pay for stops, the entry falsifier is a luxury good and Chapter 11 should say so
Two falsifiers reduce acceptance disputes Matched engagements carrying two versus one, counted over a portfolio The obvious empirical test. Nobody has run it — including us

Publishing your own attack surface is a habit worth keeping, for a reason that is selfish rather than noble: I’d rather be falsified precisely than believed vaguely. The first one teaches something.

One boundary is deliberate rather than accidental. This book has said nothing about what an engagement leaves behind for the next engagement — what the supplier learns, what that is worth, and how it changes the economics. That is a different argument with a different instrument, and it does not belong to a book about verification semantics.

One asymmetry, and one question

If you only take one operational thing from this: the exit falsifier can be retrofitted badly and still help. The entry falsifier cannot be retrofitted at all. By the time you want it, the observation is stale, the money is spent, and the only honest version of it is a post-mortem.

The question this book has actually been about

At the end of your next engagement, will acceptance be a matter of looking — or a matter of who is in the room?

Tests tell you whether the machine worked. Falsifiers tell you whether the commitment survived.

REF
Sources & Evidence

References & Sources

The evidence base behind every claim — primary research, industry analysis, and technical specifications

Research Methodology

This ebook draws on primary research from standards bodies, independent research firms, enterprise technology vendors, and consulting firms. Statistics cited throughout have been cross-referenced against primary sources.

Frameworks and interpretive analysis developed by Scott Farrell / LeverageAI are listed separately below — these represent the practitioner lens through which external research is interpreted, and are not cited inline to avoid self-promotional appearance.

Primary Research & Standards Bodies

SEBoK (Guide to the Systems Engineering Body of Knowledge), BKCASE / INCOSE / IEEE-CS / Stevens Institute — System Verification [1]

Verification is the confirmation, through the provision of objective evidence, that specified requirements have been fulfilled

https://sebokwiki.org/wiki/System_Verification

SEBoK (Guide to the Systems Engineering Body of Knowledge), BKCASE / INCOSE / IEEE-CS / Stevens Institute — System Validation [2]

Validation is the confirmation, through the provision of objective evidence, that the requirements for a specific intended use or application have been fulfilled

https://sebokwiki.org/wiki/System_Validation

U.S. Food and Drug Administration (11 January 2002, §3.1.2) — General Principles of Software Validation; Final Guidance for Industry and FDA Staff [3]

"Software testing is one of many verification activities intended to confirm that software development output meets its input requirements"; "a developer cannot test forever, and it is hard to know how much evidence is enough"

https://www.fda.gov/media/73141/download

John Rushby, SRI International, 2015 — The Interpretation and Evaluation of Assurance Cases (SRI-CSL-15-01) [4]

Modern airplanes are extraordinarily safe and no serious airplane incident has been traced to faulty software; some have been traced to faulty requirements for systems implemented in software, but that is outside the remit of DO-178C

https://www.csl.sri.com/users/rushby/papers/sri-csl-15-1-assurance-cases.pdf

Stephen Thornton — Karl Popper (Stanford Encyclopedia of Philosophy) [5]

The asymmetry between verification and falsification; corroboration counts only as the positive result of a genuinely "risky" prediction; and the caveat that a single counter-instance is never methodologically sufficient in practice

https://plato.stanford.edu/entries/popper

U.S. Food and Drug Administration — Use of Data Monitoring Committees in Clinical Trials: Guidance for Industry (draft, 2024) [6]

A DMC reviews accumulating data and recommends whether to continue, modify or stop; it is established by the sponsor but should be independent of the sponsor and the trial conduct; it may recommend stopping because the product "is not effective"

https://www.fda.gov/media/176107/download

National Institute of Standards and Technology, January 2023 — Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 [7]

MANAGE 1.1 — a determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed

https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf

ICAEW — Limited assurance vs reasonable assurance [8]

The conclusion in a limited assurance engagement is framed in a negative sense: "Based on the procedures performed, nothing came to our attention to indicate that the management assertion on XYZ is materially misstated"

https://www.icaew.com/technical/audit-and-assurance/assurance/process/scoping/assurance-decision/limited-assurance-vs-reasonable-assurance

England and Wales Court of Appeal (Civil Division) — Mears Ltd v Costplan Services (S.E) Ltd & Ors [2019] EWCA Civ 502, approved judgment (Coulson LJ at [41]) [9]

"The clause simply provides a mechanism by which a breach of contract can be indisputably identified"; materiality was introduced only in relation to room size and not in relation to the resulting breach

https://s47657.pcdn.co/wp-content/uploads/2019/04/mearscaj.pdf

Gatehouse Chambers — Case note on Mears Ltd v Costplan Services (South East) Ltd & Ors [2019] EWCA Civ 502 [10]

It is entirely possible for parties to deem a particular clause as a condition, or that the consequence of breach would preclude practical completion — but if that option is taken, they should do so clearly

https://gatehouselaw.co.uk/mears-ltd-v-costplan-services-south-east-ltd-ors-2019-ewca-civ-502

International Standard on Auditing, published by IAASA — ISA (Ireland) 500, Audit Evidence (updated October 2022), paras 5(b) and 5(f) [11]

Sufficiency is the measure of the quantity of audit evidence; appropriateness is the measure of its quality — its relevance and reliability in providing support for the conclusions

https://iaasa.ie/wp-content/uploads/2022/11/ISA-500_Oct_2022.pdf

Rapita Systems — DO-178C Guidance: Introduction to RTCA DO-178 certification [12]

DO-178B was a total re-write to move away from the prescriptive process approach and define a set of activities and associated objectives that a design assurance process must meet, allowing flexibility in the development approaches followed

https://www.rapitasystems.com/do178

Elliot Kim, Avi Garg, Kenny Peng, Nikhil Garg — Correlated Errors in Large Language Models (arXiv:2506.07962, ICML 2025) [13]

On one leaderboard dataset models agree 60% of the time when both models err; larger and more accurate models have highly correlated errors even with distinct architectures and providers

https://arxiv.org/abs/2506.07962

Elliot Kim, Avi Garg, Kenny Peng, Nikhil Garg — Correlated Errors in Large Language Models, §3.2 (arXiv:2506.07962) [14]

Each judge systematically inflates the accuracy of models that are less accurate than itself, due to correlated errors — the judge marks incorrect answers as correct if both models agree on the incorrect answer

https://arxiv.org/html/2506.07962v1

Shashwat Goel, Joschka Struber, Ilze Amanda Auzina, Karuna K Chandra, Ponnurangam Kumaraguru, Douwe Kiela, Ameya Prabhu, Matthias Bethge, Jonas Geiping — Great Models Think Alike and this Undermines AI Oversight (arXiv:2502.04313) [15]

As model capabilities increase it becomes harder to find their mistakes, and model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures

https://arxiv.org/abs/2502.04313

Lianmin Zheng et al. — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv:2306.05685) [16]

Strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80% agreement — the same level of agreement between humans

https://arxiv.org/abs/2306.05685

Linford & Company LLP — SOC Report Types: Type 1 vs Type 2 SOC Reports/Audits [17]

A Type 1 SOC report is as of a point in time and covers only the design effectiveness of internal controls; a Type 2 report covers a period of time and the operating effectiveness of those controls

https://linfordco.com/blog/soc-report-types-1-vs-2

U.S. Government Accountability Office, September 2005 — Defense Management: DOD Needs to Demonstrate That Performance-Based Logistics Contracts Are Achieving Expected Benefits (GAO-05-966) [18]

Only 1 of the 15 program offices had performed a business case update; in that single case it determined that the performance-based logistics contract did not result in expected cost savings and the weapon system did not meet established performance requirements

https://www.govinfo.gov/content/pkg/GAOREPORTS-GAO-05-966/html/GAOREPORTS-GAO-05-966.htm

Government Outcomes Lab, Blavatnik School of Government, University of Oxford — Outcomes-based contracting [19]

Outcomes-based contracts risk "cherry picking", "creaming" and "parking" — behaviours often referred to as gaming — and it can be difficult to set simple, measurable outcomes that align effectively with complex social problems

https://golab.bsg.ox.ac.uk/the-basics/outcomes-based-contracting

Charles Goodhart (1975); Marilyn Strathern (1997) — Goodhart's law [20]

Goodhart's original: "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." Strathern's formulation: "When a measure becomes a target, it ceases to be a good measure."

https://en.wikipedia.org/wiki/Goodhart%27s_law

David Manheim, Scott Garrabrant — Categorizing Variants of Goodhart's Law (arXiv:1803.04585) [21]

The importance of Goodhart effects depends on the amount of power directed towards optimizing the proxy, and so the increased optimization power offered by artificial intelligence makes it especially critical for that field

https://arxiv.org/abs/1803.04585

FDA-NIH Biomarker Working Group — BEST (Biomarkers, EndpointS, and other Tools) Resource, Glossary [22]

A surrogate endpoint does not measure the clinical benefit of primary interest in and of itself, but rather is expected to predict that clinical benefit or harm

https://www.ncbi.nlm.nih.gov/books/NBK338448

Thomas R. Fleming and John H. Powers — Biomarkers and Surrogate Endpoints in Clinical Trials (Statistics in Medicine, 2012; PMC3551627) [23]

"As indicated by Fleming and DeMets, 'A correlate does not a surrogate make.'" — evidence about a biomarker can be unreliable regarding true clinical efficacy even when strongly correlated

https://pmc.ncbi.nlm.nih.gov/articles/PMC3551627

Cardiac Arrhythmia Suppression Trial (CAST) Investigators — Preliminary report: effect of encainide and flecainide on mortality in a randomized trial of arrhythmia suppression after myocardial infarction (NEJM 1989;321(6):406-12) [24]

1,727 of 2,309 recruited patients (75 per cent) had initial suppression of their arrhythmia and were randomised; over an average of 10 months, total mortality relative risk 2.5 (95% CI 1.6 to 4.5) against placebo, and the encainide/flecainide arm was discontinued

https://pubmed.ncbi.nlm.nih.gov/2473403

U.S. Food and Drug Administration — FDA Qualifies Total Hip Bone Mineral Density (BMD) as Surrogate Endpoint for Osteoporosis Drug Development (19 December 2025) [25]

The qualified tool is the percentage change from baseline at 24 months in total hip BMD assessed by DXA, usable as a validated surrogate endpoint for post-menopausal women with osteoporosis at risk for fracture

https://www.fda.gov/drugs/drug-safety-and-availability/fda-qualifies-total-hip-bone-mineral-density-bmd-surrogate-endpoint-osteoporosis-drug-development

AI News Weekly — McKinsey Ties 25% of Fees to Outcomes as AI Erodes Billable Hours (summarising Wall Street Journal reporting, "Inside Consultants' Messy Shift From Hourly Billing") [26]

About 25% of McKinsey's global fees now come from outcome-based pricing, with clients arriving with a result they want rather than a defined scope; the figures come from firms' own characterisations rather than audited disclosures

https://aiweekly.co/alerts/mckinsey-ties-25-of-fees-to-outcomes-as-ai-erodes-billable-hours

LeverageAI / Scott Farrell — Practitioner Frameworks

The interpretive frameworks, architectural patterns, and practitioner analysis in this ebook were developed through enterprise AI transformation consulting. The articles below are the underlying thinking behind those frameworks. They are listed here for transparency and further exploration — not cited inline, as this is the author's own analytical voice.

Scott Farrell — AI-Native Service Architecture

The outcome is what was sold; the test is the instrument by which we can legitimately say the outcome occurred — and the semantics of falsification deferred to a separate treatment (ch6 #e99a49)

https://leverageai.com.au/wp-content/media/articles/226-ai-native-service-architecture.html

Scott Farrell — Governance Barbell

Control that lives only as reading the middle will always lose to the volume of the middle; senior attention changes outcomes at the ends (ch1 #15e711)

https://leverageai.com.au/wp-content/media/articles/140-governance-barbell.html

Scott Farrell — Boundary Mutation, Not Change Request

Classify possible surprises before the engagement starts, so arriving at the answer is a lookup rather than an argument; three dispositions with different decision rights (ch3 #a732ac)

https://leverageai.com.au/wp-content/media/articles/228-boundary-mutation-not-change-request.html

Scott Farrell — AI-Constituted Services

Disposition load as a build/no-build detector: when dispositions approach the number of units, the old cost curve has been rebuilt with extra steps (ch14 #cba51e)

https://leverageai.com.au/wp-content/media/articles/202-ai-constituted-services.html

Scott Farrell — Pre-Thinking Prompting

First-signal tests — define what observable evidence could change the chosen frame before substantial execution begins, falsifiable inside a short horizon

https://leverageai.com.au/wp-content/media/articles/05-pre-thinking-prompting.html

Scott Farrell — Buy Certainty First

The option resolving to zero exercise; correct non-exercise, the mirror standard, and "your incentive design will show which one you are" (ch6 #3f825e)

https://leverageai.com.au/wp-content/media/articles/204-buy-certainty-first.html

Scott Farrell — Knowledge Base Kills Projects

Kill / Fix / Double-Down as the disposition vocabulary, and the kill registry recording why rejected, how close, and what would reopen it (ch6 #3a28d6)

https://leverageai.com.au/wp-content/media/articles/208-knowledge-base-kills-projects.html

Scott Farrell — Preparedness Is the Product

A product that cannot fail its own test is only an attractive story; kill conditions written into the proof charter before the work starts, each with its named evidence (ch13 #28d96b)

https://leverageai.com.au/wp-content/media/articles/214-preparedness-is-the-product.html

Scott Farrell — The North Star Prompt

"The thing to be precise about is intent. The thing to stop over-specifying is procedure. A north star is high-information about purpose and low-information about method." — and the three cases that stay prescriptive (ch7 #0aa4c6)

https://leverageai.com.au/wp-content/media/articles/70-north-star-prompt.html

Scott Farrell — Designing Loops, Not Prompts

Running several copies of the same judge and treating agreement as verification is an echo, not a check; the failure mode is a green tick over work that never happened, and the closer is not a tax on autonomy but what buys it (ch6 #428c20)

https://leverageai.com.au/wp-content/media/articles/64-designing-loops-not-prompts.html

Scott Farrell — Same Session Supervision

If "done" is defined as whatever the session already believes done looks like, the session can pass its own test; green means the narrative is consistent with itself (ch8 #8f670c)

https://leverageai.com.au/wp-content/media/articles/180-same-session-supervision.html

Scott Farrell — Witness, Not Oracle

An oracle returns conclusions; a witness returns conclusions attached to exhibits — the four-part contract of claim, exhibit, resolvable pointer and confession of what could not be verified (ch5 #189a2b)

https://leverageai.com.au/wp-content/media/articles/93-witness-not-oracle.html

Scott Farrell — Hidden Gates

Intent stays visible because it cannot be gamed, only pursued; the rubric stays hidden, held by an orchestrator — a gate the worker can't see is a gate it can't game, and the reviewer's information asymmetry converts review into a real test (ch4 #82a1e4)

https://leverageai.com.au/wp-content/media/articles/94-hidden-gates.html

Scott Farrell — AI-Native Successor Offer

"We sell outcomes, not hours" is often a category error wearing a premium; the safer formulation is hours versus a bounded state, decision, deliverable or commitment — the successor-unit vocabulary of verified decisions, maintained states, protected periods and assessed estates (ch5 #252640)

https://leverageai.com.au/wp-content/media/articles/213-ai-native-successor-offer.html

Scott Farrell — Discussed Is Not Deployed

The eight-rung evidence ladder — discussed, proposed, coded, committed, tested, deployed, observed, accepted; tests answer demonstrated behaviour under the harness, not operational truth, and green dashboards are not acceptance (ch5 #0d59e6)

https://leverageai.com.au/wp-content/media/articles/192-discussed-is-not-deployed.html

Scott Farrell — Proof-Carrying Transformation

Contract falsifying tests as success; declaring victory at presentation acceptance is the original sin of the incomplete product (ch12 #f34e74)

https://leverageai.com.au/wp-content/media/articles/164-proof-carrying-transformation.html

Scott Farrell — Friction Thesis Compiler

The falsification meeting — a third posture between discovery theatre and website arrogance: state the thesis, show the receipts, surface the alternatives, name the known absences, state the kill condition aloud, and leave a pack that travels without you in the room (ch11 #c86f0f)

https://leverageai.com.au/wp-content/media/articles/215-friction-thesis-compiler.html

About This Reference List

Compiled August 2026. All URLs verified at time of compilation. Regulatory documents and standards specifications are subject to revision — check primary sources for the most current versions.

Some links to academic papers and vendor research may require free registration. Government and standards body publications are freely accessible.