Post-engagement learning · The trace-first instrument
The Perturbation Review
What the project taught you and never wrote down — and how to find it without inventing it.
TL;DR
- “Unrecorded” is four different conditions, and three of them cannot be reached by an interview. Most firms sample one and call it a retrospective.
- Freeze the machine’s trace-only account before anyone speaks, then credit the conversation only for the delta against it. Without that baseline you cannot tell discovery from fluent restatement — which is why no post-mortem process in your career has ever been retired on evidence.
- The paid product is the client’s closure, not your IP harvest — and the whole programme ships with its own kill conditions, written before the first review runs.
Your AI reads everything now. The proposal, the SOW variations, the architecture decision records, eighteen months of project chat, the tickets, the meeting transcripts, the final pack. It reconstructs the engagement with a completeness no human post-mortem has ever achieved, and it does it on the Friday the project closes rather than the following March. If you have installed this, you already know it works.
And here is the thing nobody says out loud at the steering committee: the firm has not changed. The wiki is bigger. The next proposal is faster. The delivery patterns are tidier. But if I asked you which commercial decision your firm made differently because of the last four engagements, you would tell me a story rather than name a change.
That gap is not effort and it is not model quality. It is instrumentation. And the reason it persists is embarrassingly simple: post-engagement learning has no customer this quarter. I have made this argument about the underwriting equivalent — the loss history row that decides whether anything compounds — and it holds here without modification: no client asks for it, no partner is measured on it, and it produces nothing this quarter, so it is always the thing that will be started properly next year.1 Everything else in the review pays out immediately. This one pays out in engagement two, and only if you built it right.
Meanwhile the vendors have arrived at the answer they were always going to arrive at. Microsoft now tells every enterprise it needs to build “Owned Intelligence—institutional know-how that compounds over time, is unique to the firm, and is hard to replicate”, on the grounds that as agents take on more work, “they also generate valuable signals: what worked, what failed, where outcomes drifted”.2 Read that second sentence carefully. Every noun in it is a recorded signal. It is a superb description of the harvest, and a complete description of its ceiling.
Push that ceiling forward five years and you get the outcome I keep warning firms about: the world’s best institutional memory for operating yesterday’s consultancy. Perfect recall of a business model that is being repriced underneath you. I have written the terminal-value argument elsewhere and I am not going to re-run it here — the one line worth importing is that declining terminal value and strong harvest returns can coexist, and the failure is not harvesting, it is harvesting while calling it a strategy.3
This piece is about the layer above the harvest. Not a better retrospective. An instrument.
1. The record is not the knowledge, and we have known it for twenty-four years
Before designing anything, look at what happened the last time a serious institution tried to industrialise project learning.
In January 2002 the US General Accounting Office audited NASA’s Lessons Learned Information System. The finding was blunt: “our survey revealed that lessons are not routinely identified, collected, or shared by programs and project managers”.4 Forty-three per cent of programme and project managers had not submitted a lesson in two years. Twenty-seven per cent had never heard of the system before the auditors asked.
But the sentence that should stop you is the one about where the learning actually lived:
“Instead, managers identified program reviews and informal discussions with colleagues as their principal sources for lessons learned.”
— GAO-02-195, NASA: Better Mechanisms Needed for Sharing Lessons Learned4
The knowledge was moving. It was moving through conversation, and the repository was not a participant. Twenty-four years later we have built a much better repository and the conversation is still where the material is.
Two more findings from the same audit, because they pre-specify two chapters of this method. On volume, one respondent said it was “difficult to weed through all the irrelevant lessons to get to the few ‘jewels’ that you need to find”.4 Capture without discrimination produces a haystack and calls it institutional memory. And on candour, a NASA programme manager put the reciprocity problem in one sentence: “People are never rewarded for telling about how they screwed-up and caused a problem/mistake…. This will continue to be a problem until a way is found to allow and encourage people to talk about their mistakes without feeling that they are risking their careers.”4
Now the structural reason a modern AI harvest inherits the same ceiling. A 2024 systematic review of two decades of lessons-learned system design found that of all the artefacts researchers had built, “all artefacts were intended to manage text-based explicit knowledge”.5 The review is sharper still about the assumption underneath: the omission of tacit knowledge “suggests an underlying assumption that tacit knowledge held by the staff can be easily explicated and converted into text-based formats effectively and accurately”.5
That assumption is false, and it has been known to be false since 1966, when Polanyi opened The Tacit Dimension by proposing to reconsider human knowledge “by starting from the fact that we can know more than we can tell”.6 The operational form of that, from the medical-education literature, is the one that matters for a project review: experts “are often unaware of the full scope of their knowledge when performing tasks”, and tacit transfer “requires sustained close interaction”.7
The inference chain, stated as an argument rather than a finding
Nobody has run the study that proves AI project-harvesting misses unrecorded knowledge. I am not going to pretend otherwise. What exists is three verified pieces that compose: (a) lessons-learned systems capture explicit, declarative, text-based material only;5 (b) a substantial share of operating expertise is, by definition, not in text;6,7 and (c) when a large institution measured where lessons actually came from, the answer was conversation, not the repository.4 That is an argument you can check. It is not a measurement, and I will not dress it as one.
Figures I went looking for and dropped
This section wanted numbers and mostly could not have them, so here is what got cut and why. The widely circulated figure for what poor knowledge sharing costs Fortune 500 companies each year traces only to an endlessly repeated secondary citation, never to a named primary source. A frequently quoted claim that most R&D projects are never reviewed after completion could not be opened at either the conference paper carrying it or the underlying study. A statistic attributed to PMI about lessons-learned practice appears only on vendor blogs. And every “percentage of professionals using an AI notetaker” figure in circulation resolves to SEO content with no research organisation behind it.
None of them appear here. Saying so is stronger than leaving the gaps silent, because the whole argument of this piece is about refusing to promote a claim above the evidence that supports it. It would be a poor look to breach that rule in the first section.
2. “Unrecorded” is four conditions, not one
Here is the design error that makes most post-project reviews low-yield, and it happens before anyone books a room. We treat “the stuff that wasn’t written down” as a single substance with a single extraction method — ask people. It is four different conditions, and they need four different sensors.
| What is missing | Best sensor | Defensible claim |
|---|---|---|
| Compiled out Intent, rejected options, uncertainty and rationale absent from the final deliverable |
Contemporaneous deliberation, decision records, work commits | “This was the stated reasoning at the time.” |
| Performed but never articulated Workarounds, branch conditions, habitual checks |
Behavioural traces, replay, observation | “This is what people actually did.” |
| Recognised only afterwards Changed judgment, political meaning, unexpected organisational consequence |
Separate post-project interviews | “This participant now interprets the episode this way.” |
| Not yet knowable Whether the intervention persisted, transferred, or produced economic value |
Delayed outcome evidence | “The claimed lesson survived contact with the world.” |
Walk each row, because the whole method falls out of them.
Compiled out. Every finished artefact structurally erases three things: the dated intent in the worker’s own words, the graveyard of considered-and-dropped approaches, and the work that was planned in detail and never built. There is no file for the third one — that is the point, and it is why the repository literally cannot contain it.8 Knowledge work adds three more: uncertainty, because publication rewards false closure; friction, because status reports reward green; and inheritance, because a document does not colour-code what survived from last year. The sensor here is not an interview. It is the deliberation record itself, which is why a firm that emits a structured work commit at each substantial session has already had most of this conversation.8
Performed but never articulated. Documents record what people say they do; event streams record what they do; the gap between the two is tacit knowledge.9 Ask someone to describe the process and you get the procedure. Watch them five times and the steps present in every trace are the procedure — and the variations are either personal style or, far more interestingly, undocumented branches nobody ever wrote down.9 No interview reaches this, because the holder cannot narrate it. That is not evasion. It is what tacit means.
Recognised only afterwards. This is the one an interview is genuinely for, and it is a narrow band: changed judgment, political meaning, consequence the participant only understood in hindsight. Notice how little of the total this is. Notice also that this is the only row where the defensible claim is about the participant’s present interpretation, not about what happened.
Not yet knowable. Adoption, persistence, avoided cost, transferability. These are future facts. They must be null at close, and a review that fills them in anyway is not reporting, it is forecasting with the confidence of a receipt.
The move that makes this urgent rather than tidy
If your engagements are AI-assisted, much of row one is already solved and you should stop paying for it. Useful AI interaction forces people to externalise purpose, alternatives, corrections and blockers while the work is happening — the interview is happening now, prompted, in real time, and the transcript is the record.10 The document is optimised for its audience; the session is optimised for getting help, and help requires the worker to say what is hard.10
Which retires most of the archaeological case for a post-project interview and leaves exactly one job: use the recorded project to discover precisely what still requires human sensing. That residual is the whole subject of this piece.
3. Freeze the account before anyone speaks
Now the move. Everything else is downstream of it.
Before you interview a single person, the machine produces a trace-only account and you freeze it. Not a summary. A reconstruction, from the record alone, covering: the intended outcome and initial assumptions; the intervention actually delivered; deviations, exceptions and contested decisions; evidence of adoption and consequence; unresolved questions; candidate explanations; and — the field that does the work — what the documents and traces cannot establish.
From that last field you compile the negative-space map: decisions whose rationale is thin; moments where the recommendation changed; repeated workarounds or exceptions; disagreements left unresolved; planned work never undertaken; assumptions with no outcome evidence; mismatches between the declared process and observed behaviour; and apparent learning candidates supported by only one source.
The output is not “the lessons”. It is a map of what the project record cannot yet explain.
Then, and only then, you talk to people — and you score the conversation on one thing: what changed relative to the frozen account.
Why the sealing is not bureaucracy
Because the alternative is unmeasurable, and because the human memory you are about to consult has already been rewritten by the outcome.
Baruch Fischhoff established this in 1975 and named it creeping determinism. Outcome knowledge increased the postdicted likelihood of reported events “regardless of the likelihood of the outcome and the truth of the report”, and — the part that ends the argument — “Judges were, however, largely unaware of the effect that outcome knowledge had on their perceptions.”11 When he and Beyth asked people to recall the probabilities they had personally assigned before an event, subjects “remembered having given higher probabilities than they actually had to events believed to have occurred and lower probabilities to events that hadn’t occurred”.11
Read that as a specification. The pre-outcome state of belief does not survive in anyone’s head. If it exists anywhere, it exists in the record — the dated intent, the option that was killed, the estimate before it was revised. The frozen account is not paperwork. It is the only surviving witness to what the project believed before it knew how it ended.
I have argued the engineering version of this before without noticing it was the same problem. When you replay a system against its own history to test whether a policy would have caught something, the rule that makes it science rather than cosplay is the Future-Leakage Rule: do not replay the final period against a corpus that already knows what happened during that period — otherwise you are not testing discovery, you are testing whether the system can read a spoiler.46 Future leakage is the silent invalidation mode for historical replay: name it, forbid it.46
A post-project interview is a replay, and outcome knowledge is the spoiler. The transposition is exact: do not interview a participant against an account that already knows how it turned out. You cannot un-know it for them — Fischhoff settled that — so instead you freeze what was knowable beforehand and measure everything against it.
Fischhoff also supplies the sentence that should be printed on the cover of every lessons-learned programme:
“the very outcome knowledge which gives us the feeling that we understand what the past was all about may prevent us from learning anything from it.”
— Fischhoff, Hindsight ≠ Foresight, 197511
The frozen account is a preregistration
The cleanest way to understand what you are doing is to borrow from science. Preregistration exists because “the distinction between postdiction and prediction is appreciated conceptually but is not respected in practice”.12 A frozen trace-only account converts the interview from postdiction into a test of a stated prior. That is the whole mechanism, and the vocabulary maps one to one: postdiction and prediction become the frozen account and the delta.
And the payoff is measurable in the domain that pioneered it. Comparing the standard psychology literature with Registered Reports — where peer review and the decision to publish happen before results are known — researchers found 96% positive results in standard reports but only 44% in Registered Reports.13
Hold that number next to your own canon
A system that decides what counts as a finding after seeing the result confirms itself at roughly twice the rate of one that pre-commits. If your firm has a dense, well-loved body of frameworks and every engagement review comes back confirming it, you now know the base rate you are being compared against. That is not an accusation. It is an instrument reading.
4. Questions that stress an account, rather than harvest a lesson
A conventional retrospective asks for a shared story: what went well, what went badly, what should we do differently. That format rewards tidy hindsight and organisationally safe consensus. The perturbation review does the opposite — it preserves incompatible views long enough to learn from them.
So the frozen account is not context you send in advance. It is the stimulus. You put the machine’s reconstruction in front of someone and ask them to break it.
- What is wrong in this reconstructed account?
- Where did actual behaviour depart from the stated process?
- What changed because of the intervention — and what would probably have happened anyway?
- Which apparent success depended on an exceptional person or a local condition?
- What failed despite the deliverable being accepted?
- What judgment still could not be compiled, and why?
- What appears general here but would fail in another environment?
- Who sees a different edge of this outcome?
- What next observation would kill our preferred explanation?
That is not oral history. It is causal-account stress testing.
Alongside those, a second set that reconstructs episodes rather than inviting general lessons. The difference is not stylistic. “What did we learn?” asks a person to perform synthesis under hindsight, which is precisely the operation Fischhoff showed is corrupted. “At what exact moment did you stop believing the original plan?” asks for an episode with a timestamp, which you can then check against the record.
- At what exact moment did you stop believing the original plan?
- What cue did you notice that someone reading the final pack would miss?
- Which option looked attractive, and why was it killed?
- What did people actually do when the designed process failed?
- What became possible only after the project?
- What now looks causal but might merely have accompanied the outcome?
- What evidence would make you reverse this lesson?
The one question that manufactures new evidence
Most review questions retrieve. One class of question creates evidence that was never in anyone’s head, because the question had not existed before. My favourite, asked of the client sponsor:
“What part of this engagement would you now refuse to pay another consultancy to do, because you’ve seen how it can be done?”
Nothing in the project record contains that answer. It is not a delivery question — it is an industry-economics question wearing a delivery costume, and it crosses an altitude in a single sentence. Two more in the same family: what did we do that you thought you were purchasing, and what did we do that you didn’t realise you needed until you saw it? And to the delivery lead: which piece of work felt like delivery but was really compensation for a bad commercial boundary? Nothing in Jira says that.
The instrument is not neutral, and you must design for that
Loftus and Palmer showed in 1974 that the verb inside the question rewrites the memory. Asking how fast the cars were going when they smashed into each other produced higher speed estimates than collided, bumped, contacted or hit — a nine mile-per-hour spread from a single word. A week later, the subjects who got smashed were more likely to say yes to “did you see any broken glass?” — and there was no broken glass in the film.14
This is not a curiosity. It is the mechanism of the self-confirming flywheel, demonstrated experimentally fifty years before anyone built a canon dense enough to supply the verb. Your frameworks are the verb. Which is why the record keeps observed fact, exact phrase and marked inference in separate fields — memory launders inference into fact, and the discipline is separation, not completeness theatre.15 I am not going to re-teach that capture object here; it is published, and this review imports it unchanged.15
5. Four positions, never averaged
“Interview the client, the users, delivery and management” is directionally right and structurally wrong. It sorts people by org chart. Sort them by what they were in a position to know instead:
| Epistemic position | The question it can answer that no other can |
|---|---|
| Authority / funder | What decision, budget or obligation actually moved? |
| Operator / participant | What changed in lived work — including the workarounds and the unrecorded burden? |
| Delivery | Where did the intended intervention meet reality, exceptions or residual scarcity? |
| Adjacent sceptic / control owner | What consequence appeared outside the project’s declared success perimeter? |
The fourth position is the one firms skip, and it is where the transferable material concentrates. “Management” reproduces the authorised narrative. A downstream recipient, a risk owner, a team that inherited a new manual step, a non-adopter — those people saw the externality. Safety investigation has been fighting this exact battle for decades, and the Australian Transport Safety Bureau names both the payoff and the resistance in the same passage: “the most effective findings for safety enhancement are often the most difficult to justify”, and “ATSB often encounters attitudes such as ‘why is the ATSB looking at that, it had nothing to do with the accident’”.16 If nobody in your firm ever asks why you are interviewing the team next door, you are not running this method.
Run them separately wherever hierarchy could suppress disagreement. The goal is not consensus. The goal is to give disagreement an address. And it is worth knowing that the suppression is measurable rather than folkloric: a meta-analysis of 89 studies, 124 samples and 303 effect estimates on sensitive-question techniques found that changing the structure in which a question is asked significantly improves response validity — and that the effect is heterogeneous, meaning you cannot correct for it with a constant factor.17 You have to change the room.
Hunt the off-diagonal
Once you have four accounts, do not average them. Compare formal importance against lived attention and go looking for the cells where they disagree — the disagreement between the two rankings is diagnostically more valuable than either alone.18 Applied to a completed engagement:
| Cell | What it probably means | What to do about it |
|---|---|---|
| Formally successful, high lived friction | The success was produced by heroics | Find who; ask what they did that wasn’t in the SOW; that is a delivery-machinery gap, not a compliment |
| Formally important, strangely little attention | Mature and stable — or ritualised and hollow | Ask the owner to explain the failure mode without reading the procedure. Their answer discriminates |
| Low formal importance, repeated operator attention | An unofficial dependency is emerging with no dashboard identity | Name it, give it an owner, or record why you are leaving it alone |
| Delivery confidence, no observable client consequence | The artefact closed; the outcome did not | Open an outcome-backfill date. Do not write a lesson yet |
One discipline governs all four cells, and it is the same one that governs weak-signal sensing generally: chatter should be a prior, never a verdict; the signal must perturb, not command.19 High friction does not equal a real problem. Consensus does not equal correctness. A cell is a nomination for investigation, not a finding.
6. Some of the most important fields must be null
Adoption. Persistence. Avoided cost. Transferability. Whether the operating-model change actually
removed the bottleneck. These are future facts, and at project close the honest value is
null.
The cross-domain model here is accident investigation, and it is worth being precise about why. A flight recorder preserves sequence and system state. Witnesses contribute perception and intent. Later physical evidence tests both. None of the three is an oracle. That is exactly the relationship between your trace-only account, your four interviews, and the outcome you have not observed yet.
The investigative bodies are explicit about how they weight the human class. The NTSB investigator’s manual instructs that witness interviews be obtained as soon as feasible because “long delays between the witnesses’ observations and the interviews increase inaccuracies in their statements”, and that witnesses “should be urged to relate only their personal observations and refrain from passing on information they have derived from some other source”.20 Both rules transfer directly. Run the interviews close to close. And when someone tells you what the client thought, that is hearsay in your record, not evidence.
The engineering version of the null field is a design I have argued for in attention systems, and it is the field almost nobody implements: a first-class nullable later outcome that a subsequent process fills, by human label or delayed evidence join. The line that justifies it is uncomfortable and correct — without backfill, review is only aesthetics; with it, learning becomes a supervised problem.21
Which produces one concrete change to how an engagement ends:
Close-out should schedule evidence events, not merely finish documentation.
“Project completed” and “outcome closed” are different states. Put the second date in the calendar while the first one is still being celebrated, because the measurement window closes at project close and does not reopen. And note the environmental squeeze: CEO planning effort has compressed — half of it now sits inside a horizon of less than a year, up from 43% the year before22 — while the outcome clock on your engagement is exactly as long as it always was. That widening gap is the thing outcome backfill exists to close.
7. Type the delta, then let the claim climb — slowly
Every difference between what the conversation produced and what the frozen account already contained gets a type. Six of them, and the list is deliberately unflattering:
| Delta type | What it earns |
|---|---|
| New fact | Correction to the frozen account. Highest-value routine output |
| Contradiction | Two accounts disagree. Preserve both; do not resolve by seniority |
| Causal hypothesis | A proposed mechanism. Needs a falsifying next observation attached before it moves |
| Local lesson | True here, stays here. This is a success, not a shortfall |
| Transferable candidate | Enters the nomination lane. Nothing more |
| Noise | Recorded and closed. If this bucket is empty, your typing is broken |
Then the ladder. “Captured” is not a terminal state; a claim has an altitude, and the only real question is whether it has earned the right to climb.
- Episode evidence — what happened, was said, or was observed in this engagement.
- Client operating prior — a client-approved interpretation that should alter the client’s next decision.
- Transfer hypothesis — a de-identified, rights-cleared possibility that the pattern recurs elsewhere.
- Field pattern — seen across suitable cases, or strong enough to justify an explicit test.
- Firm doctrine — a transferable distinction that has survived challenge, changed future action, and earned human admission.
Most material should stop at levels one or two. That is not failed compounding. It is correct discrimination.
To see why the ladder is not bureaucracy, walk one observation up it. A data reconciliation took much longer than expected. At practice level that says: we need a better reconciliation harness. At commercial level: clients shouldn’t be paying implementation rates while we discover their data estate. At offer level: perhaps certainty about the estate is itself the product. At industry level: maybe much of “consulting delivery” is monetised uncertainty created by an inability to cheaply inspect reality. At doctrine level: cheap cognition may move professional services from producing answers to bounding and proving decisions.
Same project. Six altitudes. A competent harvesting agent will find the first two increasingly well and will keep getting better at them. The scarce capability is recognising when an observation deserves to climb — and, far more often, when it does not. As I put it when I first stumbled on this: the altitude is what lets you see what’s new, and what’s trivial. The second half is the harder half.
Evidence ceilings, borrowed from people who publish theirs
An unusual n=1 can move fast to a transfer hypothesis when its consequence is large.
It still carries its evidence ceiling. “This may expose a general mechanism” is a
different claim from “this is how firms work”, and a dense canon is exactly the thing
that makes the second one feel earned.15
Safety investigation solved this by splitting two ideas most firms never separate: the standard of proof (how sure you must be) from the standard of evidence — “the quantity or quality of evidence required before a decision maker can be satisfied that the relevant standard of proof has been met”.16 The ATSB then does something almost nobody in professional services does: it publishes the threshold. The term “probably”, it says, means “more than 66 per cent likelihood”, and evidence burden scales with consequence — “the more significant the issue to be determined… the clearer and more persuasive the evidence needed”.16 It also notes, drily, that “most organisations that conduct safety investigations do not clearly specify the standard of proof (or standard of evidence) they use”.16
Neither does your firm. Write yours down before the next review, not during the argument about a particular finding.
The governing rule of the whole promotion layer
Recurrence nominates; it does not promote.23 A pattern seen three times may still be local-only, configurable, an internal delivery primitive, a supported capability, or a reject — and the promotion test is a counterfactual, not a count: would the next engagement actually be easier or safer because this was promoted?23 If you cannot answer that, you do not have a field pattern. You have hope with a release tag.
Preserve the rejects. A rejected candidate teaches the boundary, and deleting it means the next review re-litigates it blind.23
Two more constraints, both imported rather than invented. Nomination is not publication, and publication is not automatic — parking a recurrence is a success mode, and contested facts should stay contested rather than be flattened into a false consensus.24 And a durable write is justified only when understanding changes in a way that should affect future judgment; if the author cannot state what understanding changed, the write belongs on a case trajectory or nowhere.25
8. Three artefacts, kept apart
A review that produces one document has already collapsed the architecture. Three outputs, three owners, three different burdens of proof.
1. The Outcome Closure Receipt — client-owned, and the thing they actually paid for
The schema is not new and I am not going to rebuild it: proposal, execution, observation, learning, kept as separate blocks.26 What matters here is what the review has to put in it, which is mostly things a normal close-out report is designed to remove:
- contradictions, rather than a smoothed retrospective;
- unresolved attribution, rather than invented causality;
- remaining operating risks;
- the decisions and the help the client still needs;
- a later observation date wherever the outcome cannot yet be known.
Without those fields you get folklore — “that vendor is difficult”, “that playbook is outdated” — none of it addressable, none of it replayable, none of it safe to promote.26 With them you have something the client can take to its own board without you in the room, which is the actual test.
2. The Field-Pattern Nomination — the firm’s separate, rights-cleared record
Anonymised operating context; the local intervention and the value the client received now, not after productisation; the domain prior and the process prior learned; the exception and failure shape; the repeated invariant versus the local variation; evidence pointers including an explicit confession of what was not verified; confidentiality and reuse status; the transfer test; and a named promotion owner with a decision, a date and a revisit trigger.27
If the only artefact is “we built a connector for Client X”, you have a services receipt, not product discovery.27 And the failure mode on the other side is just as expensive: promoting one client’s workflow into the shared platform after a single deployment is not discovery, it is importing one client’s interface politics into everyone else’s operating system.23
3. The Machinery Change Receipt — the one that decides whether any of this was real
“Feed it back into the machinery” is where most learning programmes go to die, because the sentence has no object. Give it seven fields. For every mutation record: what changed; which evidence caused it; which prior it replaces or qualifies; the applicability conditions; the owner; the next transfer test; and the rollback or demotion condition.
A legitimate change is concrete — a new standing question, an altered offer-eligibility rule, a new acceptance test, an exception taxonomy, a delivery scaffold, a risk threshold, a reusable evaluation, a software primitive, or a documented rejection rule.
The measure is not whether IP was written. It is whether future behaviour changed.
9. The dangerous version: a canon marking its own homework
Now the part I have to argue against myself, because it is the strongest case against everything above and it is not hypothetical.
The failure chain runs like this. I arrive with a dense canon. That canon determines which anomalies look interesting and which questions I ask. Participants reconstruct events already knowing the outcome. The AI finds several of my frameworks that explain the resulting story. I recognise the fit and approve promotion. The new “learning” enters the same canon that shaped its discovery.
The loop then compounds its own priors rather than reality. And here is the vicious part: greater corpus density and better synthesis make the error more coherent, not less. That is not compounding judgment. It is compounding confidence.
The comfortable answer is a human gate. The comfortable answer is wrong. If the same taste proposes the frame, interprets the evidence and admits the result, then “human approved” can only mean “the founder remained persuaded”. And before anyone reaches for experience as a control, the 2025 study of patient-safety incident investigators is worth sitting with: worsening outcomes increased judgements of staff responsibility and produced more punitive recommendations from the same scenario, and while those with safety expertise showed the effect “to a lesser extent”, they still showed it.28
Expertise reduces the bias. It does not remove it.
Which means “we’re senior enough to notice” is not a control. It is a preference for not having one.
Epistemic separation of powers
The remedy is structural, because personal virtues fail under pressure — especially the pressure of an idea you are already in love with.29 Seven passes, seven artefacts:
- The harvesting pass reconstructs, and is not permitted to promote.
- The interviewer records testimony, and does not silently translate it into doctrine.
- A challenger generates rival explanations — and its list must include “client-local”, “outcome bias”, “no reusable learning” and “this contradicts existing doctrine”.
- The client confirms, conditions or rejects claims about its own world.
- A rights and transferability review decides what may leave that world.
- Someone other than the interviewer authorises promotion.
- Later outcomes retain the power to supersede all of it.
A two-person practice will have one person walking several of those passes. That is fine — but the passes and the artefacts must still be separated, and the system must be able to show where observation ended, interpretation began, the client spoke, the challenge occurred and promotion was authorised. This is the same structural instinct as separating epistemic access from causal authority in AI systems generally,30 and it inherits the same warning from the audit side: the machinery that compressed the ambiguity cannot independently certify that no important ambiguity was lost.31
Four dispositions, and the one almost nobody has
Every claim about the canon gets one of four dispositions: confirms, qualifies (survives only under narrower conditions), contradicts (the existing pattern must be challenged or demoted), and alien — the observation that fits no current frame and must not be discarded merely because it is difficult to classify.
Alien is the disposition that breaks the flywheel, and it needs a budget rather than good intentions. Reserve a fixed allocation for high-significance observations with zero intersection with your existing map, and then measure how often those later become canon — if the answer is never, the lane is wrong, not the idea of the lane.32 Pair it with a scheduled suppression audit: sample the candidates the system chose not to surface and score the misses, because learning only from what you kept is how echo chambers become code.32 And never accept a fluent paragraph as causal proof — if you cannot re-run the judgment from stored inputs, you do not have a judgment history, you have a story.32
The reason a dense canon needs all three is that the density is simultaneously the asset and the hazard: diffing new evidence against your accumulated map is the power and the trap, because the map can become a suppression shield around unfamiliar fields.33
The kill path has to win sometimes
Somewhere in the promotion process there must be a live, first-class option to promote nothing — and it cannot be a token seat. If the kill-path never wins, it isn’t a guard; it’s set dressing.29 The test you can run on yourself this afternoon is the one I use: when did your review process last kill a lesson you liked? If you cannot remember, you have not been running an evaluation. You have been running a ratification ceremony, and the confident yes it produces is worth exactly what a yes is worth from something that was never able to say no.29
What survives when you give up the flattering claim that human origination is irreplaceable is narrower and much more durable. A panel cannot supply the criterion it is being scored against, and it cannot carry the risk — “It’s not a gavel at the end. It’s a star overhead and a name on the line.”34 Owning the evaluation function and the consequence boundary is the job. Everything between those two poles is delegable, and getting more so.
A falsifier you can run on your own corpus this quarter
If successive engagements always confirm the canon, rarely demote anything and generate no genuinely alien categories, the system is probably assimilating evidence rather than learning from it. Contradiction is not a defect in an IP system. It is one of its highest-value outputs.
10. Reciprocity is a production constraint, not an ethic
I want to be precise about why this section exists, because it is usually written as a values statement and it is not one. If the return path fails, the sensor degrades. That is an engineering claim, and the evidence for it is the NASA manager’s sentence about risking careers.4 Fear of blame does not make people lie. It makes them stop volunteering, which is worse, because the record still looks complete.
The patient-safety literature says the same thing in the same shape: fear of punishment or blame “often discourages healthcare professionals from reporting errors and near-misses, leading to missed opportunities for improvement”.35 Aviation’s answer was to design for it — a just culture, meaning an atmosphere of trust in which people are encouraged, even rewarded, for providing essential safety-related information.36 Note that the NTSB manual insists the format is an interview, not an interrogation, conducted “on a basis of courtesy, cooperation, and neutrality”.20 That is a design decision about signal quality, not manners.
Fischhoff, again, from 1975, on why hindsight makes this worse: “When second-guessed by a hindsightful observer, his misfortune appears to have been incompetence, folly, or worse.”11 A review that knows the outcome will read every in-flight judgement as a mistake unless it is designed not to.
The rule
Every signal extracted from people must return either help, corrected authority, reduced friction, a better decision, or acknowledged learning. Without that return path, the process is vendor extraction under a collaborative costume.
Six commitments, each checkable by a participant rather than by you:
- the client receives the Outcome Closure Receipt and the resulting actions;
- participants receive a short synthesis showing what their contribution changed;
- client evidence stays client-controlled;
- reusable learning is de-identified and promoted only under agreed rights;
- a refusal to permit reuse does not reduce the client’s closure benefit;
- participant statements never become individual performance evidence.
That last one is contract-grade, not a setting. The equivalent rule in behavioural capture is the sharpest version I have written: the system records procedure shape, never individual performance; it exists to scaffold you, not to rate you; no productivity metrics, ever, contractually. The day a manager asks for time-per-transaction by staff member is the day the product has failed its own contract — and the boundary is not a limitation on the feature, the boundary is the feature.9
Consent is not theoretical either. A July 2026 Pollfish survey of 500 US workers commissioned by Kolmogorov Law found that one in three employed Americans say an AI notetaker or transcription bot has been present in their work meetings, while only 34.7% say they were always asked for permission.37 It is a single commissioned panel survey and I would not build a strategy on the exact figures — but as a description of the room you are walking into, it is directionally undeniable. People already suspect they are being recorded and already suspect nobody asked.
The membrane control that makes this implementable
One practical move does more than any policy: where possible, promote a synthetic reproducer of the failure shape rather than the original case. The shared kernel gets a regression test; the client’s evidence never travels. And keep the direction of travel honest — local delivery can be a gravel road, correct for this workflow and this approval role, while promotion is a paved road with an entirely different burden. Automatic upward movement from client exhaust into shared IP is not compounding; it is a breach of the architecture.38
All of which is downstream of one commercial fact, and it is the line I would hold in a negotiation: the paid product is the client’s closure, not your IP harvest. Get that wrong and the pitch becomes “pay me to use your projects to make my IP smarter”, which is both unattractive and, as a sensor design, self-defeating.
The test for whether you have actually handed something over is one I set for engagement close generally, and it is unforgiving: if the client cannot operate the world without your private kernel as a hidden dependency, you did not deliver a world — you delivered a leash.48 The same chapter supplies the rule for the gate on the other side, which is the one most firms quietly fail: approve and reject are both success outcomes of a working gate; a queue that only ever approves is not a gate, it is a conveyor.48
11. The strongest case against this — including the parts that argue against me
The strongest counter-case is not that interviews produce nothing. It is that trace-first analysis may already produce nearly all the valuable learning, more accurately and more cheaply.
Post-project interviews can introduce recency and hindsight bias; polite agreement with the interviewer’s thesis; attribution of systemic outcomes to the visible project; management narrative dominating operator experience; your existing frameworks forcing alien evidence into familiar categories; attractive abstractions that never change delivery; and client fatigue with unpaid participation. Every one of those is real, and several of them are experimentally demonstrated in the sources above rather than merely plausible.
Now the honest counterweight, which cuts the other way. Properly conducted structured debriefs do work: a meta-analysis of 46 samples covering 2,136 participants found that “organizations can improve individual and team performance by approximately 20% to 25% by using properly conducted debriefs”, with effects holding across teams and individuals, simulated and real settings, medical and non-medical.39
Which is exactly the right shape of evidence
Retrospection is not the problem. Unstructured, outcome-contaminated, hierarchy-averaged retrospection is the problem. The effect lives in the design of the instrument — which means a badly designed instrument should be killed rather than defended, and a well-designed one has a documented base rate to beat.
The evidence from my own history, which does not flatter this idea
I have been doing versions of this for a long time, and the record is instructive in the wrong direction. A 2006 project retrospective outline of mine proposed exactly the conventional thing — hear the client’s specifics, ask what went well and what should improve, reset the process. That establishes precedent for review. It does not establish a compounding IP system.
More sharply: in 2003 a client refused further review meetings until the product had reached a point where it added value. And feedback on a portal project around the same period rated “learning as you go” as expected quality rather than a premium selling point, while emphasising the importance of listening to the workers rather than only the boss.
The correction those three records force is the design constraint I hold today:
Clients will not value a conversation because it improves the supplier. They may value a bounded closure product that improves their own operating system.
And one dataset that argues against the whole posture
The ANA and the 4As found that clients without mandatory review periods averaged 8.1-year agency relationships, while those with frequent reviews ran as low as 3.8 years.40 Read plainly: removing the periodic falsification test lengthens the relationship. Which is the opposite of what a killable review programme does on purpose.
I am not going to reinterpret that into support. What it measures is relationship survival, not relationship value — two different quantities that are easy to confuse because one of them is easy to count. Running a review that can kill itself is a bet that a relationship designed to be falsifiable is worth more per year than one designed to persist. It might be wrong, and those are the numbers that would show it.
Meanwhile the buyer side is moving the other way. The UK National Audit Office is unambiguous that consultants “should only be used where they represent best value for money and not to replace capability required inside the civil service”, and recommends “ensuring that civil servants learn from consultants while they are working together by building knowledge transfer agreements into contracts”.41 When one of the largest buyers of advisory services in a G7 economy writes transfer into the paperwork, an instrument whose output is a client-owned closure asset stops being a nice gesture.
12. Run it as a falsification programme, with the kill conditions written first
This is the part that makes the difference between a method and a ritual, and it has to be pre-committed — before the first review, not after the first disappointing one.
The causal chain under test is:
interview → material semantic delta → transferable mechanism → machinery change → better subsequent outcome
Four links. Each can fail separately, and knowing which one failed is most of the value.
The protocol — seven steps, for the next few engagements
- Freeze the trace-only account before any interview. Timestamp it. Nobody edits it afterwards.
- Run the role-separated perturbation reviews — four positions, separate rooms.
- Record only what changed relative to the frozen account. Restatement earns nothing.
- Classify each delta into the six types.
- Have someone other than you disposition promotion.
- Test promoted candidates against an unrelated later engagement.
- Return later to observe whether the client outcome actually persisted.
What to measure
- material contradictions found only through conversation;
- candidate patterns that survive an unrelated transfer test;
- machinery changes with evidence and version history attached;
- client decisions or operating improvements caused by the closure asset;
- downstream movement in rework, speed, risk or scarce-expert demand;
- participant burden and refusal rate;
- whether routine closure becomes possible without you.
Directions and comparisons across cycles — never fabricated percentages. Shape, not invented magnitudes. And if you cannot measure a row at all, that is itself the finding: the process is not instrumented enough to claim what it is claiming.42
Kill or redesign the review if…
- interviews mostly restate the compiled trace;
- promoted patterns repeatedly collapse into client-specific exceptions;
- the client does not use the returned synthesis;
- participants experience it as surveillance or vendor R&D;
- machinery changes cannot be tied to any downstream improvement;
- your personal involvement remains necessary on known terrain.
Six conditions. Publish them to the client before the first review. A programme that cannot be killed by its own evidence is not a programme — it is a subscription.
The freeze protocol carries one more borrowed test, and it is the sharpest kill condition in the set. When you freeze a replay you are supposed to record the as-of timestamp of world knowledge, which events are allowed after it, where the decision traces are stored, and what would count as contamination — and then: “If any answer is ‘the model just knew,’ you do not have replay. You have a demo.”46 Transposed to this method: if the interviews restate the frozen account, you do not have a sensor. You have a ceremony.
And the review of the deltas should look like a suppression audit rather than a highlights reel. The four-bucket structure I specified for attention systems transposes almost unchanged: what the frozen account already claimed (precision); what it nearly explained — thin rationales, changed recommendations, unresolved disagreements (boundary calibration); what one position confidently classified as fine that another experienced as friction; and the observations with no place in the current frame at all.47 The third bucket is the one nobody builds, and it has the best name in the corpus: a case dismissed not because the evidence was weak but because the worldview had no place to hang the significance does not show up as a near miss. It shows up as silence with a clean conscience.47 That is exactly what a confident trace-only reconstruction produces.
The vocabulary for what “good” looks like already exists, and it is unforgiving: falsification throughput is the rate at which you eliminate options, each elimination carrying the named evidence that killed it. Not candidates generated. Not initiatives launched. Eliminations, with receipts.43 Applied here: a review that produces findings without eliminations is generating candidates and calling it learning. Activity does not move this clock. Evidence moves it.43
13. One review, compressed
Designed specimen — composite, not an executed client engagement
What follows is a designed walkthrough, fused from patterns rather than drawn from one client. No metric below is a measurement, and none is presented as one. Its job is to show the shape of the outputs, not to prove an effect.
The engagement. A reporting-and-reconciliation programme for a mid-sized enterprise. Delivered, accepted, invoiced, closed. Everyone reasonably pleased.
Stage one — the frozen account. The machine reconstructs the engagement from proposal, SOW variations, decision records, project chat, tickets and the final pack. Frozen and timestamped. The negative-space map that falls out of it has, among others, four entries worth following: a design recommendation that reversed in week six with a thin rationale; a repeated manual reconciliation step appearing in delivery chat but in no document; a scope item planned and silently dropped; and an adoption assumption with no outcome evidence whatsoever.
Stage two — four rooms. Authority, operator, delivery, and one adjacent sceptic: the team downstream who inherited the reports and was never in a project meeting.
Stage three — the delta table. Only what the frozen account did not already contain:
| Delta | Position | Type | Disposition |
|---|---|---|---|
| The week-six reversal was driven by a compliance conversation that never entered the project record | Authority | New fact | Confirms — corrects the frozen account |
| The manual reconciliation step is still being run weekly, months after go-live | Adjacent | Contradiction | Qualifies — “automated” is conditional |
| Delivery believed the estate was the hard part; the operator says the hard part was agreeing what the numbers meant | Operator / Delivery | Causal hypothesis | Transfer candidate, with a falsifier attached |
| The sponsor would not now pay another firm to discover the data estate | Authority | Transferable candidate | Nomination only — n=1, ceiling recorded |
| Two team members independently rebuilt the same lookup because neither knew the other had | Operator | Local lesson | Local-only. Correct outcome |
| A general observation that “communication could have been better” | Delivery | Noise | Recorded, closed |
Read the second row again, because it is the one that pays for the review. Everything about that engagement was formally successful. The report shipped, acceptance passed, the invoice cleared. And a downstream team is still doing a weekly manual step that the closure documentation says was automated. That is the delivery confidence, no observable client consequence cell, and no amount of trace analysis would have surfaced it, because the trace ends where the project ends.
Stage four — the alien lane. One observation fits nothing. The operator mentions, almost as an aside, that the team stopped trusting a particular dashboard not because it was wrong but because it had once been right in a way that embarrassed someone. There is no frame in my canon for “accuracy as a political liability”. The temptation is to file it under change management and move on. Instead it goes into the alien lane with its exact phrase preserved, and it stays there until either a second instance arrives or the lane’s scheduled review retires it. Not classifiable is a classification.
Stage five — the three artefacts. The client gets an Outcome Closure Receipt with the contradiction preserved, the unresolved attribution stated as unresolved, and an observation date set six months out. The firm gets one Field-Pattern Nomination — the “agreeing what the numbers mean” hypothesis, de-identified, marked as a transfer hypothesis with an evidence ceiling and a named test. And one Machinery Change Receipt: a new eligibility question added to the qualification checklist, with the evidence, the prior it qualifies, the applicability condition, an owner, and a demotion trigger if it fails on the next two engagements.
Stage six — the backfill that reverses a lesson. Six months later, somebody returns to the observation date. The promoted eligibility question has fired twice. Once it caught a genuine problem. Once it disqualified an engagement that a competitor took, delivered without difficulty, and that showed none of the predicted failure shape. The transfer hypothesis is demoted from field pattern back to transfer hypothesis, and the demotion is recorded with its evidence.
That last paragraph is the product
Not the nomination. The demotion — with a date, an evidence pointer and a named person who authorised it. A learning system that has never demoted anything has never been tested, and the “lessons” in it are indistinguishable from things everybody already believed.
14. Monday
Five things, and none of them requires a budget line.
One. Place your current process on the four-condition table. Which of the four do you actually sample? Most firms sample condition three, badly, and assume it covers the other three. If you cannot name your sensor for “performed but never articulated”, you do not have one.
Two. On the next engagement that closes, freeze the account before anyone talks. It does not have to be sophisticated. A dated reconstruction and an explicit list of what the record cannot explain is enough to start measuring. You cannot retrofit this — the baseline has to exist before the conversation.
Three. Add the fourth chair. Find one person who saw a consequence outside the declared success perimeter and was never in a project meeting. That single interview will tell you more about whether this method is worth running than any amount of internal reflection.
Four. Put a date in the calendar for the outcome you cannot see yet, and put it in the closure document where the client can see it too. Close-out schedules evidence events.
Five. Write your kill conditions down and give them to the client. Six of them are drafted above. Publishing them costs you nothing if the method works, and saves you a year if it does not.
Where this leaves my own role, and yours
Not every project earns this. Most closures should be run by the client’s own people, the delivery team and the machinery. A project earns the intervention when it has some combination of material strategic consequence, an unresolved contradiction, a high likelihood of transferable learning, an unusual failure shape, evidence that formal success and lived operation disagree, or a live question in the strategic portfolio.
Which means the role has to erode on purpose. The honest renewal test is the one I hold myself to: is the organisation getting better at closing known classes of outcome without me, while I keep finding consequential new classes? If engagement ten requires the same routine interview as engagement one, nothing compounded — I just installed another founder dependency with a recurring invoice attached. That is the same discipline as transfer the known, earn renewal at the frontier,42 and the same proof standard as everywhere else in this body of work: engagement two is the only proof that matters.44 The compounding test is one sentence long — did we leave merely another codebase, or a sharper corpus that makes the next project easier and better?45
One adjacency I will name and leave alone, because it belongs to a different argument with a different instrument: whether I personally get more capable while my firm becomes no more transferable is a distinct failure with its own ablation test. Different question, different apparatus — named here so you know I have not confused the two.
Here is where I have landed, after arguing myself out of the flattering version twice. AI harvesting projects, materials and documents is awesome. It does a fantastic job. But it doesn’t seem to go the next step — and the next step is not a smarter model or a bigger wiki. It is a frozen baseline, four sensors instead of one, a claim lifecycle with a ceiling, a separation of powers, and a written-down reason to stop.
AI can harvest what the project recorded. The higher-value act is discovering what the project taught and never recorded. But the version of that sentence I will actually defend is longer and less quotable: AI reconstructs what the engagement made legible; human testimony and behavioural evidence probe what remained tacit, contested or unasked; and later outcomes decide what was genuinely learned.
Every project should leave two things behind: the result the customer bought, and a better organisation than the one that started it. Most firms are now excellent at proving the first. The second one needs an instrument, and the instrument has to be able to tell you when it isn’t working.
One ask
Run one perturbation review with a frozen account and a written kill condition, then count the deltas that were genuinely absent from the trace. If the number is near zero, tell me — that is the most useful result this method can produce, and it is the one I would most like to hear about.
References
- Scott Farrell, LeverageAI. “Fixed Price Is Underwriting.” — The eighth underwriting job, engagement actuals: “Without it, engagement twenty is exactly as risky as engagement two, and every hard-won lesson lives in the head of whoever was on the engagement”; “Loss history has no customer at all — no client asks for it, no partner is measured on it, and it produces nothing this quarter. So it is always the thing that will be started properly next year”; “The blank rows are the finding. Not the filled ones.” https://leverageai.com.au/wp-content/media/articles/233-fixed-price-is-underwriting.html
- Microsoft. “2026 Work Trend Index: Agents, human agency, and opportunity” (5 May 2026; survey of 20,000 knowledge workers across 10 countries, fielded 18 February – 7 April 2026). — “Every Frontier Firm needs to build Owned Intelligence—institutional know-how that compounds over time, is unique to the firm, and is hard to replicate”; “As agents take on more, they also generate valuable signals: what worked, what failed, where outcomes drifted.” https://www.microsoft.com/en-us/worklab/work-trend-index/agents-human-agency-and-the-opportunity-for-every-organization
- Scott Farrell, LeverageAI. “The Terminal Value Doctrine: Professional Services.” — “Declining terminal value and strong harvest returns can coexist”; “The failure is not harvesting. It is harvesting while calling it a strategy.” https://leverageai.com.au/wp-content/media/articles/231-terminal-value-doctrine-professional-services.html
- US General Accounting Office. “NASA: Better Mechanisms Needed for Sharing Lessons Learned,” GAO-02-195, 30 January 2002. — “our survey revealed that lessons are not routinely identified, collected, or shared by programs and project managers”; “Instead, managers identified program reviews and informal discussions with colleagues as their principal sources for lessons learned”; “it is difficult to weed through all the irrelevant lessons to get to the few ‘jewels’ that you need to find”; “People are never rewarded for telling about how they screwed-up and caused a problem/mistake.... This will continue to be a problem until a way is found to allow and encourage people to talk about their mistakes without feeling that they are risking their careers.” https://www.govinfo.gov/content/pkg/GAOREPORTS-GAO-02-195/html/GAOREPORTS-GAO-02-195.htm
- Benjamin X. Hou and Omar El-Gayar. “Two decades of design research on lessons-learned systems: a systematic review and future research agenda,” Issues in Information Systems, Vol. 25 Issue 4, pp. 373–391, 2024. — “we found that research has primarily focused on the design of artifacts that support explicit and declarative knowledge. Despite the importance of tacit knowledge on organizational performance, we find that all artifacts were intended to manage text-based explicit knowledge”; “The omission of support for the management of tacit forms of knowledge also suggests an underlying assumption that tacit knowledge held by the staff can be easily explicated and converted into text-based formats effectively and accurately.” https://iacis.org/iis/2024/4_iis_2024_373-391.pdf
- Michael Polanyi. The Tacit Dimension, University of Chicago Press (1966; Chicago edition 2009). — “I shall reconsider human knowledge by starting from the fact that we can know more than we can tell.” (Quoted verbatim on the publisher’s page; the commonly cited page number was not verified against the book.) https://press.uchicago.edu/ucp/books/book/chicago/T/bo6035368.html
- “Insights From Michael Polanyi: Tacit Knowledge and Its Critical Importance in Medical Education,” PubMed Central. — “experts are often unaware of the full scope of their knowledge when performing tasks; tacit transfer requires sustained close interaction.” https://pmc.ncbi.nlm.nih.gov/articles/PMC12927663
- Scott Farrell, LeverageAI. “The Deliberation Is Source.” — “The most valuable things in a session are the ones the repository can’t hold: the thing you considered and rejected, and the thing you planned and never built”; the knowledge-work specials — uncertainty (“Publication rewards false closure”), friction (“Status reports reward green”) and inheritance; and the Knowledge Work Commit schema: “Every substantial AI-assisted work session should emit a structured Knowledge Work Commit.” https://leverageai.com.au/wp-content/media/articles/189-the-deliberation-is-source.html
- Scott Farrell, LeverageAI. “Capture Was Never the Bottleneck.” — “Documents record what people say they do. Clickstreams record what they do. The gap between the two is tacit knowledge”; “The steps that appear in every trace are canonical. The variations are personal style — or, more interestingly, undocumented branches”; “The system records procedure shape, never individual performance; it exists to scaffold you, not to rate you; no productivity metrics, ever, contractually”; “The boundary is not a limitation on the feature. The boundary is the feature.” https://leverageai.com.au/wp-content/media/articles/84-capture-was-never-the-bottleneck.html
- Scott Farrell, LeverageAI. “Born Structured Exhaust.” — “The interview is happening now — prompted, in real time — and the transcript is the record”; “The document is optimised for its audience. The session is optimised for getting help. Help requires the worker to say what is hard.” https://leverageai.com.au/wp-content/media/articles/191-born-structured-exhaust.html
- Baruch Fischhoff. “Hindsight ≠ Foresight: The Effect of Outcome Knowledge on Judgment Under Uncertainty,” Journal of Experimental Psychology: Human Perception and Performance, 1975, Vol. 1 No. 3, pp. 288–299. — “Judges were, however, largely unaware of the effect that outcome knowledge had on their perceptions”; “subjects remembered having given higher probabilities than they actually had to events believed to have occurred and lower probabilities to events that hadn’t occurred”; “the very outcome knowledge which gives us the feeling that we understand what the past was all about may prevent us from learning anything from it”; “When second-guessed by a hindsightful observer, his misfortune appears to have been incompetence, folly, or worse.” https://web.mit.edu/curhan/www/docs/Articles/15341_Readings/Behavioral_Decision_Theory/Fischhoff_1975_Hindsight_is_not_equal_to_foresight.pdf
- Brian A. Nosek, Charles R. Ebersole, Alexander C. DeHaven and David T. Mellor. “The preregistration revolution,” PNAS 115(11), 2600–2606, 2018. — “This distinction between postdiction and prediction is appreciated conceptually but is not respected in practice.” https://pubmed.ncbi.nlm.nih.gov/29531091/
- Anne M. Scheel, Mitchell R. M. J. Schijen and Daniël Lakens. “An Excess of Positive Results: Comparing the Standard Psychology Literature With Registered Reports,” Advances in Methods and Practices in Psychological Science 4(2), 2021. — “Analyzing the first hypothesis of each article, we found 96% positive results in standard reports but only 44% positive results in RRs”; “In Registered Reports (RRs), peer review and the decision to publish take place before results are known.” https://journals.sagepub.com/doi/10.1177/25152459211007467
- Elizabeth F. Loftus and John C. Palmer. “Reconstruction of Automobile Destruction: An Example of the Interaction Between Language and Memory,” Journal of Verbal Learning and Verbal Behavior 13, 585–589, 1974. — “The question, ‘About how fast were the cars going when they smashed into each other?’ elicited higher estimates of speed than questions which used the verbs collided, bumped, contacted, or hit in place of smashed. On a retest one week later, those subjects who received the verb smashed were more likely to say ‘yes’ to the question, ‘Did you see any broken glass?’, even though broken glass was not present in the film.” https://www.demenzemedicinagenerale.net/images/mens-sana/AutomobileDestruction.pdf
- Scott Farrell, LeverageAI. “The Perturbation Network.” — “Do not rely on memory afterward — memory launders inference into fact”; “A card with only the first four fields filled is still useful; a beautiful hypothesis with no observed fact is fiction. Discipline is separation, not completeness theatre”; and the n=1 discipline: “One good sentence can be enough when the canon is dense enough to recognise what it contains — not ‘usually enough.’” https://leverageai.com.au/wp-content/media/articles/207-the-perturbation-network.html
- Australian Transport Safety Bureau. Analysis, Causality and Proof in Safety Investigations, ATSB Transport Safety Research Report AR-2007-053, 2008. — “Associated with the concept of standard of proof is the concept of standard of evidence, or the quantity or quality of evidence required before a decision maker can be satisfied that the relevant standard of proof has been met”; “The term ‘probably’ was defined as being equivalent to ‘likely’ and meaning more than 66 per cent likelihood”; “the more significant the issue to be determined… the clearer and more persuasive the evidence needed”; “The dilemma is that the most effective findings for safety enhancement are often the most difficult to justify”; “ATSB often encounters attitudes such as ‘why is the ATSB looking at that, it had nothing to do with the accident’”; “most organisations that conduct safety investigations do not clearly specify the standard of proof (or standard of evidence) they use.” https://www.atsb.gov.au/sites/default/files/media/27767/ar2007053.pdf
- Ivar Ehler, Felix Wolter and Justus Junkermann. “Sensitive Questions in Surveys: A Comprehensive Meta-Analysis of Experimental Survey Studies on the Performance of the Item Count Technique,” Public Opinion Quarterly 85(1), 6–27, 2021. — “In total, 89 research articles with 124 distinct samples and 303 effect estimates are analyzed”; “The results show (1) a significantly positive pooled effect of ICT on the validity of survey responses compared with DQ; (2) a pronounced heterogeneity in study results.” https://kops.uni-konstanz.de/bitstreams/44bdfb71-b3b3-4bd2-9cb0-3f97180303bf/download
- Scott Farrell, LeverageAI. “Institutional Failure Radar” (the Disagreement Instrument). — “Disagreement between formal importance and lived attention is diagnostically more valuable than either ranking alone. The off-diagonal cells are where the radar should hunt”; “Output is a cell classification plus a nomination, not a single organisational risk score. Scores invite Goodhart. Classifications invite investigation.” https://leverageai.com.au/wp-content/media/articles/138-institutional-failure-radar.html
- Scott Farrell, LeverageAI. “Institutional Failure Radar” (weak signals, strong nominations). — “Chatter should be a prior, never a verdict. The signal must perturb, not command”; “high chatter does not equal high risk / low chatter does not equal safety / sentiment does not equal truth / disagreement does not equal dysfunction / consensus does not equal correctness”; “The system nominates; people decide.” https://leverageai.com.au/wp-content/media/articles/138-institutional-failure-radar.html
- National Transportation Safety Board. NTSB Investigator’s Manual, Volume III — Regional Investigations (undated compendium, archived at Embry-Riddle Aeronautical University library; not current NTSB policy). — “Long delays between the witnesses’ observations and the interviews increase inaccuracies in their statements”; “Witnesses should be urged to relate only their personal observations and refrain from passing on information they have derived from some other source”; “the questioning of witnesses should not be conducted as an ‘interrogation.’ The interview should be conducted on a basis of courtesy, cooperation, and neutrality.” http://libraryonline.erau.edu/online-full-text/books-online/1181.3.pdf
- Scott Farrell, LeverageAI. “Attention Flight Recorder.” — Field 7, later outcome (backfilled): “The field most systems never implement… Outcome is not available at decision time. It must be a first-class nullable field that later process fills… Without backfill, review is only aesthetics; with it, negative cognition becomes a supervised problem”; “A decision without versions cannot be replayed.” https://leverageai.com.au/wp-content/media/articles/200-attention-flight-recorder.html
- Oliver Wyman Forum. The CEO Agenda 2026. — CEOs now devote half of all planning effort to horizons of less than one year, up from 43% the previous year; and the accompanying caution that the CEOs taking the longest view see the most opportunity, so compressed horizons may come at the cost of strategic clarity. https://www.oliverwymanforum.com/ceo-agenda/how-ceos-navigate-geopolitics-trade-technology-people.html
- Scott Farrell, LeverageAI. “The FDE as Paid Product Discovery” (the five dispositions and the worked ledger). — “Every candidate ends in exactly one of these final dispositions. Ambiguous ‘maybe later’ is not a disposition; it is a deferred decision with a revisit trigger”; “Counterfactual: would the next deployment actually be easier or safer because this was promoted?… If any of those are missing, you do not have a supported platform capability. You have hope with a release tag”; “preserve the reject. Do not vanish it as a deleted backlog ticket”; “That is not product discovery. That is importing one client’s interface politics into everyone else’s operating system.” https://leverageai.com.au/wp-content/media/articles/172-the-fde-as-paid-product-discovery.html
- Scott Farrell, LeverageAI. “Born Structured Exhaust” (the canon promotion path). — “Most sessions stay project history. That is correct”; “Nomination is not publication. Publication is not automatic”; “Parking is a success mode: not every recurrence deserves institutional blessing”; “Promotion is not a way to launder politics into timeless truth.” https://leverageai.com.au/wp-content/media/articles/191-born-structured-exhaust.html
- Scott Farrell, LeverageAI. “Three Clocks of a Learning System.” — “A gold write is justified when understanding changes in a way that should affect future judgement”; “If the author of a gold mutation cannot state what understanding changed, the write belongs on a case trajectory or nowhere”; “A bad gold write teaches the wrong lesson with compounding authority.” https://leverageai.com.au/wp-content/media/articles/194-three-clocks-of-a-learning-system.html
- Scott Farrell, LeverageAI. “Executable Worldview” (outcome closure). — The four blocks — proposal, execution, observation, learning; “Without those fields you get folklore: ‘that email worked,’ ‘that vendor is difficult,’ ‘that playbook is outdated’ — none of it addressable, none of it replayable, none of it safe to promote into canon”; “The healthiest living wiki is not one with no disagreement. It is one where disagreement has an address.” https://leverageai.com.au/wp-content/media/articles/159-executable-worldview.html
- Scott Farrell, LeverageAI. “The FDE as Paid Product Discovery” (the field-pattern ledger). — “If the only artefact is ‘we built a connector for Client X,’ you have a services receipt. You do not have product discovery”; the ledger fields including “confession of what was not verified”; “That is the field-pattern ledger. Not a mood board. A decision record.” https://leverageai.com.au/wp-content/media/articles/172-the-fde-as-paid-product-discovery.html
- “Investigators are human too. Outcome bias and perceptions of individual culpability in patient safety incident investigations,” BMJ Quality & Safety, published online 10 February 2025 (accepted manuscript via the University of Cambridge repository). — “212 participants completed the online survey. Worsening patient outcome was associated with increased judgements of staff responsibility for causing the incident as well as greater motivation to investigate. More participants selected punitive recommendations when patient outcome was worse”; “Those with patient safety expertise demonstrated these associations but to a lesser extent, when compared to other participants.” https://api.repository.cam.ac.uk/server/api/core/bitstreams/1c9c6937-dc2c-4c2e-88d8-b4e85cbb11db/content
- Scott Farrell, LeverageAI. “Chairman, Not Judge” (the one guard). — “discipline is a personal virtue and personal virtues fail under pressure — especially the pressure of an idea you’re already in love with”; “And it can’t be a token seat. It has to actually win sometimes. If the kill-path never wins, it isn’t a guard; it’s set dressing”; “you haven’t been chairing a decision process. You’ve been running a ratification ceremony.” https://leverageai.com.au/wp-content/media/articles/104-chairman-not-judge.html
- Scott Farrell, LeverageAI. “Separation of Powers for Cognition.” — The architecture that separates epistemic access from causal authority: sensor sees, model reasons, human authorises, compiler executes. https://leverageai.com.au/wp-content/media/articles/203-separation-of-powers-for-cognition.html
- Scott Farrell, LeverageAI. “The Engagement Auditor Is Not the Janitor.” — The structural separation between the machinery that tidies a knowledge graph and the machinery that certifies its claims; audit runs read-only, findings stay on a separate plane, and humans dispose. https://leverageai.com.au/wp-content/media/articles/174-the-engagement-auditor-is-not-the-janitor.html
- Scott Farrell, LeverageAI. “The Moat Is the Memory” (three guards against self-poisoning). — “Reserve a fixed budget for high-significance items with zero wiki intersection. Measure how often those later become canon. If the answer is never, the lane is wrong — not the idea of the lane”; “Learning only from what you kept is how echo chambers become code”; “If you cannot re-run the decision from stored inputs, you do not have a judgment history — you have a story.” https://leverageai.com.au/wp-content/media/articles/149-the-moat-is-the-memory.html
- Scott Farrell, LeverageAI. “Signal Case Queue” (case boundaries through time). — “Diffing against your canon is the power and the trap. A dense map can become a suppression shield around unfamiliar fields”; “Boundary quality is a first-class benchmark — ahead of summarisation quality.” https://leverageai.com.au/wp-content/media/articles/143-signal-case-queue.html
- Scott Farrell, LeverageAI. “Chairman, Not Judge” (the two things that can’t move). — “A panel is an evaluation function: it scores options against a criterion. But it cannot supply the criterion it is being scored against”; “The panel can carry the analysis; it cannot carry the risk”; “It’s not a gavel at the end. It’s a star overhead and a name on the line.” https://leverageai.com.au/wp-content/media/articles/104-chairman-not-judge.html
- A. Kumah. “Adverse event reporting and patient safety: the role of a just culture,” Frontiers in Health Services, 2025. — “Fear of punishment or blame often discourages healthcare professionals from reporting errors and near-misses, leading to missed opportunities for improvement”; just culture as “A culture that emphasises accountability and learning over punitive measures.” https://pmc.ncbi.nlm.nih.gov/articles/PMC12405316/
- SKYbrary Aviation Safety. “Just Culture.” — “an atmosphere of trust in which people are encouraged — even rewarded — for providing essential safety-related information.” (The formulation originates with James Reason, Managing the Risks of Organizational Accidents, 1997; quoted here from SKYbrary as the accessible secondary source, since the book was not consulted directly.) https://skybrary.aero/articles/just-culture
- Kolmogorov Law / Pollfish. Survey of 500 employed U.S. adults, fielded 8 July 2026, margin of error ±4.4% at 95% confidence; distributed via Stacker. — “One in 3 employed Americans (33.4%) say an AI notetaker or transcription bot has been present in their work meetings”; “only 34.7% say they were always asked for permission.” (A commissioned single-vendor panel survey, not peer-reviewed research.) https://abc17news.com/stacker-business-economy/2026/07/17/ai-notetakers-have-sat-in-on-1-in-3-us-workers-meetings-but-only-a-third-say-they-were-asked-first/
- Scott Farrell, LeverageAI. “The FDE as Paid Product Discovery” (gravel road and paved road). — “Automatic upward movement is not a feature; it is a breach of the architecture”; the field produces “domain priors” and “process priors” at once. https://leverageai.com.au/wp-content/media/articles/172-the-fde-as-paid-product-discovery.html
- Scott I. Tannenbaum and Christopher P. Cerasoli. “Do team and individual debriefs enhance performance? A meta-analysis,” Human Factors 55(1), 231–245, 2013 (46 samples, N = 2,136; overall d = .67). — “organizations can improve individual and team performance by approximately 20% to 25% by using properly conducted debriefs.” https://pubmed.ncbi.nlm.nih.gov/23516804/
- Association of National Advertisers and American Association of Advertising Agencies. “New ANA and 4As Report Reveals Client-Agency Relationship Tenure Has Doubled Since 2016,” 30 April 2025. — “Clients without mandatory review periods (60% of respondents) have significantly longer relationships (8.1 years) than those with frequent reviews (as low as 3.8 years).” https://www.ana.net/content/show/id/pr-2025-04-tenure
- UK National Audit Office. “Government lacks a clear picture on how much it spends on consultants,” 21 November 2025. — “Consultants should only be used where they represent best value for money and not to replace capability required inside the civil service”; “ensuring that civil servants learn from consultants while they are working together by building knowledge transfer agreements into contracts.” https://www.nao.org.uk/press-releases/government-lacks-a-clear-picture-on-how-much-it-spends-on-consultants/
- Scott Farrell, LeverageAI. “The Evolution Mandate.” — “Transfer the known; earn renewal at the frontier”; the dependency gradient and its inversions, including “I have become the escalation desk of a larger firm”; “direction and comparison across cycles, never fabricated percentages. Shape, not invented magnitudes. And if a firm cannot yet measure a row at all, that is itself the finding.” https://leverageai.com.au/wp-content/media/articles/234-the-evolution-mandate.html
- Scott Farrell, LeverageAI. “Fog Is a Race Between Two Clocks.” — “Falsification throughput: the rate at which your firm eliminates strategic options, each elimination carrying the named evidence that killed it. Options closed per quarter. Not candidates generated. Not initiatives launched. Not decisions minuted. Eliminations, with receipts”; “Activity does not move this clock. Evidence moves it.” https://leverageai.com.au/wp-content/media/articles/232-fog-is-a-race-between-two-clocks.html
- Scott Farrell, LeverageAI. “Forward-Deployed Practice OS.” — “What success looks like mid-install… That is progress. It is still not proof of practice transfer. Chapter 6 is the only proof that matters: engagement two.” https://leverageai.com.au/wp-content/media/articles/167-forward-deployed-practice-os.html
- Scott Farrell, LeverageAI. “Experience Is Compressed Priors” (the compounding test). — “Did we leave merely another codebase — or a sharper dual-stream corpus that makes the next project easier and better?”; “A loop that piles findings into chat history is linear activity. A loop whose outputs become structured substrate the next loop reads as priors is compounding.” https://leverageai.com.au/wp-content/media/articles/150-experience-is-compressed-priors.html
- Scott Farrell, LeverageAI. “Attention Flight Recorder” (counterfactual calibration under freeze). — “The protection that makes this science rather than cosplay is the Future-Leakage Rule: do not replay the final period against a wiki that already knows what happened during that period. Otherwise you are not testing discovery. You are testing whether the system can read a spoiler”; “Future leakage is the silent invalidation mode for historical replay — name it, forbid it”; “If any answer is ‘the model just knew,’ you do not have replay. You have a demo.” https://leverageai.com.au/wp-content/media/articles/200-attention-flight-recorder.html
- Scott Farrell, LeverageAI. “Attention Flight Recorder” (four review buckets — the suppression audit specified). — “Near-miss review is boundary sampling… It does not reveal news that sat miles below the threshold because the worldview lacked the vocabulary to recognise it”; “A case can be suppressed not because evidence was weak, but because the worldview had no place to hang the significance — ‘already known’ misread as ‘not consequential’… Those failures do not appear as near misses. They appear as silence with a clean conscience”; “Learning only from what the radar selected is how an echo chamber becomes code.” https://leverageai.com.au/wp-content/media/articles/200-attention-flight-recorder.html
- Scott Farrell, LeverageAI. “Engagement World” (Seed to Close). — “Close is not a calendar event. It is a governed procedure that preserves evidence, transfers ownership, gates promotion, and expires the volatile”; “If the client cannot operate the world without your private kernel as a hidden dependency, you did not deliver a world. You delivered a leash”; “Approve and reject are both success outcomes of a working gate. A queue that only ever approves is not a gate; it is a conveyor.” https://leverageai.com.au/wp-content/media/articles/170-engagement-world.html
A note on numbers. Several claims in this piece have no number attached, and that is deliberate. There is no credible measurement of how often post-project retrospectives change organisational behaviour, no verified adoption figure for AI project-harvesting in professional services, and no public data on how many candidate lessons survive an unrelated transfer test. Where a passage wanted a figure the sources do not have, it states the shape of the claim instead. The falsification programme in section 12 exists precisely because those measurements do not yet exist.
