Leverage AI
LeverageAI Β· Assurance & Governance

Elastic Assurance

Compute Broadly, Disclose Narrowly

Governance dashboards are lossy compressions, sized to scarce management attention. A second, AI-native assurance plane lets you examine everything β€” and disclose only what deserves human judgment.

A worked example on the public record of Transgrid and the Australian Energy Regulator.

You will leave able to:

By Scott Farrell Β· LeverageAI

01
Part I β€” Why the Dashboard Lies by Omission

The Green Light Is Not Reality

A board pack can be entirely, defensibly green and still miss the one thing that mattered β€” not because anyone lied, but because the page had nowhere to put it.

Picture the pack that lands before a board or an audit committee on a Tuesday. Every mandatory KPI is met. Every control is attested. The status is green, the commentary box is calm, and nothing has crossed a threshold that would force an escalation. On the evidence in front of them, the directors are right to be reassured. And yet, months later, the thing that actually mattered turns out to have been happening the whole time β€” visible to people on the floor, invisible on the page. Nobody concealed it. The page simply had no column for it.

This is not a failure of honesty. It is a failure of resolution.

The green light is not reality. It is a deliberately low-resolution representation of reality, designed around the amount of attention available at the management layer.

The governance compression problem

A large organisation contains thousands of local facts, judgements, exceptions, uncertainties and trade-offs. Management cannot read them all β€” nobody can β€” so governance compresses. It squeezes the whole high-dimensional operating reality down into a handful of artefacts that fit inside a scarce attention budget:

  • a small number of KPIs
  • a few control attestations
  • an amber, green or red status
  • a short commentary box
  • an escalation, but only when a predefined threshold is crossed

That compression is necessary. It is also lossy. And the loss is not random noise β€” it is precisely the nuance that did not fit the chosen categories.

Green means the area passed the limited tests chosen for escalation. It does not mean nothing else interesting, risky or consequential is happening.

Key Insight

Assurance is a claim about what you examined β€” not a claim that something passed a predefined test. A dashboard confuses the two.

So what does green actually certify?

Consider a single process. It can be genuinely green on timeliness, completion, budget and its known safety controls β€” every box the reporting model knows how to tick. Meanwhile, the soft-data exhaust around that same process β€” the reports, the correspondence, the meeting records, the operating chatter β€” can be carrying whole dimensions the model has no field for.

What the report measures β€” and what the exhaust carries

The formal projection
  • β€’ Timeliness
  • β€’ Completion
  • β€’ Budget
  • β€’ Known, named safety controls
The soft-data exhaust
  • β€’ Implementation differs materially between two sites
  • β€’ The same control consumes radically different human effort
  • β€’ Field staff quietly qualify a number the report states with confidence
  • β€’ One site has real independent challenge; another has ritual approval
  • β€’ Workarounds have become normal practice
  • β€’ The team's human reserve margin is collapsing
  • β€’ A risk category is emerging that governance has no name for

None of these is captured by asking whether the control is green. Each of them is a legitimate assurance signal. And each is the kind of thing that, after a failure, everyone agrees "was obvious in hindsight" β€” because it was there in the exhaust all along, just never routed anywhere a decision-maker would see it.

The specific shapes that soft data takes β€” how organisational divergence looks under static analysis, how behavioural failure announces itself, how a collapsing reserve margin shows up in the numbers β€” are their own disciplines, and they are the subject of the companion pieces in this series. This book is not about the sensors. It is about the plane those sensors report into.

Too narrow to be the whole truth

It is tempting to frame all of this as deception β€” a green light that is secretly a lie. Usually it is not. The problem is more ordinary and more dangerous than dishonesty.

The green light isn't necessarily false. It is too narrow to be the whole truth β€” too rolled up to carry the nuance. The organisation is not hiding these facts. Its reporting structure simply has nowhere to put them, so they never enter the record that governance actually reads.

That distinction matters, because it tells you where the fix lives. If the problem were dishonesty, you would reach for more controls, more attestations, more audit. But the problem is compression, and you cannot fix a compression problem by demanding a more honest compression. You fix it by changing the order of operations β€” by examining the estate before you throw most of it away.

The belief this book fights

That a green dashboard is assurance. It isn't. It is a receipt that a small, fixed set of questions was answered acceptably β€” nothing more. The belief persists because attention has always been scarce at every level, the traffic light is genuinely useful, and until very recently no organisation could afford the alternative.

That last clause is the whole opening. For as long as governance has existed, it has been forced to compress first and review second, because humans cannot inspect an organisation at full resolution and there was no economical machine that could do it for them. That constraint has just changed. The next chapter is about what becomes possible when you can finally afford to examine everything the dashboard threw away β€” without drowning the people at the top.

02
Part I β€” Why the Dashboard Lies by Omission

Compute Broadly, Disclose Narrowly

There is an order-of-operations bug in governance. Fix it, and you can examine everything while still handing management a compact surface.

Traditional governance has to compress first and review second. The sequence looks reasonable until you write it down:

The old order of operations

complex reality  β†’  selected metrics  β†’  traffic light  β†’  management attention

The nuance is discarded at step two β€” before anybody has examined it.

That order was forced by a hard constraint: humans cannot inspect the underlying estate at full resolution, so the estate must be reduced to something a person can hold before any reviewing happens. The consequence is that the most important review β€” the reading of the raw material β€” is the one step that never occurs. We compress on faith and review the compression.

The compression inversion

Abundant machine cognition lets you invert the order:

The inverted order of operations

complex reality  β†’  broad AI review across many axes  β†’  selective findings  β†’  management attention

The estate is examined first. Only then is it compressed.

The system can read at high resolution and report at low resolution. That is the entire move, and it is easy to misread, so let me be exact about what it is not. It is not a proposal to flood executives with more detail. Human attention is exactly as scarce as it was. What changes is that the organisation no longer throws most of its nuance away before anyone has looked at it. Management still receives a compact representation β€” but that representation is produced after the soft-data estate has been examined, not instead of examining it. It is progressive resolution applied to assurance: coarse where a human looks, fine where the machine reads.

Compute broadly. Disclose narrowly. Drill down on demand.

Two assurance planes

The instinct, when people hear "AI review," is to imagine a machine that grades the organisation and overrules the humans. That is precisely what this is not. You keep the current governance process exactly as it is, and you run a second plane beside it.

The two planes

Formal Assurance Plane
  • β€’ Mandatory KPIs and approved procedures
  • β€’ Regulatory reports and control attestations
  • β€’ Formal green / amber / red status
  • β€’ Recognised escalation paths
  • β€’ Keeps its regulatory meaning β€” unchanged
Exploratory Assurance Plane
  • β€’ Reads reports, correspondence, meeting records, chatter
  • β€’ Compares implementation across teams and sites
  • β€’ Hunts contradictions, unusual attention, known absences
  • β€’ Tests alternative interpretations; names categories the formal model lacks
  • β€’ Reports findings with evidence β€” without changing formal status

The exploratory plane cannot move the traffic light. It cannot stop operations, approve expenditure or alter a control. What it can do is produce a sentence the old system could never say:

A result that used to be impossible

Formal status: Green. Independent findings: three items warranting management consideration.

That is a far better outcome than forcing every weak signal into amber. Amber has a cost β€” it triggers process, attention and often blame β€” so a governance system that can only express concern by turning amber will systematically suppress weak signals. Two planes let a weak signal travel as a finding rather than a status change. The regulatory meaning of the traffic light stays intact, and the organisation gets to admit it has discovered something outside its own projection.

The Assurance Finding Card

Management does not need another giant dashboard; that would just be the compression problem again, wearing a machine's clothes. It needs progressively disclosed findings. The primitive is a small card.

Field What it carries
FindingWhat unusual shape, divergence or omission was identified
SignificanceWhy it may matter despite the formal status
EvidenceReceipts from the soft-data estate
UncertaintyWhat is known, inferred, or still missing
QuestionWhat management should ask next
ResponseReview, resourcing, procedural change, monitoring β€” or no action

A worked one makes the tone concrete. Suppose the same network-security control is formally green at two sites.

Finding card β€” one control, two sites

Formal control: Green at both sites.

Finding: The same control is implemented very differently at Sites A and B.

Evidence: Site A shows broad technical review, independent challenge and stable after-hours activity. Site B shows roughly three times the exception chatter, repeated reliance on one specialist, shorter review comments, and a growing concentration of approvals near shift handover.

Interpretation: This does not establish non-compliance. It indicates the two green statuses may not represent equivalent assurance depth or sustainable operating capacity.

Question: What local conditions explain the divergence, and does Site B need additional resources or procedural support?

Notice what the card refuses to do. It does not accuse. It does not claim the AI has proven wrongdoing. It surfaces a shape, attaches the receipts, states its own uncertainty, and hands the judgement back to a human. That discipline β€” findings that carry evidence and confess their limits β€” is what makes the whole architecture safe, and it is the subject of the next chapter.

Why call it elastic?

Definition β€” Elastic Assurance

A massively parallel soft-data review layer that examines organisational activity across many more dimensions than formal governance can carry, then routes only evidence-backed, non-routine findings to human attention.

Human assurance works through scarce teams. Internal audit samples. Management picks a few KPIs. Specialists investigate known risks. Committees focus on what is already escalated. Sites get compared only when somebody commissions a review. Every one of those is rate-limited by people.

Machine assurance fans out. It can run across every site, every control family, every relevant reporting period, many independent questions, many comparison axes, historical and current implementations, formal and informal artefacts β€” at the same time. It is not literally unlimited; it is bounded by budget, access, model quality, security and evaluation. But set beside human review capacity, it is elastic: once the ingestion and governance plumbing exists, adding another site, another lens or another standing question is an incremental compute cost, not another multi-month audit programme.

So the bottleneck moves. It stops being the question that has governed assurance forever β€”

"How many things can we afford to examine?"

β€” and becomes a fundamentally better question:

Takeaway

Which findings deserve scarce human attention, and which questions should the institution keep asking?

Why now, and not five years ago?

Because the economics only just tipped. The value here is not a cheaper version of work you already do β€” it is doing the thing that was previously uneconomic: reading everything before you throw most of it away. When the cost of cognition was high, exhaustive review of an organisation's entire soft estate was a fantasy priced in consultant-years. As the cost of cognition falls, additional review that is high-volume, parallelisable and time-flexible approaches a marginal compute cost once the platform exists. A re-review that once meant standing up an audit programme becomes a job you can run across many axes in a short window, repeatedly.

The honest version of the cost claim states the shape rather than inventing a figure: a headcount-bound audit scales with people and calendar; a compute-bound review scales with budget and the falling price of tokens. The exact ratio depends on your estate. The direction does not.

That inversion is real, and it is affordable. But an un-governed AI review layer is also genuinely dangerous β€” it is one bad incentive away from becoming a machine that manufactures accusations or quietly overrules governance. What stops that is a set of boundaries, and they are strict enough to deserve their own chapter.

03
Part I β€” Why the Dashboard Lies by Omission

Un-invested, Not Objective

A second review layer is only worth building if it is safe by construction. Six boundaries turn a dangerous idea into a deployable one.

Start with an uncomfortable truth about green. Staff and management are not, as a rule, dishonest. But green has organisational value, and that value bends behaviour.

What green buys the people who report it

  • β€’ Green avoids escalation, and the executive attention that comes with it
  • β€’ Green protects delivery schedules, budgets and performance narratives
  • β€’ Green spares a team already under pressure from receiving even more scrutiny
  • β€’ Amber can read as personal or managerial failure β€” not as a request for help

So people can rationally work harder, compress uncertainty, shorten reports and absorb problems locally in order to keep the status green. The formal control ends up measuring one thing β€”

"Was the obligation met?"

β€” while quietly failing to measure the thing that actually predicts whether it will still be met next quarter:

"Was the obligation met through the designed operating model, or through unsustainable human compensation?"

Is an un-invested reviewer trustworthy?

An AI reviewer sits in a different incentive position. It receives no bonus for keeping a project green. It does not dread an uncomfortable steering committee. It is not defending a decision it made last year. That is a real and useful difference β€” but it is tempting to overclaim it, so here is the correction that keeps the whole architecture defensible.

The AI is not objective. It is un-invested.

The machine is biased β€” toward what was recorded, toward what it can access, toward the questions it was told to ask, toward the lens it was given. The difference from human bias is not that the machine has none. It is that the machine's biases can be disclosed and changed, while human investment in a narrative is often invisible even to the person holding it. An un-invested reviewer with declared, adjustable bias is a better instrument than an "objective" one whose bias you cannot see.

This is why a single review should carry multiple declared lenses, and why they should be allowed to disagree with each other inside the same package: a delivery lens, an engineering-risk lens, a workforce-capacity lens, a regulatory-prudency lens, a landholder-and-community lens, a contractor-behaviour lens, a "what would the regulator challenge?" lens, a "what would a future incident inquiry ask?" lens. Disagreement between lenses is not a defect to resolve. It is the most honest thing the review produces.

The six boundaries

Six rules make this both radical and deployable. Together they are the constitution of the exploratory plane.

Boundary What it forbids, and why
Un-invested, not objectiveAlways disclose the lens. The machine has no stake β€” not perfect sight.
Finding, not verdictIt surfaces a shape to consider. It never asserts proven wrongdoing.
Derived, not primary evidenceThe review ranks below its exhibits and cannot cite itself as proof.
Soft-first, hard-on-demandSoft data raises the question; hard data is queried only to test it.
Parallel, not replacementThe formal plane still decides. The exploratory plane only informs.
Batch, never live authorityOff the operational hot path β€” no switching, protection or real-time control.

Two of these deserve unpacking, because they are where naΓ―ve versions of this idea go wrong.

Derived evidence, and the rule that governs it

The AI review is itself worth keeping. It is evidence that a particular question was asked, that a particular evidence estate was examined, that certain patterns were visible at that time, that particular contradictions or gaps were identified, and that management accepted, rejected or ignored the finding. Retrospectively, that can be decisive: a future review might establish that an issue first surfaced in the March build, stayed below escalation thresholds, reappeared in June, and was formally investigated in September. That is valuable institutional history.

But the review must remain derived evidence. It cannot cite an earlier AI conclusion as proof that a current conclusion is true; it must always retain links back to the underlying emails, reports, decisions and system records.

Key Insight

Record the review, but never let the review outrank its exhibits.

This is the evidence-package discipline we have described elsewhere as being a witness rather than an oracle: every claim carries its verbatim exhibit, a resolvable pointer, and a confession of what could not be verified. Checkable is not the same as truthful β€” and the package is built to make the difference visible, not to hide it.

Soft-first, hard-on-demand

The review begins in the soft layer because that is where the emerging shape appears first β€” as unease in engineering correspondence, repeated qualification around one assumption, an unusual amount of coordination, unresolved disagreement, shortened review notes, repeated temporary exceptions, different interpretations across sites, staff describing overload, or confidence rising without matching evidence. Only once the soft layer has raised a question does the AI interrogate the hard systems selectively: did overtime increase, did variation volume rise, did approval latency change, are complaints clustering, did the number of independent reviewers fall, did the status change after the soft signal appeared?

Soft data generates the question. Hard data helps test it.

The warehouse stays authoritative for figures. The compiled soft layer holds meaning and pointers. The AI crosses between them only when the investigation requires it β€” never the other way around.

Batch the review; let existing governance decide

This is a near-perfect deployment of what we call the Lane Doctrine β€” the discipline of placing AI where its operating constraints are favourable, batching the thinking outside the hot path, and shipping a reviewable artefact rather than taking live authority. On one side of the line, the AI can build the ingestion and analysis software, compile the soft layer, generate and run standing questions, conduct retrospective project reviews, compare sites, produce weekly review builds, assemble evidence packages, and test new lenses in parallel. On the other side it must stay out entirely: real-time grid control, protection actions, live contingency response, automatic formal status changes, autonomous regulatory attestations, binding project or workforce decisions.

The deployment pattern

Batch the review. Ship the package. Use existing governance to decide.

There is a reason this mirrors why AI coding already works. A model can generate large volumes of code safely because the output passes through existing review, testing, versioning and approval infrastructure before it means anything. The Soft Attestation Package earns an organisation the same governance arbitrage for institutional review: the AI produces a reviewable artefact, and existing risk, engineering, audit and executive structures remain the only authorities that can act on it.

The evidence room, not the prosecution

Contrast this with how retrospective review works today. It means creating a review team, negotiating scope, locating years of records, interviewing people who know the findings may affect them, reconstructing chronology by hand, and having lawyers and executives manage the language while participants defend their historical decisions β€” producing a polished report months later. As much of that process is about covering yourself as about learning.

Abundant machine cognition changes the economics. The AI can read the complete accessible record repeatedly, reconstruct different chronologies, apply multiple lenses, and examine every project rather than only the spectacular failures. It is not perfectly objective β€” nothing is β€” but it is patient, scalable, repeatable, comparatively un-invested, able to retain contradictory accounts, able to re-run when new evidence appears, and able to produce exhibits rather than lean on retrospective testimony.

The AI does not prosecute the project team. It builds the evidence room.

Humans decide what the evidence means. That is the posture, and those are the boundaries. With them in place, we can watch the whole architecture work on a real, public, thoroughly contested case: a transmission monopoly whose entire business is arguing why to its regulator.

04
Part II β€” The Transgrid Wedge

The Two-Project Wedge

A worked example, read entirely from the public record of Transgrid and the Australian Energy Regulator. No engagement is implied β€” this is an outside-in reading of documents anyone can download.

Transgrid runs the high-voltage transmission network of New South Wales β€” roughly 11,500 kilometres of line and 136 substations and switching stations,1 and it already uses digital twins, AI and drones to manage them. On the public record it is doing at least six hard things at once: finishing a badly over-budget megaproject, building another under intense social-licence scrutiny, redesigning the technical foundations of the grid as coal retires, fielding extraordinary new data-centre demand, operating ageing assets under climate and supply-chain pressure, and asking regulators and consumers to fund several overlapping investment programs. It is close to a perfect environment for a second assurance plane.

The infrastructure may be physically complete, but the argument over why it cost what it cost β€” and who should pay β€” is only beginning.

The usual AI conversation in a business like this is drones for line inspection and copilots for document processing β€” the eighteen-project deck, all traffic lights, zero terminal value. That is horse optimisation: making the existing motion a little faster. The higher-value move does not begin with an enterprise platform. It begins with a two-project wedge.

Pilot 1 β€” a completed but contested project as a labelled corpus

Project EnergyConnect is the ideal starting corpus precisely because it is finished and disputed. On the regulator's like-for-like basis, Transgrid is seeking an additional $1.142 billion in 2022–23 dollars, on top of the previously approved $2.121 billion for the NSW component β€” about $3.263 billion in real terms β€” a request the AER says would add roughly $173 million to revenue recovered in 2027–28 and about $18 to a residential bill that year.2 Meanwhile the physical project has raced ahead of the argument: Transgrid announced in June 2026 that its 700-kilometre NSW section was complete and that Stage 2 was energising ahead of network testing.3

The regulatory question is no longer "did the project exceed its budget?" It is a chain of causal questions: which events were genuinely unforeseeable; which costs arose from contractor failure; which risks were contractually transferred; which should shareholders have carried; which management decisions were reasonable at the time; when the likely cost trajectory first became visible internally; and whether formal green reporting lagged the operational exhaust. What did Transgrid know, when, and through which evidence?

The warehouse holds the blowout. The why lives in decades of engineering decisions, variation records, consultation threads and landholder negotiations.

The product this actually calls for is not a faster response-drafting copilot. It is a causal record of the project β€” contract decisions, variations, claims, risk registers, engineering judgements, contractor correspondence, weather disruptions, schedule changes, executive and board decisions and regulatory representations, joined chronologically and causally, with receipts. Use the completed, contested project as a labelled historical corpus and ask a single question: what was the emerging shape of the overrun, when did it become visible, and which organisational signals preceded formal recognition? The output is a reference library of how project risk actually manifests in this organisation's own exhaust β€” causal patterns, warning shapes, known blind spots, standing questions and evidence expectations. It does not decide the regulator's dispute. It gives the organisation a more honest, inspectable, internally challengeable account of its own history.

Pilot 2 β€” the same questions on a live, still-moving programme

Then apply the learned questions prospectively, to a programme still in motion. Transgrid's system-strength preferred portfolio was estimated at about $6.3 billion; in April 2026 it formally notified the AER of a material change because synchronous-condenser procurement costs had risen by more than 30%.4 By July it had published a revised portfolio giving grid-forming batteries a larger role, as condensers became slower and dearer, coal-retirement timing shifted and flexibility gained value.5 Separately, it is seeking approval of a $1.185 billion System Strength Project for 2026–31, where the regulator's preliminary assessment has isolated risk costs and labour and indirect costs as its two focus areas β€” the first hybrid revenue determination combining contestable and monopoly-delivered components under this framework.6

This is not a stable process waiting to be automated. It is a living decision system in which technology credibility, prices, retirement dates, procurement outcomes and regulatory frameworks all move β€” and in which the assumptions that justified yesterday's portfolio may be obsolete tomorrow. A soft-data layer could maintain an assumption dependency graph: which conclusions depend on which equipment prices, coal-retirement dates, battery capabilities, delivery schedules and regulatory interpretations. Then every external change triggers standing questions β€” which previous conclusions no longer hold, which rejected alternatives should be reopened, which consultation statements are now stale, which current decisions still cite superseded assumptions. Today's material-change review runs when a formal trigger is crossed; this makes it continuous.

The world loop

Use the completed failure to create the questions; use the live project to test whether the questions prevent recurrence.

Where do the findings actually live?

HumeLink shows the answer: in the joins between reporting categories. The AER approved $3.965 billion in Stage 2 capital expenditure β€” cutting $314.4 million from the original application β€” and explicitly tied its decision not only to consumer cost but to landholders, local communities and Transgrid's obligation to build and maintain social licence throughout delivery.7 The formal structure reports cost, schedule, safety, environment, land access, community engagement, workforce and construction quality as separate lanes. The emerging failure lives between them: are repeated landholder complaints later associated with access delays or expensive design changes; do contractor workarounds keep creating environmental or community issues; does schedule pressure make consultation language more certain while the underlying issue stays unresolved; are commitments made in engagement channels represented consistently in delivery instructions; are two delivery partners interpreting equivalent obligations differently? No ordinary dashboard owns those questions, because each one crosses an organisational and contractual boundary.

Their own admission

The strongest validation is first-party. Transgrid's System Security Roadmap operational-technology case says the transforming grid is causing a substantial increase in information and analysis requirements, and its public material warns that without better capability, operators may need to run the grid more conservatively and constrain low-cost renewables more often, may become overburdened having to access and confirm information across multiple sources, may be less able to act within the available time during a contingency, and may increase expected unserved energy. It is seeking $163.5 million for OT upgrades to help.8

That is an explicit statement of the management-attention constraint this whole architecture addresses. The obvious response β€” better screens, integrated tools, improved analytics β€” is useful and still horse optimisation. The larger question is the one the OT project cannot ask about itself:

What important review and institutional vigilance cannot occur at all, because operators and planners must spend their attention keeping the live system running?

The OT project improves what operators can see while operating. A soft-data layer, sitting off the hot path, examines what the organisation is failing to notice about how it operates β€” the two are complementary, not competing.

Demand credibility, not connection processing

Nearby sits a different soft-data problem. Since late 2024, Transgrid has received data-centre connection enquiries totalling 14 gigawatts within 12 kilometres of Sydney West β€” roughly the whole of NSW's peak winter load concentrated in one area β€” while Western Sydney's transmission capacity is largely exhausted, and it says households and ordinary businesses should not bear the associated cost or risk.9

One number, two very different questions

The hard-data question
  • β€’ How many gigawatts have applied?
The soft-data questions
  • β€’ Which enquiries are speculative land-banking versus credible projects?
  • β€’ Which proponents actually have capital, land, approvals and realistic plans?
  • β€’ Which applications quietly rely on the same scarce capacity or the same hidden assumptions?
  • β€’ Where does executive-forecast certainty exceed the conditional language in engineering analysis?

The value of answering the second column is measured in billions of avoided augmentation built for a queue whose semantic quality is invisible in its rows. This is a demand-credibility engine, not a connection-processing tool.

Pull the wedge together and the commercial logic is plain. A transmission monopoly's entire business is arguing why to a regulator, across a widening burden of proof that now spans the EnergyConnect reopener, the system-strength hybrid, the OT project, the 2028–33 revenue determination and the rate-of-return review.

Bottom Line

Regulated entities live and die on the defensible record. Institutional memory and causal trace are a compounding asset β€” not a productivity toy.

So what does one filed review actually contain, and what does it let a board finally ask? That is the artefact at the centre of the whole system, and it has its own chapter.

05
Part II β€” The Transgrid Wedge

The Soft Attestation Package

One filed review β€” its contents, how it is refiled as evidence, and the single question it lets a board finally ask: not whether the project was green, but how green was produced.

Going back through a completed project today is expensive and defensive. It means standing up a review team, negotiating scope, locating years of records, interviewing people who know the findings may affect them, reconstructing chronology by hand, and letting lawyers and executives manage the language while participants protect their historical decisions β€” a polished report emerging months later. The output is real, but the process is as much about covering yourself as about learning. The alternative is not a better postmortem. It is a contemporaneous artefact that already existed at every reporting cycle, before anyone knew the outcome: the Soft Attestation Package.

One technical distinction makes the whole thing safe, and it must be stated before anything else.

The package attests that a defined review occurred against a known evidence state. It does not attest that the AI's conclusion is true.

What is inside a package

A practical package is a fixed, reproducible structure. Another authorised reviewer should be able to open the exhibits, inspect the hard-data queries, and understand exactly why an issue surfaced.

Section What it records
SubjectProject, control, site, reporting period and formal status reviewed
Evidence stateDocuments, correspondence, reports and systems examined β€” with timestamps and source pointers
Review specificationStanding questions, lenses and AI-system versions used
FindingsRanked issues or hypotheses β€” not asserted root causes
ExhibitsVerbatim evidence and resolvable source pointers
Contradictions & absencesConflicting accounts, missing voices, and evidence that should exist but does not
Hard-data checksStructured data queried to validate or falsify a soft-data finding
Counter-caseAlternative interpretations, and what evidence would disprove the finding
Workforce implicationWhether the pattern suggests under-resourcing, concentration risk, shallow review or hidden human subsidy
DispositionAccepted, rejected, investigated, deferred β€” or promoted into a standing question
Seal & versionCorpus snapshot, review version, timestamp, and diff from the previous review

Refile it β€” as typed, derived data

Once the panel has considered a package, it does not disappear into a folder. It is refiled into the project's soft estate as a new, dated, derived record β€” the same write-back move Andrej Karpathy made explicit for knowledge bases in April 2026: useful analyses and discovered connections should be filed back rather than lost to chat history.10 It is filed not as truth, but as a precise, dated claim: this is what an independent AI review could see, from the information available, at this reporting date.

The write-back needs strict epistemic typing, or the system starts believing its own past opinions.

soft_attestation.record
type: soft-attestation
status: derived
project: HumeLink
formal-report-period: 2027-06
observed-as-at: 2027-06-30
formal-status: green
review-status: three-findings
supports: [ source-pointers, report-versions, correspondence, meeting-records, structured-queries ]
supersedes: [ prior-soft-attestation ]
human-disposition: [ accepted | rejected | investigate | monitor ]

Key Insight

Record the review, but never let the review outrank its exhibits β€” the package ranks below primary evidence, cannot cite itself as proof, and goes stale when its supporting evidence changes.

So the review compounds β€” but it never promotes its own earlier opinion into ground truth. Every finding points to exhibits; missing and inaccessible evidence is recorded; later reviews may strengthen, split, reject or supersede it; and the human panel's response is attached and preserved. The project wiki does not eat its own tail.

How was green produced?

Attach a package to every formal report and the governing question changes. It stops being the binary the dashboard already answers, and becomes the question that actually predicts the future.

Not:

Was the project green?

But:

How was green produced?

A green state can have several very different underlying compositions. The taxonomy below is a green-provenance record β€” not an accusation field. It describes shapes, not guilt.

Green formation What the soft record shows
DesignedMet via planned staffing, normal hours, intended controls and adequate independent review. Genuinely healthy green.
CompensatedHeld green through overtime, borrowed staff, key-person heroics, deferred work and compressed review. The obligation was met; the operating design did not support it.
CoercedSchedule or executive pressure; objections quietly disappearing; dissent narrowing after senior intervention. Formal confidence exceeds the confidence visible in the deliberation.
ReclassifiedScope quietly narrowed, thresholds reinterpreted, work moved outside the boundary. The number stayed green, but what it represented moved.
NarrativeCaveats and disagreement in the source; a much cleaner final report. An amber conversation that generated a green summary.

Each mode has a recognisable vignette. Designed green is the control that held steady with stable effort and real challenge in the record. Compensated green is the site staying green on one specialist's after-hours heroics. Coerced green is the amber concern that goes quiet the week after a steering-committee push. Reclassified green is the KPI that stayed green because the definition of "complete" quietly moved. Narrative green is the correspondence full of qualifiers that became a single confident sentence in the board pack. The package's job is to compare the semantic shape of the source deliberation with the shape of the formal report β€” and to say, with receipts, which kind of green this was.

Which clock are we reading?

Because each package is dated and derived, the project record becomes bitemporal β€” it carries two clocks at once.

Two clocks on every reporting period

What appeared true at the time

What the evidence, policies, costs, schedules and operating conditions indicated for that reporting period.

What the organisation knew at the time

What the formal report stated, what the AI review surfaced, what management considered, and what actions were taken.

Separating the two clocks β€” a discipline we develop at length in The Answer Depends on the Date β€” lets a board ask very precise past-tense questions. What did formal governance say in June 2027, and what did the soft attestation identify in the same month? Which evidence available then supported or contradicted each view? Which later information changed our understanding only retrospectively? Was the project's actual condition unknowable, or visible but not escalated? That last distinction is the whole point.

That is the difference between uncertainty and organisational blindness.

And it cuts both ways, which is why honest teams should want this more than anyone.

Takeaway

The system does not merely catch failure. It preserves evidence of reasonable contemporary judgement.

A contemporaneous record can show that a later failure genuinely was not reasonably foreseeable, that a concern was considered and rationally rejected on the evidence then available, that management funded additional capacity the moment the human reserve margin fell, or that today's obvious explanation simply was not obvious at the time. It protects the honest team as reliably as it exposes the blind one β€” which is exactly what you want from an evidence room rather than a prosecution.

The living review loop

Put together, a bounded implementation runs as a cycle rather than a postmortem β€” and it never waits for a project to turn amber or fail:

  1. The formal team produces its existing report.
  2. The report and the reporting-period soft exhaust are frozen as a review snapshot.
  3. The AI runs the standing project questions.
  4. It selectively queries hard systems where a soft finding needs testing.
  5. It produces ranked findings, counter-cases, exhibits and gaps.
  6. A human assurance panel accepts, rejects or redirects each finding.
  7. The package and the panel's disposition are attached to the formal report.
  8. Both are refiled into the project's institutional memory.
  9. The next cycle diffs the formal state, the soft state and the previous findings.
  10. Useful new questions are promoted into standing review lanes.

Two of those steps β€” the panel, and the promotion of questions into standing lanes β€” are where the whole system either compounds or stalls. They are the subject of the next chapter.

06
Part III β€” Running the Plane

Standing Questions

The strongest part of the architecture: a human panel that does not read the exhaust, but teaches the institution what to keep watching β€” turning good questions into permanent capability.

A director asks, once: "Why do two sites implementing the same control require radically different amounts of discussion and exception handling?" The first time, that is an investigation β€” a walk through the estate, an answer, a disposition. What makes this architecture compound rather than merely repeat is what happens next.

From a good question to a standing one

Once validated as useful, the question is promoted and generalised: "Across all sites implementing this control, identify material divergence in attention, exception handling, evidence depth, independent challenge and human reserve margin." It stops being a one-off and becomes a versioned organisational asset.

A Standing Question carries… …so that
The question, and why it mattersIts purpose survives the person who asked it
The entities and source types it coversIts scope is explicit and auditable
Its normal baseline and the signals that perturb itA change is legible as a change
Evidence requirements and escalation criteriaA finding has a defined bar before it reaches a human
Cadence and a responsible human ownerIt runs on a schedule and someone owns it
Conditions for revision or retirementIt can be improved or killed, not left to rot

It is, in effect, a Question Ledger entry promoted into an operational review lane. And the progression it enables is the mechanism that makes the whole institution smarter over time.

Question compounding

A human asks a good question  β†’  the AI conducts a deep walk  β†’  humans review the result  β†’  the useful question is promoted  β†’  it runs repeatedly across the estate  β†’  new findings and variations emerge  β†’  the panel improves the question  β†’  the system's assurance capability compounds.

Traditional consulting answers a question and leaves. This system turns a good question into a permanent institutional capability.

These are not dashboard KPIs. They are persistent lines of institutional inquiry β€” and they accumulate.

What is the panel actually for?

The panel's role is often misunderstood as validation β€” reading every AI observation and signing it off. It is not. It is closer to chairmanship and system stewardship.

The panel decides β€” the AI supplies

The panel provides
  • β€’ Which questions matter; the North Star and boundaries
  • β€’ Significance, accountability, organisational context, proportionality
  • β€’ The challenge to evidence and to alternative interpretations
  • β€’ Promotion of valuable questions; retirement of noisy ones
  • β€’ The authority to act β€” and to keep the system pro-human
The AI provides
  • β€’ Breadth: every site, control family and period
  • β€’ Persistence: it re-reads without tiring
  • β€’ Parallelism: many questions and lenses at once
  • β€’ Exhibits: receipts, not testimony

This is human-over-the-loop supervision, not humans rubber-stamping machine output β€” and crucially, the panel must keep the system pointed at process and team-level assurance, never letting it mutate into individual performance surveillance. The panel's questions are not overhead on top of the review. They are a new ingestion stream into it.

Here is why. Each serious review produces two assets: the finding, and the route through the organisational graph that produced it. The route itself reveals source relationships that should be encoded, places where the AI repeatedly needed more evidence, dead ends, missing links between policy and operation, comparisons users keep requesting, areas that are formally important but never traversed, and emergent categories that deserve their own review lane. So the system improves not only when more documents arrive, but when humans interrogate it β€” a discipline we call filing back the walk. Questioning the map is how the map gets better.

The standing-question board

What does a mature board of standing questions look like? Drawn from the kind of pressures visible on any large regulated portfolio, the lanes are persistent lines of inquiry no dashboard owns:

  • Where is a formally green project being sustained by rising human effort? (The capacity and reserve-accounting that answers this in detail is its own discipline β€” a companion piece in this series.)
  • Where have risk statements become more confident without new evidence?
  • Which obligations are implemented differently across equivalent sites?
  • Which project assumptions have changed without downstream documents changing?
  • Where do cost explanations conflict with contemporaneous operational discussion?
  • Which community commitments have no clear implementation owner?
  • Which recurring exceptions indicate a temporary workaround has become the real process?
  • Where are several independent assurance functions relying on one evidence source?
  • Which large-demand connection assumptions are treated as firm despite conditional proponent language?
  • What is receiving abnormal discussion but has no formal risk category β€” and what is formally critical but no longer receiving substantive challenge?

Keeping it alive: scheduled builds

Standing questions only compound if they actually run. The operating rhythm is a scheduled build β€” nightly or weekly β€” that re-runs each question, diffs the findings against last cycle, detects new divergences and updates the evidence. This is the same CI/CD hygiene we apply in Nightly AI Decision Builds: versioned, diffable outputs, and human review layered into frontline, system-quality and governance tiers, each seeing what the others miss, with observations flowing back into the system. The diff is the product. A finding that was there last month and is gone this month is as informative as a new one.

Positioned this way, the panel is doing genuine board-level work rather than reviewing a deck. It is challenging the search β€” the questions and the evidence path β€” with better questions, using inspectable artefacts. And it treats the question ledger the way you treat a navigation instrument in permanent fog: refreshed continuously, never settled once.

All of which can sound like a decade-long enterprise program. It is not β€” provided you refuse to build the enterprise version first. The last chapter is about how to start small enough that it actually happens.

07
Part III β€” Running the Plane

Your First Plane

How to stand up a bounded second assurance plane you could start this month β€” one control, two sites, one year of exhaust, a small panel β€” and grow it radially into something an institution keeps.

The traditional way to build a capability like this would try to design everything up front: the complete enterprise taxonomy, every control relationship, one universal risk ontology, all the data integrations, the final reporting model, the governance committee structure and a multi-year implementation roadmap. It would almost certainly fossilise before completion β€” obsolete by the time it shipped, and too expensive to change once it had. The AI-native approach does the opposite. It begins with one bounded vertical slice.

1

operational control

2

comparable sites

1 yr

of reports, emails, exceptions, reviews

1

small panel who knows the work

Where do you actually start?

The purpose of the first slice is not to prove that the system can audit the whole enterprise. It is to discover what the system needs to know before it can be trusted at scale:

  • which source material actually contains meaningful signal
  • what semantic grain is useful
  • which comparisons produce insight, and which findings are noise
  • what evidence humans require before they will take a finding seriously
  • which questions are worth running repeatedly
  • what privacy and access boundaries the work needs

Answer those on one control at two sites, and you have earned the right to grow. From there, scaling is radial β€” you widen one dimension at a time, on a substrate built so that adding a new axis is cheap rather than another project.

Radial scaling
  1. Same control, more sites
  2. Related controls within the same operational domain
  3. The same question across different control families
  4. Additional soft-data sources
  5. Additional independent review lenses
  6. Enterprise-wide standing questions

You do not design every useful axis in advance. You build a substrate that makes new axes cheap to add.

The operating loop, as a checklist

A bounded implementation is ten concrete steps β€” not a transformation program:

  1. Select one project, control, or two comparable sites.
  2. Ingest the relevant soft estate with source permissions preserved.
  3. Define five to ten standing review questions.
  4. Run parallel AI reviews outside the operational path.
  5. Let reviews query structured systems when corroboration is needed.
  6. Produce Soft Attestation Packages with findings, exhibits, counter-cases and gaps.
  7. Have an engineering, risk and workforce panel accept, reject or redirect each finding.
  8. Promote useful questions into recurring review lanes.
  9. Record the panel's disposition as part of the package.
  10. Re-run periodically and diff what changed.

The product of that loop is not a new green light. It is an expanding collection of evidence-bearing alternative views arranged around the green light β€” never replacing it.

What to expect β€” and what to refuse

Running your first plane

βœ“ Do

  • β€’ Expect noise first β€” learning which findings are noise is half the point of a pilot
  • β€’ Hold a real evidence bar at the panel; make findings earn their way to a human
  • β€’ Keep the review off the operational hot path
  • β€’ Keep it pro-human: process and team level, never individual surveillance
  • β€’ File every package as derived data and diff it next cycle

βœ— Don't

  • β€’ Let the AI change the traffic light or take live authority
  • β€’ Let it become an accusation engine β€” findings, never verdicts
  • β€’ Try to boil the ocean with an enterprise taxonomy first
  • β€’ Cite the review as its own proof, or skip the human disposition

What changes for a board or regulator

The payoff is a fundamentally stronger position. Instead of a retrospective narrative β€” "the extent of the cost pressure was not fully apparent until late 2027" β€” the organisation can show the formal status issued at each point, the contemporaneous soft review beside it, the evidence available to both, which issues were accepted, rejected or monitored, which assumptions later changed, the exact date a pattern became strong enough to escalate, and whether management acted proportionately on what was knowable then. That is a far more defensible record β€” and, as we saw, it protects honest teams as surely as it exposes blind ones. For an organisation whose entire business is arguing why to a regulator, that record is a compounding asset, not a productivity toy.

The larger point

Existing governance assumes attention is scarce at every level β€” workers to do the job, reviewers to examine it, managers to understand it, executives to govern it β€” so the organisation samples, rolls up, compresses and assigns traffic lights. AI does not remove the executive attention constraint. But it removes much of the intermediate cognition constraint. It can read before compressing, compare before sampling, and ask many questions before selecting which one deserves escalation. It can revisit the organisation without commissioning another audit team. And it can discover that the current categories are incomplete, rather than merely checking whether the current categories are green.

So governance no longer has to be a small number of predefined questions answered repeatedly. It can become a continuously expanding field of machine-scale review, shaped by human questions and compressed back into the few findings that deserve human judgment.

Stand up one plane

Pick one control, two sites and a year of exhaust. Run a batch review outside the hot path, file the first Soft Attestation Package beside the formal report, and let a small panel decide what deserves a standing question. Then diff it next cycle.

That is not a faster dashboard. It is a new institutional organ.

REF
Sources & Evidence

References & Sources

The evidence base behind every claim — primary research, industry analysis, and technical specifications

Research Methodology

This ebook draws on primary research from standards bodies, independent research firms, enterprise technology vendors, and consulting firms. Statistics cited throughout have been cross-referenced against primary sources.

Frameworks and interpretive analysis developed by Scott Farrell / LeverageAI are listed separately below — these represent the practitioner lens through which external research is interpreted, and are not cited inline to avoid self-promotional appearance.

LeverageAI / Scott Farrell — Practitioner Frameworks

The interpretive frameworks, architectural patterns, and practitioner analysis in this ebook were developed through enterprise AI transformation consulting. The articles below are the underlying thinking behind those frameworks. They are listed here for transparency and further exploration — not cited inline, as this is the author's own analytical voice.

Scott Farrell, LeverageAI — Maximising AI Cognition and AI Value Creation

Cost of cognition and elastic cognitive capacity: previously-uneconomic exhaustive review becomes an incremental compute cost

https://leverageai.com.au/wp-content/media/articles/27-maximising-ai-cognition.html

Scott Farrell, LeverageAI — Look Mum No Hands

Progressive disclosure via proposal cards; deep overnight parallel cognition feeding a concise daytime decision surface

https://leverageai.com.au/wp-content/media/articles/43-look-mum-no-hands.html

Scott Farrell, LeverageAI — The Lane Doctrine: Deploy AI Where Physics Is on Your Side

Governance arbitrage: batch the thinking outside the hot path and ship reviewable artefacts rather than take live authority

https://leverageai.com.au/wp-content/media/articles/47-the-lane-doctrine.html

Scott Farrell, LeverageAI — Witness, Not Oracle

The evidence-package contract: claim, verbatim exhibit, resolvable pointer, confession β€” checkable is not truthful

https://leverageai.com.au/wp-content/media/articles/93-witness-not-oracle.html

Scott Farrell, LeverageAI — Differently Sighted, Not Objective

Un-invested, not objective; disclose the lens, show the exhibit, keep humans terminal; declared lenses may disagree

https://leverageai.com.au/wp-content/media/articles/135-differently-sighted-not-objective.html

Scott Farrell, LeverageAI — The Answer Depends on the Date

Bitemporal record: separate what appeared true at the time from what the organisation knew; return the version, applicable window and evidence path

https://leverageai.com.au/wp-content/media/articles/101-the-answer-depends-on-the-date.html

Scott Farrell, LeverageAI — File Back the Walk

File syntheses as typed derived cache ranked below sources and invalidated when evidence changes

https://leverageai.com.au/wp-content/media/articles/80-file-back-the-walk.html

Scott Farrell, LeverageAI — Nightly AI Decision Builds

CI/CD hygiene for AI: versioned diffable outputs and layered micro/macro/meta human review

https://leverageai.com.au/wp-content/media/articles/45-nightly-ai-decision-builds.html

Scott Farrell, LeverageAI — The Cognition Dimension Ladder

Boards challenge the search with better questions; refresh the question ledger continuously in permanent fog

https://leverageai.com.au/wp-content/media/articles/62-cognition-dimension-ladder.html

Scott Farrell, LeverageAI — Your Organization Has Source Code (And You Can Finally Read It)

BI for Soft Data: the exhaust holds the causal why-layer; compile it into a navigable estate with receipts

https://leverageai.com.au/wp-content/media/articles/86-your-organization-has-source-code.html

Scott Farrell, LeverageAI — The Soft Join: SQL Discipline for Soft Data

Joining the compiled soft estate to systems of record so soft context and hard figures become one queryable substrate

https://leverageai.com.au/wp-content/media/articles/88-the-soft-join.html

Scott Farrell, LeverageAI — BI Tells You Where; the Wiki Tells You Why

Hard figures stay in systems of record; soft context generates joinable explanations returned with receipts

https://leverageai.com.au/wp-content/media/articles/106-bi-where-wiki-why.html

Industry Analysis & Vendor Research

Infrastructure Partnerships Australia — 2026 O&SP finalist β€” Transgrid Network Asset Strategy [1]

Network scale: ~11,500km of HV lines and 136 substations and switching stations; existing use of digital twins, AI and drones

https://infrastructure.org.au/tools-resources/2026-o-and-sp-finalist-transgrid-network-asset-strategy/

Australian Energy Regulator — AER begins consultation on Transgrid's application to reopen its 2023-28 determination β€” Project EnergyConnect [2]

Additional $1.142bn (2022-23 dollars) on approved $2.121bn NSW component; ~$173m added to 2027-28 revenue; ~$18 residential bill impact; submissions received from AGL, consumer advocates, infrastructure groups

https://www.aer.gov.au/news/articles/communications/aer-begins-consultation-transgrids-application-reopen-2023-28-transmission-revenue-determination-project-energyconnect

Transgrid — EnergyConnect powers up as clean energy transition forges ahead [3]

700km NSW section construction complete, Stage 2 energising ahead of AEMO inter-network testing, June 2026

https://www.transgrid.com.au/media-publications/news-articles/energyconnect-powers-up-as-clean-energy-transition-forges-ahead/

Australian Energy Regulator — Transgrid β€” system strength material change of circumstances [4]

~$6.3bn preferred portfolio; April 2026 material-change notice as synchronous-condenser procurement costs rose >30%; AER required updated analysis and public consultation

https://www.aer.gov.au/industry/registers/determinations/transgrid-system-strength-material-change-circumstances

Transgrid — Batteries elevated in optimised system strength plan for NSW grid [5]

Revised portfolio (14 July 2026) elevates grid-forming batteries while retaining proven synchronous equipment and fallback hydro/gas

https://www.transgrid.com.au/media-publications/news-articles/batteries-elevated-in-optimised-system-strength-plan-for-nsw-grid/

Australian Energy Regulator — AER consults its preliminary position paper β€” Transgrid's hybrid system strength project revenue determination [6]

$1.185bn System Strength Project proposal 2026-31; AER isolates risk costs and labour/indirect costs; first hybrid revenue determination under the NSW framework

https://www.aer.gov.au/news/articles/communications/aer-consults-its-preliminary-position-paper-transgrids-hybrid-system-strength-project-revenue-determination

Australian Energy Regulator — AER approves reduced costs β€” HumeLink Stage 2 [7]

$3.965bn Stage 2 capex approved, $314.4m cut from application; decision tied to landholders, communities and social licence; 365km corridor with independent environmental audits and complaints registers

https://www.aer.gov.au/news/articles/news-releases/aer-approves-reduced-costs-humelink-stage-2

AEMO / Australian Energy Regulator — Transgrid PACR β€” System Security Roadmap Operational Technology upgrades [8]

Operators may be overburdened confirming information across multiple sources, less able to act in-time during a contingency; $163.5m sought for OT upgrades

https://www.aemo.com.au/consultations/current-and-closed-consultations/transgrid-pacr-system-security-roadmap-operational-technology-upgrades

Transgrid — Data centres and electricity capacity in NSW [9]

14 GW of data-centre enquiries within 12km of Sydney West since late 2024 (~whole NSW peak winter load); Western Sydney capacity largely exhausted; households should not bear cost/risk

https://www.transgrid.com.au/about-us/network/network-connections/data-centres-and-electricity-capacity-in-nsw/

Andrej Karpathy — LLM knowledge-base note and gist [10]

April 2026 write-back move: useful analyses and discovered connections filed back into the knowledge base rather than lost to chat history

https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f

About This Reference List

Compiled July 2026. All URLs verified at time of compilation. Regulatory documents and standards specifications are subject to revision — check primary sources for the most current versions.

Some links to academic papers and vendor research may require free registration. Government and standards body publications are freely accessible.