AI That Survives Audit
What an FDE must carry so a useful system can pass Australian enterprise approval—and still be defensible after it reaches production.
TL;DR
- A demo proves capability. An approval path asks whether this organisation may safely and accountably operate it.
- “Survives audit” means you can replay why the system was allowed to act and what actually occurred. Logs alone do not prove authority.1
- The scarce last-mile artefact is an approval-ready deployment package—one joined object that answers each forum’s decision question with evidence, owners, gates, and receipts. “Approval-ready” means ready for organisational review, never regulator-certified.
The pilot looked fine. The model classified claim documents, summarised the mess, and suggested a routing path. Someone recorded a Loom. Someone pasted prompt screenshots into Confluence. Then the real questions started: architecture wanted data flows; cyber wanted identity and model egress; privacy wanted purpose, accuracy, and human oversight; risk wanted eval thresholds and residual ownership; operations wanted rollback; the accountable owner wanted to know whose name was on the outcome.
That is not a documentation problem. It is a delivery problem. In regulated Australian enterprise, the scarce last mile is often not connecting a model. It is producing an intervention that can pass multi-forum review with its evidence still attached.
This article is for the forward-deployed engineer and the delivery sponsor who have to walk that path. It does not re-teach the whole Governance Stack or Proof-Carrying Transformation.2,3 It owns the concrete package those journeys require when the destination is production inside an Australian organisation that will later ask hard questions.
Demo capability is not operational permission
Traditional advisory often sells borrowed authority: the board can approve because a prestigious firm said so. That authority expires the moment an Architecture Review Board, cyber team, or privacy office asks for exhibits instead of brand.3
Generalized / Observed On a large Australian insurer programme, prestigious recommendations arrived as motherhood statements on an ill-defined problem. Implementation staff could not map the advice back to organisational findings. Governance and cyber were not designed in; they were left as someone else’s later problem. The project then tried to reconcile inherited artefacts—strategy slides, an outsourced “Agile way of working,” and poorly expressed governance—by drawing three circles and calling the middle the holy grail. That is political settlement, not synthesis. As one field formulation puts it: they began with three documents and tried to infer a reality; the opposite move is to begin with reality and determine which documents survive.
Evidence bound: the case is generalised from source field observation. It does not prove every forum name, rejection letter, condition text, or elapsed-day metric at any named firm. Where those receipts do not exist, this article leaves them blank rather than inventing them. EVIDENCE GAP for a complete public longitudinal approval journey with interviews and measured rework.
The FDE answer is not a prettier slide. It is rubber on the road: receipts, reasons you can follow, governance as an input, security considered early, and an executable path through deployment and evaluation.
One bounded Australian deployment (not a universal checklist)
Universal compliance matrices create false confidence. Obligations are deployment-specific. For this article, lock one bounded case:
Deployment: an APRA-regulated Australian general insurer’s internal claims-document assist agent. It proposes classification, summary, and routing recommendations over claim documents. It does not autonomously determine settlements or communicate adverse decisions to customers without a named human authority.
Why this shape: it is high-volume enough to matter, personal-information-heavy enough to trigger real privacy and data-risk questions, operationally material enough for operational-risk and security forums, and still controllable with human gates—so we can show earned autonomy without pretending the first release is fully autonomous claim adjudication.
Obligation classes relevant to this deployment
Distinguish four classes carefully. This is orientation for package design, not legal advice.
| Class | What applies here | What it is |
|---|---|---|
| Binding | Privacy Act 1988 / Australian Privacy Principles for APP entities handling claim personal information4; APRA CPS 230 Operational Risk Management for APRA-regulated entities including general insurers — current reissue in force from 1 July 2026; for insurers, claims processing is a minimum critical operation unless justified otherwise5; APRA CPS 234 Information Security — in force 1 July 2019, applies to general insurers (roles, commensurate capability and controls, systematic testing/assurance, incident notification)6; ASIC RG 271 enforceable IDR standards when an expression of dissatisfaction meets the complaint definition (conditional — not every claim)7 | Law / prudential standard / enforceable RG standards |
| Regulator guidance / report | APRA CPG 230 / CPG 235 (operational and data risk practice guidance)5,8; OAIC guidance on privacy and commercially available AI products (Updated 21 October 2024)9; ASIC REP 802 (Dec 2024) on general-insurance complaints handling and RG 271 identification/recording expectations10 | How regulators interpret practice |
| Voluntary / framework | Guidance for AI Adoption (21 Oct 2025) — six essential practices; voluntary industry guidance that updates/evolves the earlier Voluntary AI Safety Standard; not law11; ASD Information Security Manual (ISM) as an optional technical control catalogue — not the binding cyber basis for this APRA insurer (that is CPS 234)12 | Risk-based / voluntary frameworks |
| Internal control | ARB, cyber, privacy office, model/risk committee, operations, accountable owner risk acceptance, IDR/complaint systems | Entity governance workflow |
For this claims-assist agent, reviewers will typically force concrete answers that those instruments imply without naming your product:
- Privacy / APP: What personal information enters the model path? For what primary purpose? Who (vendor or otherwise) can access prompts and outputs? What accuracy steps exist for inferred personal information? How is human oversight embedded where effects on individuals could become significant?9
- Data risk: Is lineage, quality, and fitness-for-purpose of claim features known and owned?8
- Operational risk (CPS 230): Claims processing is a minimum critical operation for a general insurer unless justified otherwise. What fails closed, what are disruption tolerances, how is change controlled—including model/vendor (service provider) paths—and how does this assist path sit inside claims-processing continuity?5
- Cyber (CPS 234 — binding): Are information-security roles defined? Is capability commensurate with threats to claim and model assets? Are controls tested systematically with assurance? How will material incidents be detected, responded to, and notified to APRA within required timeframes? Third-party model endpoints are still the entity’s information assets under CPS 234.6 ISM patterns may support design but do not replace CPS 234.12
- Complaints / IDR (conditional): If a document is (or contains) an expression of dissatisfaction meeting RG 271’s complaint definition, can the system identify it, record it, and preserve it for the firm’s IDR process—rather than treating it as ordinary claim noise?7,10
Australian Government AI policy and assurance materials (e.g. Policy for the responsible use of AI in government v2.0 effective 15 December 2025) bind agencies, not private insurers.13 They are useful only as parallel vocabulary for designated accountability and use-case risk actions—not as obligations on this deployment.
What “survives audit” actually means
Most programmes can answer: was the model validated? was the data broadly governed? Almost none can answer: who authorised this action, under what mandate, with what admissible evidence, at decision time?1
That gap is why more logs fail as a strategy. Observability says what happened. Provenance and authority say who was allowed to drive. Policies and dashboards that cannot prevent an unauthorised action are compliance cosplay: useful theatre until the accountability moment.1,15
So the FDE’s job is twofold and simultaneous:
- Discover the governance workflow as seriously as the business workflow—which forums decide, which evidence each accepts, who owns residual risk, what must be technically enforceable.
- Build one joined package that compiles first-generation views for those forums from the same ground truth—not a telephone chain of rewritten decks.
Sibling work on Proof-Carrying Transformation owns the full professional-services transaction; the Governance Stack owns the missing authority layer; production-ready engineering owns eval and release reliability.2,3,14 This article joins those ideas at the point of organisational permission: the approval-ready deployment package.
The approval-ready deployment package
Designed / Synthetic Specimen for the claims-document assist agent. Field names are real engineering objects; values are illustrative. Not a certified template. Not any named insurer’s live pack.
The package is not a PDF dump after the build. It is the working system plus the artefacts that make permission and operation defensible:
| Package component | What it joins | Owner (role) | Receipt pointer (example) |
|---|---|---|---|
| 1. Working system slice | Deployable service + config + model route + feature flags | FDE / platform eng | git:svc-claims-assist@sha… |
| 2. Authority model | What the agent may propose vs execute; human roles; delegation limits | Accountable owner + risk | policy:authority-v3.yml |
| 3. Evidence / data boundary | Admissible inputs, PII handling, retention, vendor access, purpose map; record-class tags (claim file vs complaint/IDR) and retention rules per class | Privacy + data owner + IDR owner | pia:claims-assist-2026-… + data-flow + record-class:schema |
| 3b. Complaint preservation (conditional) | When input is an expression of dissatisfaction meeting RG 271: identify, record, route to IDR system; do not silently reclassify as “feedback” or ordinary claim mail7,10 | Claims ops + IDR / complaints owner | idr:complaint-handoff-spec |
| 4. Eval suite + thresholds | Green path, edges, incomplete docs, high-risk actions; pass bars | FDE + model risk | eval:report-r0.9.json |
| 5. Threat model + CPS 234 control pack | Prompt injection, data exfil, confused deputy, model supply chain; roles; capability statement; systematic control testing/assurance plan; APRA incident notification path6 | Cyber (CPS 234 accountable roles) | sec:tm-claims-assist-v2 + cps234:control-test-plan |
| 6. Human gates | When PAUSE is mandatory; dual control; UI proposal card | Ops + owner | gate:human-authority-matrix |
| 7. Release metadata | Version, eval metrics, config hash, deployment timestamp, environment | Platform | release:r1.0.0-meta.json |
| 8. Observability | Traces for proposals, gate outcomes, tool calls—not only app logs | SRE | obs:dashboard-claims-assist |
| 9. Incident + rollback | Kill switch, prior version, data repair notes, comms path | Ops | runbook:rb-claims-assist |
| 10. Operating receipts | Production observations written back into evals; residual risk log | Owner + FDE | receipts:weekly-… |
| 11. Approval manifest | Forum map: question, evidence object, owner, status, conditions, receipt | FDE (compiler) / owner (accepts) | See table below |
Build, release, and run stay separated: offline evals gate the build; the release is tagged with metrics; production runs with online monitoring and write-back into the next eval set.14
Approval manifest (specimen)
Designed / Synthetic Status values illustrate rework visibility. They are not observed outcomes from a named firm.
| Forum | Decision question | Evidence object | Accountable owner | Status / disposition | Conditions | Receipt pointer |
|---|---|---|---|---|---|---|
| Architecture (ARB) | May this sit in our estate with these interfaces and data flows? | Architecture pack + data-flow + dependency register | Enterprise architect (named role) | Conditional | No direct write to policy admin; read-only claims store v2 API only | arb:2026-…-cond |
| Cyber / InfoSec (CPS 234) | Are roles, capability, controls, systematic testing/assurance, and incident notification acceptable for these information assets? | Threat model + CPS 234 control map + test/assurance evidence + incident playbook | CISO delegate | Rework then Conditional | Private model endpoint only; block public generative tools for claim PII9; secret scan in CI; annual control-test sufficiency review | cyber:rework-02 → cond |
| Privacy | Is personal information handling proportionate, purpose-bound, accurate, and secured? | PIA + APP map + retention by record-class + vendor access matrix | Privacy officer | Conditional | Human review before any customer-facing use of model text; APP 10 accuracy sampling weekly; complaint-class retention separate from ordinary claim drafts | priv:pia-approved-cond |
| IDR / complaints (if triggered) | If dissatisfaction is present, is the matter identified and recorded as a complaint under RG 271? | Complaint-identification rules + IDR handoff + record log | IDR / complaints owner | Conditional (path only when definition met) | Mandatory PAUSE + record when complaint definition met; no silent “feedback” reclassification7 | idr:gate-cond |
| Risk / model | Is residual model risk acceptable at this autonomy level? | Eval report + failure taxonomy + limitation statement | Model risk / operational risk | Rejected → resubmitted | Initial reject: no threshold for “incomplete document.” Resubmit with mandatory PAUSE rule | risk:reject-01 → accept-r1 |
| Operations | Can we run, observe, roll back, and recover? | Runbooks + SLO draft + rollback drill record | Ops lead | Conditional | Rollback drill before production traffic; on-call roster named | ops:drill-… |
| Accountable business owner | Who owns residual risk and production outcomes? | Risk acceptance + success metrics + RACI | Claims operations executive | Accepted with conditions | Autonomy ceiling: propose/route only; no settlement execution | owner:risk-accept-r1 |
| Finance (if material) | Is spend justified against measured value? | Business case + unit cost per assisted claim | Finance partner | Not required at r1 / deferred | — | fin:deferred |
Rework and rejection are first-class rows. Hiding them is how organisations recreate the empty Venn: everyone pretends the middle is green because no one recorded the conditions.
One end-to-end decision trace
Designed / Synthetic Illustrative run for a single claim document assist. Stages are the minimum spine that ties engineering to authority.
- Admissible input — Claim PDF and structured claim ID enter via the approved claims-store API. Source system, record-class (ordinary claim document vs potential complaint), retention class, and sensitivity tag are attached. Out-of-scope channels (personal email forward into a public chatbot) are architecturally blocked, not merely discouraged.9
- Model proposal — The model emits a structured proposal: document type, summary, suggested work-queue, confidence, cited field IDs, reason codes, and a complaint-signal flag when the text appears to be an expression of dissatisfaction. It does not call the write API and does not alone “close” a complaint.
- Deterministic policy gate — Versioned rules check: schema validity, confidence floor, prohibited actions (e.g. settlement), data-scope match, model/release allow-list, and complaint-class rules (if complaint-signal or human-marked dissatisfaction: force PAUSE, set record-class=complaint, require IDR record path). Outcome: ALLOW propose to human queue, PAUSE for mandatory review, or DENY with reason.15
- Human authority where required — For incomplete documents, low confidence, or complaint-class items, a named claims officer (and IDR staff where required) accepts, edits, rejects, or confirms complaint recording. Identity, time, and decision are recorded. Settlement-affecting actions always PAUSE. Expressions of dissatisfaction that meet RG 271 must be recorded as complaints, not re-labelled as ordinary claim notes.7
- Executed action — Only the permitted side effect runs (e.g. route to queue Q17, or create/update IDR complaint record with preservation of original document). The agent never self-escalates its authority.
- Immutable result — Append-only record joins: input hash, record-class, proposal ID, policy version, gate outcome, human decision (if any), complaint-record pointer (if any), action result, release metadata. Later audit is a replay, not an archaeological dig through chat logs.1
Operating rule for the FDE practice: be tight on intent, loose on method, and hard on verification. Prove the green path first; then edge cases, failure modes, recovery, and human-in-the-loop. The pipeline—not the model’s memory—decides which checks must run.
Earned autonomy: conditions, production evidence, next gate
Controlled release is not a vibe. It is a sequence of explicit ceilings:
- Shadow / recommend-only — model proposals logged; humans work as today.
- Assist with human gate on every action — proposal cards in the claims UI.
- Assist with auto-route only on high-confidence green path — still no settlement execution.
- Later autonomy — only if production observations are written back into the eval suite, residual risk is re-accepted, and the next forum gate is explicitly approved.
Each stage earns the right to the next. A proposed target for confidence thresholds or pass rates is not an observed result. When you lack receipts, leave the number blank.
Measures (honest blanks)
| Measure | Observed (with receipt) | Designed target (not observed) | Receipt slot |
|---|---|---|---|
| Approval elapsed time (first complete package → conditional production) | — | Designed entity-specific; track calendar days and wait states per forum | [blank] |
| Unanswered review questions at first forum submission | — | Designed count open questions by forum; drive toward zero before resubmit | [blank] |
| Rework cycles (reject or major condition → resubmit) | — | Designed record each cycle; specimen table shows risk reject-then-accept pattern | [blank] |
| Production incidents attributable to the assist agent | — | Designed severity-tagged; feed failure cases into evals within one sprint | [blank] |
EVIDENCE GAP: no public, receipt-backed before/after series for this exact Australian journey is claimed here. Inventing numbers would recreate the prestige problem the package is meant to end.
What the FDE actually does differently
Stop treating governance as a decorative circle on the kickoff slide. Treat each prior artefact—strategy pack, operating model, control standard—as an input with a disposition: retain, modify, reject, or requires evidence.
Compile radial views from one ground truth: the ARB pack, cyber submission, privacy pack, risk pack, ops runbooks, and owner brief should be first-generation translations of the same system, not serial rewrites that drift apart.
Put authority in the execution path: model proposes; deterministic gates and human authority decide; immutable results make audit a replay.
And when someone asks for “just more logging,” answer cleanly: logs are necessary and insufficient. Surviving audit means the team can show why the system was allowed to act—and what it did—without reconstructing a story under pressure.
Next step. Before your next demo graduates to “production,” write the approval manifest with empty status fields and the six-stage decision trace for one real action. Fill only what you can point to. The blanks are the work.
References
- Scott Farrell / LeverageAI. “The Governance Stack” — authority gap / three layers. — Observability says what happened; provenance/authority say who was allowed to act. Layer 3 is not “more logging.” https://leverageai.com.au/wp-content/media/articles/57-governance-stack.html
- Scott Farrell / LeverageAI. “The Governance Stack” — three-layer diagnostic (data / model / authority). Cameo only in this article. https://leverageai.com.au/wp-content/media/articles/57-governance-stack.html
- Scott Farrell / LeverageAI. “Proof-Carrying Transformation” — incomplete consulting transaction; borrowed vs transferable authority. https://leverageai.com.au/wp-content/media/articles/164-proof-carrying-transformation.html
- Commonwealth of Australia. Privacy Act 1988 (Cth); Australian Privacy Principles. Binding for APP entities. https://www.legislation.gov.au/Series/C2004A03712
- Australian Prudential Regulation Authority (APRA). Prudential Standard CPS 230 Operational Risk Management — current reissue in force from 1 July 2026. For insurers (general, life, private health), claims processing is a minimum critical operation unless justified otherwise; service-provider and BCP/tolerance requirements apply. https://www.apra.gov.au/standards/cps-230
- APRA. Prudential Standard CPS 234 Information Security — in force 1 July 2019; applies to general insurers. Requires defined information-security roles and responsibilities; capability commensurate with threats; controls commensurate with criticality/sensitivity plus systematic testing and assurance; notification of material information security incidents (including no later than 72 hours in specified cases). https://www.apra.gov.au/standards/cps-234 · APRA cyber orientation: https://www.apra.gov.au/cyber-security
- Australian Securities and Investments Commission (ASIC). Regulatory Guide 271 Internal dispute resolution (issued 2 September 2021). Enforceable standards for in-scope firms. Complaint definition RG 271.27 (expression of dissatisfaction…); firms must deal with expressions meeting the definition (RG 271.28); firms must record all complaints received (RG 271.179). https://www.asic.gov.au/regulatory-resources/find-a-document/regulatory-guides/rg-271-internal-dispute-resolution/
- APRA. CPG 235 Managing Data Risk — prudential practice guide on data risk for regulated entities. https://www.apra.gov.au/practice-guides/cpg-235
- Office of the Australian Information Commissioner (OAIC). “Guidance on privacy and the use of commercially available AI products.” Updated 21 October 2024. Privacy obligations on personal information inputs/outputs; due diligence; human oversight; APP 6/10; high-risk decision care; best practice against public generative tools for personal/sensitive information. https://www.oaic.gov.au/privacy/privacy-guidance-for-organisations-and-government-agencies/guidance-on-privacy-and-the-use-of-commercially-available-ai-products
- ASIC. REP 802 Cause for complaint: Complaints handling in general insurance (released 5 December 2024). Supervisory review of general insurers against select RG 271 obligations; emphasises identification and recording of complaints (RG 271.28; RG 271.179). https://www.asic.gov.au/regulatory-resources/find-a-document/reports/rep-802-cause-for-complaint-complaints-handling-in-general-insurance/
- National AI Centre / Department of Industry. Guidance for AI Adoption (published 21 October 2025) — six essential practices for safe and responsible AI governance; voluntary industry guidance that updates/evolves the earlier Voluntary AI Safety Standard; not law. https://www.industry.gov.au/publications/guidance-for-ai-adoption · Implementation guidance: https://www.ai.gov.au/staying-safe-and-responsible/essential-ai-practices/guidance-ai-adoption-implementation-guidance · PDF: https://www.ai.gov.au/sites/default/files/2026-05/Guidance-for-AI-adoption-implementation-guidance_0_0.pdf
- Australian Signals Directorate / ACSC. Information Security Manual (ISM) — optional cyber framework organisations may apply using their risk management framework; not the binding cyber basis for this APRA-regulated private insurer (see CPS 234). https://www.cyber.gov.au/business-government/asds-cyber-security-frameworks/ism
- Australian Government. Policy for the responsible use of AI in government (v2.0 effective 15 December 2025) and Pilot AI assurance framework materials — agency-facing; not applied here as private-insurer law. https://www.digital.gov.au/policy/ai · https://www.digital.gov.au/policy/ai/pilot-ai-assurance-framework
- Scott Farrell / LeverageAI. “Production-Ready LLM Systems” — demo-to-production gap; systematic evaluation; evaluation-driven build/release/run quality gates, release metadata, and continuous evaluation write-back. https://leverageai.com.au/wp-content/media/articles/02-production-ready-llm-systems.html
- Scott Farrell / LeverageAI. “Compliance Cosplay” / Decision Authority Infrastructure — model proposes; in-path ALLOW/PAUSE/DENY; governance that cannot prevent unauthorised execution is theatre. https://leverageai.com.au/wp-content/media/articles/55-compliance-cosplay.html · https://leverageai.com.au/wp-content/media/articles/53-governance-as-code.html
