Leverage AI

Knowledge Architecture · Assurance · Engagement World

The Engagement Auditor Is Not the Janitor

A polished project knowledge graph can be clean, cited, and confidently wrong. Structural maintenance and epistemic warrant are different jobs — and combining them creates a conflict of interest.

Scott Farrell · LeverageAI · Extender of Engagement World · ~12 min read

Takeaways

The architecture page is pristine. Duplicates are gone. Edges are typed. Every consequential sentence ends with a citation. The retrieval score looks healthy. Someone is about to export the proposal.

Then you notice the security approval it cites was signed before Thursday’s boundary redesign. The page is clean. The claim is not justified. That is not a documentation failure. It is an assurance-boundary failure — and it is exactly what happens when teams treat the Janitor as if it were the Auditor.

This piece extends Engagement World without re-teaching the parent model.1 Parent doctrine already places a private, project-bounded, provenance-bearing world between institutional kernels and ephemeral task rooms, and already names three maintenance roles: Scribe, Janitor, Auditor. What this extender owns is the assurance boundary: the audit contract, the finding taxonomy, the separation of findings from canonical truth, the Janitor-only baseline, and the hard rule that the Auditor never becomes approval authority.

Two jobs, two questions

The Scribe absorbs new material and proposes claims with source pointers. That is integration. The remaining split is where most teams go soft.

RoleMaintainsPrimary questionFailure if alone
Janitor Shape Is this world still lean and coherent? Pretty structure, unjustified claims
Auditor Warrant Is this world still justified? Findings with nowhere durable to land — or worse, findings silently rewritten as truth

The Janitor merges duplicates, splits overloaded nodes, marks supersession, repairs indexes, enforces schema, manages expiry, and — critically — can preserve contradictions as edges rather than averaging them into false consensus. That work prevents the knowledge graveyard: an append-only corpus that grows larger and dumber because nothing subtracts.2

The Auditor does something else. It reconstructs each consequential claim from its support path. It challenges derivation. It compares dates and authority. It tests mandatory coverage. It emits findings. It does not tidy.

The operation

Not “Are you sure?” — Reconstruct the current claim from its evidence path and report any mismatch.

That reconstruction stance is not a taste preference. Provenance standards exist because the history of how information was produced is “crucial in deciding whether information is to be trusted.”3 Having a citation field filled is not the same as having a warrant you can still defend.

Why combining the roles fails

Put both jobs in one process and you invent a conflict: the same machinery that compressed ambiguity is asked to certify that no load-bearing ambiguity was lost.

A merge that removes a live release-gate disagreement makes the graph easier to walk. It also destroys the disagreement the Auditor needed to see. A supersession that collapses two authority boundaries makes retrieval cleaner. It also launders an inference into a fact-shaped node. A lint pass that eliminates orphans and stale links can raise every cosmetic quality metric while leaving circular opportunity chains untouched.

Public LLM-wiki practice already includes lint for contradictions, stale claims, and orphans — cousin work to the Janitor.1 Useful. Incomplete. Lint can tell you the graph is messy. It cannot tell you the graph is still justified.

Industry language often confuses a related pair: data integrity versus data quality — structural consistency versus whether the data is actually fit for the decision you are about to make.4 The analogy is imperfect and still useful. The Janitor is closer to structural integrity and navigability. The Auditor is closer to whether consequential claims remain justified by current evidence under declared authority rules. Equating either with “the RAG score looks fine” is how teams sleepwalk into confident export.

There is a second failure mode once assurance is present but correlated. Paraphrasing the Institutional Linter: reviews that share one upstream evidence feed are not multi-line assurance—only a shared root in different coats.5 Engagement Worlds call a cousin of this circular glory: opportunity slides that all depend on the same unverified rail, marketed as independent confirmations.1 The Auditor’s job includes making that shared root visible. The Janitor’s job does not.

Specimen audit contract

Designed The contract below is a specimen for an Engagement World audit run. Adapt the scope; keep the separation of powers.

Audit contract — specimen

1. Scope. Consequential claims only: proposal commitments, architecture assertions, security boundary claims, deployability statements, client-confirmed priorities used as decision inputs, and any derived node that gates release, commercial export, or handover.

2. Selection rule. Start from the export surface (proposal / ARB pack / production cut / close pack). Walk inbound supports[]. Expand to any node whose authority is derived or inferred-unconfirmed and that appears on a critical path. Cap depth by risk, not by comfort.

3. Source and authority rules. Every load-bearing node must expose origin, authority, supports, as-at, and status. Derived ranks below source-backed. Inferred-unconfirmed never equals client-confirmed. Secondary research never silently becomes client bronze. Empty supports on derived claims fail closed.

4. Tests (minimum).

5. Severities. Critical (blocks named export), High (must dispose before export), Medium (time-boxed), Low (hygiene for later), Info (coverage note).

6. Dispositions. open → accepted_risk | remediate | reject_claim | reclassify_authority | split_node | supersede | needs_human_evidence. Only a human (or explicitly authorised human process) may apply a disposition that changes canonical nodes.

7. Re-test. After remediation or after any material source/design mutation in the impact set, re-run the affected claim bundle and attach a re-test receipt to the finding thread.

8. Non-authority rule. The Auditor is read-only against canonical truth during the run. It may write only to a findings plane. It may recommend. It may not approve, merge, or auto-close a gate.

That last clause is load-bearing. Academic and industry audit-trail work is converging on chronological, tamper-evident, context-rich ledgers that let organisations reconstruct what changed, when, and who authorised it — process transparency as a durable object, not a model monologue.6 Internal algorithmic auditing frameworks similarly treat audit integrity as a lifecycle process, not a confidence theatre at the end.7 None of that requires handing the Auditor the keys to production truth.

Elastic Assurance makes the same separation at organisational scale: an exploratory plane produces evidence-backed findings, not silent edits to the formal traffic light.8 The Engagement Auditor should behave the same way inside one project world.

How an Auditor run should speak

A validation run has a fixed shape:

  1. Select a claim or proposed external output.
  2. Traverse its supports graph.
  3. Restore relevant source versions / as-at snapshots.
  4. Confirm cited passages or artefacts exist.
  5. Check mandatory evidence classes for this claim type.
  6. Compare present synthesis with evidence.
  7. Identify unsupported extension, stale support, circular derivation, missing coverage, or predating receipt.
  8. Emit a proof-carrying validation report — claim, exhibit, resolvable pointer, and a confession of what could not be verified.9

The report should sound like structure, not vibe:

What it should not contain: the model’s private chain-of-thought as a substitute for path. Governance traces earn trust by externalising the version-pinned path through admissible knowledge — wiki snapshot, observed pages, rejected candidates — not by asking the model to narrate hidden reasoning.10

And it should not be self-graded by the same agent instance that authored the claim. Hidden Gates doctrine is blunt about that pattern: once the worker can see the exact rubric, the rubric becomes a target rather than a measure; review from outside, without turning every check into a gameable specification.11 Independence is part of warrant.

Synthetic case: a seeded Engagement World

Synthetic The following case is designed and seeded for illustration. It is not a client engagement, not a measured pilot, and not a claim about production false-positive rates. Product, consultancy, and client names are generalised on purpose.

World: EW-SYN-Payment-Modernisation-2026 — a multi-week pursuit world for a payments-modernisation opportunity. Federated nodes include firm capability pages, a client landing-zone absence, architecture decisions, opportunity hypotheses, and a proposal draft.

As-of snapshot: T0 seed → T1 workshop inferences → T2 security boundary redesign → T3 Janitor hygiene pass → T4 Auditor run before proposal freeze.

What the Janitor-only baseline did

At T3, a Janitor-only pass:

Janitor report (shape): 0 blockers. Navigability up. No warrant tests run.

What the Auditor found (findings plane only)

At T4 the Auditor ran read-only against canonical nodes and wrote only finding.* objects. Dispositions remain open until a human acts. Canonical claim status is unchanged by the run.

IDClassSeverityFinding (summary)
F-01 Predating receipt / stale support Critical Architecture asserts “boundary verified” with receipt sec-approval-T1; security boundary v2 landed at T2. Receipt predates material design change.
F-02 Circular derivation High opp.marketplace-rail supports value.cycle-time supports opp.marketplace-rail. No client or firm bronze root.
F-03 Unsupported extension + authority mismatch Critical Proposal promotes “client can accept Marketplace private offers” as firm-ready capability. Origin is consultant inference from T1; never client-confirmed; no known-absence ticket.
F-04 Missing mandatory coverage High Deployability claim lacks landing-zone evidence class; only a known-absence node exists, and the proposal text does not preserve the absence.
F-05 Merge hazard High Janitor-proposed merge of AI-triage vs deterministic-routing would erase a live release-gate disagreement still open with the client security owner.
F-06 Change-impact revalidation set Medium After boundary v2, dependents requiring revalidation: architecture summary, test plan pointer, proposal §Security, handover residual-risk note.
Evidence label: All six findings are synthetic fixtures designed to cover the brief’s required failure classes. They demonstrate what an Auditor should catch, not what a production system measured on a real client estate.
finding.F-01
  claim: arch.boundary-verified
  class: predating_receipt
  severity: critical
  evidence:
    - receipt: sec-approval-T1 @ T1
    - design_change: security-boundary-v2 @ T2
  canonical_write: false
  disposition: open
  owner: human:engagement-lead
  retest_required_after: remediate | accept_risk

Notice the shape: the Auditor did not “fix” the architecture page. It did not auto-supersede the receipt. It did not approve the proposal with conditions. It emitted a finding with evidence pointers and left disposition to a human.

Janitor-only versus Auditor (comparison)

QuestionJanitor-only at T3Auditor at T4
Is the world navigable? Improved Not the primary question
Are citations present? Yes Insufficient — dates and authority checked
Stale / predating support? Missed F-01 Critical
Circular derivation? Missed (may even look “well linked”) F-02 High
Unsupported extension / authority laundering? Missed F-03 Critical
Missing mandatory region? Missed if prose is fluent F-04 High
Merge erases live disagreement? May propose the merge F-05 blocks merge closure
Post-change revalidation set? Not produced F-06 impact list

The comparison is the thesis in operational form: structural maintenance can succeed while epistemic assurance fails. If your only green report is the Janitor’s, you are measuring the wrong green for export readiness.

Change-impact revalidation

When a source mutates or a design decision lands, the Auditor’s job is not only to flag the local mismatch. It is to list the dependent claim bundle that must be re-justified before the next external commitment.

Designed In the synthetic world, security-boundary v2 at T2 generates an impact set: every active derived claim whose supports include the pre-v2 boundary or the T1 approval. The revalidation receipt is not “we still believe it.” It is a re-run of reconstruction over that set, with new findings or a clean re-test attached to the disposition thread.

This is the Engagement-scale cousin of governance regression thinking: change is cheap to declare and expensive to re-justify if you refuse to name dependents. Teams that skip the impact set quietly recycle old receipts as political cover — the “spirit of the test still holds” argument the Auditor is hired to refuse.

Findings stay findings until a human says otherwise

Three rules keep the system honest:

  1. Read-only canonical during the run. The Auditor may traverse and report. It may not flip claim status as a side effect of scanning.
  2. Findings plane is first-class. Severity, evidence pointers, owner, disposition, and re-test receipt live on the finding, not as a silent footnote on the claim.
  3. Disposition is human-owned. Accept risk, remediate, reject, reclassify, split, supersede — each is an accountable act. The Auditor can recommend; it cannot authorise.

That is also why the Auditor must not become an autonomous approval authority. Approval is a different speech act. It binds the organisation. Reconstruction improves the quality of what a human is asked to bind. It does not replace the binding.

Knowledge-graph research is increasingly explicit that much of what systems treat as “data” is really asserted, interpreted material — and that provenance of who asserted what under which authority is essential to assessing reliability.12 Federation fields on Engagement World nodes exist for that reason. The Auditor enforces them under pressure. The Janitor keeps them walkable.

What to do before the next export

  1. Write the one-page audit contract for your next proposal freeze or ARB pack. Name scope, tests, severities, dispositions, and the non-authority rule.
  2. Separate the planes in tooling if you can: Janitor PRs for structure; Auditor findings for warrant. Do not let one agent instance grade its own claims.
  3. Seed at least the five failure classes as fixtures — stale support, circular derivation, unsupported extension, missing coverage, predating receipt — so “green” means something mechanical.
  4. On every material design change, demand the revalidation set before anyone reuses old receipts.
  5. Keep humans on disposition. If your system can close a Critical finding without a named owner, you did not build an Auditor. You built an unsupervised editor with better vocabulary.

North star

Stop asking whether the graph looks maintained. Ask whether its consequential claims can still be independently reconstructed and challenged — without letting the challenger silently rewrite the world.

The Engagement World is how multi-person, multi-agent delivery compounds on shared understanding.1 The Janitor keeps that world from rotting into a graveyard. The Auditor keeps it from hardening into a lie. They are partners. They are not the same job. And only one of them is allowed to rearrange the furniture.

References

  1. Scott Farrell / LeverageAI. “Engagement World: The Project Reality the Slide Deck Pretended to Hold.” — Parent doctrine for the project-bounded Engagement World and Scribe/Janitor/Auditor roles; this article extends the assurance boundary without re-teaching the lifecycle. https://leverageai.com.au/wp-content/media/articles/170-engagement-world.html
  2. Scott Farrell / LeverageAI. “Designing Loops, Not Prompts.” — Knowledge-graveyard / Janitor doctrine: append-only loops grow larger and dumber without subtractive consolidation; navigability is not warrant. https://leverageai.com.au/wp-content/media/articles/64-designing-loops-not-prompts.html
  3. W3C. “PROV-DM: The PROV Data Model.” W3C Recommendation, 30 April 2013. — “Provenance is information about entities, activities, and people involved in producing a piece of data or thing, which can be used to form assessments about its quality, reliability or trustworthiness.” Also: provenance of information is “crucial in deciding whether information is to be trusted.” https://www.w3.org/TR/prov-dm/
  4. Datafold. “Data integrity vs. data quality.” — Industry distinction often confused in practice: integrity concerns structural consistency; quality concerns fitness for use / correctness for decisions. Used here as analogy only for Janitor (structure) vs Auditor (warrant). https://www.datafold.com/blog/data-integrity-vs-data-quality/
  5. Scott Farrell / LeverageAI. “The Institutional Linter.” — Independence as a graph property; correlated checkers that share one upstream feed create the appearance of multi-line assurance without independent warrant (paraphrase, not a source-turn quotation). https://leverageai.com.au/wp-content/media/articles/137-institutional-linter.html
  6. Ojewale et al. “Audit Trails for Accountability in Large Language Models.” arXiv:2601.20727, 2026. — LLM audit trails as chronological, tamper-evident, context-rich ledgers linking technical provenance with governance records so organisations can reconstruct what changed, when, and who authorised it. https://arxiv.org/abs/2601.20727
  7. Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., et al. “Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing.” ACM FAccT, 2020. — End-to-end internal algorithmic auditing across the development lifecycle to support audit integrity and close accountability gaps. https://arxiv.org/abs/2001.00973
  8. Scott Farrell / LeverageAI. “Elastic Assurance.” — Exploratory assurance produces evidence-backed findings rather than silent formal status changes; findings plane stays off the operational hot path. https://leverageai.com.au/wp-content/media/articles/136-elastic-assurance.html
  9. Scott Farrell / LeverageAI. “Witness Not Oracle.” — Evidence packages return claim + exhibit + resolvable pointer + confession of what could not be verified; witnesses you can check, not oracles you must trust. https://leverageai.com.au/wp-content/media/articles/93-witness-not-oracle.html
  10. Scott Farrell / LeverageAI. “The Model Is Not the Memory.” — Governance traces externalise version-pinned paths through admissible knowledge rather than relying on model chain-of-thought as the audit object. https://leverageai.com.au/wp-content/media/articles/68-the-model-is-not-the-memory.html
  11. Scott Farrell / LeverageAI. “Hidden Gates.” — Share intent, keep diagnostic rubrics from becoming gameable targets; independent review without self-grading. https://leverageai.com.au/wp-content/media/articles/94-hidden-gates.html
  12. “Provenance-Enhanced Statements in Knowledge Graphs.” arXiv HTML 2606.15246, 2026. — Documented, verifiable provenance as fundamental to KG quality assurance; authority and epistemic status of asserted material (capta) matter for reliability assessments. https://arxiv.org/html/2606.15246