Knowledge Architecture · Assurance · Engagement World
The Engagement Auditor Is Not the Janitor
A polished project knowledge graph can be clean, cited, and confidently wrong. Structural maintenance and epistemic warrant are different jobs — and combining them creates a conflict of interest.
Takeaways
- The Janitor asks whether the world is still lean and coherent; the Auditor asks whether it is still justified.
- Audit is read-only against canonical truth during a run: findings stay separate until a human disposition changes the world.
- Required tests catch stale support, circular derivation, unsupported extension, missing coverage, and post-receipt design change.
- Janitor-only hygiene can improve navigability while missing every load-bearing warrant failure.
- The Auditor must never become autonomous approval authority.
The architecture page is pristine. Duplicates are gone. Edges are typed. Every consequential sentence ends with a citation. The retrieval score looks healthy. Someone is about to export the proposal.
Then you notice the security approval it cites was signed before Thursday’s boundary redesign. The page is clean. The claim is not justified. That is not a documentation failure. It is an assurance-boundary failure — and it is exactly what happens when teams treat the Janitor as if it were the Auditor.
This piece extends Engagement World without re-teaching the parent model.1 Parent doctrine already places a private, project-bounded, provenance-bearing world between institutional kernels and ephemeral task rooms, and already names three maintenance roles: Scribe, Janitor, Auditor. What this extender owns is the assurance boundary: the audit contract, the finding taxonomy, the separation of findings from canonical truth, the Janitor-only baseline, and the hard rule that the Auditor never becomes approval authority.
Two jobs, two questions
The Scribe absorbs new material and proposes claims with source pointers. That is integration. The remaining split is where most teams go soft.
| Role | Maintains | Primary question | Failure if alone |
|---|---|---|---|
| Janitor | Shape | Is this world still lean and coherent? | Pretty structure, unjustified claims |
| Auditor | Warrant | Is this world still justified? | Findings with nowhere durable to land — or worse, findings silently rewritten as truth |
The Janitor merges duplicates, splits overloaded nodes, marks supersession, repairs indexes, enforces schema, manages expiry, and — critically — can preserve contradictions as edges rather than averaging them into false consensus. That work prevents the knowledge graveyard: an append-only corpus that grows larger and dumber because nothing subtracts.2
The Auditor does something else. It reconstructs each consequential claim from its support path. It challenges derivation. It compares dates and authority. It tests mandatory coverage. It emits findings. It does not tidy.
The operation
Not “Are you sure?” — Reconstruct the current claim from its evidence path and report any mismatch.
That reconstruction stance is not a taste preference. Provenance standards exist because the history of how information was produced is “crucial in deciding whether information is to be trusted.”3 Having a citation field filled is not the same as having a warrant you can still defend.
Why combining the roles fails
Put both jobs in one process and you invent a conflict: the same machinery that compressed ambiguity is asked to certify that no load-bearing ambiguity was lost.
A merge that removes a live release-gate disagreement makes the graph easier to walk. It also destroys the disagreement the Auditor needed to see. A supersession that collapses two authority boundaries makes retrieval cleaner. It also launders an inference into a fact-shaped node. A lint pass that eliminates orphans and stale links can raise every cosmetic quality metric while leaving circular opportunity chains untouched.
Public LLM-wiki practice already includes lint for contradictions, stale claims, and orphans — cousin work to the Janitor.1 Useful. Incomplete. Lint can tell you the graph is messy. It cannot tell you the graph is still justified.
Industry language often confuses a related pair: data integrity versus data quality — structural consistency versus whether the data is actually fit for the decision you are about to make.4 The analogy is imperfect and still useful. The Janitor is closer to structural integrity and navigability. The Auditor is closer to whether consequential claims remain justified by current evidence under declared authority rules. Equating either with “the RAG score looks fine” is how teams sleepwalk into confident export.
There is a second failure mode once assurance is present but correlated. Paraphrasing the Institutional Linter: reviews that share one upstream evidence feed are not multi-line assurance—only a shared root in different coats.5 Engagement Worlds call a cousin of this circular glory: opportunity slides that all depend on the same unverified rail, marketed as independent confirmations.1 The Auditor’s job includes making that shared root visible. The Janitor’s job does not.
Specimen audit contract
Designed The contract below is a specimen for an Engagement World audit run. Adapt the scope; keep the separation of powers.
Audit contract — specimen
1. Scope. Consequential claims only: proposal commitments, architecture assertions, security boundary claims, deployability statements, client-confirmed priorities used as decision inputs, and any derived node that gates release, commercial export, or handover.
2. Selection rule. Start from the export surface (proposal / ARB pack / production cut / close pack). Walk inbound supports[]. Expand to any node whose authority is derived or inferred-unconfirmed and that appears on a critical path. Cap depth by risk, not by comfort.
3. Source and authority rules. Every load-bearing node must expose origin, authority, supports, as-at, and status. Derived ranks below source-backed. Inferred-unconfirmed never equals client-confirmed. Secondary research never silently becomes client bronze. Empty supports on derived claims fail closed.
4. Tests (minimum).
- Stale support — support superseded or content-changed; dependent still active.
- Circular derivation — mutual supports with no external bronze or kernel root.
- Unsupported extension — new commitment without supports or known-absence handling.
- Missing mandatory coverage — claim type lacks required evidence class (e.g. deployability without landing-zone evidence).
- Predating receipt / post-receipt design change — evidence timestamp precedes a design mutation that invalidates what the receipt was attached to.
- Authority mismatch — client-local inference presented as firm capability or client fact.
- Merge hazard — proposed structural merge would erase a live disagreement relevant to a gate.
5. Severities. Critical (blocks named export), High (must dispose before export), Medium (time-boxed), Low (hygiene for later), Info (coverage note).
6. Dispositions. open → accepted_risk | remediate | reject_claim | reclassify_authority | split_node | supersede | needs_human_evidence. Only a human (or explicitly authorised human process) may apply a disposition that changes canonical nodes.
7. Re-test. After remediation or after any material source/design mutation in the impact set, re-run the affected claim bundle and attach a re-test receipt to the finding thread.
8. Non-authority rule. The Auditor is read-only against canonical truth during the run. It may write only to a findings plane. It may recommend. It may not approve, merge, or auto-close a gate.
That last clause is load-bearing. Academic and industry audit-trail work is converging on chronological, tamper-evident, context-rich ledgers that let organisations reconstruct what changed, when, and who authorised it — process transparency as a durable object, not a model monologue.6 Internal algorithmic auditing frameworks similarly treat audit integrity as a lifecycle process, not a confidence theatre at the end.7 None of that requires handing the Auditor the keys to production truth.
Elastic Assurance makes the same separation at organisational scale: an exploratory plane produces evidence-backed findings, not silent edits to the formal traffic light.8 The Engagement Auditor should behave the same way inside one project world.
How an Auditor run should speak
A validation run has a fixed shape:
- Select a claim or proposed external output.
- Traverse its supports graph.
- Restore relevant source versions / as-at snapshots.
- Confirm cited passages or artefacts exist.
- Check mandatory evidence classes for this claim type.
- Compare present synthesis with evidence.
- Identify unsupported extension, stale support, circular derivation, missing coverage, or predating receipt.
- Emit a proof-carrying validation report — claim, exhibit, resolvable pointer, and a confession of what could not be verified.9
The report should sound like structure, not vibe:
- “The proposal claims the client can deploy into its AWS environment, but no client-cloud landing-zone evidence has been opened.”
- “Three opportunity recommendations independently depend on the unverified assumption that procurement can use Marketplace private offers.”
- “The architecture page cites a test receipt produced before the security boundary changed.”
- “This ‘client priority’ originated as a consultant inference and has never been confirmed by the client.”
What it should not contain: the model’s private chain-of-thought as a substitute for path. Governance traces earn trust by externalising the version-pinned path through admissible knowledge — wiki snapshot, observed pages, rejected candidates — not by asking the model to narrate hidden reasoning.10
And it should not be self-graded by the same agent instance that authored the claim. Hidden Gates doctrine is blunt about that pattern: once the worker can see the exact rubric, the rubric becomes a target rather than a measure; review from outside, without turning every check into a gameable specification.11 Independence is part of warrant.
Synthetic case: a seeded Engagement World
Synthetic The following case is designed and seeded for illustration. It is not a client engagement, not a measured pilot, and not a claim about production false-positive rates. Product, consultancy, and client names are generalised on purpose.
World: EW-SYN-Payment-Modernisation-2026 — a multi-week pursuit world for a payments-modernisation opportunity. Federated nodes include firm capability pages, a client landing-zone absence, architecture decisions, opportunity hypotheses, and a proposal draft.
As-of snapshot: T0 seed → T1 workshop inferences → T2 security boundary redesign → T3 Janitor hygiene pass → T4 Auditor run before proposal freeze.
What the Janitor-only baseline did
At T3, a Janitor-only pass:
- Merged three near-duplicate “cloud landing” paragraphs into one tidy architecture section.
- Converted repeated prose into typed edges.
- Archived two cold research pages.
- Proposed merging
decision.ai-triage-wedgewithdecision.deterministic-routingbecause both mentioned “intake path” and retrieval treated them as near-duplicates. - Left all citations in place. The graph looked excellent.
Janitor report (shape): 0 blockers. Navigability up. No warrant tests run.
What the Auditor found (findings plane only)
At T4 the Auditor ran read-only against canonical nodes and wrote only finding.* objects. Dispositions remain open until a human acts. Canonical claim status is unchanged by the run.
| ID | Class | Severity | Finding (summary) |
|---|---|---|---|
| F-01 | Predating receipt / stale support | Critical | Architecture asserts “boundary verified” with receipt sec-approval-T1; security boundary v2 landed at T2. Receipt predates material design change. |
| F-02 | Circular derivation | High | opp.marketplace-rail supports value.cycle-time supports opp.marketplace-rail. No client or firm bronze root. |
| F-03 | Unsupported extension + authority mismatch | Critical | Proposal promotes “client can accept Marketplace private offers” as firm-ready capability. Origin is consultant inference from T1; never client-confirmed; no known-absence ticket. |
| F-04 | Missing mandatory coverage | High | Deployability claim lacks landing-zone evidence class; only a known-absence node exists, and the proposal text does not preserve the absence. |
| F-05 | Merge hazard | High | Janitor-proposed merge of AI-triage vs deterministic-routing would erase a live release-gate disagreement still open with the client security owner. |
| F-06 | Change-impact revalidation set | Medium | After boundary v2, dependents requiring revalidation: architecture summary, test plan pointer, proposal §Security, handover residual-risk note. |
finding.F-01
claim: arch.boundary-verified
class: predating_receipt
severity: critical
evidence:
- receipt: sec-approval-T1 @ T1
- design_change: security-boundary-v2 @ T2
canonical_write: false
disposition: open
owner: human:engagement-lead
retest_required_after: remediate | accept_risk
Notice the shape: the Auditor did not “fix” the architecture page. It did not auto-supersede the receipt. It did not approve the proposal with conditions. It emitted a finding with evidence pointers and left disposition to a human.
Janitor-only versus Auditor (comparison)
| Question | Janitor-only at T3 | Auditor at T4 |
|---|---|---|
| Is the world navigable? | Improved | Not the primary question |
| Are citations present? | Yes | Insufficient — dates and authority checked |
| Stale / predating support? | Missed | F-01 Critical |
| Circular derivation? | Missed (may even look “well linked”) | F-02 High |
| Unsupported extension / authority laundering? | Missed | F-03 Critical |
| Missing mandatory region? | Missed if prose is fluent | F-04 High |
| Merge erases live disagreement? | May propose the merge | F-05 blocks merge closure |
| Post-change revalidation set? | Not produced | F-06 impact list |
The comparison is the thesis in operational form: structural maintenance can succeed while epistemic assurance fails. If your only green report is the Janitor’s, you are measuring the wrong green for export readiness.
Change-impact revalidation
When a source mutates or a design decision lands, the Auditor’s job is not only to flag the local mismatch. It is to list the dependent claim bundle that must be re-justified before the next external commitment.
Designed In the synthetic world, security-boundary v2 at T2 generates an impact set: every active derived claim whose supports include the pre-v2 boundary or the T1 approval. The revalidation receipt is not “we still believe it.” It is a re-run of reconstruction over that set, with new findings or a clean re-test attached to the disposition thread.
This is the Engagement-scale cousin of governance regression thinking: change is cheap to declare and expensive to re-justify if you refuse to name dependents. Teams that skip the impact set quietly recycle old receipts as political cover — the “spirit of the test still holds” argument the Auditor is hired to refuse.
Findings stay findings until a human says otherwise
Three rules keep the system honest:
- Read-only canonical during the run. The Auditor may traverse and report. It may not flip claim status as a side effect of scanning.
- Findings plane is first-class. Severity, evidence pointers, owner, disposition, and re-test receipt live on the finding, not as a silent footnote on the claim.
- Disposition is human-owned. Accept risk, remediate, reject, reclassify, split, supersede — each is an accountable act. The Auditor can recommend; it cannot authorise.
That is also why the Auditor must not become an autonomous approval authority. Approval is a different speech act. It binds the organisation. Reconstruction improves the quality of what a human is asked to bind. It does not replace the binding.
Knowledge-graph research is increasingly explicit that much of what systems treat as “data” is really asserted, interpreted material — and that provenance of who asserted what under which authority is essential to assessing reliability.12 Federation fields on Engagement World nodes exist for that reason. The Auditor enforces them under pressure. The Janitor keeps them walkable.
What to do before the next export
- Write the one-page audit contract for your next proposal freeze or ARB pack. Name scope, tests, severities, dispositions, and the non-authority rule.
- Separate the planes in tooling if you can: Janitor PRs for structure; Auditor findings for warrant. Do not let one agent instance grade its own claims.
- Seed at least the five failure classes as fixtures — stale support, circular derivation, unsupported extension, missing coverage, predating receipt — so “green” means something mechanical.
- On every material design change, demand the revalidation set before anyone reuses old receipts.
- Keep humans on disposition. If your system can close a Critical finding without a named owner, you did not build an Auditor. You built an unsupervised editor with better vocabulary.
North star
Stop asking whether the graph looks maintained. Ask whether its consequential claims can still be independently reconstructed and challenged — without letting the challenger silently rewrite the world.
The Engagement World is how multi-person, multi-agent delivery compounds on shared understanding.1 The Janitor keeps that world from rotting into a graveyard. The Auditor keeps it from hardening into a lie. They are partners. They are not the same job. And only one of them is allowed to rearrange the furniture.
References
- Scott Farrell / LeverageAI. “Engagement World: The Project Reality the Slide Deck Pretended to Hold.” — Parent doctrine for the project-bounded Engagement World and Scribe/Janitor/Auditor roles; this article extends the assurance boundary without re-teaching the lifecycle. https://leverageai.com.au/wp-content/media/articles/170-engagement-world.html
- Scott Farrell / LeverageAI. “Designing Loops, Not Prompts.” — Knowledge-graveyard / Janitor doctrine: append-only loops grow larger and dumber without subtractive consolidation; navigability is not warrant. https://leverageai.com.au/wp-content/media/articles/64-designing-loops-not-prompts.html
- W3C. “PROV-DM: The PROV Data Model.” W3C Recommendation, 30 April 2013. — “Provenance is information about entities, activities, and people involved in producing a piece of data or thing, which can be used to form assessments about its quality, reliability or trustworthiness.” Also: provenance of information is “crucial in deciding whether information is to be trusted.” https://www.w3.org/TR/prov-dm/
- Datafold. “Data integrity vs. data quality.” — Industry distinction often confused in practice: integrity concerns structural consistency; quality concerns fitness for use / correctness for decisions. Used here as analogy only for Janitor (structure) vs Auditor (warrant). https://www.datafold.com/blog/data-integrity-vs-data-quality/
- Scott Farrell / LeverageAI. “The Institutional Linter.” — Independence as a graph property; correlated checkers that share one upstream feed create the appearance of multi-line assurance without independent warrant (paraphrase, not a source-turn quotation). https://leverageai.com.au/wp-content/media/articles/137-institutional-linter.html
- Ojewale et al. “Audit Trails for Accountability in Large Language Models.” arXiv:2601.20727, 2026. — LLM audit trails as chronological, tamper-evident, context-rich ledgers linking technical provenance with governance records so organisations can reconstruct what changed, when, and who authorised it. https://arxiv.org/abs/2601.20727
- Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., et al. “Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing.” ACM FAccT, 2020. — End-to-end internal algorithmic auditing across the development lifecycle to support audit integrity and close accountability gaps. https://arxiv.org/abs/2001.00973
- Scott Farrell / LeverageAI. “Elastic Assurance.” — Exploratory assurance produces evidence-backed findings rather than silent formal status changes; findings plane stays off the operational hot path. https://leverageai.com.au/wp-content/media/articles/136-elastic-assurance.html
- Scott Farrell / LeverageAI. “Witness Not Oracle.” — Evidence packages return claim + exhibit + resolvable pointer + confession of what could not be verified; witnesses you can check, not oracles you must trust. https://leverageai.com.au/wp-content/media/articles/93-witness-not-oracle.html
- Scott Farrell / LeverageAI. “The Model Is Not the Memory.” — Governance traces externalise version-pinned paths through admissible knowledge rather than relying on model chain-of-thought as the audit object. https://leverageai.com.au/wp-content/media/articles/68-the-model-is-not-the-memory.html
- Scott Farrell / LeverageAI. “Hidden Gates.” — Share intent, keep diagnostic rubrics from becoming gameable targets; independent review without self-grading. https://leverageai.com.au/wp-content/media/articles/94-hidden-gates.html
- “Provenance-Enhanced Statements in Knowledge Graphs.” arXiv HTML 2606.15246, 2026. — Documented, verifiable provenance as fundamental to KG quality assurance; authority and epistemic status of asserted material (capta) matter for reliability assessments. https://arxiv.org/html/2606.15246