Engagement World
The Project Reality the Slide Deck Pretended to Hold
After Reading This Ebook, You Will:
- ✓ Name the mesoscale Engagement World between institutional kernels and task rooms
- ✓ Federate claims with origin, authority, and supports—without false synthesis
- ✓ Run Scribe / Janitor / Auditor and catch stale, circular, unsupported, and predating-receipt failures
- ✓ Close with client ownership, expiry, and human-gated promotion—without contamination
TL;DR
- • Chat is too transient, folders too flat, firm wikis too permanent. Multi-person multi-agent work needs an Engagement World.
- • Stack: source → institutional truth → engagement understanding → task room → model attention.
- • Federate territories; keep origin/authority; derived joins stay subordinate to supports.
- • Scribe integrates, Janitor structures, Auditor asks: is this world still justified?
-
•
Executable proof suite under
ew_proto/: 10/10 tests passed this session (executed, not invented).
The Deck That Pretended to Hold Reality
Day one. Three circles. A holy grail that cannot answer why the project exists.
On day one of a serious delivery programme, the opening slide is sometimes a Venn diagram.
Three circles. One holds inherited strategy recommendations. Another holds an outsourced “way of working.” The third gestures at governance without quite saying control, authority, or obligation. The overlap is labelled the holy grail. The room is invited to believe that the middle is a solution.
It is not.
Those three artefacts are not sets in a shared universe. One is advice. One is an operating model. One is a control environment. You cannot overlap them and discover a project. To create a meaningful intersection you would first need to translate each artefact into claims and criteria: what problem is asserted, what organisational evidence supports it, which outcomes are meant, where the models conflict, what is retained or rejected, what first test could falsify the theory. Without that work, the centre of the diagram is an empty patch of slides used as a substitute for synthesis.
The most charitable reading still fails. “We need to reconcile the recommendation, the existing way of working, and governance” is a reasonable starting brief. It is not an answer. Implement only what all three already agree on, and the intersection may be trivial or empty. Combine all three as a conjunction, and you may smuggle direct contradictions. Use all three as inputs to design something new, and you need trade-offs, a causal model, and evidence—none of which a Venn supplies.
So the middle is usually performing a political function. The expensive recommendation remains authoritative. The existing operating model remains valid. Governance has been “acknowledged.” Nobody’s prior work has to be declared wrong. Nobody has to reopen whether the recommendation survives contact with the organisation. The project has been constituted around reconciling documents, not solving a demonstrated problem. That is the opposite of adult work.
What the deck pretends to be
People use slide decks as if they held organisational thinking. In practice they mostly hold assertions arranged spatially, fixed snapshots, conclusions with the reasoning flattened away, and weak provenance. A deck can communicate a point in time. It cannot naturally answer:
- Which source supports this claim?
- What changed after the workshop?
- Which alternative was rejected, and why?
- What other decisions rely on this assumption?
- Is this still current?
- Who disagreed?
- Which code implements it?
- What test proved it?
- What would invalidate it?
- How should this look for a salesperson versus an architect?
When those questions have no home, teams invent social substitutes: the latest deck, the loudest voice, the prestige of the firm that “said so,” the folder full of PDFs nobody can join. Advice becomes a person speaking a document. There are no necessary receipts behind it. There is no delivery mechanism. There is no test against the real world.
Implementation teams then inherit an unresolved obligation: interpret the vague recommendation, reconstruct missing evidence, push it through governance, design the actual system, manage organisational resistance, implement it, and discover whether the original idea was any good—after the recommendation has already been monetised. The consultant completed an engagement by handing over a document. The client inherited the expensive half.
We think that bargain is broken. Rubber is supposed to hit the road. Receipts are supposed to exist. Governance is supposed to be an input, not a surprise later. Rejected alternatives are supposed to remain visible. Success is supposed to mean tested change in the real world—not acceptance of a presentation.
Evidence label — observed
The day-one Venn and “person speaking a document” pattern are field observations from source delivery narrative (generalised client and firm names). They are not a controlled multi-site study. Treat them as diagnostic pattern, not population statistics.
Not only a consulting pathology
AI delivery fails in related ways. Independent interview research from RAND finds that, by some estimates, more than eighty percent of AI projects fail—about twice the rate of non-AI IT projects—and that the leading root causes are organisational, not mystical model limits.1 Stakeholders often misunderstand or miscommunicate what problem needs to be solved; models get optimised for the wrong metrics or do not fit the workflow and context. Miscommunications about intent and purpose are reported as the most common reason projects fail.1
That matters for this book because a project that cannot hold a stable, inspectable account of problem, alternatives, constraints, and evidence is not going to be rescued by a larger context window or a more fluent agent. It will fail in the same place the Venn failed: at the shared understanding of what is true, what is provisional, and what is allowed. The deck is simply the older, more prestigious form of the same disease—a recommendation that externalises verification to the people who must live with it.
The reader question
What knowledge object should a human–AI project team share when a chat is too transient, a document repository too flat, and the institutional wiki too permanent?
That is the question this book answers. Not “which PM tool?” Not “which model?” Not “how do I staff a forward-deployed practice?” The object: a project-bounded world that can be defined, governed, rendered, audited, and closed without mistaking project inference for canonical truth.
Thesis
Multi-person, multi-agent delivery compounds only when the project owns a persistent, federated, provenance-bearing Engagement World between durable institutional kernels and ephemeral task worlds—so every act can change shared understanding without mistaking project inference for canonical truth.
After this book you should be able to:
- define the Engagement World and place it in a five-layer memory stack;
- federate claims with origin and authority instead of erasing joins;
- coordinate humans and agents through a claim/evidence/decision graph;
- run Scribe, Janitor, and Auditor as distinct jobs;
- compile task rooms from engagement state for different acts;
- close with client ownership, expiry, and human-gated promotion;
- and read honest evidence labels: designed, executed, observed, hypothesized.
What this book will not do
It will not re-derive the durable institutional wiki, the Intent-Conditioned Task World, the Third Substrate, or File Back the Walk in full. Those appear as cameos where the boundary matters.
It will not treat every chat restatement as permanent knowledge.
It will not recommend automatic promotion of client or task-local inference into a firm or vendor kernel.
It will not become a project-management platform feature list.
It will not sell you a consultancy go-to-market for a forward-deployed practice. The knowledge object is the product of this book; the practice packaging is someone else’s brief.
How proof will work here
Later chapters report an executable prototype under ew_proto/. Where tests ran this session, claims are labelled executed. Where the narrative comes from delivery experience in the source material, claims are labelled observed. Where effects are plausible but unmeasured—for example onboarding time versus a slide/file trail—claims are labelled hypothesized. Where structure is specified for software that must still be run, claims are labelled designed. Invented pass rates and fictional case metrics are not allowed.
Key Takeaways
- The slide deck often performs political settlement among documents, not synthesis.
- Adult questions need a durable home for supports, rejects, absences, and tests.
- AI project failure research points first at intent and problem miscommunication—not model magic.
- This book owns the mesoscale Engagement World between kernels and task rooms.
- Evidence language is part of the doctrine: no invented results.
Chapter 2 names the missing layer those documents were pretending to be—and why chat, folders, and firm wikis cannot substitute for it.
The Missing Layer
Chat dies. Folders lie flat. Firm wikis freeze the wrong things.
Chapter 1 diagnosed the deck that pretends to hold project reality. The natural next move is to ask where understanding should live instead. Teams already try three homes. All three fail—for different structural reasons.
Three wrong homes
Chat is too transient. Retail conversations compact, drop tool fidelity, and die when the window or the product decides they should. Multi-agent systems make the problem worse: coordination, evaluation, and reliability become first-class concerns, and stateful work needs resume-from-checkpoint rather than hope-the-thread-survives.2 Multi-agent research systems also consume far more tokens than ordinary chat—on the order of fifteen times in published engineering analysis—so architecture must buy continuity and compression, not pure parallelism theatre.2 Chat is a doorway. It is not a memory authority.
The project folder is too flat. It holds files. It does not hold claims, supports, rejected alternatives, known absences, or authority levels. Onboarding means reading a trail of decks and tickets and reconstructing a world that was never compiled. Contradiction detection is social—someone remembers that slide four no longer matches the architecture decision. Change impact is heroic archaeology: open twelve PDFs and guess which assumption still binds.
The institutional wiki is too permanent. Durable firm or client truth is real and necessary. Project inference is not institutional truth. If every workshop hypothesis is written into the firm kernel, you get zombie narratives: sharp one-off interpretations that outlive the deal, the partnership, or the architecture they described. If nothing from the project is allowed near memory, every engagement starts cold and every agent re-pays comprehension forever.
The stack is missing a mesoscale layer: persistent enough to outlive sessions and tools, temporary enough not to contaminate durable kernels by accident.
The corrected stack
| Layer | Role | Persistence |
|---|---|---|
| Bronze / sources | Immutable records: documents, code, transcripts, client data | Permanent source |
| Institutional kernels | Durable firm, client, or capability truth | Permanent, governed |
| Engagement World | Private evolving synthesis for this pursuit or delivery | Project lifecycle |
| Task World | Temporary intent-shaped room for one act | Default death |
| Active context | What one model invocation holds in attention | Invocation |
Source → institutional truth → engagement understanding → task-specific room → model attention
That ordering is load-bearing. The Engagement World can live for months without pretending its claims are permanent organisational truth. A Task World can still expire after a meeting, design review, or analysis. The active context can remain clean and disposable. Selected learning can later promote upward into a consultancy or capability kernel—after review, never by accident.
Misplace a claim and the failure mode is predictable. Put engagement inference in the firm kernel and you freeze a hypothesis as doctrine. Put durable policy only in chat and you re-derive it every Tuesday. Put task-local colour schemes in the engagement forever and the world becomes a junk drawer. The stack is a set of economic and epistemic rules, not a diagram for a slide appendix.
Definition
An Engagement World is a private, project-bounded, provenance-bearing wiki assembled from several authoritative knowledge territories and continuously enriched by the work of a human–AI team for the life of an engagement.
It is not a longer chat. It is not a smarter SharePoint. It is not “the firm wiki with a project tag.” It is not a vector store with a pretty name.
It asks a different question than a Task World. A Task World asks: what needs to be present for this particular decision? An Engagement World asks: what have we collectively learned, proposed, rejected, built, and verified about this client pursuit so far?
It accumulates, among other things:
- client entities and relationships;
- current hypotheses and discovered opportunities;
- stakeholders and constraints;
- architecture and security choices;
- decisions and rejected alternatives;
- code and deployment artefacts;
- tests, receipts, open questions, known absences;
- and outcome evidence.
Each new Task World is compiled from this richer engagement state plus the relevant authoritative kernels. A sales conversation, architecture review, governance review, and implementation sprint do not start from the same static client dump. They inherit the current engagement state and project it through different intents.
Four words that must stay distinct
If you blur these, the architecture collapses into a mood:
- Kernel — what the organisation knows durably.
- Engagement World — what this project team currently understands about this pursuit.
- Task World — what this actor needs present for this particular act.
- Context window — what the current model can attend to right now.
Prior doctrine already owns Task Worlds as temporary, provenance-bearing rooms that default to death, and the wiki as institutional kernel. This book owns the middle object: persistent for one bounded project, federated across authorities, explicitly promotable and expirable at close.
What you are not building
You are not building a second fabric that pretends to be the firm. You are not asking agents to “remember the project” inside weights. You are not hoping retail product memory will coordinate humans and tools. You are building a world-building layer that compiles several organisations’ truths into a shared, auditable project reality—then lets humans and AI work inside it together for the life of the engagement.
Evidence labels
Observed (doctrine): stack and definition crystallised in source collaborative session and brief. Executed (later): the ew_proto suite implements engagement objects, task compilation, and lifecycle states—reported in Parts II–III, not claimed as field production metrics.
Key Takeaways
- Chat, folder, and firm wiki each fail a different persistence requirement.
- Engagement World is the mesoscale between institutional truth and task rooms.
- Keep Kernel / Engagement World / Task World / Context Window distinct.
- Promotion upward is gated; task rooms still default to death.
- Retail chat is a doorway into the world, not the memory authority.
Chapter 3 makes federation precise: every claim keeps an origin and an authority, and derived joins stop lying about who said what.
Federation Without False Authority
One opportunity. Three territories. No erased joins.
Chapter 2 placed the Engagement World in the stack. Placement alone does not stop a project from lying. The most dangerous sentence in a project is the one that sounds integrated and cites nothing.
“The client needs a soft-data decision layer on top of their payment modernisation, and we can deliver it with our SAP capability using our BI doctrine.” That sentence may be brilliant. It may also be three authorities blended into one voice so thoroughly that nobody can later say which part was client fact, which part was firm claim, and which part was licensed method. When the recommendation fails a governance review, the room argues about tone instead of supports.
An Engagement World is federated on purpose. It does not pretend the project has one omniscient narrator.
Many territories, one graph
A pursuit is assembled across territories such as capability or practitioner doctrine, the consulting firm kernel, client institutional knowledge, public research, code repositories, meeting and conversation transcripts, live operational sources, and the engagement’s own accumulating outputs.
Therefore the engagement is not a filter over one mega-wiki. It is a federated synthesis graph. Filters remove. Construction selects, joins, interprets, preserves dissent, names absences, and records provenance. That is the same constructive move Task Worlds make for a single act—applied continuously at project scale.
Canonical fields
Every load-bearing node carries at least:
| Field | Meaning |
|---|---|
origin | Which territory produced this node |
authority | How much weight it may carry (source-backed, firm-canonical, licensed-doctrine, derived, inferred-unconfirmed, secondary-source, client-confirmed) |
supports[] | Explicit links to prior nodes or bronze pointers (required for derived) |
as_at | When this claim was justified |
status | active / superseded / rejected / expired |
Authority is not a vibe. It is a type. Derived material ranks below source-backed material. Inferred-unconfirmed never pretends to be client-confirmed. Secondary public research never quietly becomes client bronze.
Worked join
[[client.payment-modernisation]]
origin: client-kernel
authority: source-backed
[[consultancy.sap-capability]]
origin: firm-kernel
authority: firm-canonical
[[framework.bi-for-soft-data]]
origin: capability-kernel
authority: licensed-doctrine
[[opportunity.soft-data-decision-layer]]
origin: engagement-world
authority: derived
supports:
- client.payment-modernisation
- consultancy.sap-capability
- framework.bi-for-soft-data
The opportunity node is not “in” any source wiki. It exists because the engagement joined the worlds. That is where much of the value lives—and where most of the fraud-adjacent risk lives if you erase the lineage. The join is the product. The flattened sentence is the failure mode.
W3C’s provenance model states the general rule in domain-agnostic language: provenance is information about entities, activities, and people involved in producing a piece of data or thing, usable to assess quality, reliability, or trustworthiness.3 Derivation is the construction of a new entity based on pre-existing entities.3 Professional delivery still ships high-stakes recommendations with less structure than that. The Engagement World refuses that free pass.
What federation forbids
Erased joins. If a proposal states a landing-zone assumption without a client source, the graph must show a hole—not a confident paragraph. Known absences are first-class facts, not awkward silences.
Silent elevation. A consultant inference that “the client prioritises Marketplace private offers” remains inferred-unconfirmed until a client-confirmed receipt exists. Promoting it to firm doctrine because it was useful in a workshop is contamination with good storytelling.
Circular glory. Three opportunity slides that all depend on the same unverified procurement rail are not three independent confirmations. They are one assumption with marketing copies. Chapter 7’s Auditor treats circular derivation as a machine-checkable fault.
Operational habit
Before any external commitment—proposal, ARB pack, production cutover—walk the load-bearing claims and ask: can I name origin, authority, supports, and as-at for each? If not, you are about to export a Venn diagram with better fonts. Federation is not bureaucracy. It is the minimum structure that lets multi-person multi-agent work compound without laundering authority.
Key Takeaways
- Federation keeps origin and authority typed on every load-bearing claim.
- Derived joins are engagement-local value—only if supports survive.
- Prototype executed: three-origin join + reject empty supports.
- Known absences and inferred-unconfirmed are honest states, not defects to hide.
Chapter 4 assigns the jobs that keep the graph honest: Scribe, Janitor, and Auditor.
Scribe, Janitor, Auditor
Federation without maintenance is a graph that rots. “Are you sure?” is not an operation.
Chapter 3 made origin, authority, and supports first-class. That schema only matters if someone keeps writing it, cleaning it, and testing whether it still holds. When a claim looks shaky, most teams ask a model—or a person—to sound more confident. That is vibes-based auditing. You trade structural governance for a story invented after the fact.
“Are you sure?” is not an operation. “Reconstruct this claim from its evidence path and report any mismatch” is an operation. An Engagement World needs three distinct jobs that match that distinction.
Coordination without mutual memory
Multi-agent systems introduce hard problems of coordination, evaluation, and reliability. Agents are stateful; errors compound; restart-from-scratch is often the wrong recovery.2 Humans and agents cannot keep complete memory of one another’s sessions and still scale. The reliable move is not cleverer inter-agent chat. It is durability: place state somewhere that outlives any single actor, and coordinate through that medium.
The Engagement World is that medium at professional-services scale. Account lead, architect, security specialist, research agent, coding agent, governance agent, and retail chat clients read and write the same externalised graph. A meeting adds outcomes. A research agent adds a public signal. An architect revises a boundary. A coding system attaches test receipts. An auditor flags unsupported assumptions. The world survives each participant’s session.
The unit of continuity is not a file. It is the evolving network of claims, relationships, evidence, decisions, absences, and unresolved tensions—the same fields Chapter 3 required, kept alive by roles rather than by heroics.
Role 1 — Scribe (integrator)
The Scribe absorbs new material: meetings, files, research, code, client comments, tool outputs, decisions, tests. It proposes claims and edges with source pointers. It does not silently promote inference to fact. It does not tidy away disagreement to make the deck prettier.
Operational habits for the Scribe:
- Every new claim gets
origin,authority, andas_atat write time—not later when memory is soft. - Client-confirmed and inferred-unconfirmed are different types; never collapse them for “clarity.”
- Known absences are written as nodes, not left as awkward gaps in a summary.
- Rejected alternatives keep a reason, so the next actor does not re-open a closed door without new evidence.
- Derived opportunity nodes only appear after supports exist—never as free-floating synthesis prose.
Scribe success metric: new evidence is reachable as typed structure, not trapped as another PDF in the folder. If the graph cannot answer “what did Tuesday’s workshop actually confirm?” without rereading a transcript, the Scribe failed.
Role 2 — Janitor (structure)
The Janitor keeps the world lean and navigable: merge duplicates, split bloated pages, mark supersession, turn repeated prose relationships into typed edges, identify cold material, preserve contradictions as edges rather than averaging them into false consensus.
Public LLM-wiki practice already includes a lint pass for contradictions, stale claims, orphans, and missing cross-references—cousin work to the Janitor.4 Useful. Incomplete. Lint can tell you the graph is messy. It cannot tell you the graph is still justified. That boundary matters: a tidy world can still be wrong.
Operational habits for the Janitor:
- Supersession is explicit: old nodes move to
superseded; dependents become the Auditor’s problem if left active. - Cold pages are marked or archived, not left as false scent for the next agent walk.
- Contradictions stay as typed edges; averaging them into a “balanced” paragraph is how dissent dies.
- Duplicate opportunity nodes are merged without erasing divergent supports.
Janitor question: Is this world still lean and coherent?
Role 3 — Auditor (justification)
The Auditor asks: Is this world still justified? That is not cleaning. That is epistemic integrity:
- Does every derived claim have valid supports?
- Were those sources actually opened?
- Do cited receipts still resolve?
- Have source changes invalidated dependents?
- Did an agent ignore a mandatory region?
- Do several “independent” conclusions share one upstream assumption?
- Does a claimed implementation exist?
- Did the test exercise the path being claimed?
- Has task-local inference been relabelled as client fact?
- Is the as-at date still valid?
Example findings the Auditor is for:
- “The proposal claims deployability into the client AWS environment, but no landing-zone evidence has been opened.”
- “Three opportunity recommendations depend on an unverified Marketplace private-offer assumption.”
- “The architecture page cites a test receipt produced before the security boundary changed.”
- “This ‘client priority’ originated as consultant inference and has never been confirmed.”
These are not style notes. They are the same failure families Chapter 3 forbade under federation—stale supports after supersession, circular glory, unsupported extension, predating receipts. Chapter 7 turns four of them into machine-checkable fixtures that the prototype executed this session.
When to run which role
| Trigger | Primary role | Why |
|---|---|---|
| Workshop, research drop, code merge | Scribe | New evidence must enter typed |
| Weekly hygiene / graph bloat | Janitor | Navigability before the next walk |
| Proposal, ARB, production cut, close | Auditor | External commitment requires justification |
| Support supersession event | Janitor then Auditor | Structure first; dependents re-justified |
Collapse modes
| Collapse | What breaks |
|---|---|
| Scribe only | Graph fills with unvalidated exhaust |
| Janitor only | Pretty structure, unjustified claims |
| Auditor only | Findings with nowhere durable to write |
| Human hero only | Single point of failure; no multi-tool continuity |
The Engagement World is richer through use when participation produces verified relationships—not merely more documents. That is semantic accretion: more meaningful actors improve the graph instead of only raising coordination tax. Collapsing the three roles recreates Chapter 1’s deck problem with more tools: confident surfaces, weak reconstruction.
Evidence labels
Observed: role triad and Auditor question from source doctrine. Executed: multi-role graph writes and four Auditor fault classes in ew_proto (Parts II–III). Role schedules and human operating rhythm remain designed process guidance, not a measured staffing study.
Key Takeaways
- Coordinate through a durable claim graph, not mutual session memory.
- Scribe integrates; Janitor structures; Auditor justifies—three questions, not three job titles on one person’s forehead.
- Vibes-based “are you sure?” is not reconstruction from supports.
- Run Auditor before external commitments; run Scribe when evidence arrives; run Janitor before the graph becomes unwalkable.
Part II puts five actors on one world and forces the Auditor to earn its name in sequence.
Five Actors, One World
Chapter 4 named the roles. This chapter runs them in sequence on one engagement.
The fantasy of modern delivery is that everyone will stay aligned because they are all “on the project.” They are not in the same room. They are not in the same tool. They are not in the same context window. The account lead runs discovery in a retail chat. A research agent walks public signals overnight. An architect revises boundaries in a design session. A coding agent opens a repository on a different machine. An auditor—if you have one—arrives late and is asked to bless a proposal that has already been emotionally accepted.
If continuity lives in heads and decks, you get telephone. If continuity lives in an Engagement World, you get the semantic accretion Chapter 4 promised: each meaningful act adds verified structure the others can inherit without replaying every session. Scribe writes first. Janitor would tidy if the graph bloated. Auditor closes the loop before external commitment.
This chapter is the multi-role case. The story is designed as a generalised pursuit. The ordered graph build is executed in ew_proto (meeting seed → research → architect → coding → auditor on a clean world).
Case frame (generalised)
Engagement ID: northline-paymod-2026. Pursuit: help a mid-market payments firm decide and deliver a soft-data decision layer beside an existing modernisation programme. Territories: client kernel facts, firm capability claims, licensed doctrine, public research, code, meetings.
No real client names. No brand theatre. The point is the graph, not the logo. The same federation rules from Chapter 3 apply on every write.
Beat 1 — Meeting (human Scribe input)
A discovery workshop produces:
- confirmed pain: exception-heavy payment exceptions queue;
- disputed interpretation: whether “AI triage” or “deterministic routing first” is the first wedge;
- stakeholder map and a known absence: no documented cloud landing-zone readiness;
- rejected alternative: full ERP replacement (out of scope, high political cost).
The Scribe writes claims with origin: meeting-transcript, authorities split between client-confirmed and inferred-unconfirmed, and an explicit known-absence: landing-zone-evidence. The disputed wedge is not forced into false consensus. The ERP rejection keeps a reason so later actors do not re-open it without new evidence.
Beat 2 — Research agent
A research agent opens public signals about payment-ops automation patterns and procurement constraints. It adds nodes with origin: public-research, authority: secondary-source, linked to open questions—not smuggled in as client fact.
The critical discipline: it must not “fill” the landing-zone absence with a generic cloud best-practices page and call the absence closed. Known absence remains first-class until client bronze appears. Secondary research can inform probes; it cannot mint client-confirmed deployability.
Beat 3 — Architect
A senior architect revises the system boundary between exception queue and core ledger, and a security constraint: no agent write-path to settlement without human gate. Derived architecture claims keep supports to client systems inventory and firm patterns.
If the architect’s claim depends on a usable landing zone, the node must still show the absence or a new client-backed support. Architecture opinion is not bronze. This is Chapter 3’s authority typing under pressure: fluency does not upgrade authority.
Beat 4 — Coding agent
A coding agent implements a vertical slice: intake classification harness with tests. It attaches code pointers, test receipts with timestamps, a claim that the green path works under fixture set A, and a walk log of what it opened. Answer and path are both assets.
The coding agent does not get to declare the business case proven. It declares what the harness demonstrated. Scope discipline here prevents the classic failure: a passing unit test used as political proof of a strategy slide.
Beat 5 — Auditor
Before a proposal gate, the Auditor reconstructs load-bearing claims: opportunity node supports, landing-zone assumption, security boundary versus test receipt dates, any circular cluster of “independent” recommendations.
On the clean multi-role world, the executed Auditor report passes. That is intentional. Faults—stale support, circular derivation, unsupported extension, predating receipt—are injected in separate fixtures in Chapter 7 so the clean case proves accretion without smuggling defects, and the fault suite proves detection.
What the sequence proves operationally
Nobody needed complete memory of everyone else’s session. They needed a world that accepted heterogeneous writes, preserved authority, kept absences visible, and allowed challenge without social theatre. That is the shared blackboard at work. Multi-agent coordination lessons from production research systems—plan persistence, checkpointing, evaluation of outcomes rather than one true path—apply here as professional requirements, not lab curiosities.2
Inspectability after each beat matters. If only the final proposal is visible, you cannot tell whether the research agent closed an absence illegally, whether the architect skipped supports, or whether the coding agent overclaimed. The Engagement World makes those intermediate states first-class.
Evidence labels
Designed: Northline story and generalised actors. Executed: ordered graph mutations and clean Auditor pass in ew_proto. Not claimed: real client outcomes, onboarding-time savings, or field A/B against slide trails.
Key Takeaways
- Five writers, one
engagement_id, inspectable state after each beat. - Research must not invent client bronze to close absences.
- Coding proves harness behaviour; Auditor guards external commitments.
- Clean multi-role pass and fault fixtures are separate proofs—do not conflate them.
Chapter 6 shows how each act compiles a different temporary room from the same evolving world—sales versus architecture, without dumping the entire graph into every context window.
Compile the Room You Need
Chapter 5 left five actors writing one graph. Now each act needs a different room from that same graph.
A sales conversation and an architecture review can both be “about Northline” and still need different operative realities. The sales act needs commercial position, opportunity status, blockers, relationship risks, and what has been promised. The architecture act needs systems, boundaries, dependencies, threat notes, and which decisions are already locked.
Dumping the entire Engagement World into both contexts is not honesty—it is attention vandalism. Starting both from a static client folder is amnesia with better branding. The multi-role case only compounds if each subsequent act can inherit engagement state without either drowning in it or restarting cold.
The Engagement World is the persistent project understanding. Each act compiles a Task World: a temporary, intent-shaped room that defaults to death when the act ends.
Compilation, not vibes
Prior doctrine already defines an Intent-Conditioned Task World: intent acts as a relevance and meaning filter that constructs a provenance-bearing temporary world from durable sources, shaped by role, date, access, and attention. Filters only remove. Construction selects, joins, interprets, raises resolution, preserves dissent, identifies absence, and establishes provenance.
This book’s addition is the input set. The Task World is compiled from:
- the current Engagement World state (claims with origin, authority, supports);
- relevant kernel slices (firm / client / capability) under access scope;
- the parent intent, role, and lens for this act;
- the as-at date and expiry policy.
So strategy, architecture, governance, build, and handover do not restart from the same cold dump. They inherit engagement understanding and project it. Origin and authority travel into the room; they are not flattened into undifferentiated “context.”
Manifest: inspectable or it did not happen
If a walk leaves only a chat transcript and a final answer, you have an outcome without a world. Promote the residue into a declared structure. At minimum a manager who was not in the room should be able to answer:
- What was the parent intent?
- Which role and lens conditioned the room?
- As-at which date?
- What access scope applied?
- Which pages and edges were included, and why?
- Which minority paths were preserved?
- What known absences remained?
- Which sources actually opened?
- When does this room expire?
If those answers do not exist, you had a lucky dump that fit in a window. That is the same failure mode as Chapter 1’s deck, moved into agent tooling.
Parent intent
Role and declared lens
As-at date
Access scope
Boot profile / entry
Questions / probes
Pages and edges included (+ why)
Convergence hotspots
Important minority paths
Counter-lens
Known absences
Sources opened
Derived interpretations
Expiry / promotion policy
Same engagement, different rooms
| Act | Intent shape | Emphasised regions |
|---|---|---|
| Sales / pursuit | Value, fit, blockers, next ask | Opportunity, account, unknowns |
| Architecture review | Boundaries, dependencies, decisions | Architecture, constraints, rejects |
| Security review | Controls, authority, data movement | Security edges, absences, receipts |
| Build sprint | Green path, harness, failure modes | Artefacts, tests, open implementations |
| Handover | Ownership, runbooks, residual risk | Decisions, receipts, open questions |
On Northline, a sales compile should surface the soft-data opportunity, rejected ERP path, and landing-zone absence as a commercial blocker. An architecture compile should surface the exception-queue boundary, settlement write gate, and the same absence as a technical risk. Same engagement_id. Different parent intents. Different working sets. Shared authority typing.
Answer, path, and what not to persist
When an agent walks the world to answer a question, it produces two assets: the answer (compiled claim) and the path (what was opened, abandoned, never touched). Summaries delete negative space. Dead ends are informative. Hard multi-hop answers may be filed back as derived cache with supports; trivial lookups should not litter the permanent fabric.
Inside an Engagement World, that rule keeps the project from paying repeatedly to rediscover the same join—and from filling the world with conversational restatements. Task-local colour for a workshop board is not an engagement claim. A multi-hop join between client pain, firm capability, and doctrine is.
Compiler contract
The executable prototype exposes the equivalent of:
compile_task_world(
engagement_id,
intent,
role,
lens,
as_at,
access_scope
) -> TaskWorldManifest + working_set
It must pull from engagement claims with origin and authority intact. It must expire. It must not write task-local interpretation back into firm or capability kernels without a human promotion gate (Chapter 8). Generated artefacts—proposal sections, diagrams, backlogs, test reports—are views over these rooms and the parent world, not rival sources of truth.
Failure modes of bad compilation
- Whole-graph dump: attention diffusion; model re-litigates settled rejects.
- Cold folder restart: loses multi-role accretion from Chapter 5.
- Authority flattening: secondary research looks like client confirmation.
- Immortal task rooms: temporary colour becomes permanent doctrine.
- No manifest: cannot audit what the agent thought it inhabited.
Evidence labels
Cameo: Task World and File Back the Walk doctrine (not re-derived). Executed: two-intent compilation with distinct manifests in ew_proto. Designed: full field catalogue of production manifests beyond the prototype minimum. Hypothesized: token savings versus whole-graph dumps in production retail clients (not measured here).
Key Takeaways
- Compile task rooms from engagement + kernels + intent; do not dump or restart cold.
- Manifests make rooms reviewable; lucky dumps do not.
- Prototype executed two different intent compilations from one world.
- Answer and path are both assets; only hard joins earn persistence.
- Task rooms default to death so engagement and kernels stay clean.
Chapter 7 attacks the room and the parent world with the only question that matters before you ship a commitment: is this still justified?
Is This World Still Justified?
Chapter 6 compiled the room. Before you export from that room, reconstruct the path.
The architecture page says the security boundary was redesigned on Thursday. The test receipt attached to “boundary verified” was produced on Tuesday. Everyone is busy. The proposal is due. Someone will argue that the spirit of the test still holds. The Auditor’s job is to refuse spirit and demand path.
That is the same discipline Chapter 4 named and Chapter 5 deferred on the clean multi-role world: justification is not cleanliness, not confidence, and not a green unit test used as political cover.
The operation
Not:
Are you sure?
But:
Reconstruct the current claim from its evidence path and report any mismatch.
A validation run has a fixed shape:
- Select a claim or proposed external output (proposal paragraph, architecture assertion, deployability claim).
- Traverse its
supportsgraph. - Restore relevant source versions / as-at snapshots.
- Confirm cited passages or project artefacts exist.
- Check mandatory evidence classes for this claim type.
- Compare present synthesis with evidence.
- Identify unsupported extension, stale support, circular derivation, or predating receipt.
- Emit a proof-carrying validation report.
Tool calls and opened pages are evidence of cognition. If the agent used the world, the trace can show it. If it freelanced, the trace can show that too. Asking the model to narrate hidden reasoning is a different—and weaker—sport. Reconstruction is mechanical enough to be repeated; judgment remains where authority requires a person.
Four failures the Auditor must catch
These are the brief’s minimum epistemic tests. Fixtures fail closed. The prototype executed each class.
1. Stale support
Shape: Support node superseded or content-changed; dependent derived claim still active with old confidence.
Fixture sketch: region policy moves from single-region to dual-region; a DR claim still asserts the old policy as current. Auditor must flag dependency invalidation.
Executed PASS — auditor_stale_support · ew_proto/reports/fault_stale_support.json
2. Circular derivation
Shape: Claims support each other with no external bronze or kernel root.
Fixture sketch: opportunity A supports B, B supports value-case C, C supports A. No client or firm source-backed root. Auditor must report the circular cluster with node ids.
Executed PASS — auditor_circular_derivation · fault_circular.json
3. Unsupported extension
Shape: New commitment appears without supports or explicit known-absence handling—often presented as fact.
Fixture sketch: proposal text claims “client can accept Marketplace private offers” with no support edge, no known-absence ticket, authority mis-typed. Auditor must fail the claim as unsupported extension.
Executed PASS — auditor_unsupported_extension · fault_unsupported.json
4. Predating receipt
Shape: Evidence timestamp precedes a design change that invalidates what the receipt was attached to.
Fixture sketch: security boundary v2 lands at T2; receipt stamped T1 still linked as proof of v2. Auditor must flag predating receipt. This is the Thursday/Tuesday story made machine-checkable.
Executed PASS — auditor_predating_receipt · fault_predating_receipt.json
How the Auditor should speak
When the Auditor speaks, it should sound like structure—not like a vibe:
- “Derived claim X lists supports […]; support S2 was superseded at …; claim remains active.”
- “Claims A↔B↔C form a cycle with no source-backed root.”
- “External text introduces commitment Q with zero supports and no known-absence.”
- “Receipt R timestamp predates design change D linked as proven.”
What the Auditor is not
- It is not the Janitor. A world can be neat and unjustified.
- It is not a style guide. Tone is not a support.
- It is not a human’s social veto alone—though humans still own consequential acceptance. Structure makes disagreement inspectable; it does not replace judgment.
- It is not “run the model over the folder and hope.”
- It is not a substitute for Chapter 6’s manifest. You still need to know which room was compiled; the Auditor then tests whether claims in that room (and the parent world) remain justified.
When to run it
Run the Auditor before external commitments: proposal send, ARB pack, production cut, client workshop that asserts deployability, and engagement close. After support supersession, run Janitor first to mark structure, then Auditor to re-justify dependents. On the clean multi-role path (Chapter 5), expect green. On fault fixtures, expect red with typed findings. Do not narrate fictional pass rates beyond the machine-checkable suite.
Evidence labels
Executed: all four fault fixtures PASS in ew_proto/reports/test_results.json (re-runnable). Designed: broader mandatory evidence classes for enterprise claim types beyond the four fixtures. Not claimed: false-positive rates in production, or human Auditor staffing ratios.
Key Takeaways
- Reconstruction beats confidence theatre.
- Four failure classes are machine-checkable fixtures, not slogans.
- All four auditor tests executed green in the prototype suite.
- Clean multi-role green and fault red are complementary proofs.
- Humans still accept consequential exports; the Auditor makes the path inspectable.
Chapter 8 takes a justified—or honestly incomplete—world to the only finish line that matters: close without amnesia or contamination, with the client owning what was learned.
Seed to Close
Chapter 7 asked whether the world is still justified. Close asks who owns it when the pursuit ends.
On the last day of an engagement, two bad outcomes are common.
Amnesia: the team ships code and a slide pack; understanding stays in people’s heads and chat logs. The next team rebuilds the world from folklore. Every multi-role write from Chapter 5 evaporates. Every compiled room from Chapter 6 leaves no durable residue. The Auditor’s findings become hallway memory.
Contamination: every workshop inference is shovelled into the firm wiki as if it were doctrine. Six months later agents treat a dead partnership narrative as current institutional belief. Task-local colour becomes firm canon. Client-confidential specifics leak into a practitioner kernel. The zombie lens from Chapter 2’s warning becomes institutional software.
An Engagement World exists to avoid both. It persists for the pursuit—and it ends on purpose. Close is not a calendar event. It is a governed procedure that preserves evidence, transfers ownership, gates promotion, and expires the volatile.
Lifecycle
Seed
Assemble the initial world from client kernel slices under access rules, firm kernel slices, capability or licensed doctrine, public research where relevant, initial intent, and known absences. Declare engagement_id, access scope, and as-at. Do not pretend the seed is complete. Seed honesty is a known-absence list—the same first-class absences Chapter 3 and Chapter 5 required so research agents cannot invent client bronze.
Grow
Every substantive interaction may add evidence, hypotheses, decisions, artefacts, paths, and receipts. Growth is not “more files.” Growth is more verified relationships. Scribe integrates; humans judge; agents contribute with typed authority. This is the multi-role accretion of Chapter 5 applied across the engagement calendar.
Stabilise
Janitor and Auditor run as continuous hygiene—not a pre-demo panic. Coherence and justification remain separate bars (Chapter 4). A pretty graph that is wrong is still wrong. A justified graph that cannot be navigated will be re-derived by the next agent at cost.
Branch
Workstreams may hold lens-specific derived material—commercial, architecture, security, data, change, delivery—while sharing canonical entities and locked decisions. Branches are not parallel fictions; they are lenses with explicit divergence edges where they disagree. Task rooms (Chapter 6) compile from the shared trunk plus the lens they need, then die.
Close
At the end of the engagement, execute a governed close:
- Preserve the complete private engagement archive, including walks and receipts.
- Hand the client its owned project world—the living semantic model of what was built and why—not only code and PDFs.
- Promote firm-reusable learning into the consultancy kernel only under human gate.
- Promote appropriately abstracted doctrine into capability kernels only after de-identification.
- Expire volatile task-local interpretations so they are not queryable as current.
- Retain receipts and outcomes for audit and later comparison.
The economic rule, applied at engagement scale: persist the expensive, reusable, and integrative; generate or expire the volatile and task-local.
Client ownership is not a slogan
If the client cannot operate the world without your private kernel as a hidden dependency, you did not deliver a world. You delivered a leash. That violates the takeaway of the brief: the reader must be able to close an Engagement World while preserving origin, authority, disagreement, and evidence under client ownership.
Client package should include, at minimum:
- compiled problem and capability map;
- decisions and architecture with supports;
- known constraints and rejects;
- test and evaluation evidence;
- operational playbooks where they exist;
- open questions and residual risks;
- governance receipts and Auditor reports that still resolve.
Support, extension, audit, onboarding, and change-impact analysis get cheaper when the semantic model travels with the software. Whether that onboarding is faster than a slide/file trail remains hypothesized until measured; the ownership requirement does not wait on that measurement.
Promotion is a gate, not a vibe
Promotion candidates enter a queue with:
- what is proposed for which kernel;
- evidence that it is integrative and reusable;
- de-identification checklist;
- human owner and decision;
- reject ledger when “no” is the right outcome.
Automatic promotion of client-confidential specifics into a practitioner kernel is a contract and trust failure. Automatic promotion of task-local narrative into firm doctrine is how zombie lenses are born. Approve and reject are both success outcomes of a working gate. A queue that only ever approves is not a gate; it is a conveyor.
Run the Auditor (Chapter 7) on promotion candidates before human decision. A derived join that no longer has valid supports should not enter a kernel because it felt useful in a workshop.
Close exercise — executed
| Step | Required demonstration |
|---|---|
| Archive | Immutable snapshot of engagement graph + walks |
| Client package | Export client-owned world without firm-private nodes as current truth |
| De-identify | Strip or abstract client markers from promotion candidates |
| Expiry | Task-local nodes marked expired; not queryable as current |
| Promotion | At least one human-approved promote and one human-rejected promote |
| Receipts | Close report lists what moved where and what died |
The finish line
The visible product may be software, an AI capability, a governance system, or a process change. The Engagement World is also a delivered asset. You do not leave only code and documentation. You leave a living semantic model of the capability and why it exists—and you leave the firm smarter only where a human gate allowed learning to rise.
They began with three documents and tried to infer a reality. At close, the documents that survive should be generated views over a world the client owns, with supports that still reconstruct, and with task-local vapour expired on purpose.
Evidence labels
Executed: close archive, client export, one approve, one reject, expiry in ew_proto. Designed: full enterprise de-identification checklists and multi-tenant legal packaging. Hypothesized: onboarding-time and contradiction-detection gains versus slide/file trails. Observed: amnesia and contamination failure modes as field patterns in source narrative (generalised).
Key Takeaways
- Close is a governed procedure, not a calendar event.
- Client owns a living semantic model alongside software—not a leash to your private kernel.
- Promotion approve and reject are both success outcomes of a working gate.
- Expire the volatile; archive the evidence; promote only what is integrative and de-identified.
- Auditor runs before promotion decisions, not after regret.
Chapter 9 inventories the human and AI views, the full prototype contract, and the honest evidence labels that keep this book falsifiable.
Views, Prototype, Proof Contract
Chapter 8 closed ownership. This chapter names how the world is rendered—and what was actually proven.
A substrate nobody can render is a graveyard of claims. A renderer with no substrate is a slide deck with ambitions. Human and AI actors need different morphologies over the same Engagement World. They do not need separate truths. Views are how Decision and Unknowns stay human-legible while Map and Manifest stay agent-operable—without reintroducing Chapter 1’s rival documents as sources of truth.
Mandatory views — executed
| Audience | View | Must show |
|---|---|---|
| Human | Decision | Proposal, alternatives, evidence, approval status, supports |
| Human | Unknowns | Open questions, known absences, required probes |
| AI | Map | Compact typed map of current engagement neighbourhood |
| AI | Manifest | Task-world manifest for the active intent (Chapter 6) |
These four are the minimum proof pair set. They force both “what are we deciding?” and “what do we not know?” for humans, and both “where am I?” and “what room was compiled?” for agents. Without Unknowns, absences disappear into optimism. Without Manifest, agents cannot be audited for the room they inhabited.
Full catalogue (reference, not re-proved)
Humans may also use account, opportunity, architecture, delivery, timeline, and receipts views. AI may also need access scope, version state, boot profile, open hypotheses, and authority-filtered neighbourhoods. Build what the engagement needs; prove the four mandatory ones first. Additional views are generated projections over the same federated graph from Chapter 3—not new sources of truth.
Prototype inventory
The working suite lives under ew_proto/ (not rewritten in this revision):
ew_proto/
schema/claim.json
tools/ew_core.py
tools/run_all.py
case/multi_role_script.md
reports/test_results.json
reports/execution_log.txt
reports/derived_join.json
reports/task_world_*.json
reports/view_*.json
reports/fault_*.json
reports/close_*.json
Executed acceptance (this session)
| Requirement | Chapter | Result | Label |
|---|---|---|---|
| Origin/authority + derived joins | 3 | PASS | executed |
| Task compilation (two intents) | 6 | PASS | executed |
| Decision + Unknowns + Map + Manifest | 9 | PASS | executed |
| Multi-role case | 5 | PASS | executed |
| Auditor four faults | 7 | PASS (4/4) | executed |
| Close / promote / expire / archive | 8 | PASS | executed |
| Onboarding vs slide-trail metrics | 8–9 | Not measured | hypothesized |
Command: python3 ew_proto/tools/run_all.py → exit 0; 10/10 tests passed (test_results.json). Re-running the suite must not require rewriting fixtures for this book’s proof claim to hold.
Evidence-label protocol
| Label | Meaning |
|---|---|
| designed | Specified structure, test, or process not yet run, or beyond current fixtures |
| executed | Command or assertion ran with machine-checkable report |
| observed | Field pattern in source narrative or external research with provenance |
| hypothesized | Plausible effect without measurement (e.g. onboarding-time vs slide trail) |
Honesty is part of the doctrine. Invented percentages and fictional client ROI are out of scope. External stats stay tied to research sources (RAND, W3C PROV, Anthropic engineering notes, Karpathy LLM Wiki lint pattern). Self-frameworks appear as author voice with leverageai REFs for the references list, not as fake third-party proof.
Monday checklist
- Is there an
engagement_idand a living graph, or only a folder? - Do load-bearing claims carry origin, authority, and supports?
- What are the top unknowns and known absences?
- Which task world is active, and when does it expire?
- What did the last Auditor run find?
- What is queued for promotion vs expiry at close?
- Can the client own this world without our private kernel?
If you cannot answer these, you are back in deck cosplay—three circles and a holy grail with no reconstructable path.
What this book refuses
It refuses to restate the thesis as a slogan in every chapter. Part I named it; Parts II and III applied it.
It refuses FDE go-to-market cosplay, PM feature laundry lists, and automatic promotion of inference.
It refuses to hide the difference between designed, executed, observed, and hypothesized. A world that cannot fail a test is not a world—it is a brand. This one failed four Auditor fixtures on purpose and passed ten suite tests on the record.
Closing
The slide deck promised a shared project mind. An Engagement World is how you stop pretending:
- private to the pursuit;
- honest about authority;
- usable by humans and agents through different views;
- maintained by Scribe, Janitor, and Auditor;
- closable without amnesia or contamination.
They began with documents and tried to infer a reality. Build the world—then decide which documents survive.
Key Takeaways
- Four views are the minimum human/AI morphology pair set.
- Proof burden for the brief’s executable requirements: 10/10 executed in
ew_proto. - Slide-trail comparison remains hypothesized—do not invent metrics.
- Monday checklist returns the team to reconstructable work, not deck theatre.
References & Sources
The evidence base behind every claim — primary research, industry analysis, and technical specifications
Research Methodology
This ebook draws on primary research from standards bodies, independent research firms, enterprise technology vendors, and consulting firms. Statistics cited throughout have been cross-referenced against primary sources.
Frameworks and interpretive analysis developed by Scott Farrell / LeverageAI are listed separately below — these represent the practitioner lens through which external research is interpreted, and are not cited inline to avoid self-promotional appearance.
Primary Research & Standards Bodies
RAND Corporation — The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed [1]
>80% AI project failure estimate; 2x non-AI IT
https://www.rand.org/pubs/research_reports/RRA2680-1.html
W3C — PROV-DM: The PROV Data Model [3]
Provenance for quality, reliability, trustworthiness
https://www.w3.org/TR/prov-dm/
Industry Analysis & Vendor Research
Anthropic Engineering — How we built our multi-agent research system [2]
Multi-agent coordination, evaluation, reliability; stateful resume
https://www.anthropic.com/engineering/multi-agent-research-system
Andrej Karpathy — LLM Wiki [4]
Lint for contradictions, stale claims, orphans
https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f
LeverageAI / Scott Farrell — Practitioner Frameworks
The interpretive frameworks, architectural patterns, and practitioner analysis in this ebook were developed through enterprise AI transformation consulting. The articles below are the underlying thinking behind those frameworks. They are listed here for transparency and further exploration — not cited inline, as this is the author's own analytical voice.
Scott Farrell — Intent-Conditioned Task World
Temporary provenance-bearing world; default death
https://leverageai.com.au/wp-content/media/articles/160-intent-conditioned-task-world.html
Scott Farrell — File Back the Walk
Answer and path are separate assets
https://leverageai.com.au/wp-content/media/articles/80-file-back-the-walk.html
About This Reference List
Compiled July 2026. All URLs verified at time of compilation. Regulatory documents and standards specifications are subject to revision — check primary sources for the most current versions.
Some links to academic papers and vendor research may require free registration. Government and standards body publications are freely accessible.