Project memory · Human–AI delivery
Engagement World: The Project Reality the Slide Deck Pretended to Hold
Multi-person, multi-agent delivery compounds only when the project owns a persistent, federated, provenance-bearing Engagement World between durable institutional kernels and ephemeral task rooms — so every act can change shared understanding without mistaking project inference for canonical truth.
In brief
- Chat is too transient, a file dump too flat, and the firm wiki too permanent — the missing layer is an Engagement World.
- Stack: source → institutional truth → engagement understanding → task room → model attention.
- Federate territories without erasing origin or authority; derived joins stay subordinate to supports.
- Coordinate through a claim/evidence/decision graph; Scribe integrates, Janitor structures, Auditor asks: is this world still justified?
- Lifecycle ends in client ownership and gated promotion — not silent elevation of task-local guesses into doctrine.
On day one of a delivery programme, the opening slide is sometimes a Venn diagram. Three circles: inherited strategy recommendations, an outsourced “way of working,” and something that gestures at governance. The overlap is labelled the holy grail. Nobody can say what is changing, for whom, by what mechanism, with what evidence. The diagram is not synthesis. It is a political settlement among documents.
That scene is not rare. It is what happens when teams treat the slide deck as if it held organisational thinking. Decks hold assertions arranged spatially, fixed snapshots, conclusions with the reasoning flattened away, and weak provenance. They can communicate a point in time. They cannot answer the adult questions: which source supports this claim; what changed after the workshop; which alternative was rejected; who disagreed; which test proved it; what would invalidate it.
If you have sat through that kickoff, you already know the pain. Independent research keeps finding a related pattern in AI delivery: projects fail less because models are weak and more because problem intent is miscommunicated, metrics are wrong, or the solution does not fit the workflow and context.3,4,5 The deck is simply the older, more prestigious form of the same disease — a recommendation that externalises verification to the people who must live with it.
The reframe
They began with three documents and tried to infer a reality. An Engagement World begins with evidence and determines which documents survive.
The missing layer
Most teams already have pieces of a memory stack. They have raw sources — transcripts, tickets, code, contracts. They may have a durable wiki or knowledge base for firm doctrine. They have chats and agent sessions. What they lack is an honest place for this:
What have we collectively learned, proposed, rejected, built and verified about this pursuit so far?
That question is not answered by institutional memory (too permanent, wrong authority boundary) and not answered by a single meeting’s context window (too temporary, wrong scope). Between them sits a mesoscale object:
The corrected stack
| Layer | Role |
|---|---|
| Bronze / sources | Immutable records: documents, code, transcripts, client data |
| Institutional kernels | Durable firm, client, or capability truth |
| Engagement World | Private, evolving synthesis for this pursuit or delivery |
| Task World | Temporary intent-shaped room for one act |
| Active context | What one model invocation can attend to now |
Source → institutional truth → engagement understanding → task-specific room → model attention
An Engagement World is a private, project-bounded, provenance-bearing wiki assembled from several authoritative knowledge territories and continuously enriched by a human–AI team for the life of the engagement. It can live for months without pretending its claims are permanent organisational truth. Task rooms still die after a design review. The active context stays disposable. Selected learning promotes upward only under review.
Name the four words carefully, or the whole stack collapses into mush:
- Kernel — what the organisation knows durably.
- Engagement World — what this project team currently understands about this pursuit.
- Task World — what this actor needs present for this particular act.
- Context window — what the current model can attend to right now.
The Task World is already a named object in prior work: intent compiles a temporary, provenance-bearing room from durable sources, inspectable through a manifest, dead by default.11 The Engagement World does not replace that. It is the evolving substrate from which many task worlds are compiled — sales call, architecture review, security review, build sprint, handover — each with a different intent, each inheriting the current engagement state rather than restarting from a static client folder.
What the slide deck was pretending to be
People use decks as if they were the team’s shared mind. Relational databases, by contrast, hold known entities and predefined relationships well, but struggle with emergent meaning, unresolved interpretation, contradictions, and provisional hypotheses. Natural-language claims with typed edges sit in the middle: humans can correct a claim directly; the next AI session behaves differently; the same object can be working surface and audit route. That is the practical force of a legible “third substrate” for organisational learning — not a philosophy lecture, a coordination requirement.observed pattern
So the Engagement World is not “AI-enabled PowerPoint.” It is closer to the executable semantic model that the slide deck, meeting notes, project folders and databases were all incomplete views of. Proposals, architecture diagrams, executive briefs, backlogs, test reports and attestations become generated views over that world — not rival sources of truth.
W3C’s provenance model already states why this matters for trust: provenance is information about entities, activities, and people involved in producing a piece of data or thing, used to assess quality, reliability or trustworthiness.1 Professional delivery still ships million-dollar recommendations with less structure than a weather-data pipeline. Derivation — constructing a new entity from prior ones under influence — is a first-class relation in PROV;2 in most project folders it is a paragraph that begins with “we believe.”
Federation without false authority
An Engagement World is not a filter over one mega-wiki. It is a federated synthesis graph. It assembles claims from capability doctrine, firm knowledge, client facts, public research, code, conversation, and live sources — while each node retains origin and authority.
[[client.payment-modernisation]]
origin: client-kernel
authority: source-backed
[[consultancy.sap-capability]]
origin: firm-kernel
authority: firm-canonical
[[framework.bi-for-soft-data]]
origin: leverageai-kernel
authority: licensed-doctrine
[[opportunity.soft-data-decision-layer]]
origin: engagement-world
authority: derived
supports:
- client.payment-modernisation
- consultancy.sap-capability
- framework.bi-for-soft-data
That last node is not “in” any source wiki. It exists because the engagement joined the worlds. That is where much of the value lives — and where most of the risk lives if you erase the joins. A derived claim must remain typed as derived, linked to supports, invalidated when supports change, and ranked below source-backed material. That discipline is the same economic boundary already used for task-local synthesis: persist the expensive, reusable and integrative; generate or expire the volatile and task-local.11,12
Worked failure the auditor exists to catch
The proposal assumes a usable cloud landing zone. No client source has established it. Three “independent” opportunities all hang on an unverified procurement rail. A test receipt predates a security-boundary change. A “client priority” originated as consultant inference and was never confirmed. These are not style issues. They are epistemic defects.
The team’s shared blackboard
Multi-agent systems introduce hard problems of coordination, evaluation and reliability.6 Production multi-agent work is stateful; errors compound; restarting from scratch is often the wrong recovery model — systems need to resume from durable checkpoints.8 Token cost is real: multi-agent research systems can consume on the order of fifteen times a chat interaction’s tokens, so architecture must buy compression and continuity, not pure parallelism theatre.7
The Engagement World is how a professional-services team buys continuity. Account consultant, architect, security specialist, research agent, coding agent, governance agent, and retail chat clients do not need perfect memory of one another’s sessions. They coordinate by reading and writing the same externalised world. A meeting adds outcomes. A research agent adds a public signal. An architect revises a boundary claim. A coding system attaches test receipts. An auditor flags unsupported assumptions. A proposal run records rejected alternatives.
The unit of continuity is not a file. It is the evolving network of claims, relationships, evidence, decisions, absences and unresolved tensions.
Public patterns already point this direction. Karpathy’s LLM Wiki pattern argues that RAG re-derives knowledge on every question, while a maintained wiki compounds: knowledge is compiled once and kept current.9 Engagement World is that instinct under professional constraints — project bounds, multi-authority federation, client handover, and an auditor that goes beyond “lint for orphans.”
Human views and AI views
Same substrate, different morphologies. Humans rarely want a thousand-node graph. They want:
- Account — people, history, commercial position, next moves
- Opportunity — candidates, value, fit, evidence, blockers
- Architecture — systems, boundaries, decisions
- Decision — proposal, alternatives, evidence, approval status
- Timeline — how understanding evolved
- Receipts — what has been demonstrated
- Unknowns — open questions, absences, required probes
AI views need compact maps, typed page bodies, source authority, neighbourhoods, access scope, version state, parent intent, open hypotheses, and a task-world manifest for the current act. The important design claim: humans and AI are not running separate knowledge systems. They are rendering the same learned state differently. Retail chat remains a doorway into that state — not the memory authority itself.
Three maintenance roles
Ingestion alone is not enough. An engagement needs three distinct jobs.
1. Scribe (integrator)
Absorbs meetings, files, research, code, client comments, tool outputs, decisions and tests. Proposes new claims and edges with source pointers. Does not silently promote inference to fact.
2. Janitor (structure)
Keeps the world lean and navigable: merge duplicates, split bloated pages, mark supersession, turn repeated prose relationships into typed edges, identify cold material, preserve contradictions rather than averaging them into false consensus. The public “lint the wiki” instinct — contradictions, stale claims, orphans, missing cross-references — lives here.10
3. Auditor (justification)
Asks a different question than the Janitor:
Is this world still justified?
Not “are you sure?” — reconstruct the current claim from its evidence path and report any mismatch. Check that supports resolve, sources were actually opened, receipts still match the claim’s as-at date, mandatory regions were not skipped, “independent” conclusions do not share one unverified upstream assumption, and task-local inference has not been re-labelled as client fact.
Four failure classes the auditor must catch in any serious implementation designed:
| Failure | What it looks like |
|---|---|
| Stale support | Support changed or was superseded; dependent claim still asserted |
| Circular derivation | Claims support each other with no external bronze or kernel root |
| Unsupported extension | New commitment appears without supports or known-absence handling |
| Predating receipt | Evidence timestamp precedes a design change that invalidates it |
That is structural governance, not vibes-based auditing where the model narrates confidence after the fact.
Lifecycle: seed, grow, stabilise, branch, close
A Task World defaults to death. An Engagement World needs a governed project lifecycle.
Seed. Assemble the initial world from client kernel, firm kernel, capability doctrine, public research, and initial intent. Declare access scope and as-at dates.
Grow. Every substantive interaction may add evidence, hypotheses, decisions, artefacts, paths and receipts. More meaningful participation can create semantic accretion rather than only coordination tax — if and only if contributions land on the shared graph.
Stabilise. Janitor and Auditor runs continuously. Coherence and justification are separate quality bars.
Branch. Commercial, architecture, security, data and delivery workstreams may keep lens-specific derived material while sharing canonical entities and decisions.
Close. At the end of the engagement:
- preserve the private engagement archive;
- hand the client its owned project world — a living semantic model of what was built and why;
- promote firm-reusable learning into the consultancy kernel only under human gate;
- promote appropriately abstracted doctrine into capability kernels only after de-identification;
- expire volatile task-local interpretations;
- retain receipts and outcomes.
Automatic promotion of client or task-local inference into a firm kernel is not learning. It is contamination with good marketing copy.
Minimum operating checklist (next engagement)
- Create
engagement_id; seed entities for client, constraints, open questions. - Require origin + authority on every claim; type derived joins with
supports[]. - Compile task worlds for each major act from the engagement state + intent manifest.
- Expose at least two human views (e.g. decision + unknowns) and two AI views (map + manifest).
- Run Auditor before any external commitment (proposal, ARB pack, production cut).
- At close: client package, expiry pass, promotion queue with human sign-off only.
What must later be proven — without inventing results
A framework that cannot fail a test is a slogan. The full ebook’s proof burden is real; Stage A states it honestly. None of the following are claimed as already shipped here:
| Proof item | Status |
|---|---|
| Executable prototype: origin/authority fields, derived joins, task-world compilation, ≥2 human and ≥2 AI views | designed |
| Multi-role case: meeting, research agent, architect, coding agent, auditor over time on one world | designed |
| Auditor tests for stale, circular, unsupported, and predating-receipt failures | designed |
| Close/promotion exercise: client ownership, archival, de-identification, expiry, human-gated upward learning | designed |
| Comparison vs the project’s actual slide/file trail on onboarding time, contradiction detection, change impact | hypothesized |
| Field pattern: recommendations and Venn “synthesis” without organisational grounding or verification path | observed |
The engagement itself should become the evidence that the method works — but only with receipts, not with a prettier narrative of success.
What this is not
This is not a project-management feature list. It is not a request to dump every chat into the firm wiki. It is not a replacement for institutional kernels or for temporary task rooms. It is not a consultancy go-to-market plan. Adjacent work covers durable wikis, intent-conditioned task worlds, walk telemetry and third-substrate learning in depth; this piece owns the project-bounded world that makes multi-person multi-agent delivery compound without lying about authority.
If you already have a “project wiki,” ask four questions: Do claims carry origin and authority? Can an auditor reconstruct supports? Do task rooms compile from current engagement state rather than from a zip file of decks? At close, can the client own the world without your private kernel remaining a hidden dependency? If the answers are no, you have a document pile with a search bar.
Start with one claim
Pick the next external commitment your team will make — a proposal assertion, an architecture boundary, a “client priority.” Write it as a node. Attach origin, authority, supports, and known absences. Ask whether a hostile auditor could reconstruct it from evidence that still resolves. If not, you have discovered why the deck felt solid and the implementation felt like fog.
The slide deck promised a shared project mind. An Engagement World is how you stop pretending and start building one — private to the pursuit, honest about authority, usable by humans and agents, and closable without either amnesia or contamination.
Next step
On your current engagement, open an engagement_id and migrate five load-bearing claims into origin/authority/supports form. Run one audit pass before the next steering committee. If you want the fuller blueprint — schema, roles, views, and the executable proof plan — follow the Engagement World ebook as it lands.
Scott Farrell · leverageai.com.au · scott@leverageai.com.au
References
- W3C. “PROV-DM: The PROV Data Model.” W3C Recommendation. — “Provenance is information about entities, activities, and people involved in producing a piece of data or thing, which can be used to form assessments about its quality, reliability or trustworthiness.” https://www.w3.org/TR/prov-dm/
- W3C. “PROV-DM: The PROV Data Model.” — “A derivation is a transformation of an entity into another, an update of an entity resulting in a new one, or the construction of a new entity based on a pre-existing entity.” https://www.w3.org/TR/prov-dm/
- RAND Corporation (Ryseff, De Bruhl, Newberry). “The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed.” — “By some estimates, more than 80 percent of AI projects fail—twice the rate of failure for information technology projects that do not involve AI.” https://www.rand.org/pubs/research_reports/RRA2680-1.html
- RAND Corporation. Same report. — “First, industry stakeholders often misunderstand—or miscommunicate—what problem needs to be solved using AI. Too often, trained AI models are deployed that have been optimized for the wrong metrics or do not fit into the overall business workflow and context.” https://www.rand.org/pubs/research_reports/RRA2680-1.html
- RAND Corporation. Same report. — “Misunderstandings and miscommunications about the intent and purpose of the project are the most common reasons for AI project failure.” https://www.rand.org/pubs/research_reports/RRA2680-1.html
- Anthropic Engineering. “How we built our multi-agent research system.” — “Systems with multiple agents introduce new challenges in agent coordination, evaluation, and reliability.” https://www.anthropic.com/engineering/multi-agent-research-system
- Anthropic Engineering. Same article. — “agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats.” https://www.anthropic.com/engineering/multi-agent-research-system
- Anthropic Engineering. Same article. — “Agents are stateful and errors compound… When errors occur, we can't just restart from the beginning… Instead, we built systems that can resume from where the agent was when the errors occurred.” https://www.anthropic.com/engineering/multi-agent-research-system
- Andrej Karpathy. “LLM Wiki.” — “Instead of just retrieving from raw documents at query time, the LLM incrementally builds and maintains a persistent wiki… The knowledge is compiled once and then kept current, not re-derived on every query.” https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f
- Andrej Karpathy. “LLM Wiki.” — “Lint. Periodically, ask the LLM to health-check the wiki. Look for: contradictions between pages, stale claims that newer sources have superseded, orphan pages with no inbound links…” https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f
- Scott Farrell / LeverageAI. “Intent-Conditioned Task World.” — Temporary provenance-bearing world compiled by intent; inspectable manifest; default expiry; persist integrative meaning. https://leverageai.com.au/wp-content/media/articles/160-intent-conditioned-task-world.html
- Scott Farrell / LeverageAI. “File Back the Walk.” — Answer and path are separate assets; hard multi-hop answers filed as derived cache with supports. https://leverageai.com.au/wp-content/media/articles/80-file-back-the-walk.html