Route-Invariant Grounding
Many Paths, Same Genba
How do you know your agentic knowledge base is robust — or that the model is answering from its priors and the graph is decoration?
After reading this ebook, you will:
- ✓ Measure four quantities across a family of walks — path variance, evidence invariance, answer invariance, traversal efficiency
- ✓ Read the variance table and know when path variance is healthy route resilience
- ✓ Run the omission test that detects model-prior substitution
- ✓ Apply a weekly operating battery with honest labels for executed vs proposed work
TL;DR
- • Path identity is the wrong score. Evidence invariance across many walks of the same question is the right one.
- • Three instruments, three units. Path Testing grades the agent; telemetry mutates the map; Route-Invariant Grounding grades the substrate.
- • Loose front, strict back. Genchi genbutsu is a terminal attractor; capstones intensify the pull.
- • Five runs: 27–57% pairwise citation overlap; six stable doctrines. Paths varied; the genba held.
- • Omission is the sharp test. Confident same answer with no evidence is model-prior substitution.
The Graph or the Prior?
Path identity is the wrong quality score. The question is whether the substrate is doing load-bearing work — or whether the model is answering from its priors and the graph is decoration.
Imagine two agents handed the same hard business question. Same graph. Same tools. Same model family, if you like. Both return a sharp, plausible answer that would pass a casual human skim. Grade the final string and you cannot tell them apart. That is the trap this book is built to make expensive.
Walk with agent A
Agent A does not begin by answering. It begins by orienting. A map or search call lands a framework page that names the territory. An edge leads to a concept that sharpens the question. Another edge, or a source-descent rule, pulls the walker into the actual ebook chapter that carries the load-bearing claim. The log shows the canonical page opening before synthesis. Claims in the draft answer can be matched, line by line, to passages that were actually read. When a qualification appears in the source, it appears in the answer. When a contradiction sits on an opened page, the answer either holds it or says why it prefers one side.
None of this is theatrical. The walk may still be slightly messy. Agent A may open one wrong neighbour, backtrack, and re-enter through a better edge. Path length is not the prize. What matters is that the answer was built after the genba was visited — the actual place where the doctrine lives as source text, not as a rumour in a secondary mention.
An observer reading agent A’s log can answer inspection questions without guessing: Did the agent open the canonical source before synthesising? Which edges were available and used? Which abandoned branches were left without use? Those questions are how Path Testing makes the good path visible.
Walk with agent B
Agent B looks busier and, for a while, more impressive. It fires a broad search. It opens three loosely related pages, none of them the chapter that should carry the claim. It skims a later article that alludes to the idea without restating the warrant. It never follows the edge that would have forced source descent. Then it writes.
The prose is fluent. The structure is clean. The conclusion happens to match what a well-read practitioner would have said, because the model’s pre-training already contains a compressed version of the public shape of the idea. The graph was present. The graph was not load-bearing. It was scenery.
What would an observer have had to notice to catch agent B? Not the final paragraph. The final paragraph is the camouflage. The tells sit in the walk:
- Canonical page never opened, or opened after the answer was already drafted in the model’s working set.
- Citations, if any, point at secondary mentions that cannot support the strongest claims made.
- Qualifications that live only in the source chapter do not appear in the answer, because the chapter was never read.
- Search thrash and abandoned branches dominate the log; typed edges that would have led home sit unused.
- The same answer would likely survive if the pivotal page were removed tomorrow — a foreshadowing of the omission test in Chapter 8.
Grade the path and the two agents land in different cells of a matrix we already own: good or bad path crossed with right or wrong answer. Agent B occupies the dangerous cell — bad path, right answer; unsupported lucky success that works until priors stop holding. Output-only suites celebrate it. Path-aware suites treat it as a failure of architecture dressed up as a win.
From one walk to a family of walks
Path Testing answers a per-walk question: was this journey sound? That is necessary and still not enough for the question this book owns. You can have one beautiful walk and a fragile substrate. You can have path variance that looks noisy on a dashboard and a substrate that is actually healthy. You need a third cut: many walks of the same question under deliberately varied wording, routing parameters and traversal order, scored not for path identity but for whether the load-bearing evidence region and the material answer still hold.
Route-Invariant Grounding
A knowledge substrate is robust when reasonable variations in question wording, routing parameters and traversal order still lead the agent to an equivalent load-bearing evidence region and a materially consistent answer — without depending on one prescribed path or an ungrounded model prior.
Many paths, same genba.
Genba here is not poetry for its own sake. It is the Japanese business term for the actual place — the source chapters and claims that must be visited if the answer is to be grounded rather than performed. Many paths may enter the territory. The same genba must still be reached.
Why this is urgent now
Agent evaluation platforms have converged on capturing full trajectories: tool selection, arguments, loops, retrieval paths, latency and cost — not only final strings.1 That is necessary infrastructure. It also creates a new temptation: treat path identity as the score.
If last week’s “good” walk opened pages A→B→C, this week’s walk that opened A→D→C looks like a regression on a trajectory dashboard. It may not be. It may be healthy route resilience — two bush tracks into the same valley. If your instrumentation is optimising path identity, you are training the system to look like last Tuesday’s favourite model and one prompt style. You will punish legitimate entrances and miss lucky priors that never needed the graph at all.
The quantity worth measuring is evidence invariance across many walks, not path identity within one.
That sentence is the spine of the book. Everything else — the four quantities, the variance table, the terminal attractor, the five-run specimen, the omission test — is apparatus for making the sentence operational.
The reader question
How do you tell whether your agentic knowledge base is genuinely robust, or whether the model is answering from its priors and your graph is decoration?
After this book you should be able to measure four separate quantities across a family of walks, and to run the one test that distinguishes a grounded substrate from a well-dressed model prior. You should also be able to look at a green answer in your own logs the way an observer looks at agent B: not “does this sound right?” but “what was actually opened, and would the answer survive if that evidence disappeared?”
Where we are going
Part I names the third instrument beside Path Testing and walk telemetry, then splits groundedness into four quantities and a diagnostic table. It explains why routes can converge without being identical: genchi genbutsu as a terminal attractor with loose navigation at the front and strict epistemics at the back, intensified by capstone ebooks and boot profiles.
Part II runs the flagship five-run study — executed evidence with real overlap figures — and shows how claim-level support comparison prevents you from fooling yourself with prose similarity.
Part III turns omission, claim-level support and class-level variation into protocols specified tightly enough to run, keeps honest labels on what has and has not been executed, then leaves you with a weekly battery and a hard kill list of neighbouring doctrines this book refuses to steal.
The graph must be load-bearing. The rest of the book is how you prove it — and how you catch the walk that never went to the genba at all.
Three Instruments, Three Units
One walk record supports three different questions. If you do not name the unit of analysis, you build a dashboard and call it architecture.
Every instrument in this chapter can be built from the same walk log: which pages opened, which edges fired, what was abandoned, what was cited, what the final answer claimed. That is the problem and the opportunity. The opportunity is that you do not need three capture systems. The problem is that a team that sets out to measure grounding can end up measuring thrift, or path sameness, or green strings — and not notice until the dashboard stops producing anything actionable.
File Back the Walk already draws the boundary for four instruments that read the walk record. Path Testing grades one walk and returns a verdict on the agent. Dispute tickets label failures with frozen evidence. Observability and provenance audit one execution. Walk telemetry aggregates many walks and returns a graph mutation — missing edges, dead ends, cold pages. Those four instruments are real. This book does not replace them. It adds a fifth column that none of them fully owns.
| Instrument | Question | Unit of analysis | Output |
|---|---|---|---|
| Path Testing | Was this walk sound? | One walk | Verdict on the agent |
| Walk telemetry | What should the map become? | Many walks, in aggregate | Graph mutation |
| Route-Invariant Grounding | Is the substrate robust under variation? | Many walks of the same question under deliberate variation | Verdict on the substrate |
Dispute tickets and execution provenance remain neighbours. They matter. They are not this book’s spine. The three-way contrast above is the one that keeps teams from using a path-quality score as if it were a knowledge-base score.
Path Testing grades the agent
We already own the 2×2: good or bad path crossed with right or wrong answer. The four cells route to different repairs. Good path, right answer is real success — keep it as a regression golden. Good path, wrong answer is synthesis failure — fix judgement, not the map shape. Bad path, wrong answer is navigation failure, often with missing content. Bad path, right answer is the invisible cell: a guess from model priors that never touched the canonical source. Output-only suites celebrate it. Path-aware suites treat it as a failure of architecture dressed up as a win.
This book will not re-derive that matrix. It will stand on it. Path Testing remains the per-walk honesty device. When Chapter 1’s agent B looks green, Path Testing is what makes the walk expensive to ignore. What Path Testing is not, by itself, is a property of the knowledge base under reworded questions, altered routing parameters, and a second model that never saw your house vocabulary.
Per-walk dashboards still matter: path length in steps, search calls, canonical-source reach before synthesis, abandoned branches, time and token cost, answer quality, and the explicit 2×2 cell label. Without the cell label, teams optimise path length into silence or accuracy into luck. With only the cell label, they still cannot say whether the substrate is robust when the question is rephrased tomorrow.
Telemetry grades the map
Walk telemetry asks a different question of many walks: what should change in the graph? Its output is a candidate mutation, not a green checkmark. Repeated dead ends propose edges. Cold pages propose attention or retirement. Missing joins propose structure the corpus already implied but never compiled.
A perfect path over a bad map still leaves a missing-edge ticket. A bad path over a perfect map still fails the 2×2. Neither instrument can produce the other’s finding. That boundary is load-bearing. If you feed family-of-walks invariance numbers into a janitor queue, you will start mutating the graph for the wrong reason. If you feed janitor candidates into a robustness score, you will start calling connectivity a property it is not.
Route-Invariant Grounding grades the substrate
Fix the question — or a small family of deliberate rewordings of the same question. Vary the conditions: search parameters, result depth, prompt style, sometimes the model. Record every walk. Ask whether the load-bearing evidence region stays equivalent and whether the answer stays materially consistent. That is not “did one agent thrash?” and not “should we add an edge?” It is “does the knowledge substrate still do work when the walk is allowed to move?”
Path Testing grades one walk. Telemetry aggregates many walks into a graph mutation. Route-Invariant Grounding grades a family of walks of the same question and returns a verdict on the substrate.
The unit of analysis is the family, not the single trajectory. The deliberate variation is not noise; it is the experimental condition. Without variation you may only be measuring one model learning its own interface. Without a fixed question family you are comparing different tasks and calling the scatter a substrate failure.
Practically, that means you design the family before you run it. Write the core question and the rewordings down. Pin the graph and the tools. Decide which axes are allowed to move — parameters, wording, natural pathing — and which are frozen for this battery. Then score the family as one object. A single green walk filed under “looks robust” is still Path Testing. Five walks of the same question, read together, is the substrate instrument.
Teams that skip that design step usually discover they mixed three questions, two graph versions, and one warm chat context, then argued about the scatter. The instrument did not fail. The unit was never fixed.
Wrong-fix map
If you conflate the instruments you ship the wrong repair. The cost is not academic. It is weeks spent rearranging the wrong surface:
- Agent thrash on a good map → fix tools, boot order, contracts, and synthesis rules. Path Testing territory. Telemetry that “adds more edges” will not teach a model to open the canon first.
- Repeated dead ends across questions → fix the graph. Telemetry territory. A family-of-walks score that only scolds the agent will leave the missing edge in place for the next question.
- Stable answers with no shared evidence under variation → substrate thinness or model-prior luck. This book’s territory. Path Testing on a single lucky walk will not catch the family-level pattern; omission will (Chapter 8).
- Shared evidence with contradictory conclusions → synthesis or judgement instability. Not a map mutation and not automatically a prior. Fix how conflicting claims are presented and scored.
Name the unit before you celebrate the metric. A green string is not a unit. A short path is not a unit. A new edge is not a unit. Agent, map, and substrate are three different objects that happen to share a log format.
One more practical discipline follows from the three units: report them separately even when they are computed from the same log. A weekly ops review that shows only average path length will retrain the organisation toward path identity. A review that shows only open janitor tickets will retrain it toward connectivity. A review that shows only answer green will retrain it toward priors. Put the family row for route-invariant grounding on the same page as the Path Testing cell summary and the telemetry candidate count — three columns, three questions, three next actions. When one column moves, you will know which instrument owns the fix.
If your tooling cannot yet produce a family row, start with a handwritten sheet for one question family. The instrument is the unit of analysis, not the sophistication of the dashboard. A crude family score that names the substrate is healthier than a polished path chart that pretends the substrate was measured.
The next chapter splits the substrate metric itself into four quantities people routinely glue into one word — groundedness — and gives you a diagnostic table for reading them together.
Four Quantities, One Diagnostic Table
Groundedness is a blob. Split it into four measures and read them through a four-row variance table — or you will smuggle priors past the instrumentation.
External evaluation literature already knows that “correct against reality” and “faithful to the retrieved context” are different failures. Faithfulness — groundedness in the strict sense — is defined against observed context, not parametric memory.2 Citation correctness is not genuine reference use; systems can post-rationalise attributions that align with prior beliefs. One study reports that up to 57% of citations can lack faithfulness even when they are correct.3
Hold that as a neighbour. Do not glue it to the five-run figures in Chapter 6. Those are a different quantity: pairwise citation overlap across runs, not faithfulness of a citation within a run. Merging the two “57%s” in conversation is how a careful book becomes a sloppy slide.
The four quantities
| Measure | Meaning |
|---|---|
| Path variance | How different were the pages, edges and ordering across runs? |
| Evidence invariance | How much overlap existed in the load-bearing source pages, chapters and claims? |
| Answer invariance | Did independently decomposed claims reach the same conclusions and qualifications? |
| Traversal efficiency | How much useful territory and source grounding were obtained for the time and payload spent? |
Path variance is not a failure mode by itself. It is what free agents do when the territory has more than one entrance. Compute it from sets of page ids, edge ids, and ordered sequences across the family. High path variance with low evidence variance is often health. Low path variance can mean robustness — or overfit routing that only one model will tolerate.
Evidence invariance is the load-bearing property. Prefer overlap of load-bearing sources and claims over raw citation lists. A page opened and unused is not evidence. A chapter that carries the warrant is. Chapter 7 tightens this to claim-level support; Chapter 6’s first cut uses citation overlap plus doctrine survival because that is what the executed study measured.
Answer invariance is necessary and insufficient. Answers can agree while floating free of evidence. Score conclusions and qualifications: two answers that share a headline but drop opposite caveats are not invariant. Prose similarity is a weak proxy; claim decomposition is the honest one.
Traversal efficiency is an engineering constraint, not the quality score. Useful territory and source anchoring per unit of attention can rise or fall for good reasons. Do not promote it into a single number that replaces the other three.
Key Insight
Call count is not a scoreboard. A better agent may make more calls because deeper investigation is available; a more mature map may need fewer broad searches because typed edges replace guessing. Measure useful territory and source anchoring per unit of attention.
The variance table — worked row by row
Once you have four quantities you can diagnose a family of walks without collapsing them into a vibe. The table is not a poster. It is a decision surface.
| What varies? | What remains stable? | Interpretation |
|---|---|---|
| Path varies | Evidence and answer stable | Healthy route resilience |
| Path and evidence vary | Answer stable | Possibly useful redundancy — or model-prior luck |
| Evidence stable | Answer varies | Synthesis or judgement instability |
| Evidence and answer both vary | — | Retrieval, graph or question ambiguity |
Row one — healthy route resilience
What it looks like in a log: page sequences differ; citation sets only partially overlap; the same load-bearing chapters or doctrines keep appearing; answers share conclusions and the important qualifications. An observer sees different bush tracks and the same valley.
What it means: the substrate is reachable from more than one entrance, and source-descent is still doing terminal work. Path variance is not noise to be erased.
What you do next: keep the question family as a regression. Resist the urge to force one golden path. Improve efficiency only if source anchoring stays intact. Promote this family as a positive control when you change routing.
Row two — the ambiguous cell (and the whole point of this book)
What it looks like in a log: paths differ and the sets of opened sources differ in ways that matter. Load-bearing pages come and go across runs. Yet the final answers still sound the same — same headline conclusions, similar confidence, tidy structure.
What it might mean (two rival stories):
- Useful redundancy / alternate entrances. Different pages legitimately support the same claims. The territory has multiple valid warrants. Claim-level support comparison (Chapter 7) would show shared underlying assertions even when page ids differ.
- Model-prior luck. The answer is stable because the model already believed it. The graph is decoration. Different weak pages are window dressing. Remove a pivotal source (Chapter 8) and the answer barely flinches, without alternate grounded route or honest confession.
This is the cell where prose-level answer invariance becomes actively dangerous. It is the family-level cousin of Path Testing’s bad path + right answer cell. Chapter 1’s agent B, repeated five times with different thrash patterns and the same fluent conclusion, would land here until omission forces a verdict.
What you do next: do not celebrate. Run claim-level support. Then run omission on a page that should be pivotal. If alternate entrances carry the same claims, document the redundancy as reachability (and leave the full taxonomy to a neighbouring doctrine). If the answer survives without evidence, you have model-prior substitution — fix grounding discipline, source-descent enforcement, and synthesis gates, not the path-identity dashboard.
Row three — synthesis or judgement instability
What it looks like in a log: the family keeps opening the same load-bearing region. Citations cluster. Yet conclusions or qualifications swing: one run hedges hard, another states the claim as unconditional; one includes a boundary condition, another erases it.
What it means: navigation is not the bottleneck. Judgement is. Rearranging edges will not fix contradictory synthesis over the same evidence.
What you do next: fix how conflicting claims are presented, how qualifications are required, and how answer schemas force boundary conditions into the open. Score answer invariance at claim level so the swing is visible. Do not punish path variance that was not the problem.
Row four — retrieval, graph or question ambiguity
What it looks like in a log: different sources, different conclusions, little stable core. The family does not agree on where the genba is or what the answer is.
What it means: either the question family is not one task, the graph does not yet have a navigable territory for this intent, or retrieval is too noisy to converge.
What you do next: stabilise the question, inspect whether load-bearing pages exist at all, and only then measure robustness. You cannot score route-invariant grounding on a task that is not yet a task.
Plumbing: cognitive provenance
You cannot score evidence invariance from a model’s post-hoc story about what it “must have read.” You need a record of which pages, claims and edges were actually observed at decision time — external to the model, version-pinned, restorable. That is cognitive provenance: the ability to reconstruct the knowledge world that governed the run.
Final-answer accuracy alone cannot explain how an output was produced, which evidence supported each claim, or whether tool calls were justified.4 The four quantities are how we operationalise that warning for a navigable knowledge graph.
Chapter 4 supplies the causal engine: why paths can vary while evidence still holds when source descent is doing its job at the end of the walk.
Loose Front, Strict Back
Genchi genbutsu is not mere citation hygiene. It is a terminal attractor — and the reason varied paths can still share a genba.
Chapter 3 told you how to read a family of walks. This chapter answers a different question: why would evidence invariance ever appear when path variance is high? Luck is one answer. Architecture is the better one. The architecture has a shape.
Loose navigation at the front, increasingly strict epistemics at the back.
Genchi genbutsu means go and see the actual place. In this stack it is operating doctrine: descend to source chapters before you synthesise. The measurement story needs more than the slogan. It needs to explain how free early movement can coexist with convergent late grounding — and what breaks when you invert that order.
What is allowed to vary early
The early route can enter the territory through many doors. The walker may land first on a framework page that names the region. Or on a concept that restates the principle in different vocabulary. Or on a later update that already carries enough of the earlier shape for a current question. Or on a capstone ebook that consolidates the region. Or on an adjacent project page that made the idea operational in another domain.
That looseness is not sloppiness if the graph is dense with entrances. It is how a reworded question still finds a foothold. A practitioner who asks in house jargon and a vocabulary-naive agent who asks in public language should both be able to enter somewhere. Vocabulary matches wherever it can. Search parameters and bloom settings can change which neighbours are visible first. Traversal order can shuffle. None of that should be treated as automatic failure on a path-identity dashboard.
When we reviewed walk paths under different search strategies, different parameters, and slightly different questions, the path rarely repeated. The outcome quality barely moved. The pivotal chapters still got read. Reword the question a bit and the walker still hits the right material. Genchi genbutsu keeps pulling the walk into the correct final-ebook chapters even when the path in differs — and that rule is directing the walk more than people notice.
That observation is the empirical seed of the attractor claim. It is not yet the five-run specimen with magnitudes (Chapter 6). It is the qualitative pattern that made the property worth naming: path variance without answer collapse, under a source-descent discipline that keeps dragging the walker home.
What is not allowed to stay loose late
A source-descent rule eventually funnels those routes toward actual ebook chapters and source material before synthesis. The model is not free to stop at a gist, a map descriptor, or a secondary mention and call the job done. Those early artefacts are routing signals. They are not evidence. The final answer is not allowed to inherit all that navigational looseness.
Several bush tracks enter the same valley; one marked trail leads down to the site. Use that metaphor once and mean it operationally: early tools may explore; late tools must ground. If your stack is inverted — rigid routing early and soft synthesis late — you get brittle paths and fluent priors. Agents that cannot find the one prescribed entrance thrash; agents that skip grounding still write confidently. Both failure modes look like “the AI is bad” on a string-only eval. Both are design failures under this shape.
Terminal attractor
In this measurement sense, a terminal attractor is a rule or structure that allows free movement early and forces convergence onto load-bearing sources late. Genchi genbutsu is the rule. Capstone ebooks (Chapter 5) are high-density structures that intensify the pull.
Strict late epistemics is not the same as forbidding exploration. It is a gate on synthesis. The walker may wander. The answer may not. That distinction is what lets path variance and evidence invariance coexist. Without the gate, exploration becomes prior-shopping: open anything, then write what the model already believed.
Why one canonical path is the wrong fix
When path variance shows up on a dashboard, the instinct is to prescribe the route: always open these three pages in this order. That instinct is understandable and usually wrong. You do not want every agent on one prescribed sequence of pages. That would mean you over-specified the route for one model and one prompt style.
Reflexive Agent Design’s variation protocol exists because optimising for one session’s habits is overfitting with good PR. Tool order, vocabulary and boot profiles can lock in one model’s quirks until the next generation makes the design look stupid overnight. Over-prescribing the path is not grounding. It is governance cosplay that collapses the moment a utility model or a vocabulary-naive agent arrives.
Natural movement is allowed. Ungrounded synthesis is not. The attractor protects the second rule without destroying the first. If your only tool for “making walks more reliable” is path prescription, you will trade away the very multi-entrance resilience that makes reworded questions work.
How the attractor shows up in scores
If the attractor is working, row one of the variance table becomes common: path varies, evidence and answer hold. Different page sequences still open the same load-bearing chapters before synthesis. Citation sets may only partially overlap; doctrines and claims still recur. That is healthy route resilience, not noise.
If the attractor is missing or optional, you get more of row two and row four. Answers float free of shared sources, or the family never converges on a territory at all. Path Testing still matters per walk — agent B from Chapter 1 is still a bad walk even when the attractor exists for other walks. The attractor is what makes a family of good-enough walks land in the same genba without being clones.
Inspection questions that surface a broken attractor:
- Do walks that enter from different vocabulary still open the same source chapters before synthesis?
- Is source descent enforced, or merely encouraged in a prompt the model can ignore?
- Do secondary pages act as entrances that lead home, or as stopping points that satisfy the model early?
- When a walk never reaches a capstone or canonical chapter, does the answer still claim full confidence?
- If you remove the source-descent instruction for a controlled replay, does path variance stay while evidence invariance collapses?
That last question is a cousin of the omission test in Chapter 8. Omission removes a page. Weakening the attractor removes the duty to visit pages that matter. Both probes ask whether the answer was earned from the graph.
What genchi genbutsu is not
It is not a citation style guide. Decorating an answer with footnote numbers after the fact is not descent. It is not “prefer long documents.” Length without warrant is not genba. It is not a ban on intermediate pages. Framework and concept pages are legitimate entrances; they become a problem only when the walk stops there and still answers with full force.
Established doctrine in this stack already treats source descent as go-to-the-actual-place before synthesis. This chapter does not redefine the Japanese term from first principles. It explains why that doctrine produces route-invariant grounding when the front of the walk is allowed to move: the terminal rule is what keeps the family honest while the entrances stay plural.
Cost order, not continuous panic
You do not need to replay every experiment continuously. The Progressive Evaluation Ladder climbs only when a cheaper rung fails to discriminate: inspect existing walks first, then pinned counterfactual replay, then active generation. Family-of-walks scoring starts as inspection of walks you already have. Omission and multi-model batteries are later rungs. The attractor explains why inspection often already shows convergence; the later rungs test whether that convergence is honest.
Teams that skip inspection and jump straight to multi-model generation often discover they never instrumented opens, never enforced descent, and never distinguished path variance from evidence collapse. The ladder is not bureaucracy. It is cost-ordered honesty.
Chapter 5 tightens the structural half of the attractor: why capstone ebooks and boot profiles make many entrances land in one territory without forcing path identity.
Capstones and Entrances
Final ebooks pull harder than their page count suggests. Boot profiles guarantee reachability without forcing one path. Together they make many entrances land in one territory.
Chapter 4 described the rule: source descent as terminal attractor. Rules need structures to pull against. Two structures matter disproportionately for route-invariant grounding: capstone ebooks and boot profiles. Neither requires every walker to take the same sequence. Both make it more likely that different sequences still share a genba.
Why final ebooks get opened so often
Under a source-descent rule, final ebooks are not merely another copy of the IP. They are high-density, source-oriented attractors. They consolidate a region so every walker does not have to rediscover the full publication history of an idea — the original article, the correction, the application note, the side argument that never got merged.
That consolidation is why they matter more than a page-count would suggest. A walker that enters through a concept page can still be pulled into a capstone that carries the current claim and the pointers back to bronze sources. A walker that enters through an update can still land in a chapter that holds the warrant rather than stopping at the update’s summary. The capstone is not a cache of identical text. It is a compile: structure and density that raw sources do not have on their own.
Key Insight
Capstones matter disproportionately because they sit at the end of the marked trail. Loose navigation can wander; strict epistemics still wants a place dense enough to ground a synthesis.
For measurement, the implication is practical. When you score evidence invariance, weight load-bearing chapters and claims, not every intermediate page. A family of walks may differ on which entrance pages they open and still agree on the capstone region that actually supports the answer. That is healthy path variance with evidence-region stability — row one of the table in Chapter 3.
An observer reading logs should therefore separate three layers of pages: entrances (how the walk found the territory), intermediates (orientation and relationships), and load-bearing sources (where the warrant lives). Optimising for identical entrances is how path-identity dashboards go wrong. Ignoring whether load-bearing sources were opened is how agent B from Chapter 1 stays green.
Boot profiles: many starts, one field
Most organisations eventually invent named kernels — a brand voice document, a house style guide, an operating doctrine — and then watch them drift apart as each restates the others. In a graph, a named kernel becomes a boot profile: a hub page plus a selection bias, not a competing document. It carries almost no content of its own. What it carries is a starting point and a bias about what to surface first, with edges into the regions a given kind of work tends to need.
The deeper shift is what “boot” means. With a monolith, booting is loading the document — a single fixed act. With a graph, you can enter at any page and still reach everything the task needs, because the edges guarantee reachability. The entry point sets where you start and what you see first; it never limits what you can reach.
The kernel stopped being a thing you load and became a field you stand in.
That is why path identity is the wrong success criterion. If every legitimate boot profile must produce the same page sequence, you have turned a field back into a script. What should be stable is reachability into the load-bearing region, not the order of footsteps. A writing profile may surface voice and examples first; a governance profile may surface authority boundaries first; both should still be able to walk to the same capstone claim when the question requires it.
If a flat file is required by some tool at startup, generate it from the graph rather than hand-authoring a rival monolith. Hand-authored kernels reintroduce drift. Generated kernels remain disposable outputs of a field you can stand in from any page.
Entrances without forcing path identity
Multiple semantic entrances — the same principle appearing in different contexts with different vocabulary — are one reason path variance can coexist with evidence invariance. Walk it a different way, hit different pages, and you can still get a clear answer. That is not automatically bloat; sometimes it is how the territory stays reachable from questions that do not share house jargon.
Typed edges orient without forcing one path. Edges that mark development, application, update or supersession tell the walker why a page exists in the neighbourhood. They do not require every walk to traverse the same chain. Canonicality belongs partly to the claim and the question, not simply to one sacred document: for an origin question the original may be required; for a present-doctrine question the update or capstone may be enough; for a change-history question both may be required.
This book will not taxonomise redundancy into entrances, progressive development, cross-domain restatement and accidental duplication. That sort is neighbouring doctrine. The measurement claim here is narrower: if your graph has only one brittle entrance to a load-bearing claim, route-invariant grounding will fail the moment the question is phrased wrong. Capstones and typed edges reduce that brittleness without requiring every walker to take the same sequence.
How structure fails in practice
Three failure patterns show up repeatedly when you read families of walks with structure in mind:
- Entrance without descent. The walker finds a related page and stops. The answer is fluent. The capstone was never opened. Path Testing cell: bad path, right-looking answer.
- Capstone without warrant. The walker opens a dense ebook but never follows pointers to the specific claim that should support the conclusion. Density becomes theatre.
- Rival kernels. Two boot documents restate doctrine differently and drift. Walks that start in different profiles no longer share a genba because the field has split into two monoloths again.
None of these failures is fixed by making the path more identical. They are fixed by making the field more reachable and the late gate stricter: descent before synthesis, one field rather than rival documents, entrances that lead somewhere load-bearing.
What to look for in your own logs
- Do walks that enter from different boot profiles still open the same capstone or source chapters before synthesis?
- Are entrance pages acting as doors or as dead ends that satisfy the model without descent?
- When path sequences differ, does the set of load-bearing chapters still overlap more than the set of peripheral pages?
- If you force one canonical path in tooling, does a second model or a reworded question break immediately?
- When answers agree, is agreement coming from shared capstone claims or from shared model priors with different decorative pages?
Those questions are how structure becomes measurable without inventing a new percentage. They prepare the eye for Chapter 6, where partial citation overlap coexists with doctrine survival — the empirical shape you should expect when entrances are plural and genba is real.
Bridge to proof
If the attractor is real, a family of walks of the same question should show path variance with a stable load-bearing region. Chapter 6 is that specimen: five runs, citation sets that only partially overlap, doctrines that survive every run. Doctrine ends here. Evidence begins.
Five Runs, Same Genba
Primary executed evidence: five runs of one business question, 27–57% pairwise citation overlap, six doctrines that never dropped out. This is the chapter a sceptical reader decides the book on.
Route-invariant grounding is a property statement. Properties need specimens. This chapter is the specimen that was actually run — not a proposed protocol, not a thought experiment. Everything in Part I was scaffolding for reading these numbers without panicking at the wrong one.
Where the walks came from
On a dated replay session we reviewed live walk history from AskUI — Scott’s AWS Marketplace application that instruments agent walks over a knowledge substrate. This is not the third-party UI-automation vendor of a similar name. The figures below are this system’s instrumented telemetry, not industry benchmarks and not a multi-company survey. Production instrumentation from a shipped product is stronger evidence than a lab toy; it is still not a claim about the industry at large.
The work used Claude Code as a co-analyst on stored walks: compare repeated runs, write analysis scripts, inspect what the walker cited and what doctrines survived. No raw session identifiers belong in this book; the point is the measurement, not the ticket number.
What was held fixed, and what was allowed to move
The fixed object was the business question itself — one task, not five different tasks dressed as one family. That is the experimental hinge. If the question drifts, scatter in citations is uninteresting; you compared different intents. The study’s claim is stronger: the same question, walked more than once, still did not replay one golden path.
What moved between runs was the natural variance of a real walker under a live stack: different pathing through the graph, different citation sets, different emphasis in the final synthesis. The wider harness work around the same substrate had already been deliberately varying search strategies, result-depth parameters, and slight rewordings of questions; those qualitative patterns (path shifts, pivotal chapters still read) are the backdrop. The five-run battery is the place where magnitudes land on citation overlap and doctrine survival for one fixed business question.
What was not claimed as the experimental lever for these five runs is a full multi-model class battery. That remains a proposed extension (Chapter 9). What was executed is repeated walks of the same question under the instrumented stack, with enough natural path variance to make the overlap figure interesting.
What the five runs showed
Across five runs of the same business question
Survived in every run despite path variance
Five runs of the same business question produced only 27–57% pairwise citation overlap. The same six main doctrines survived in every run. Paths varied. The load-bearing evidence region held.
The answers were not photocopies of each other. Emphasis still differed usefully: one run was sharpest on pilot mechanics, another on the authority boundary, another on governance structure. The shared doctrinal core did not abandon ship. That is material consistency with room for angle — not robotic identity of prose.
Why 27–57% is the interesting number, not the disappointing one
A path-identity dashboard would read low pairwise citation overlap as failure. That is exactly the wrong reflex. If five walks of the same question produced near-100% citation identity, you might have overfit a route, over-constrained the walker, or simply lucked into one model’s habit. The interesting regime for route-invariant grounding is partial citation overlap combined with stable load-bearing content.
Read the range carefully. Pairwise overlap between runs sat between twenty-seven and fifty-seven percent. Some pairs of runs shared more of their citation set; some shared less. None of that prevented the same six main doctrines from appearing every time. The citation set is a path-sensitive surface. Doctrine survival is a first cut of evidence-region stability.
In the language of Chapter 3: path variance is high; evidence invariance is positive at the doctrine level; answer invariance is material consistency with useful differences of emphasis; traversal efficiency is not the hero metric of this specimen. On the variance table, the family sits near row one — healthy route resilience — pending finer claim-level instrumentation (Chapter 7).
How to read those numbers
- 27–57% citation overlap is path variance made visible, not a failure of grounding.
- Doctrine survival is the first cut of evidence invariance.
- It is not claim-level support comparison — that protocol is proposed in Chapter 7, not reported as completed on this battery.
- Do not merge this range with Wallat et al.’s figure that up to 57% of citations can lack faithfulness even when correct. Different quantity, different study.
What “six stable doctrines” means — and what it does not
The source record states that the same six main doctrines survived every run. It does not hand us a published roster of the six names for this book to reprint as if they were a standard checklist. Honesty requires that limit. What we can say operationally is what doctrine survival is as a measurement move:
- You name the load-bearing doctrines or frameworks that a correct answer to this business question must touch.
- For each run, you mark which of those doctrines appear in the answer and which supporting pages were opened.
- A doctrine “survives” when it is present across the family, not when every run cites the identical page id for it.
- Emphasis may still differ: pilot mechanics, authority boundary, governance structure — different runs can lead with different facets of the same core.
That is why doctrine survival is a stronger signal than prose similarity and a coarser signal than claim-level support. It is the right first cut for an executed study that measured citation overlap and core survival without completing a claim-support matrix.
Map the specimen to the four quantities
- Path variance: high. Citation sets only partially overlap; the walker is not replaying one golden sequence.
- Evidence invariance: first cut, positive. Six doctrines survive every run. Region-level stability, not yet claim-level Jaccard.
- Answer invariance: materially consistent core with useful differences of emphasis and qualification.
- Traversal efficiency: not the hero metric of this specimen. Do not retrofit a call-count victory narrative onto a study that was about region stability.
How a reader would compute the equivalent figure
If you have walk logs, you do not need our stack to compute the analogous quantities. You need discipline:
- Fix one business question (and optionally two deliberate rewordings if you want a family).
- Pin graph version, tools/boot profile and model id so the family is comparable.
- Run at least five walks under conditions that are allowed to vary (parameters, natural pathing, optional rewordings).
- For each walk, extract the set of cited page ids (or source ids) that appear as citations in the answer or as explicit citation events in the log.
- For each pair of runs, compute citation-set overlap (intersection over union, or a simpler shared-count ratio — pick one definition and keep it). The five-run study’s reported shape is a range of pairwise overlaps between twenty-seven and fifty-seven percent; your absolute numbers will differ; the shape you care about is partial overlap with stable core content.
- Separately, pre-declare the doctrines or load-bearing claims that must appear for this question. Score survival across runs.
- Place the family on the variance table. Do not stop at the overlap percentage.
That procedure is how 27–57% becomes a method rather than a magic constant. The constant is this system’s telemetry for one question family. The method is portable.
The qualitative pattern around the specimen
Earlier harness work showed the same shape under reworded questions and varied routing parameters: the path shifts; the pivotal chapters still get read; answer quality stays subjectively in band. The five-run study puts magnitudes on the citation half of that story. It is the best executed evidence this book has.
Path variance with doctrine survival is success, not noise.
Chapter 7 is how you stop fooling yourself when citation lists and final prose both look “close enough” — including how you would refine this specimen without pretending the refinement has already been run.
Claim-Level Support
Final prose similarity and raw citation lists can both mislead. Decompose answers into claims and map each to a supporting source claim — specified tightly enough to run next week.
Two answers can sound alike and rest on different warrants. Two answers can sound different and share the same warrants. Prose is the wrong layer for evidence invariance. Raw citation lists are better than nothing and still coarse: legitimate alternate pages can support the same claim, and the same page can be cited while the model leans on a prior.
That last failure is the neighbour from groundedness research: post-rationalisation, where citations align with beliefs rather than with genuine reference use.3 Final-answer accuracy alone cannot explain which evidence supported each claim.4
Chapter 6 measured pairwise citation overlap and doctrine survival. Those are the right first cuts for an executed study. They still leave open whether two runs that both “mention” a doctrine actually rest on the same warrant, on legitimate alternate warrants, or on fluent restatement without opens. Claim-level support is how you close that gap without inventing a Jaccard figure you never computed.
Purpose
Distinguish three things Chapter 6 cannot fully separate with citation overlap alone:
- legitimate alternate sources supporting the same underlying claim
- answers that merely sound alike while resting on different warrants
- claims that float free of any opened source (prior candidates)
Those distinctions map directly to the variance table. Row one should survive claim-level scrutiny. Row two often collapses into either multi-entrance health or free-floating priors once you force the map from claim to opened source. Row three becomes visible when headlines stay fixed and qualifications flip over the same opens.
Inputs (what you must have before you start)
- A fixed question family (same business question; optional rewordings recorded as variants).
- At least two answer artefacts from separate walks (five is better if you are extending Chapter 6).
- Walk logs that show pages opened, not only final prose. If the log only stores the final string, you cannot run this protocol honestly.
- Version pins: graph id, tool/boot profile, model id for each run.
- Access to the source pages themselves so a human (or independent judge) can point at supporting passages.
What gets pinned
Pin the world before you score. If two runs used different graph commits or different tool contracts, you are no longer comparing walks of the same substrate. Record pins in the score sheet. Historical replay conditioned on an old payload is still useful; just do not pretend the pins were identical if they were not.
Also pin the decomposition rules: what counts as a claim, what counts as a material qualification, and whether you allow “equivalent assertion on a different page” as a match. Changing those rules mid-battery is how you manufacture agreement.
The decomposition step
- Take answer A. Split it into atomic claims. A claim is a single conclusion or a single material qualification (boundary, condition, exception). Prefer claims a second reader would recognise without your commentary. Avoid mega-claims that smuggle three conclusions into one sentence.
- For each claim, search the walk log for opened pages that could support it. Open those pages yourself. Record the supporting source claim: page id, chapter or section, and a short quotation or assertion identifier from the source.
- If no opened page supports the claim, mark it free-floating (prior candidate). Do not accept the model’s self-report that it “used” a page it never opened. Self-report is not cognitive provenance.
- Repeat for answer B (and remaining runs).
- Align claims across runs: same conclusion with the same qualification counts as the same claim; same headline with opposite qualifications counts as disagreement. Record disagreements explicitly; they are signal, not noise.
Worked shape (no invented numbers): Run A claims “X under condition C.” Run B claims “X” with no condition. At prose level they look similar. At claim level they disagree on a material qualification. If both opened the same chapter that states C, run B is synthesis instability (row three). If neither opened a supporting page, both free-float and the shared headline is prior-shaped agreement, not evidence invariance.
What gets scored
- Support overlap: for claims present in multiple runs, how often the supporting source claims are the same or clearly equivalent (same assertion, possibly different entrance pages).
- Conclusion and qualification agreement: answer invariance at claim level, not paragraph similarity.
- Free-floating rate: share of claims with no opened support. Rising free-floating rate is a prior signal even when prose looks stable.
- Legitimate alternates: claims with different page ids but equivalent source assertions — evidence of multiple entrances, not of failure.
Key Insight
If you cannot name the supporting claim, you do not have evidence invariance — you have vibes with footnotes.
What result would falsify the property
Route-invariant grounding, under this protocol, is under pressure when:
- answer-level conclusions look stable while free-floating claims are common
- shared citations exist but claim support does not (post-rationalisation neighbour)
- qualifications flip across runs while the headline stays fixed and the opened evidence stays fixed (row three of Chapter 3)
- doctrine-level survival from Chapter 6 cannot be traced to any opened assertion for some runs
It is not falsified merely because page ids differ while source assertions match. That is often healthy multi-entrance behaviour — the structural story of Chapter 5 showing up in the score sheet.
What this would refine on the five-run study
Chapter 6 shows doctrine survival and citation overlap. Claim-level support would ask whether the six stable doctrines were supported by the same underlying assertions, by legitimate alternate sources, or by fluent restatement without opens. It would also expose qualifications that appeared in one run and vanished in another while the headline conclusion stayed put — for example, a run that leads on pilot mechanics may still need the same authority-boundary qualification as a run that leads on governance structure.
Until that matrix exists, treat doctrine survival as a strong first cut and refuse to pretend it is claim-level Jaccard. Do not invent a claim-overlap percentage for the five runs. The honesty that made Chapter 6 trustworthy is the same honesty that keeps this chapter a protocol rather than a fake result.
Cognitive provenance as the precondition
The method requires knowing what was opened, not what the model says it used. Cognitive provenance — pages, claims and edges actually observed at decision time, external and version-pinned — is the plumbing. Without it, claim-level support collapses into another post-hoc narrative.
If your tooling cannot yet export opens, build that first. A beautiful claim matrix built on guessed opens is worse than no matrix: it trains the organisation to trust a second prior.
Chapter 8 is the disambiguator when answers hold and evidence appears to wander: remove the genba and watch.
The Omission Test
Counterfactual omission is the sharp test that distinguishes a grounded substrate from a well-dressed model prior. Detection target: model-prior substitution. Specified so you can run it next week.
Row two of the variance table (Chapter 3) is the ambiguous cell: path and evidence vary while the answer stays stable. That can mean genuine alternate entrances. It can also mean the model is answering from its priors and the graph is decoration. Stable answers under path variance do not, by themselves, prove grounding. Omission is how you force the rival stories apart.
Chapter 1’s agent B is the single-walk form of the disease. Chapter 8 is the family-level counterfactual: remove the place that should have been genba and see whether the answer was ever visiting it. Chapter 7 can make free-floating claims visible; omission makes them undeniable when the answer still struts without the page that should have been load-bearing.
From duty to detection target
Reflexive Agent Design already treats deliberate missing information as a duty: drop a page, delay a canonical source, remove an edge, and see whether the walk fails honestly or luckily invents. Route-Invariant Grounding converts that duty into a named detection target.
Counterfactual omission
Temporarily remove a pivotal page, source route or edge. Replay the same question. A healthy walker takes an alternate grounded route, qualifies the answer, or confesses the absence. A walker that confidently returns the same answer with no evidence has just revealed model-prior substitution.
Inputs
- A question family you already trust enough to use as a regression (ideally one that previously showed healthy or ambiguous results under Chapter 3’s table).
- Walk logs from the baseline family identifying pages and edges that repeatedly carry load-bearing claims. Chapter 6’s doctrine survival list, or a rough Chapter 7 map, is enough to nominate candidates.
- Ability to hide or remove one element without redesigning the rest of the stack in the same experiment.
- Version control so you can restore the world after the probe and keep the omission run as an artefact.
What gets pinned
Pin graph version (with the omission applied as a named delta), tool/boot profile, model id, and the exact question text. Pin the identity of the removed element and why you believed it was pivotal. If you change two things at once — omit a page and change bloom policy — you will not know which change produced the behaviour. Single-variable probes are slower and more honest.
Procedure
- From a baseline family of walks, identify a page, route or edge that repeatedly carries load-bearing claims. Prefer something claim mapping would mark as support for a core conclusion, even if you only map roughly.
- Record baseline answers and opened-page sets for the question family. Note confidence tone and qualifications.
- Apply a single omission: hide the page, remove the edge, or block the route. Keep everything else fixed.
- Replay the same question family under the same model and boot profile.
- Score each replay against the healthy-outcome checklist below. Capture the walker’s language verbatim when it qualifies or confesses.
- Restore the omitted element. Keep the omission run as a regression artefact next to the baseline family.
What gets scored
- Whether an alternate grounded route appears (other pages that actually support the same claims).
- Whether the answer qualifies or narrows when the warrant is gone.
- Whether the walker confesses absence or unreachability in plain language.
- Whether confidence stays high while support is empty — the prior signature.
- Path Testing cell on the omission walk: bad path + right answer is the red flag form.
- Whether free-floating claims (Chapter 7) increase under omission while prose similarity to baseline stays high.
Healthy outcomes
- Alternate grounded route: another entrance reaches equivalent source claims; the answer still rests on opens. This is Chapter 5’s multi-entrance story earning its keep.
- Qualification: the answer narrows scope because the warrant is gone. Confidence tracks evidence.
- Confession: the walker states that the pivotal material is missing or unreachable. That is success of honesty, not failure of the product.
Healthy outcomes can still look “worse” on a string-only eval that wanted the original full answer. That is the point. A system that cannot say “I cannot ground that without X” is not robust; it is fluent.
Train the organisation to read those outcomes correctly. A confession is not a product failure to be papered over with a friendlier prompt. A qualified answer is not incomplete work to be forced full again without restoring the evidence. If your eval suite marks honest qualification as a miss and fluent invention as a hit, the suite is training model-prior substitution. Align the suite with the three healthy outcomes before you run the probe at scale.
What falsifies the property (model-prior substitution)
If it confidently returns the same answer without evidence, you have detected model-prior substitution.
That is the same disease as Path Testing’s bad path + right answer cell, caught with a counterfactual instrument on a family of walks. Output-only green becomes expensive to believe. A sceptical reader does not need a long philosophical argument after that transcript; they need the omission log.
Note what does not by itself falsify route-invariant grounding: a longer path that finds a legitimate alternate entrance. That is the attractor working. Failures that look like thrash but end in honest qualification are still healthier than fluent invention. A partial answer that drops the unsupported claims is healthier than a full answer that invents support.
Choosing what to omit
Bad omission targets produce uninformative results. Do not omit a decorative page nobody used. Do not omit the entire corpus. Prefer a single page or edge that baseline walks treated as load-bearing for a core claim. If baseline walks never opened it, omission cannot distinguish prior from path. If every claim depended on it and no alternate entrance exists, a healthy confession is the expected outcome — and that is still a successful test of honesty, plus a signal that the territory may be single-entrance brittle (Chapter 5).
What to keep as the receipt
When you run the test, keep the walker’s words. Record what was removed, what was opened instead, whether claims changed, and any confession language. This book will not invent a worked failure quote. The receipt is the point of the protocol, and inventing one would violate the honesty rule that made the five-run study trustworthy.
Store the receipt next to the baseline family sheet from Chapter 10’s battery: same question texts, same pins, omission delta named, outcome classified. Next month’s reader should be able to replay the logic without asking you what you meant by “it still looked fine.”
Chapter 9 multiplies the agent class so the substrate score is not a single vendor session in disguise.
Class-Level Variation
Route-invariant grounding must hold across the agent class, not one favourite session. Otherwise you measured a model’s private habits and called them a substrate property.
A beautiful walk on one frontier model is not class-level legibility. Tool order, vocabulary and boot profiles can lock in one model’s quirks until the next generation makes the design look stupid overnight. Reflexive Agent Design without a variation protocol is one vendor session wearing a general label.
Chapter 6’s five-run study is executed evidence of path variance with doctrine survival under the instrumented stack. It is not, by itself, a proof that every model family and every prompt style will land in the same genba. Class-level variation is how you stop overclaiming. The attractor of Chapters 4–5 is only as good as the class of agents that can find it.
Purpose
Test whether the four quantities and the variance-table row remain in a healthy regime when the agent class changes — not only when pathing varies inside one familiar session. You are asking whether the substrate is legible to the users it will actually have, or only to the session that built it.
Inputs
- The same question family used for your baseline family-of-walks score.
- At least two model classes (for example one frontier and one utility model that cannot paper over bad contracts with clever compensation).
- Ability to start fresh contexts with no warm conversational memory of house vocabulary.
- Optional alternate boot profiles and prompt styles (terse vs conversational).
- Pins and walk capture as in Chapters 6–8.
What gets pinned
For each cell of the battery, pin graph version, tools/boot profile, model id, prompt style, and whether the context was fresh. If you change the graph while changing the model, you will not know which variable moved the score. The Progressive Evaluation Ladder still applies: inspect cheaply before generating a full multi-model grid.
Write the pin table before the first run. Teams that improvise pins mid-battery always discover, too late, that they compared different worlds.
Axes that change how an agent meets your affordances
- Frontier and utility models — including a cheap model that cannot paper over bad tool contracts with clever compensation.
- Fresh contexts — no warm conversational memory of wiki vocabulary, no prior turns that taught house names.
- Terse versus conversational questions — both appear in real agent traffic.
- Vocabulary-familiar versus vocabulary-naive agents — framework names versus public language only.
- Boot profiles and tool-description order — deliberately swap order, rename tools, shorten or lengthen descriptions; measure path change.
- Deliberately delayed or missing information (Chapter 8), applied to more than one model if you can afford it.
The aim is not to make one session walk beautifully. It is to make the system’s affordances legible across the class of agents intended to use it.
Procedure (runnable form)
- Fix the question family and the graph pin.
- Run the baseline family on Model F (your usual frontier) as control. Score four quantities and variance-table row.
- Re-run the family on Model U (utility) in a fresh context. Score again.
- Optionally re-run a subset under terse vs conversational prompts on one model only.
- Optionally swap boot profiles (search-first vs canonical-tool-first) on one model only.
- For each cell, score path variance, evidence invariance, answer invariance, traversal efficiency, and variance-table row. Optionally run a light claim map on divergences.
- Where cells diverge, refuse designs that depend on one model’s search bias or vocabulary preference.
- Keep diverging walks as regression: the next redesign must not reintroduce one-model specialisation.
What gets scored
Score the substrate, not only the string. Green answers with collapsed evidence invariance are not a pass. Path Testing cells still matter per walk; the family row is the substrate verdict. Compare rows across models: a control that sits in row one while a utility model sits in row two or four is a class-level failure even if the frontier still looks pretty.
Also compare free-floating claim rates if you can. A frontier model that papers over missing opens with fluent synthesis will look more invariant than a utility model that stumbles into honesty. That is not a reason to prefer the frontier on robustness; it is a reason to distrust the frontier’s green string.
What falsifies class-level route-invariant grounding
- Only the favourite model reaches the load-bearing region; others thrash or invent.
- Answer invariance holds for the favourite model via free-floating claims, exposed when a utility model cannot compensate.
- A vocabulary-naive agent never finds the territory because entrances only exist under house jargon.
- Omission (Chapter 8) catches prior substitution on one model class while another class fails honestly — still a substrate design problem if production traffic includes both.
- Boot-profile A works and boot-profile B never reaches genba, with no documented reason why production will only ever use A.
Worked reading of divergence (shape, not a fake study)
Suppose Model F shows partial citation overlap with doctrine survival (the Chapter 6 shape). Model U, on the same question family and graph pin, never opens the capstone chapters and still produces a fluent answer. That is not “Model U is dumb.” That is class-level evidence that your affordances are not legible without frontier compensation. Fix tool contracts, labels, boot order, and entrance vocabulary. Re-run until the utility model can reach the region or fails honestly under omission.
Suppose both models reach the same load-bearing region with different paths. That is the attractor working across the class. Keep those walks as goldens of multi-entrance health.
A third shape is subtler: both models reach the region, but only the frontier preserves qualifications that live in the source. The utility model drops boundary conditions while keeping the headline. That is not path failure; it is synthesis fragility under class change. Score it as answer-invariance pressure, not as a missing edge. Fix how qualifications are required in the answer schema, then re-run both models. Class-level scoring is how you notice that the frontier was quietly compensating for a soft synthesis gate.
Historical replay is not the whole story
Historical replay is useful and conditioned on what the old walker was shown. When you change retrieval design — payload shape, bloom policy, edge exposure — the decisive evidence is new traffic under the new design, not only replaying old traces through a world they never met. Live evaluation sits beside replay: inspect family-of-walks first, replay when comparison needs pins, generate active traffic when the world has moved.
Do not let historical convenience become a class-level blind spot. Old traces are usually dominated by the models and prompt styles you already favoured. A clean replay on last month’s traffic can ratify one-model habits. Budget at least one live cell with a second model or a fresh context whenever the claim is “the substrate is robust,” not merely “last month’s favourite walker still works.”
Chapter 10 compresses everything into a battery you can run without re-reading the doctrine chapters — the page that should work standalone on a Monday morning.
The Operating Battery
The most usable page in the book. A weekly checklist with enough surrounding instruction that it works standalone — without re-reading the preceding ten chapters.
If you only keep one chapter, keep this one. The doctrine is load-bearing; the battery is how the doctrine becomes a habit. Everything below is written so a competent engineer with walk logs can execute it without reopening the theory chapters, while still knowing where to go if a step fails.
What this battery is for
You are answering one question: is the knowledge substrate doing load-bearing work under reasonable variation, or is the model answering from its priors with the graph as decoration? You answer it by scoring a family of walks of the same question, not by celebrating one green string.
You are not answering: is the model smart? Is path length short? Did call count go down? Those may matter for engineering. They are not the substrate verdict.
Preconditions (do not skip)
- Walk capture that records pages opened, edges used, abandoned branches, citations and final answer text. If you only store the final answer, stop and instrument first. Without opens, Chapters 7 and 8 cannot run honestly.
- Version pins: graph commit or export id, tool/boot profile, model id. Write them on the score sheet before run one.
- A judge who is not the answering model for claim decomposition — a human, or an independent process that cannot invent opens the log does not show.
- One business question worth regressing — hard enough that a prior-only answer would be a real risk, stable enough that you can re-run it next month.
- A place to file sheets — a directory, ticket, or wiki page where baseline and later runs sit next to each other.
Operating battery
- Fix a question family — one core business question plus two deliberate rewordings. Record all three texts verbatim. Do not improvise rewordings mid-run.
- Pin versions — graph, tools/boot, model. Do not change pins mid-family.
- Run at least five walks under deliberate variation of wording and routing parameters (and natural pathing). Save full logs, not summaries.
- Score all four quantities for the family:
- Path variance — how different were pages, edges, ordering?
- Evidence invariance — overlap of load-bearing sources/chapters/claims (citation overlap is a coarse start; doctrines or claims are better).
- Answer invariance — do conclusions and qualifications agree when decomposed?
- Traversal efficiency — useful territory and source grounding per time/payload; not raw call count as a trophy.
- Place the family on the variance table
- Path varies; evidence and answer stable → healthy route resilience.
- Path and evidence vary; answer stable → redundancy or prior luck (do not celebrate yet).
- Evidence stable; answer varies → synthesis instability.
- Evidence and answer both vary → question/graph/retrieval not ready.
- Claim-level support map — even manually once. Decompose answers; map each claim to an opened source claim; mark free-floating claims. (Proposed protocol; still run it.)
- One omission probe on a pivotal page, route or edge. Record alternate route, qualification, confession, or confident prior. (Proposed protocol; still run it.)
- Optional class check — one alternate model or fresh context. Score the four quantities again. (Proposed application; still run it when you can.)
- Classify twice — Path Testing cell per walk (good/bad path × right/wrong answer); substrate row for the family.
- File the sheet — pins, overlap notes, doctrine/claim survival, omission outcome, pass/fail. Keep it as regression after the next graph or walker change.
How long this should take
The first time, budget a working session: instrumentation check, five walks, one manual claim map, one omission. Later runs of the same family should be cheaper — re-run walks, recompute overlaps, spot-check claims, omit only when the graph or walker changed. Continuous full replay of everything is not maturity; it is often avoidance of reading the traces you already have.
If the first session is blocked on missing opens in the log, the session still produced a finding: you do not yet have the plumbing for evidence invariance. Fix capture before you invent a dashboard that only scores final strings.
Pass shape
- Path may vary freely. Do not punish path variance alone.
- Load-bearing evidence region holds (doctrines and/or claim supports).
- Answers are materially consistent, with qualifications that track the evidence.
- Omission fails honestly: alternate route, qualify, or confess.
- Traversal efficiency is monitored, not worshipped; call count is not the trophy.
- If you ran a class check, the healthy row is not unique to one favourite model.
- Proposed versus executed labels are written on the sheet so next month’s reader knows what was measured.
Fail shapes (and what they point at)
- Confident same answer under omission with no evidence — model-prior substitution. Fix grounding discipline and synthesis gates; do not celebrate path dashboards.
- Answer holds while evidence vanishes and no alternate grounded route appears — treat as prior risk until proven otherwise.
- Evidence holds while answers contradict on load-bearing claims — synthesis instability; fix judgement, not edges.
- Total scatter of evidence and answer — question or graph not yet navigable; stabilise the task before scoring robustness.
- Beautiful single-model walk that collapses under fresh context or utility model — class-level overfitting, not substrate success.
- Low path variance forced by tooling, with no proof of multi-entrance reachability — you may have overfit a route; test a rewording and a second model before calling it robust.
Minimal score sheet (copy this)
- Date; operator; question family texts (core + rewordings).
- Pins: graph / tools-boot / model / context freshness.
- For each run: path summary, citation set size, doctrines or load-bearing claims hit, Path Testing cell.
- Family: path variance notes; evidence invariance notes; answer invariance notes; efficiency notes; variance-table row.
- Claim map: free-floating count (even approximate); legitimate alternates noted.
- Omission: what removed; outcome (alternate / qualify / confess / confident prior).
- Optional class cell: second model or fresh context; row comparison.
- Verdict; next action; link to raw logs.
Standing regression
Keep the question family. Re-run after graph edits, retrieval-policy changes, boot-profile changes, or model upgrades. Compare new sheets to old ones. The battery is not a one-off audit; it is how route-invariant grounding stays a property rather than a story you told once after a good week.
When the row moves from one to two, do not “fix path variance.” Run claim map and omission. When the row moves from one to three, fix synthesis. When the row moves to four, fix the question or the territory before you touch efficiency knobs.
Key Takeaways
- Four quantities or you are not measuring route-invariant grounding.
- Omission is the sharp test for model-prior substitution.
- Label proposed versus executed work every time you report.
- If it is not on a battery, it is not a property you have — it is a story you like.
Chapter 11 draws the boundary around this doctrine so neighbouring books can own their own axes without this one collapsing into mush.
Boundaries and the Reliability Case
What we refuse to write is part of the doctrine. This book owns measurement of route-invariant grounding — cleanly.
Seasoned editorial judgment is often what does not ship. Neighbouring ideas are real and load-bearing elsewhere. Folded into this book they would turn a measurement doctrine into a hedged essay that cannot be tested. The John West principle applies: the fish you reject is part of what makes the can trustworthy.
A reader who finishes Chapter 10 has a battery. A reader who finishes Chapter 11 knows what that battery is not allowed to become: a dumping ground for every adjacent insight about wikis, redundancy, bloom payloads, edge direction, or learning that looks like rediscovery. Those insights matter. They get their own rooms. This room is the reliability scoreboard.
Why the reliability case stays unhedged
There is a true correction nearby: the same destination does not make two journeys cognitively equivalent, and answer invariance can hide derivational novelty. That correction is important. It is also somebody else’s chapter. If you import it here as a softener, the four quantities stop being decision tools and become decorative vocabulary. A reliability score that apologises for itself mid-sentence is not a reliability score.
This book states the reliability case cleanly: measure whether many paths still share a genba, and run the omission test for model-prior substitution. Learning-sensitive memory of different derivations can arrive later without sabotaging the scoreboard. First you need to know whether the graph is load-bearing at all.
The same discipline applies to bloom and context-budget debates. Payload shape can change what historical replay means. That is why Chapter 9 insists live evaluation sits beside replay. It is not a licence to turn this book into a payload design manual. Measure the substrate property under a pinned world; redesign the payload under its own brief.
What this book did claim
It claimed a third instrument: family of walks of the same question under deliberate variation, returning a verdict on the substrate. It claimed four quantities and a four-row diagnostic table, with row two given the space it deserves as the ambiguous cell. It claimed a causal engine — loose navigation at the front, strict epistemics at the back — with genchi genbutsu as terminal attractor and capstones as high-density pull. It claimed one executed specimen with real figures and an honest reading of why partial citation overlap is interesting rather than disappointing. It specified three proposed protocols tightly enough to run. It left you a battery that can stand alone.
It did not claim that path variance is always healthy. It did not claim that answer invariance is always good. It did not claim the omission test has already been run on the five-run family. It did not claim industry benchmarks from AskUI telemetry. It did not invent names for the six surviving doctrines when the source did not publish a roster. Those refusals are part of the claim set.
How the pieces fit without re-arguing them
Chapter 1 made agent B recognisable in a log. Chapter 2 named the unit: substrate, not agent, not map. Chapter 3 gave four quantities and a table that decides next action. Chapters 4 and 5 explained why row one can exist: attractor shape plus entrances and capstones. Chapter 6 was the specimen. Chapters 7 through 9 made the ambiguous cell and the class-level claim runnable without pretending they were already executed. Chapter 10 is the Monday-morning page.
If you find yourself re-deriving Path Testing’s full matrix here, stop — that instrument is reused by name. If you find yourself designing janitor mutations here, stop — that instrument belongs to File Back the Walk. If you find yourself arguing that identical answers hide valuable new proofs, stop — that is the neighbouring reliability-versus-learning tension, not this book’s scoreboard.
Honesty inventory
| Item | Status in this book |
|---|---|
| Five-run study (27–57% citation overlap; six stable doctrines) | Executed evidence |
| Claim-level support comparison | Proposed protocol |
| Counterfactual omission / model-prior substitution | Proposed protocol |
| Cross-model / fresh-context battery for RIG | Proposed application |
Keep that table in your own reports. The fastest way to lose trust in a measurement doctrine is to blur proposed and executed until a sceptical reader cannot tell which sentences are receipts. The second-fastest way is to merge external faithfulness figures with your own citation-overlap range because the digits look related.
What a reader should be able to do tomorrow
- Open a walk log and recognise agent B without waiting for a production incident.
- Score a five-walk family on four quantities and place it on the variance table.
- Compute pairwise citation overlap for their own logs without treating low overlap as automatic failure.
- Run one omission probe and keep the walker’s words as a receipt.
- Refuse to call a single-model beauty contest a substrate proof.
- File the sheet and re-run it after the next graph or model change.
The property, once more
A knowledge substrate is robust when reasonable variations in question wording, routing parameters and traversal order still lead the agent to an equivalent load-bearing evidence region and a materially consistent answer — without depending on one prescribed path or an ungrounded model prior.
You can measure four quantities across a family of walks. You can run the one test that distinguishes a grounded substrate from a well-dressed model prior. You can compute a pairwise citation-overlap range the way Chapter 6 did, without treating low overlap as automatic failure. You can refuse to merge that range with external faithfulness figures that happen to share a digit shape.
Many paths. Same genba.
Trajectory tooling made path variance cheap to record. The discipline is not to erase that variance. The discipline is to measure whether the genba still holds — and to catch the walk that never went there at all.
If you take nothing else: open the log before you trust the answer, score the family before you trust the substrate, and remove the pivotal page before you trust the story that everything is fine. Neighbouring books will teach you what to do with redundancy, payloads, edges and rediscovery. This one taught you how to know whether the graph is doing work.
Carry the kill list into the battery, not only into this closing chapter. When a sheet tempts you to annotate “also interesting for bloom” or “same answer, new proof,” park that note outside the reliability verdict. The family row, the four quantities, and the omission outcome stay unmuddied. You can open a second ticket for the neighbouring insight the same afternoon. What you must not do is let the second ticket rewrite the first score until the scoreboard no longer decides anything.
That is the last operational form of the John West principle: reject the good fish that does not belong in this can. Route-invariant grounding is a narrow, testable property. Keep it narrow enough to fail cleanly — and therefore narrow enough to trust when it passes.
References & Sources
The evidence base behind every claim — primary research, industry analysis, and technical specifications
Research Methodology
This ebook draws on primary research from standards bodies, independent research firms, enterprise technology vendors, and consulting firms. Statistics cited throughout have been cross-referenced against primary sources.
Frameworks and interpretive analysis developed by Scott Farrell / LeverageAI are listed separately below — these represent the practitioner lens through which external research is interpreted, and are not cited inline to avoid self-promotional appearance.
Industry Analysis & Vendor Research
LangChain — LangSmith Evaluation documentation [1]
Agent evaluation captures full trajectories of steps, tool calls and reasoning, not only final strings
https://docs.langchain.com/langsmith/evaluation
LeverageAI / Scott Farrell — Practitioner Frameworks
The interpretive frameworks, architectural patterns, and practitioner analysis in this ebook were developed through enterprise AI transformation consulting. The articles below are the underlying thinking behind those frameworks. They are listed here for transparency and further exploration — not cited inline, as this is the author's own analytical voice.
Scott Farrell — Reflexive Agent Design
Grade the path, not only the answer; bad path + right answer is unsupported lucky success
https://leverageai.com.au/wp-content/media/articles/146-reflexive-agent-design.html
Scott Farrell — File Back the Walk
Four instruments on one walk record: path testing, disputes, provenance, walk telemetry
https://leverageai.com.au/wp-content/media/articles/80-file-back-the-walk.html
Scott Farrell — RAG Was Built for Chatbots — Agents Need a Wiki
Path Testing 2x2: grade path independently of answer
https://leverageai.com.au/wp-content/media/articles/69-rag-was-built-for-chatbots-agents-need-a-wiki.html
Scott Farrell — The Model Is Not the Memory
Cognitive provenance: reconstruct which pages, claims and edges were observed at decision time
https://leverageai.com.au/wp-content/media/articles/68-the-model-is-not-the-memory.html
Scott Farrell — The Wiki Playbook
Boot profiles as hub plus selection bias; edges guarantee reachability from any entry
https://leverageai.com.au/wp-content/media/articles/176-the-wiki-playbook.html
Scott Farrell — Route-Invariant Grounding: Many Paths, Same Genba
Executed five-run AskUI walk study (this system's telemetry): 27–57% pairwise citation overlap; six stable doctrines — primary specimen for route-invariant grounding
Scott Farrell — Route-Invariant Grounding: Many Paths, Same Genba
Four quantities (path variance, evidence invariance, answer invariance, traversal efficiency) applied to the five-run specimen
Scott Farrell — Route-Invariant Grounding: Many Paths, Same Genba
Operating battery: family of walks, four quantities, variance table, claim-level map, omission probe
Scott Farrell — Route-Invariant Grounding: Many Paths, Same Genba
Reliability case restated: four quantities + omission test for model-prior substitution; sibling doctrines deferred
Primary Research & Standards Bodies
deepset — Measuring LLM Groundedness in RAG Systems [2]
Factual vs unfaithful hallucination; faithfulness defined against retrieved context
https://www.deepset.ai/blog/rag-llm-evaluation-groundedness
Wallat, Heuss, de Rijke & Anand (arXiv:2412.18004) — Correctness is not Faithfulness in RAG Attributions [3]
Citation correctness insufficient; post-rationalisation; up to 57% of citations lack faithfulness even when correct
https://arxiv.org/abs/2412.18004
Yiqi Wang et al., arXiv:2606.04990 — From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents [4]
Final-answer accuracy alone cannot explain which evidence supported each claim; execution provenance as typed graph of an agent execution
https://arxiv.org/abs/2606.04990
About This Reference List
Compiled July 2026. All URLs verified at time of compilation. Regulatory documents and standards specifications are subject to revision — check primary sources for the most current versions.
Some links to academic papers and vendor research may require free registration. Government and standards body publications are freely accessible.