Measurement · Agentic knowledge systems
Route-Invariant Grounding: Many Paths, Same Genba
How do you know your agentic knowledge base is robust — or that the model is answering from its priors and the graph is decoration?
In short
- Path identity within one walk is the wrong quality score. Evidence invariance across many walks of the same question is the right one.
- Four quantities are routinely conflated: path variance, evidence invariance, answer invariance, and traversal efficiency.
- Genchi genbutsu (source descent) acts as a terminal attractor: loose navigation at the front, strict epistemics at the back.
- One executed study: five runs, 27–57% pairwise citation overlap, six stable doctrines. One sharp proposed test: counterfactual omission for model-prior substitution.
Imagine two agents handed the same hard business question. Same graph. Same tools. Both return a sharp, plausible answer that would pass a casual human skim. Grade the final string and you cannot tell them apart. That is the trap. Grade the path and you can. Grade a family of walks under deliberately varied conditions and you can say something stronger still: whether the knowledge substrate is doing load-bearing work, or whether the graph is decoration.
Two walks that look the same from the outside
Agent A does not begin by answering. It orients. A search or map call lands a framework page. An edge leads to a concept. Source-descent pulls the walker into the actual ebook chapter that carries the warrant. The log shows the canonical page opening before synthesis. Claims in the draft can be matched to passages that were actually read. Qualifications in the source appear in the answer. Abandoned branches, if any, are left behind; the load-bearing region was still visited.
Agent B looks busier. It thrash-searches. It opens three loosely related pages, none of them the chapter that should carry the claim. It skims a secondary mention. It never follows the edge that would have forced descent. Then it writes. The prose is fluent. The conclusion matches what a well-read practitioner would have said, because the model’s pre-training already contains a compressed public shape of the idea. The graph was present. The graph was not load-bearing. It was scenery.
What would an observer have had to notice to catch agent B? Not the final paragraph. The tells sit in the walk: canonical page never opened (or opened after the answer was already formed); citations that cannot support the strongest claims; missing qualifications that live only in unread chapters; unused typed edges; a confidence that would likely survive if the pivotal page were removed tomorrow.
That is the Path Testing cell we already named the dangerous one — bad path, right answer; unsupported lucky success that works until priors stop holding.2 Output-only suites celebrate it. Path-aware suites treat it as a failure of architecture dressed up as a win.
Path Testing is necessary and still not enough. You can have one beautiful walk and a fragile substrate. You can have path variance that looks noisy on a dashboard and a substrate that is actually healthy. The unit this piece owns is different: not one walk’s cell, but a family of walks of the same question under deliberate variation, scored for whether the load-bearing evidence region still holds.
If you only ever grade final strings, agent B will keep shipping. If you only ever grade path identity, agent A’s legitimate alternate entrance will look like a regression. Both mistakes are cheap once trajectory tooling makes the path visible. The discipline is to score the right quantity on the right unit.
Route-Invariant Grounding
A knowledge substrate is robust when reasonable variations in question wording, routing parameters and traversal order still lead the agent to an equivalent load-bearing evidence region and a materially consistent answer — without depending on one prescribed path or an ungrounded model prior.
Many paths, same genba.
Genba is the actual place: the source chapters and claims that must be visited if the answer is to be grounded rather than performed. Many paths may enter the territory. The same genba must still be reached.
The quantity worth measuring is evidence invariance across many walks, not path identity within one.
Why path identity became the wrong temptation
Agent evaluation platforms have converged on capturing full trajectories — tool selection, arguments, loops, retrieval paths, latency and cost — not only final strings.1 That is necessary infrastructure. It also creates a new temptation: treat path identity as the score. If last week’s “good” walk opened pages A→B→C, this week’s walk that opened A→D→C looks like a regression. It may not be. It may be healthy route resilience — two bush tracks into the same valley.
If your dashboard is optimising path identity, you are training the system to look like last Tuesday’s favourite model and one prompt style. You will punish legitimate entrances and miss lucky priors that never needed the graph at all.
A third instrument on the same log
Our canon already owns related instruments that read the same walk record and are easy to confuse.
Path Testing grades one walk and returns a verdict on the agent. Its 2×2 — good/bad path × right/wrong answer — makes the dangerous cell expensive to ignore.29
Walk telemetry aggregates many walks and returns a graph mutation: missing edges, dead ends, cold pages. Its unit is the map, not the agent.3
Neither instrument answers the question this piece owns. Path Testing asks: was this walk sound? Telemetry asks: what should the map become? Route-Invariant Grounding asks: is the substrate robust enough that reasonable variation still lands in an equivalent load-bearing evidence region with a materially consistent answer?
| Instrument | Unit | Output |
|---|---|---|
| Path Testing | One walk | Verdict on the agent |
| Walk telemetry | Many walks, aggregate | Graph mutation |
| Route-Invariant Grounding | Many walks of the same question under deliberate variation | Verdict on the substrate |
Name the unit before you celebrate the metric. A green string is not a unit. A short path is not a unit. A new edge is not a unit. Agent, map, and substrate are three different objects that happen to share a log format.
If you conflate them you ship the wrong repair: more edges for an agent that never opens the canon; more path constraints for a map that is missing joins; more path dashboards for a substrate that is thin while the model is fluent.
Four quantities, not one blob called “groundedness”
Groundedness research already warns that “correct against reality” and “faithful to the retrieved context” are different failures.4 Citation correctness is not the same as genuine reference use; systems can post-rationalise attributions that align with prior beliefs rather than with what was actually used. One study reports that up to 57% of citations can lack faithfulness even when they are correct.5
Hold that as a neighbour. Do not glue it to the five-run figures below. Those are a different quantity: pairwise citation overlap across runs, not faithfulness of a citation within a run. Merging the two “57%s” in conversation is how a careful argument becomes a sloppy slide.
| Measure | Meaning |
|---|---|
| Path variance | How different were the pages, edges and ordering across runs? |
| Evidence invariance | How much overlap existed in the load-bearing source pages, chapters and claims? |
| Answer invariance | Did independently decomposed claims reach the same conclusions and qualifications? |
| Traversal efficiency | How much useful territory and source grounding were obtained for the time and payload spent? |
Path variance is not a failure mode by itself. It is what free agents do when the territory has more than one entrance. Evidence invariance is the load-bearing property. Answer invariance is necessary and insufficient: answers can agree while floating free of evidence. Traversal efficiency is an engineering constraint, not the quality score.
Call count is not a scoreboard. A better agent might make more calls because deeper investigation is available; a more mature map might need fewer broad searches because typed edges replace guessing.3 Measure useful territory and source anchoring — not raw tool calls as a trophy.
Final-answer accuracy alone cannot explain how an output was produced or which evidence supported each claim.8 That is why the four quantities must be scored together.
The variance table — read it as a decision surface
| What varies? | What remains stable? | Interpretation |
|---|---|---|
| Path varies | Evidence and answer stable | Healthy route resilience |
| Path and evidence vary | Answer stable | Possibly useful redundancy — or model-prior luck |
| Evidence stable | Answer varies | Synthesis or judgement instability |
| Evidence and answer both vary | — | Retrieval, graph or question ambiguity |
Row one is the property you want: different routes, same genba, same material answer. Keep the question family as a regression. Resist forcing one golden path.
Row two is the ambiguous cell — and the whole point of this piece. Paths differ and the sets of opened sources differ in ways that matter, yet answers still sound the same. Two rival stories compete: legitimate alternate entrances supporting the same claims, or model-prior luck with the graph as decoration. Prose-level answer invariance becomes actively dangerous here. This is the family-level cousin of bad path + right answer. Do not celebrate. Run claim-level support: decompose conclusions and qualifications, map each to an opened source claim, mark free-floating claims. Then run omission on a page that should have been pivotal. If alternate entrances carry the same warrants, document multi-entrance health. If the answer survives without evidence, you have model-prior substitution — fix grounding discipline, not the path-identity dashboard.
Row two is also where teams waste the most time “fixing” the wrong surface. They shorten paths, add more golden sequences, or punish citation diversity — and leave the prior untouched. The ambiguous cell demands a probe that removes evidence, not a probe that enforces sameness.
Row three is a synthesis problem: same evidence region, swinging conclusions or qualifications. Rearranging edges will not fix judgement. Fix how conflicting claims and boundary conditions are forced into the open.
Row four means you do not yet have a stable question or a navigable territory. Stabilise the task before you score robustness.
You cannot score evidence invariance from a model’s post-hoc story about what it “must have read.” You need cognitive provenance: which pages, claims and edges were actually observed at decision time — external, version-pinned, restorable.
Why routes converge: genchi genbutsu as terminal attractor
Genchi genbutsu — go and see the actual place; descend to source — is already operating doctrine in this stack. The measurement story needs a causal engine, not a slogan. The engine is shape:
Loose navigation at the front, increasingly strict epistemics at the back.
The early route can vary: framework first, concept first, a later update first, a capstone ebook first, an adjacent implementation first. Vocabulary matches the question wherever it can. Search parameters and result depth can shuffle which neighbours appear first. That looseness is how reworded questions still find a foothold.
But a source-descent rule eventually funnels those routes toward actual ebook chapters and source material before synthesis. The final answer is not allowed to inherit all that navigational looseness. Several bush tracks enter the same valley; one marked trail leads down to the site.
When we reviewed walk paths under different search strategies, different parameters, and slightly different questions, the path rarely repeated. The outcome quality barely moved. The pivotal chapters still got read. Reword the question a bit and the walker still hits the right material. Genchi genbutsu keeps pulling the walk into the correct final-ebook chapters even when the path in differs — and that rule is directing the walk more than people notice.
You do not want every agent on one canonical sequence. That would mean you over-specified the route for one model and one prompt style.7 Over-prescribing the path is not grounding. It is overfitting dressed as governance. Natural movement is allowed. Ungrounded synthesis is not. Strict late epistemics is a gate on synthesis, not a ban on exploration. The walker may wander. The answer may not. That distinction is what lets path variance and evidence invariance coexist. Without the gate, exploration becomes prior-shopping: open anything, then write what the model already believed.
Inspection questions that surface a broken attractor: Do walks that enter from different vocabulary still open the same source chapters before synthesis? Is source descent enforced, or merely encouraged in a prompt the model can ignore? Do secondary pages act as entrances that lead home, or as stopping points that satisfy the model early? When a walk never reaches a capstone or canonical chapter, does the answer still claim full confidence?
Final ebooks matter disproportionately under this shape. They are high-density, source-oriented attractors that consolidate a region without requiring every walker to rediscover its complete publication history. Boot profiles and typed edges do related work: a graph can be entered from many pages while still guaranteeing reachability into the territory a task needs.6 The entry point sets where you start; it should not limit what you can reach.
The five-run study (executed evidence)
Properties need specimens. On a dated replay session we reviewed live walk history from AskUI — Scott’s AWS Marketplace application that instruments agent walks over a knowledge substrate (not the third-party UI-automation vendor of a similar name). The figures below are this system’s instrumented telemetry, not industry benchmarks. Production instrumentation from a shipped product is stronger evidence than a lab toy; it is still not a claim about the industry at large.
The fixed object was the business question itself — one task, not five different tasks dressed as one family. What moved was the natural variance of a real walker under a live stack: different pathing, different citation sets, different emphasis in the final synthesis. The wider harness work around the same substrate had already been deliberately varying search strategies, result-depth parameters, and slight rewordings; those qualitative patterns (path shifts, pivotal chapters still read) are the backdrop. The five-run battery is where magnitudes land.
Five runs of the same business question produced only 27–57% pairwise citation overlap. The same six main doctrines survived in every run. Paths varied. The load-bearing evidence region held.10
The answers were not photocopies. Emphasis still differed usefully — one run was sharpest on pilot mechanics, another on the authority boundary, another on governance structure — without abandoning the shared doctrinal core. That is material consistency with room for angle.
Why 27–57% is the interesting number, not the disappointing one
A path-identity dashboard would read low pairwise citation overlap as failure. That is the wrong reflex. If five walks of the same question produced near-total citation identity, you might have overfit a route or locked in one model’s habit. The interesting regime for route-invariant grounding is partial citation overlap combined with stable load-bearing content. Some pairs of runs shared more of their citation set; some shared less. None of that prevented the same six main doctrines from appearing every time. The citation set is path-sensitive. Doctrine survival is a first cut of evidence-region stability.
What “six stable doctrines” means operationally: you pre-declare the load-bearing doctrines a correct answer to this question must touch; a doctrine survives when it is present across the family, not when every run cites the identical page id for it. The source record for this study states that six main doctrines survived; it does not hand us a published roster of six names to reprint as a standard checklist. Honesty requires that limit. Doctrine survival remains a stronger signal than prose similarity and a coarser signal than claim-level support.
Map to the four quantities: path variance high; evidence invariance positive at doctrine level; answer invariance material with useful emphasis differences; traversal efficiency not the hero metric of this specimen. On the variance table, the family sits near row one — healthy route resilience — pending finer claim-level instrumentation.
How to read those numbers
27–57% citation overlap is path variance made visible, not a failure of grounding. Doctrine survival is the first cut of evidence invariance. It is not claim-level support comparison — that protocol is proposed below. Do not merge this range with Wallat et al.’s figure that up to 57% of citations can lack faithfulness even when correct. Different quantity, different study.
If you have walk logs, compute the analogous shape: fix one question; pin versions; run at least five walks; extract citation sets; compute pairwise overlap with one consistent definition; separately score survival of pre-declared doctrines; place the family on the variance table. The 27–57% range is this system’s telemetry for one question family. The method is portable. The constant is not a universal law.
What “six stable doctrines” means operationally: you pre-declare the load-bearing doctrines a correct answer to this question must touch; a doctrine survives when it is present across the family, not when every run cites the identical page id for it. The source record for this study states that six main doctrines survived; it does not hand us a published roster of six names to reprint as a standard checklist. Honesty requires that limit. Doctrine survival remains a stronger signal than prose similarity and a coarser signal than claim-level support — which is why the next protocol exists as proposed work, not as a fake Jaccard on the five runs.
Do not compare prose — compare claim support
Protocol status: claim-level support comparison is a proposed measurement protocol. The five-run study was executed; a completed claim-level battery on those five runs is not reported.
Raw citation lists are coarse: two answers can cite different pages that support the same claim, or cite the same pages while supporting different qualifications. Method:
- Decompose each answer into atomic claims (conclusions and material qualifications).
- For each claim, record which opened source claim supports it (page, section, assertion).
- Mark claims that float free of any opened source — prior candidates.
- Score support overlap, qualification agreement, free-floating rate, and legitimate alternates (different page ids, same assertion).
What pressures the property: stable headlines with common free-floating claims; shared citations without real support; qualifications that flip while evidence stays fixed. What does not falsify the property by itself: different page ids with equivalent source assertions — often healthy multi-entrance behaviour.
If you cannot name the supporting claim, you do not have evidence invariance. You have vibes with footnotes.
The omission test for model-prior substitution
Reflexive Agent Design already treats deliberate missing information as a duty: drop a page, delay a canonical source, remove an edge — and see whether the walk fails honestly or luckily invents.7 Route-Invariant Grounding converts that duty into a named detection target.
Counterfactual omission (proposed protocol)
Temporarily remove a pivotal page, source route or edge. Pin versions. Replay the same question family. A healthy walker takes an alternate grounded route, qualifies the answer, or confesses the absence. A walker that confidently returns the same answer with no evidence has just revealed model-prior substitution.
Protocol status: proposed, not presented here as an executed catch. The five-run study is the executed evidence; omission is the sharp test you should instrument next. Do not invent a failure transcript you do not have.
This is the disambiguator for row two. Answer holds while path and evidence vary: either the territory has genuine alternate entrances, or the model is answering from its priors and the graph is decoration. Omission is how you tell which. Keep the walker’s words as the receipt.
Change only one thing at a time. If you omit a page and redesign retrieval in the same experiment, you will not know which change produced the behaviour. Score alternate routes, qualifications, confessions, and confidence-with-empty-support. A longer path that finds a legitimate alternate entrance is the attractor working — not a failure.
Variation across models and prompts
Protocol status: proposed application of established variation doctrine to this substrate score. The five-run study is not by itself a multi-model class proof.
A beautiful walk on one frontier model is not class-level legibility. Multiply models (including a utility model that cannot compensate for bad contracts), fresh contexts, terse versus conversational prompts, vocabulary-naive agents, and boot profiles.7 Score the four quantities and the variance-table row per cell. Green answers with collapsed evidence invariance are not a pass.
Worked shape of divergence (not a fake second study): Model F shows partial citation overlap with doctrine survival — the five-run shape. Model U, on the same question family and graph pin, never opens the capstone chapters and still produces a fluent answer. That is not “Model U is dumb.” That is class-level evidence that your affordances are not legible without frontier compensation. Fix tool contracts, labels, boot order, and entrance vocabulary. Re-run until the utility model can reach the region or fails honestly under omission. If both models reach the same load-bearing region with different paths, the attractor is working across the class — keep those walks as goldens of multi-entrance health.
What falsifies class-level route-invariant grounding: only the favourite model reaches the load-bearing region; answer invariance holds via free-floating claims that a utility model cannot hide; a vocabulary-naive agent never finds the territory because entrances only exist under house jargon; boot-profile A works and boot-profile B never reaches genba with no documented reason production will only ever use A.
Historical replay is useful and conditioned on what the old walker was shown. Live evaluation under a changed retrieval design remains decisive when the payload itself has changed. Inspect family-of-walks first; replay when comparison needs pins; generate active traffic when the world has moved.
What to run this week
- Fix a small question family (one core question plus two rewordings). Record the texts.
- Pin the graph version, tool/boot profile and model id.
- Run five walks under deliberate parameter and wording variation. Save full logs.
- Score all four quantities. Plot which variance-table row you are in.
- Decompose answers into claims; map support (even manually once). Mark free-floating claims.
- Pick one pivotal page. Run the omission probe. Record alternate route, qualification, confession — or confident prior.
- Optional: one alternate model or fresh context.
- Classify Path Testing cell per walk and substrate row for the family. File the sheet as regression.
Pass shape: path may vary; load-bearing region holds; answers materially consistent; omission fails honestly; call count is not the trophy; if you ran a class check, the healthy row is not unique to one favourite model.
Fail shapes: confident prior under omission; answer holds while evidence vanishes without alternate route; evidence holds while answers contradict; total scatter; beautiful single-model walk that collapses under a second class; low path variance forced by tooling with no proof of multi-entrance reachability.
File the sheet: pins, overlap notes, doctrine or claim survival, omission outcome, Path Testing cells, substrate row, next action, link to raw logs. Re-run after graph edits, retrieval-policy changes, boot-profile changes, or model upgrades. The battery is not a one-off audit; it is how route-invariant grounding stays a property rather than a story you told once after a good week.
When the row moves from one to two, do not “fix path variance.” Run claim map and omission. When the row moves from one to three, fix synthesis. When the row moves to four, fix the question or the territory before you touch efficiency knobs.
What this piece is not
It is not the redundancy taxonomy, not the context-budget argument, not edge-directionality design, not the claim that answer invariance can conceal derivational novelty, and not walk telemetry as a janitor. Those are neighbouring doctrines. This piece owns measurement of route-invariant grounding — cleanly — so the reliability case does not get hedged into mush. Learning-sensitive memory of different derivations can arrive later without sabotaging the scoreboard. First you need to know whether the graph is load-bearing at all. A reliability score that apologises for itself mid-sentence is not a reliability score; keep the neighbouring insights for their own briefs and keep this scoreboard unhedged.
Honesty inventory
| Item | Status |
|---|---|
| Five-run study (27–57% citation overlap; six stable doctrines) | Executed evidence |
| Claim-level support comparison | Proposed protocol |
| Counterfactual omission / model-prior substitution | Proposed protocol |
| Cross-model / fresh-context battery for RIG | Proposed application |
Keep that table in your own reports. The fastest way to lose trust in a measurement doctrine is to blur proposed and executed until a sceptical reader cannot tell which sentences are receipts.
The property, once more
A knowledge substrate is robust when reasonable variations in question wording, routing parameters and traversal order still lead the agent to an equivalent load-bearing evidence region and a materially consistent answer — without depending on one prescribed path or an ungrounded model prior.
Many paths. Same genba.
Trajectory tooling made path variance cheap to record. The discipline is not to erase that variance. The discipline is to measure whether the genba still holds — and to catch the walk that never went there at all.
Open the log before you trust the answer. Score the family before you trust the substrate. Remove the pivotal page before you trust the story that everything is fine. Neighbouring doctrines will teach you what to do with redundancy, payloads, edges and rediscovery. This piece taught you how to know whether the graph is doing work — and how to catch the walk that never went to the genba at all.
References
- LangChain. “LangSmith Evaluation documentation.” docs.langchain.com/langsmith/evaluation — Agent evaluation captures full trajectories (tool selection, arguments, loops, retrieval paths, latency and cost), not only final strings.
- Scott Farrell, LeverageAI. “Reflexive Agent Design.” leverageai.com.au/wp-content/media/articles/146-reflexive-agent-design.html — Grade the path, not only the answer; bad path + right answer is unsupported lucky success (Path Testing 2×2).
- Scott Farrell, LeverageAI. “File Back the Walk.” leverageai.com.au/wp-content/media/articles/80-file-back-the-walk.html — Four instruments reading the walk record; call count is not a scoreboard; path testing grades the agent, telemetry mutates the map.
- deepset. “Measuring LLM Groundedness in RAG Systems.” deepset.ai/blog/rag-llm-evaluation-groundedness — Factual vs unfaithful hallucination; faithfulness defined against retrieved context, not parametric memory.
- Wallat, Heuss, de Rijke & Anand. “Correctness is not Faithfulness in RAG Attributions.” arxiv.org/abs/2412.18004 — Citation correctness insufficient; post-rationalisation; up to 57% of citations can lack faithfulness even when correct.
- Scott Farrell, LeverageAI. “The Wiki Playbook.” leverageai.com.au/wp-content/media/articles/176-the-wiki-playbook.html — Boot profiles and edge reachability; enter from any page; edges as load-bearing asset.
- Scott Farrell, LeverageAI. “Reflexive Agent Design” (variation protocol). leverageai.com.au/wp-content/media/articles/146-reflexive-agent-design.html — Multi-model / multi-style variation; drop a page or remove an edge and see whether the walk fails honestly or luckily invents.
- Yiqi Wang et al. “From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents.” arxiv.org/abs/2606.04990 — Final-answer accuracy alone cannot explain which evidence supported each claim; execution provenance as typed graph of an agent execution.
- Scott Farrell, LeverageAI. “RAG Was Built for Chatbots — Agents Need a Wiki.” leverageai.com.au/wp-content/media/articles/69-rag-was-built-for-chatbots-agents-need-a-wiki.html — Path Testing: grade path independently of answer.
- Primary field evidence. AskUI walk-history replay (Scott’s AWS Marketplace app; this system’s telemetry), dated replay session, 26 July 2026 — five runs of one business question; 27–57% pairwise citation overlap; six stable doctrines across runs. Source transcript: content.md (editor2).
