Leverage AI

Knowledge systems · Memory architecture

Derivational Provenance: Same Answer, Different Proof

📖 This article has an expanded ebook edition — read the full ebook.

A conclusion that already sits in the canon can still be new knowledge — because the derivation carries the warrant, the domain of application, and the future affordances. Diff derivations, not propositions.

Scott Farrell · LeverageAI · July 2026

What you leave with

Your AI keeps telling you that you have already said this. Sometimes it is right — you did say it, you forgot, and the retrieval is a gift. Sometimes it is wrong in a subtler way: you are arriving at a familiar conclusion by an unfamiliar road, and that road is the part that matters. The system matches the final sentence, declares redundancy, and quietly discards the warrant, the domain of application, and the future use that only the new path opened.

That is a matching error at the wrong level. This piece names the error, gives you a diagnostic table for telling the cells apart, and specifies what a memory system should actually write down when the conclusion looks old but the proof is new.

This piece is a correction — say that early

In Route-Invariant Grounding, the reliability argument ran one way: different wiki paths reaching substantially the same answer is a desirable property. Many paths, same genba. That remains true for reliability. You want a substrate that is not brittle to one golden route.

This article is the correction that reliability argument needs.

Answer invariance can conceal derivational novelty.

The conclusion may already exist in the canon while the path by which you have just reached it is new — and that path may change what the conclusion means, how strongly it is warranted, where it applies, or how it can be reused. For reliability, you want route-invariant grounding. For learning, you also need derivation-sensitive memory. A system that only celebrates the first will erase the second.

What this piece will not do: re-argue the five-run route study, the citation-overlap figures, or the measurement battery that 182 owns. Route into that article if you need the reliability instrument. Here we own the derivation axis only.

The reader’s question

Frame the whole problem as the person living it asks it:

My AI keeps telling me I have already said this. Sometimes it is right. How do I tell the difference, and what should the system record when it is wrong?

After this article you should be able to distinguish repetition from re-derivation, and know what to write down when a familiar conclusion arrives by an unfamiliar road.

The lived experience is almost comic in its predictability. You are mid-thought. You reach a conclusion that sounds like something already indexed. The partner, helpful as ever, says you have covered this. Sometimes that is true — you forgot, the canon is large, the retrieval worked. Sometimes you are saying the same thing for a different reason: same words, different angle, different warrant. It takes a turn or two of argument for the partner to see that this is not a path through the wiki. It is a path through thought and logic. That path changes the interpretation of the outcome, the learning, and how the idea can be used next. If the partner can see the same answer, the nuance gets overlooked.

Matching at the wrong level

When a system says “you’ve already discussed this,” it is usually matching at the level of the proposition:

Existing claim ≈ current claim

That comparison can be correct and still incomplete. The same proposition can differ by any of these axes — and more than one can move at once:

None of those differences is decoration. Each one changes what the claim can do next week: which transfers are licensed, which objections are already answered, which experiments would falsify it, which neighbouring claims it now pulls into the same room.

A memory system that diffs only propositions will treat all of those as “already have it” and compress them away. That is how a helpful thinking partner performs a live version of a failure we already named for offline maintenance: hallucinated consolidation — merging distinct ideas into one falsely confident claim because they look old and adjacent.1 When a janitor does it overnight, you can at least review a diff. When a partner does it mid-conversation, it arrives as assistance. That makes it harder to notice and easier to accept.

The diagnostic table

Here is the table that should sit beside every “you already said this” detector:

Final proposition Derivation Meaning
Same Same Probably repetition
Same Different New warrant, lens, mechanism or transfer path
Different Same evidence Synthesis or interpretation changed
Different Different Genuine unresolved divergence

Walk every row for what a memory system should actually do.

Row one — same proposition, same derivation

This is genuine repetition. You said it before, for the same reasons, with the same evidence class, and nothing material moved. The correct system behaviour is almost boring: point at the existing claim, offer the prior context, and do not mint anything new. This is where “you already said this” earns its keep. Forgetting is real. Large canons make it more real. A good partner that recovers the prior turn is doing Working Fidelity’s return job — handing the thought back so you can continue rather than re-derive from a cold fossil.2

What the system records: a retrieval event, maybe a last-accessed stamp. Not a new concept. Not a rival page. Not a derivation receipt that pretends novelty where there is none.

Row two — same proposition, different derivation

This is the heart of the piece. This is the row naive deduplication compresses away.

The final sentence matches something in the canon. The road does not. Something in the derivation changed: observation, mechanism, evidence class, lens, domain, boundary, independence, or join. The claim is not “new” as a slogan. It is newly warranted, newly scoped, or newly connected.

What a naive system does: match proposition → declare redundancy → discard the new road. What a derivation-sensitive system does: keep the existing claim page as the address, and attach a receipt that records how the claim was re-reached, when, from which bronze event, and what that changes about future use.

Dwell here. The failure is not that retrieval works. The failure is that retrieval works so well that the partner can find the old proposition and stop thinking. Compression becomes the dominant interaction pattern precisely when personal and institutional canons get large enough for the model to search them. That is a new problem created by success.

Sibling pieces in this run address neighbouring compression problems: redundancy as error correction rather than pure bloat; novelty-preserving carve-outs in payload design; inbound edges as a different question from outbound ones. Those are real axes. This axis is different. It is not about how many pages say a thing, or which reverse edges arrive, or how much of the bloom payload is novel. It is about whether the thought that produced the proposition is the same thought as last time.

Same answer, different proof.

Row three — different proposition, same evidence

The exhibits did not change. The synthesis did. Two honest readers of the same bronze can diverge. Here the system should not pretend the second reading never happened, and should not overwrite the first. Contradiction-as-edge discipline applies: keep disagreement as structure rather than averaging it into a smoother paragraph. The memory action is to record that interpretation shifted on shared evidence — a synthesis note, a contested edge, a supersession candidate if one reading legitimately replaces the other.

Row four — different on both

Different proposition, different derivation: genuine unresolved divergence. Do not force a merge. Do not declare one side “already covered” by the other. Hold both, with provenance, until evidence or decision resolves them — or until they remain permanently as a mapped tension.

If your system only has two behaviours — “duplicate of existing” and “brand new concept” — rows two through four will be misclassified. Row two becomes silent discard. Row three becomes a fight. Row four becomes either noise or a fake consolidation. The table is the minimum instrument that makes those mistakes visible.

The provenance family — and the missing fifth

You already have several provenance concepts in working use. Name them carefully; do not invent near-synonyms.

None of those four records the thing this failure erases.

Idea Provenance tells you who said it and when. It does not tell you why the thought is now warranted differently. Evidence provenance tells you which exhibits support the claim. It does not store the chain of tensions, analogies and pivots that made the claim meaningful in this moment. Cognitive Provenance reconstructs what the agent observed. That is essential for auditing agent cognition, and still not the same object as a human re-deriving a familiar conclusion through a new operational loop. Walk provenance records the pages traversed. A path through the wiki is not automatically a path through thought.

What is missing is the fifth member:

Derivational Provenance

The preserved chain of observations, tensions, analogies, pivots, evidence and logical moves through which a thought became meaningful — even when its final proposition resembles an existing claim.

Or more simply: same answer, different proof.

This is already foreshadowed by Working Fidelity. Preserve only a conclusion and you produce a fossil: you can cite it; you cannot re-enter it; to extend it you must climb the hill again.2 A useful return needs not only the right conclusion but enough of the construction to continue thinking from inside it.6 Derivational Provenance is the graph-level form of that insight. The construction has to survive not only for personal re-entry, but as a typed relationship a system can reason about — so a later agent does not “helpfully” delete the part that changed.

Semantic Refraction supplies a neighbouring discipline on a different axis: one fixed bronze event can support several legitimate, lens-qualified implications without becoming several rival facts.7 This article fans out on the derivation axis instead of re-arguing that thesis. The reuse allowance is the metadata shape, not the whole argument: lens, owner, date, pointer, status, conditions. A derivation receipt should imitate that shape so it is structured like a claim, not a vibe.

Two worked pairs from a real canon

Doctrine without specimens is a slogan. Here are two pairs grounded in the same operational material: a practitioner reviewing walk paths, A/B-testing retrieval ideas, rejecting variants, mutating design, and watching a thinking partner run analysis code in the breadth of a conversation.

Pair A — genuine repetition (same proposition, same derivation)

Proposition (rough form): walk telemetry and history-backed analysis improve how you judge retrieval designs.

First occurrence: the claim already exists as written doctrine — File Back the Walk and related retrieval instrumentation — as paper doctrine in the canon. The derivation is the published argument: walks are durable artefacts; instruments read them for different questions; the map should improve from use.

Second occurrence that is truly the same: later, in conversation, you restate the same claim by rehearsing the same argument structure — same causal story, same evidence class (doctrine / prior writing), same domain (wiki walk instrumentation). No new observation loop. No new reject chain. No new operational measurement.

What the system should record: a pointer to the existing doctrine page. Optionally “retrieved on date D.” No new concept. No observed-in-operation upgrade. No independent-derivation receipt. This is row one. Saying “you already covered this” is correct behaviour.

Pair B — same conclusion, different derivation (the row that gets compressed)

Proposition (rough form): still recognisable as “walk telemetry improves retrieval design.” The final words can look like Pair A.

New derivation: you are not rehearsing the paper. You are watching the whole mechanism operate inside minutes:

idea → historical walks → analysis code → measurement → rejected idea → changed measure → mutated design → implementation → replay.

Ideas get tried. The partner says some of them stink. A-B tests prove some of them stink. Ameliorations fail. A different cut — bloom only the top convergent results; surface reverse relationships the outbound lines cannot show — survives. The loop is not “I remember I wrote about this.” The loop is “I just saw the instrument run, reject, mutate, and re-run.”

The proposition may resemble existing doctrine. Its evidence class, experiential origin and operational meaning changed. Paper doctrine became observed-in-operation. Rejected variants became part of the warrant (what failed, under which measure). The design that survived is not the first proposal; the trail is the proof.

What a naive system records: “already in canon — File Back the Walk / walk telemetry improves retrieval.” Conversation compressed. Learning discarded.

What a derivation-sensitive system records: keep the existing claim address. Attach a receipt: observed-in-operation (and likely extends-by-mechanism), dated to the session, pointed at the bronze transcript or walk-analysis record, with a short new-route summary and a note that evidence status upgraded from paper doctrine to operational observation. Optionally an independently-derived-from edge if the operational loop re-reached the claim without merely quoting the page.

That is row two, fully walked. Same answer. Different proof. The system’s job is not to invent a second concept called “walk telemetry (again).” Its job is to make the second proof queryable.

One downstream use only the new derivation opened

Why does the receipt matter beyond bookkeeping? Because derivation differences change future affordances.

In the operational pair above, the new route exposed design moves that pure paper doctrine did not force into the room in the same way: harsh exponential thresholds that sounded rigorous and failed against the actual mass distribution; a top-K bloom that treats convergence mass as an attention convenience rather than a licence to re-print everything; inbound residual as a stabilising question — what is pointing at us that we have not surfaced? Those are not hypothetical transfers. They are the contents of the derivation chain itself: rejected ideas, changed measures, mutated design, implementation, replay.

Once the claim is only stored as the old proposition, a later agent cannot answer: “Under what failed measure was the exponential idea rejected?” “What boundary did the top-K cut protect?” “Why is reverse-edge residual part of the warrant rather than a nice-to-have?” Those questions are only answerable if the derivation survived. A companion piece on novelty-preserving payload design develops the carve-out economics of not re-printing salience; another develops inbound edges as a different epistemic question. Neither replaces the need to keep the derivation that made those moves live.

The portable test: if a later transfer, boundary condition, or rejection reason would be invisible after proposition-only storage, you are looking at row two — and silent discard is an active deletion of future capability.

The receipt, not a duplicate concept page

When the system detects “I’ve said this before but the derivation differs,” do not mint a rival concept page. Attach a receipt to the existing claim. Imitate the implication metadata discipline already used for lens-qualified readings: lens/type, owner, date, pointer, status, conditions.7

Receipt schema (implementable)

receipt_type: one of independently-derived-from | observed-in-operation | extends-by-mechanism | lens-qualified-implication | derivation-note | evidence-status-upgrade

target_claim_id: the existing canon claim or concept page this receipt strengthens (not a new slug for the same idea).

source_bronze_pointer: URI or durable id of the bronze record — transcript segment, session id + turn range, commit, walk-log batch — the original source the derivation came from.

derived_at: ISO-8601 timestamp of the re-derivation event (when the thought was reached).

recorded_at: ISO-8601 timestamp when the receipt was written (may differ if written in a later review).

owner: human or agent identity standing behind the receipt.

status: candidate | accepted | superseded.

prior_route_summary: how the claim was previously warranted (one short paragraph or structured bullets).

new_route_summary: observation → mechanism → evidence class → domain/boundary in the new path.

what_changed: one or more of warrant | evidence_class | domain | boundary | transfer | independence | join.

conditions: when this receipt applies; what would reopen or falsify it.

evidence_status_after: intuition | paper_doctrine | observed_in_operation | measured | contested.

Dating is not optional. Without derived_at, chronological stacking cannot place the re-warrant in time. Without source_bronze_pointer, the receipt becomes storytelling detached from exhibits — the opposite of evidence-package discipline. The pointer must open the bronze, not merely name a gold page that summarises it. Gold addresses reality; bronze holds what happened. The receipt is the join that keeps both honest.

Edge types in practice:

Strengthened evidence status is a first-class outcome. Moving a claim from “written doctrine” to “seen operating end-to-end” is not a second concept. It is a status change with a dated trail.

The better conversational behaviour

Schema without dialogue still fails in the room where the damage happens. The portable artefact is the partner’s line:

The conclusion resembles something already in the canon, but the route is new. Previously you reached it through X; today you reached it through Y, which changes Z.

That beats both “you already said this” and falsely declaring every rediscovery a new concept. It forces the three fields that matter: resemblance (not identity), prior route, new route, and the delta (Z). Z is where the receipt’s what_changed comes from. If the partner cannot name Z, it should not claim the routes differ — and it should not claim they are the same.

The two-turn lag is an engineering smell. If it routinely takes human pushback for the system to notice path-through-thought, the default match is still proposition-only. Raise the default: when proposition similarity is high, run a cheap derivation check before emitting the collapse line. Compare observation, mechanism, evidence class, and domain tags if you have them. If they are missing, ask one clarifying question instead of declaring redundancy.

What not to do — and where this piece stops

Do not re-argue route invariance as if reliability were the enemy. It is not. Do not expand into Semantic Refraction’s full thesis about multi-lens implications over one bronze event — reuse the metadata shape; leave the axis. Do not expand into corporate deliberation capture; that is a different sub-topic in this publishing run. Do not invent false-collapse rates or industry percentages you did not measure. Write the shape of the experience and the schema of the fix.

Related published pieces you can route into without retelling them: Route-Invariant Grounding for the reliability instrument this corrects; Wiki Redundancy Is Error Correction for when multiple entrances are health rather than bloat; The Novelty-Preserving Carve-Out for payload novelty vs repeated salience; Inbound Edges Are a Different Question for reverse-edge epistemology.

How to run the table in a real week

Doctrine without a weekly habit dies. Here is a minimal operating rhythm that does not require a full graph rewrite on day one.

During conversations: when the partner emits a collapse line, pause. Ask which cell you are in. If you believe you are in row two, require X, Y and Z before anyone files the moment as redundancy. If Z is empty, accept row one and move on. If Z is real, write a candidate receipt before the session ends — even as a short structured note.

After operational loops: whenever you watch a mechanism run end-to-end — reject, mutate, replay — check whether the final proposition already existed as paper doctrine. If it did, you are in Pair B territory. Attach observed-in-operation and capture the reject trail that became part of the warrant.

Weekly lint: receipts without bronze pointers; collapse events without a derivation check; candidate receipts that never got accepted or rejected; near-duplicate concept pages that should have been receipts on one address. Thirty minutes is enough if the fields exist.

This rhythm invents no industry percentage. It sequences the work so classification and write behaviour land before automation. Automating collapse without classification only accelerates the failure mode.

What “path through thought” looks like in logs

Walk provenance will show page opens and edge traversals. Derivational provenance may show almost none of that. The expensive part of Pair B was not necessarily a long wiki walk. It was the chain of proposed ideas, failed measures, mutated design and replay — much of it in conversation and code. If your only instrument is page telemetry, you will keep missing the cheapest re-derivations: the ones that re-warrant doctrine without reopening the gold page.

That is why the receipt points at bronze session segments, not only at gold doctrine pages. Gold addresses reality. Bronze holds what happened. The receipt is the join. Without the join, later agents can find the slogan and still be unable to answer which measure killed which idea.

Building evaluation pairs without inventing a canon

Take moments when you disagreed with “you already said this.” Reconstruct honestly. If the derivation was the same, keep it as row-one training — the partner was right. If the derivation differed, write the receipt you should have written and keep the pair as row-two evaluation. Grade partners on whether they ask for Z, attach a receipt, or falsely collapse. A suite with only restatements and brand-new ideas will look green while still erasing learning.

A filled receipt for Pair B (so the schema is not abstract)

Using the operational pair above, a concrete receipt looks like this — not as invented telemetry numbers, but as structured bookkeeping over the real loop:

Example (illustrative field fill)

receipt_type: observed-in-operation; extends-by-mechanism

target_claim_id: the existing walk-telemetry / File Back the Walk doctrine claim

source_bronze_pointer: the thinking-session segment that held the idea→reject→mutate→replay loop (openable bronze, not only the gold doctrine page)

derived_at: date-time of the operational loop

recorded_at: date-time the receipt was written (may be next-day review)

prior_route_summary: paper doctrine — walks as durable artefacts; instruments improve the map from use

new_route_summary: live harness loop; exponential thresholds tried and failed; top-K bloom and inbound residual survived; implementation and replay closed the loop

what_changed: evidence_class; mechanism; boundary

evidence_status_after: observed_in_operation

status: candidate until a human accepts

If you cannot fill what_changed, you are not in row two. If you cannot open source_bronze_pointer, you have a story, not a receipt. Those two checks catch most theatre.

Close

The same destination does not make two journeys cognitively equivalent.

When your partner says you have already said this, run the table. If it is row one, thank the retrieval. If it is row two, refuse the collapse and write the receipt: type, target claim, bronze pointer, derived_at, prior route, new route, what changed. Keep the concept page. Strengthen the warrant. Leave the future affordances queryable.

Reliability wants many paths to the same genba. Learning wants the paths themselves to remain first-class. Build for both — or your memory will keep erasing the part that changed.

References

  1. Scott Farrell / LeverageAI. “The Index Is the Data” (ebook), ch.6 — Hallucinated consolidation: a janitor under a loose North Star can merge two genuinely distinct ideas because they are old and adjacent. Cite key #bdc832. leverageai.com.au
  2. Scott Farrell / LeverageAI. “The Clasp” (ebook), ch.1 — Working fidelity vs fossil: writing often preserves the destination and discards the path; re-entry requires construction, not only conclusion. Cite key #bca011. leverageai.com.au
  3. Scott Farrell / LeverageAI. “Idea Provenance” (ebook), ch.6 — Organisational idea attribution by timestamped, attributable exhibits rather than status-weighted recollection. Cite key #65d1b1. leverageai.com.au
  4. Scott Farrell / LeverageAI. “The Model Is Not the Memory” (ebook), ch.3 — Cognitive provenance: reconstruct exactly which pages, claims and edges an agent observed at decision time, external and version-pinned. Cite key #aca00a. leverageai.com.au
  5. Scott Farrell / LeverageAI. “File Back the Walk” (ebook), ch.11 — Walk records serve multiple instruments; path testing, telemetry and execution provenance ask different questions of the same artefact. Cite key #2bcb76. leverageai.com.au
  6. Scott Farrell / LeverageAI. “The Clasp” (ebook), ch.6 — Capture without timely working-fidelity return is an archive; right fidelity on the return path is re-enterable construction. Cite key #b5da47. leverageai.com.au
  7. Scott Farrell / LeverageAI. “Semantic Refraction” (ebook), ch.6 — Lens-qualified implications carry lens, owner, date/as-at, source pointer, status and conditions without forking bronze into rival facts. Cite key #d63a05. leverageai.com.au