Two Ladders, One Climb: Tier-Three Retrieval Is Rung-Three Value
The three tiers of retrieval and the three rungs of AI value turn out to be the same ladder, climbed from two sides — and once you see why, you can tell which AI features win by volume, which win by existence, and how to put a dial on the ones that matter.
The short version
- Lookup → search → unsolicited corroboration lines up, rung for rung, with Don't Compete → Augment → Transcend. The two ladders are the same ladder.
- The alignment is causal, not cute: tier-three retrieval is rung-three value because it was unpurchasable before cheap cognition — it made a nonexistent capability exist, not an old one cheaper.
- So rung-one AI invites comparison; rung-three AI has no comparator. You measure it not by a benchmark it can't have, but by frequency — tier-three events per week — and its real driver is edge density, not model quality.
The single most powerful thing I've used AI for isn't writing code or drafting proposals. It's that the system has started answering questions I didn't ask. I'll be talking through something half-formed — not querying anything, just thinking out loud — and it comes back with a document I didn't know existed. An email from years ago. A blog post I forgot I wrote. Not proving me right, not proving me wrong. Just volunteering a piece of my own past that turned out to matter.
And here's the part I want to sit with: it's happening more often. That started as a vibe and I've come to treat it as a signal, because — as I'll get to — my own architecture predicts exactly that trend, and the trend is where the metric hides.
I've already named that third kind of retrieval. In the Idea Provenance work I called it the third tier — above lookup and above search — and filed the concept as unsolicited corroboration. I'm not going to re-derive the tiers here; that piece does it. What I want to do here is something that piece deliberately left on the table: put a price on the tier, and put a dial on it. Because the day I drew a parallel between that retrieval model and a value framework I'd written years earlier, the two collided into a single claim — and the claim is useful.
1. Two ladders I drew by accident
The parallel arrived mid-sentence, which is the honest way most of these arrive. I was explaining why a RAG search feels limited. A RAG search answers a question you already know how to ask. You have to know the words, because it's matching against your words. That's genuinely useful, and it's also — let me be blunt — something we could mostly already do. Google indexed the web. We've been able to full-text search our own Gmail for fifteen years. Read the page faster, find the page faster: fine, but not new in kind.
Then I remembered a framework I'd written — the Cognition Ladder. Three rungs. Rung one, AI doing in real time what humans already do well, which is the failure-prone rung. Rung two, batch augmentation — ten to a hundred times more analysis on problems you already have. Rung three, the things we never even attempted because the coordination or cognitive overhead was prohibitive. The uneconomical tier. The stuff that was off the table entirely, not just expensive.
And the two ladders line up. Not loosely — rung for rung.
| Retrieval tier | Value rung | What it does | Its relationship to a human |
|---|---|---|---|
| Lookup — you know the artefact exists | 1. Don't Compete | Fetch the thing you already know to fetch | Duplicates a capability we already had (a filing cabinet, a URL) |
| Search — you have the question, not the answer | 1–2. Compete / Augment | Match your query against your own words, faster and wider | Google did it; full-text Gmail did it; RAG does it marginally better |
| Unsolicited corroboration — you have neither, only a belief | 3. Transcend | Walks the graph unprompted and volunteers a record you never requested | No human, at any salary, ever did this for you |
When I first noticed the alignment I thought it was a nice coincidence — two of my frameworks rhyming. It isn't a coincidence. That's the whole point of this piece.
2. The hinge is unpurchasability
Here's why the ladders are the same ladder and not just two things drawn with three boxes each.
Lookup and search are rung one because they duplicate a capability that already existed. When you can already do a thing, and AI does it a bit better or a bit faster, you are — by definition — competing with the incumbent way of doing it. Sometimes you win by thirty percent. Sometimes you lose. It's always arguable.
Tier-three retrieval is rung three for the opposite reason. It doesn't compete with anything, because the thing it does had no prior version. You could not previously hire a person to hold the entire exhaust of your business life in mind — every email, every project, every repo, every proposal — and then, mid-conversation, volunteer a forgotten record because it happens to bear on a belief you just voiced in passing. That capability wasn't expensive. It was unpurchasable. No salary bought it, because no human mind holds that much at that resolution and walks it on demand.
The cost-of-cognition collapse didn't make an old thing cheaper. It made a nonexistent thing exist.
That's the causal link. Rung three is defined as work that was structurally impossible before AI removed the ceiling — and unsolicited corroboration is precisely a capability that was structurally impossible before. The two ladders align because they're measuring the same variable from two directions: the retrieval ladder measures how much you knew to ask for, and the value ladder measures how much of the capability existed before. At the top of both, the answer is: nothing. You knew to ask for nothing, and nothing like it existed to buy. That's not two facts. It's one fact wearing two labels.
It also explains a feeling. Rung-three results always feel slightly unreasonable — freakish, even — and now I know why. Your intuitions about what's reasonable were calibrated on a world where this was impossible. Of course it feels like too much. Your sense of "too much" was trained on the old ceiling.
3. Rung one invites comparison. Rung three has no comparator.
Once you see the hinge, a diagnostic falls straight out of it, and it's the most practically useful thing in this whole essay.
Rung-one AI invites comparison; rung-three AI has none. A RAG chatbot gets benchmarked against the search box it replaced. It's thirty percent better, occasionally worse, always arguable — and the endless tuning is just the cost of fighting for marginal wins against a capability that already worked. The tier-three wiki gets benchmarked against nothing, because "the system that corroborates beliefs you voiced in passing" had no prior product, no prior price, no prior expectation to be measured against. There's nothing on the other side of the scale.
This is why the genuinely good AI projects are so hard to sell up the chain, and it's worth naming the mechanism plainly: rung three is illegible to a job description. When a company writes a role or a requirement for "AI," it stops at RAG — because rung one is legible. It maps onto a workflow someone can already name: "search our documents, but with a chatbot." You cannot write a requirement for a capability you have never seen. Rung three isn't just missing from the spec because the author is behind; it's missing by definition, because a spec is a description of a known thing, and the whole nature of rung three is that it was unknown. The vocabulary is fossilised at rung one. The value is one rung up, in a language the org chart doesn't have a word for.
4. The proof pair, with receipts
I don't want this to stay abstract, so here is the argument's whole case made of two events from a single afternoon, on a single wiki. Same tool, same session, two different rungs — and you can feel the difference in kind, not degree.
The CV rewrite. I had the system rewrite a CV against my history. It did it well and fast, drawing on more of my material than I'd have pulled by hand. That is a real win. It is also a rung-two win: a human CV writer could have done it. AI did it faster and wider — a volume win. Useful, defensible, and entirely comparable to the human way of doing it.
The Joel email. In the same conversation I was chasing a memory of a bug I'd shipped years ago — a stand-pat problem in some chess-engine work, the kind of thing I've written about separately. I asked the system to find my recollection of the bug. What it came back with instead was an email from a colleague, Joel, warning me about that exact bug before I shipped it. I didn't know the email survived. Nobody could have searched for it, because "you don't want to be failing high on every capture" and "stand-pat bug" share no vocabulary at all. There was no query that finds that. The system had to generate the connection from the neighbourhood of the graph, off a belief I'd only spoken aloud.
No human, at any price, finds a twenty-year-old warning about a bug in an archive nobody knew to search, triggered by a sentence said in passing.
That's the rung-three win. Not faster than a human — impossible for a human. An existence win, not a volume win. Two events, one afternoon, one wiki, and the felt difference between them is not "one was quicker." It's "one had a human alternative and one had none." That felt difference is the ladder.
And I'll be honest about how thick the receipts got, because it's almost funny. This conversation ran tier three on itself, twice. The Joel email was one instance. The other: I'd asked, half-embarrassed, "have I already blogged about this pattern?" — and the answer arrived as an unprompted hint at the bottom of a search result, pointing at a chapter I'd written, titled with the very pattern I was describing. The system corroborated my writing about the system corroborating me. I couldn't have staged that if I'd tried.
5. Price it — then put a dial on it
So the tier is named and the value is priced. The last move, the one I actually care about, is turning it into something you can watch on a dashboard. Because a value you can't measure loses every budget fight to a value you can, even when the measurable one is worth less.
Start with the observation I opened on: it's happening more often. My architecture predicts that, and the prediction is the metric in disguise. Tier-three's hit rate is a function of edge density, not model quality. A sparse graph — a few documents, few relationships — walking out from wherever you're currently standing reaches nothing worth volunteering. There's simply no neighbour to surface. But every package you ingest, every new edge you draw, widens the neighbourhood one hop from wherever you happen to be. So unsolicited corroboration should be rare at first and increasingly routine as the corpus fills in — which is exactly the trend I'm reporting. I appear to be living on the bendy part of the curve.
That gives you a metric, and I want to be careful about what kind of claim it is. Here is the spec:
What counts as a tier-three event
A retrieval where no query was issued for the thing returned. You supplied a belief, a topic, or an aside — not a question whose answer is the surfaced artefact — and the system volunteered a record you did not know to ask for and, often, did not know existed. If you could write the search that would have found it, it wasn't tier three. The test is falsifiable in both directions: if there's a query that retrieves it, downgrade it; if there genuinely isn't, count it.
Why edge density is the driver, not the model
A better model reasons better over what it reaches. It does not change what it reaches. Reach is a property of the graph — how many relevant things are one or two hops from your current position. You can hold the model constant and watch tier-three frequency climb purely by ingesting more of your life and drawing more edges. That's why "buy a smarter model" is the wrong lever here and "ingest more, connect more" is the right one. The dial you turn is density.
Now the honesty flag, because the whole point of this system is that it won't flatter you and I shouldn't either: I'm not going to quote you a hit rate. I don't have a measured one, and inventing a number here would betray the exact discipline the piece is arguing for. What I have is a shape and a prediction: the count of tier-three events per week, plotted over time, should bend upward as edge density rises — and if you grow the graph and that curve stays flat, my claim is wrong. That's the deal. A falsifiable prediction about my own system, offered as a proposed instrument, not a benchmark I've already run.
6. A sketch dashboard
If I were building the instrument — and this is a sketch, not a shipped product — three things would go on it.
Proposed — a tier-three dashboard
- Tier-three events per week (the headline curve). A running count of retrievals that met the no-query-was-issued test. The number matters less than the slope; you're watching for the bend.
- Edge density (the leading indicator). Nodes, edges, and edges-per-node in the graph over time. If the headline curve is the outcome, this is the cause — and it should move first. When density climbs and, a little later, tier-three events climb, you've got your mechanism on film.
- Distinct workflows served by one asset. Not a tier-three metric exactly, but the terminal-value tell: count how many different jobs the same corpus quietly did this month, none of which you built it for. In the afternoon I described, one wiki grounded a CV, settled a provenance question, surfaced a colleague's old warning, and retrieved two of my own frameworks on demand. That tally is the reuse-across-unknown-futures claim, witnessed rather than asserted.
Treat all three as proposed. I'm describing the shape of the instrument and the relationship it should reveal, not handing you calibrated numbers. The point is that the illegible tier — the one with no comparator — is not actually unmeasurable. It just has to be measured by its frequency and its driver instead of by a head-to-head it structurally can't have.
Where this leaves your AI strategy
If you take one thing from this, take the classifier, because it will save you money on the next AI decision you make. Ask of any AI feature: does it lose to a human, win by volume, or win by existence? Rung one competes with people on their home turf and mostly loses. Rung two wins by doing far more of what a human could have done — real, defensible, and where most honest ROI lives today. Rung three wins because there is no one on the other side of the table at all.
The trap is that rung one is the intuitive place to point AI — make the existing thing faster — and it's the low-value rung. That's the horse: you feed money and strategy and attention into making an existing workflow go a bit quicker, and you've built nothing that compounds. There's no terminal value in a faster horse. The compounding asset is the corpus that answers the question you didn't know to ask — and, unlike the horse, it carries over into every future workflow you haven't designed yet. This piece is really just the sharp end of my Terminal Value Doctrine: put your strategy tokens on the thing that survives, not the thing you'll turn off.
So build the asset. Feed it your whole exhaust, not a tidy slice. Draw the edges. And then, once it starts volunteering things you never asked for — and it will, more and more often — count them. Watch the curve bend. That bend is the most valuable thing your AI does, and until now nobody was even keeping score.
No horse got faster. The rider got a library.
Scott Farrell writes on AI strategy, knowledge architecture, and the economics of building compounding intellectual assets at LeverageAI. If your corpus has started answering questions you didn't ask, the thing worth doing next is instrumenting it — counting the tier-three events and watching what edge density does to the curve.
